Neural network computing system and device

Sparse data is converted into time pulses and frequency signals through digital time conversion and frequency conversion modules, and multiplication and accumulation calculations are performed in combination with phase accumulators, which solves the problems of large resource overhead and high energy consumption in sparse data processing and achieves efficient sparse data processing and energy efficiency balance.

CN120764599APending Publication Date: 2025-10-10PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510909664.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing neural network computing strategies have high resource overhead and energy consumption when processing sparse data. Traditional sparse optimization methods are difficult to adapt to changes in sparsity, resulting in reduced hardware resource utilization and low inference efficiency.

Method used

The digital time conversion module and the digital frequency conversion module are used to convert sparse data into time pulse signals and frequency signals, and multiplication and accumulation calculations are performed through the phase accumulator. Combined with sparse perception and mixed precision support, efficient processing of sparse data can be achieved.

Benefits of technology

It improves computing throughput, fully utilizes hardware resources, reduces energy consumption, and takes into account both computing accuracy and energy efficiency. It is suitable for resource-constrained edge environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120764599A_ABST
    Figure CN120764599A_ABST
Patent Text Reader

Abstract

The invention discloses a neural network computing system and device, and relates to the technical field of neural networks, the neural network computing system comprises a digital time conversion module, a digital frequency conversion module and a phase accumulator, and the output end of the digital time conversion module and the output end of the digital frequency conversion module are connected with the input end of the phase accumulator; the digital time conversion module is used for compressing and converting sparse data in the received input data to obtain a corresponding time pulse signal and outputting the time pulse signal to the phase accumulator; the digital frequency conversion module is used for converting the received weight data into corresponding frequency signals and outputting the frequency signals to the phase accumulator; and the phase accumulator is used for carrying out multiply-accumulate calculation on the time pulse signal and the frequency signal based on the phase and outputting a calculation result. According to the method, the technical problems of high resource overhead and high energy consumption of a neural network calculation strategy aiming at sparse data at present are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of neural network technology, and in particular to a neural network computing system and device. Background Art

[0002] With the rapid development of deep learning technology, the structure of neural network models has become increasingly complex, with deeper layers, larger parameter sizes, and higher computational density, placing higher demands on computing resources and storage systems. In mainstream models such as the Transformer, in particular, the massive matrix operations and nonlinear activation operations further exacerbate the demand for computing power and energy efficiency. Furthermore, neural networks generally exhibit significant sparsity during actual reasoning. This sparsity primarily stems from optimization methods such as activation functions, attention mechanisms, and pruning, manifesting as a large number of weights or activation values ​​being zero, resulting in redundant computation and storage.

[0003] Sparse data exhibits non-stationary and unpredictable characteristics. To mitigate the impact of sparse data, traditional acceleration methods that rely on static sparse graphs or predefined pruning structures are unable to effectively adapt to runtime sparsity changes, easily leading to reduced hardware resource utilization and even impacting inference efficiency and energy efficiency. Furthermore, processing based on purely digital or analog architectures incurs high computational resource overhead and energy consumption, making them difficult to apply in resource-constrained edge environments.

[0004] The above information disclosed in this Background section is only for understanding the background of the present invention and therefore it may contain information that does not constitute prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a neural network computing system and device, aiming to solve the technical problems of large resource overhead and high energy consumption in current neural network computing strategies for sparse data.

[0006] To achieve the above objectives, the present application provides a neural network computing system, which includes a digital time conversion module, a digital frequency conversion module, and a phase accumulator, wherein the output ends of the digital time conversion module and the digital frequency conversion module are respectively connected to the input end of the phase accumulator;

[0007] The digital time conversion module is used to compress the sparse data in the received input data, and convert it into a corresponding time pulse signal to output to the phase accumulator;

[0008] The digital frequency conversion module is used to convert the received weight data into a corresponding frequency signal and output it to the phase accumulator;

[0009] The phase accumulator is used to perform multiplication and accumulation calculations on the time pulse signal and the frequency signal based on the phase, and output a calculation result.

[0010] In one embodiment, the digital time conversion module includes at least a synchronous counter, a clock source connected to the synchronous counter, and a sparse detection module;

[0011] The synchronous counter and the clock source are used to provide a reference clock signal to the sparse detection module as a reference for time encoding;

[0012] The sparse detection module is used to identify sparse data with an input of 0 in the input data and compress the sparse data in the time encoding stage to output a corresponding time pulse signal.

[0013] In one embodiment, the digital frequency conversion module includes at least one digitally controlled oscillator, which is used to convert the input weight data into a corresponding frequency signal, wherein each weight data corresponds to an output frequency to maintain periodic oscillation.

[0014] In one embodiment, the digitally controlled oscillator includes at least a plurality of inverters connected in sequence, wherein the pull-up path and the pull-down path of the inverters are respectively connected in series with a controlled current source, and the digitally controlled oscillator is used to adjust the oscillation frequency by regulating voltage or digital control signal.

[0015] In one embodiment, the phase accumulator comprises at least a phase calibration unit, a multiplication unit, a phase accumulation unit, and a data readout unit connected in sequence;

[0016] The phase calibration unit is used to align the phases of the time pulse signal and the frequency signal;

[0017] The multiplication unit is used to multiply the phase-aligned time pulse signal and the frequency signal to obtain a product result, and process the phase-aligned time pulse signal and the frequency signal according to the exclusive OR logic to determine the accumulation direction of the product result;

[0018] The phase accumulation unit is used to phase-modulate the phase-aligned time pulse signal and the frequency signal according to the accumulation direction and accumulate them in time to obtain a multiplication and addition operation result;

[0019] The data readout unit is used to read the multiplication and addition operation result, and convert the multiplication and addition operation result into a calculation result in a target format through a preset digital time conversion interface.

[0020] In one embodiment, the phase accumulator further includes a signal jump protection unit, which is connected to the phase accumulator unit and is used to latch the current multiplication and addition result when the time pulse signal or frequency signal jumps until the operation is resumed after the signal stabilizes.

[0021] In one embodiment, the jump protection unit includes a first delayer, a second delayer and a logic XOR gate connected in sequence; the first delayer is used to receive the input time pulse signal and frequency signal, and output an intermediate processing signal after delay processing, and input the intermediate processing signal to the second delayer, and then the second delayer delays the intermediate processing signal and outputs the delayed signal to the logic XOR gate; the logic XOR gate is used to perform XOR processing on the input time pulse signal and frequency signal with the corresponding delay signal, and output a jump control signal, and the jump control signal is used to represent the jump interval, so as to latch the multiplication and addition operation result of the intermediate processing signal in the jump interval.

[0022] In one embodiment, the neural network computing system further includes a data segmentation module and a weight segmentation module, wherein the data segmentation module is connected to the digital time conversion module, and the weight segmentation module is connected to the digital frequency conversion module;

[0023] The data segmentation module is used to split the input data into high-order data and low-order data based on a preset data splitting rule, and input the data into the digital time conversion module for independent coding conversion to obtain a high-order time pulse signal and a low-order time pulse signal respectively;

[0024] The weight segmentation module is used to split the weight data into high-order weights and low-order weights based on preset weight splitting rules, and input them into different digitally controlled frequency oscillators of the digital frequency conversion module for independent encoding conversion to obtain high-order frequency signals and low-order frequency signals.

[0025] In one embodiment, the phase accumulator includes a rectangular computing array consisting of multiple computing units, each computing unit is respectively connected to a digital time conversion module and a digitally controlled frequency oscillator in the digital frequency conversion module, each row of computing units in the rectangular computing array shares the same time pulse signal, and each column of computing units in the rectangular computing array shares the same frequency signal, and each of the computing units is used to perform multiplication and accumulation operations based on the received time pulse signal and frequency signal.

[0026] In addition, to achieve the above-mentioned purpose, the present application also provides a neural network computing system device, which includes the neural network computing system described above.

[0027] The present application provides a neural network computing system, which includes a digital time conversion module, a digital frequency conversion module, and a phase accumulator, wherein the output ends of the digital time conversion module and the digital frequency conversion module are respectively connected to the input end of the phase accumulator; the digital time conversion module is used to compress the sparse data in the received input data, convert it into a corresponding time pulse signal and output it to the phase accumulator; the digital frequency conversion module is used to convert the received weight data into a corresponding frequency signal and output it to the phase accumulator; the phase accumulator is used to perform multiplication and accumulation calculations on the time pulse signal and the frequency signal based on the phase and output the calculation result. The neural network system of the present application adopts a digital time conversion and digital frequency conversion mechanism, which can encode and map the input data and weight data in the time domain and the frequency domain respectively, and first process and compress the sparse data in the digital time conversion module, which can increase the proportion of effective computing cycles, thereby improving the overall computing throughput. Compared with the computing scheme of pure digital or analog architecture, it fully utilizes hardware resources, not only can effectively process sparse data in neural network computing, but also overcomes the defects of large resource overhead and high energy consumption, taking into account both accuracy and energy saving requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0029] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0030] Figure 1 This is a schematic diagram of the time-frequency calculation framework corresponding to the neural network computing system in the embodiment of the present application;

[0031] Figure 2 This is a schematic diagram of the internal data flow of the digital time conversion module in the embodiment of the present application;

[0032] Figure 3 This is a statistical diagram of the conversion results of the digital time conversion module in the embodiment of the present application;

[0033] Figure 4 This is a schematic diagram of the multi-stage inverter structure of the digital oscillator in the embodiment of the present application;

[0034] Figure 5 A statistical diagram of the output frequency signal converted by the digitally controlled oscillator in an embodiment of the present application;

[0035] Figure 6 This is a structural diagram of a phase accumulator in an embodiment of the present application;

[0036] Figure 7 This is a structural diagram of a phase calibration unit in an embodiment of the present application;

[0037] Figure 8 The waveforms corresponding to the frequency signal F, the time pulse signal T, the processed signal T_cal, and the calibrated signal T×F in the embodiment of the present application are shown respectively;

[0038] Figure 9 This is a structural diagram of a jump protection unit in an embodiment of the present application;

[0039] Figure 10 The waveforms corresponding to the original signal SIGN, the intermediate processing signal SIGN_DE, and the jump control signal HOLD in the embodiment of the present application are shown respectively;

[0040] Figure 11 Schematic diagram of data flow for performing independent conversion processing on high-order weights and low-order weights in an embodiment of the present application;

[0041] Figure 12 Schematic diagram of data flow in a computing array composed of multiple computing units in an embodiment of the present application.

[0042] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0043] To make the above-mentioned purposes, features, and advantages of the present application more clearly understood, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present application without making any creative work are within the scope of protection of this application.

[0044] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0045] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0046] Traditional neural network computing systems usually use a unified precision data path, which is difficult to flexibly adapt to the precision requirements of different layers. Either the overall performance is sacrificed to ensure the calculation precision, or the calculation precision is compressed to reduce energy consumption, thereby affecting the final accuracy of model inference. This "global unified" processing method further amplifies the contradiction between energy efficiency and performance when facing end-side deployment. Although there are currently various hardware acceleration schemes to improve the execution efficiency of neural networks on the end side, there are still the following key technical bottlenecks to be solved: first, the dynamic and unpredictable data sparsity widely existing in neural networks makes it difficult for traditional static sparse optimization methods to adapt; second, existing computing schemes are mostly based on pure digital or pure analog architecture, lacking an effective compromise between precision and energy consumption, and it is difficult to meet the demand for low power consumption and high precision on the end side; finally, the requirements of each layer of neural network for calculation precision are obviously different, and traditional schemes generally use fixed precision, lacking a flexible mixed precision support mechanism, which will cause resource waste or precision loss. The above problems limit the efficient deployment and execution of deep neural networks in energy-sensitive scenarios, and it is urgent to propose a new computing system architecture that is sparse-aware, precision-adjustable, and energy-efficient.

[0047] To solve the above problems, the embodiment of the present application provides a neural network computing system, referring to Figure 1 , Figure 1 is a schematic diagram of a time-frequency computing framework corresponding to the neural network computing system of the present application. The neural network computing system comprises a digital time conversion module 100, a digital frequency conversion module 200 and a phase accumulator 300, and the output ends of the digital time conversion module 100 and the digital frequency conversion module 200 are connected with the input end of the phase accumulator 300;

[0048] The digital time conversion module 100 is used to compress the sparse data in the received input data and convert the corresponding time pulse signal output to the phase accumulator 300;

[0049] The digital frequency conversion module 200 is used to convert the received weight data into a corresponding frequency signal and output to the phase accumulator 300;

[0050] The phase accumulator 300 is used to multiply and accumulate the time pulse signal and the frequency signal based on the phase, and output the calculation result.

[0051] The neural network computing system in the embodiment of the present application can be applied to a neural network computing circuit, which aims to improve the efficiency of sparse data processing, realize efficient reuse of input and weight, and balance low power consumption and computing accuracy. The circuit adopts a time-frequency coding mechanism, and by encoding and mapping the input data and weight data in the time domain and frequency domain respectively, it can also build an efficient matrix multiplication unit in the phase accumulator to perform phase accumulation processing on sparse data and dense weight data, which can achieve the effects of sparse zero skipping, high throughput, mixed precision support, etc., and is suitable for deploying neural network models in resource-constrained edge devices.

[0052] The digital time conversion module receives external input data through the input end, and then converts and processes it to output the time pulse signal to the input end of the phase accumulator through the output end for further processing by the phase accumulator; the digital frequency conversion module receives external input data through the input end, and then converts and processes it to output the frequency signal to the input end of the phase accumulator through the output end for further processing by the phase accumulator. Specifically, the data flow direction is as follows: Figure 1 As indicated by the arrow.

[0053] Exemplarily, the above system is mainly used in the sparse input-dense weight matrix multiplication (Sparse×Dense Matrix Multiplication) scenario in neural networks, and supports multiplication and accumulation calculations of input data IN and weight data W. The main principle is to encode the input IN and weight W into time pulse signals and frequency signals respectively, and finally accumulate them into a phase result (i.e., the calculation result).

[0054] Furthermore, in a feasible embodiment, as Figure 2 As shown, the digital time conversion module 100 at least includes a synchronous counter 101, a clock source 102 connected to the synchronous counter 101, and a sparse detection module 103;

[0055] The synchronous counter 101 and the clock source 102 are used to provide a reference clock signal to the sparse detection module 103 as a reference for time coding;

[0056] The sparse detection module 103 is used to identify sparse data with an input value of 0 in the input data, compress the sparse data in the time encoding stage, and output a corresponding time pulse signal.

[0057] The input data IN is converted into a corresponding time pulse signal through the digital time converter module (DTC), such as Figure 2As shown in Figure 1, the pulse width can be used to represent the time pulse signal, i.e., IN T0. To address the sparsity of the input data, the sparse sensing module is equipped with a sparse sensing circuit that can identify data with an input of 0 and perform "zero skipping" processing during the time encoding stage to compress invalid pulses in the time pulse signal, increase the proportion of effective computing cycles, and thus improve the overall computing throughput.

[0058] It is understandable that in actual deep neural networks, especially after the activation function (such as ReLU, linear rectification function) is applied, the intermediate layer data often exhibits high sparsity characteristics, and the sparse position changes dynamically under different inputs. Traditional hardware solutions are difficult to predict or utilize this sparsity in advance, resulting in invalid calculations and waste of energy and resources.

[0059] To address these issues, the present invention encodes input data into corresponding time pulse signals through a digital-to-time conversion module. A sparse-aware zero-skipping circuit is also introduced to detect zero values ​​in real time. If zero values ​​are detected, the calculation is automatically skipped, thus shortening calculation time and improving the computational throughput of the system or chip.

[0060] Exemplarily, the digital time conversion module encodes an input digital signal IN into a time pulse signal (IN T0) proportional to the digital signal. The module includes a base clock source and a synchronous counter circuit. The base clock source provides a fixed-period reference clock signal T0, which serves as a counting basis for time conversion.

[0061] Specifically, the synchronous counter is constructed based on a JK flip-flop and adopts a unified clock drive method to ensure that the triggers at all levels flip simultaneously at the same clock edge, avoiding errors caused by signal propagation delays. The synchronous counter counts up in each clock cycle T0 until the count value reaches the target value corresponding to the input digital IN, outputs a time pulse signal (IN T0) with a fixed width, and then resets to enter the next conversion cycle. Compared with the asynchronous counter structure, the synchronous counter has a smaller sequence delay and higher time accuracy, which significantly improves the stability and accuracy of the time encoding process, and helps to ensure the accuracy of subsequent time-frequency multiplication calculations. Figure 3 As shown, the horizontal axis is the input count value (discrete numbers 0-15), and the vertical axis is the pulse width modulation and error, both in ns (nanoseconds). The conversion result (i.e., the time pulse signal) output by the digital time conversion module of the embodiment of the present application is shown, which has good linearity, high output accuracy and small error.

[0062] In a feasible embodiment, the digital frequency conversion module 200 includes a digitally controlled oscillator, which is used to convert the input weight data into a corresponding frequency signal, wherein each weight data corresponds to an output frequency to maintain periodic oscillation.

[0063] Exemplarily, the weight data W is usually a dense matrix. In the embodiment of the present application, a digitally controlled oscillator (DCO) is used to convert the weight data into a corresponding frequency signal, namely WF0. Each weight data corresponds to a frequency output, maintaining periodic oscillation to cooperate with the time pulse signal of the input data to form a time window for multiplication operation.

[0064] Furthermore, in a feasible embodiment, as Figure 4 As shown, the digitally controlled oscillator 200 includes at least a plurality of inverters 201 connected in sequence, wherein the pull-up path and the pull-down path of the inverter 201 are respectively connected in series with a controlled current source. The digitally controlled oscillator is used to adjust the oscillation frequency by regulating the voltage or the digital control signal.

[0065] exist Figure 4 In this circuit, the inverters are connected end-to-end, comprising N stages. The pull-up and pull-down paths of each inverter are connected in series with a controlled current source, which provides Ibias (controlled current). The current source includes VDD (Voltage Drain Drain) and GND (Ground).

[0066] In order to achieve effective conversion of digital signals into frequency signals, an embodiment of the present application provides a frequency-adjustable digitally controlled oscillator DCO. The DCO adopts a multi-stage inverter ring structure as the oscillation core, and its output frequency is adjusted by a control signal. Specifically, the inverter adopts a current-starved structure, that is, a controlled current source is connected in series in the pull-up and pull-down paths of the inverter. By adjusting the control voltage or the digital control signal to regulate the conduction capability of the current source, fine control of the inverter delay time is achieved, and then the frequency of the entire oscillation ring is adjusted. The adjustment result is as follows: Figure 5 As shown (the horizontal axis x is the digitally controlled oscillator input signal (for example, discrete integer value 0-15), the vertical axis y is the output frequency, in MHz). The corresponding curve function is y=73.528x+6.5634, R 2 =0.9995, R 2 is the determination coefficient, which is a linear reliability index. The closer it is to 1, the stronger the linear controllability of the frequency control.

[0067] The structure of the aforementioned digital controlled oscillator has the following advantages: First, the frequency is highly adjustable: by finely controlling the current size of the current source, the oscillation frequency can be continuously adjusted, which is suitable for the frequency mapping requirements in multi-precision computing scenarios; second, it has good linear response characteristics: because the current-starved inverter has a relatively linear response relationship to the control signal, the mapping relationship between the digital control signal and the output frequency is more predictable, which is convenient for subsequent system calibration and control; third, it has a simple structure and low area overhead: the inverter used has a simple structure and is easy to integrate, and can be deployed on a large scale under low power conditions; fourth, it has high compatibility with CMOS (Complementary Metal Oxide Semiconductor) process: the DCO is constructed based on a standard CMOS process, which is convenient for coordinated integration with other digital circuits to achieve efficient on-chip deployment. Therefore, the frequency-adjustable DCO provided in the embodiment of the present application can efficiently complete the conversion task from digital to frequency domain, provide key support for the multiplication calculation process based on time × frequency in the embodiment of the present application, and is particularly suitable for frequency encoding of dense weighted signals, while ensuring accuracy and reducing overall computing power consumption.

[0068] In a possible embodiment, Figure 6 As shown, the phase accumulator 300 at least includes a phase calibration unit 301, a multiplication unit 302, a phase accumulation unit 303 and a data readout unit 304 connected in sequence;

[0069] The phase calibration unit 301 is used to align the phases of the time pulse signal and the frequency signal;

[0070] The multiplication unit 302 is used to multiply the phase-aligned time pulse signal and the frequency signal to obtain a product result, and process the phase-aligned time pulse signal and the frequency signal according to the exclusive OR logic to determine the accumulation direction of the product result;

[0071] The phase accumulation unit 303 is used to phase-modulate the phase-aligned time pulse signal and the frequency signal according to the accumulation direction and accumulate them in time to obtain a multiplication and addition operation result;

[0072] The data reading unit 304 is used to read the multiplication and addition operation result, and convert the multiplication and addition operation result into a calculation result in a target format through a preset digital-to-time conversion interface.

[0073] In the embodiment of the present application, the time pulse signal and the frequency signal act together on a phase accumulator. During the valid period of the input pulse, the corresponding frequency signal drives the phase to grow, realizing the multiplication calculation of "time × frequency". The formula is expressed as Phase = ∑IN T0 × WF0, where Phease is the result. The result is accumulated in situ in the form of phase to complete the multiplication and accumulation operation. This method can maintain basic multiplication accuracy while maintaining low power consumption, which is suitable for the power consumption-accuracy trade-off requirements of edge scenarios.

[0074] Specifically, this application proposes a phase-aware multiply-accumulate unit (PA-MAC) as a phase accumulator. This calculation structure implements the multiplication and addition calculation in sparse-dense matrix multiplication based on the phase relationship between the input time pulse and the frequency signal. The specific structure is as follows: Figure 6 As shown, the input coded time pulse (i.e., time pulse signal) and weighted coded frequency signal are first input into the phase calibration unit of the phase accumulator, and then calculated by the multiplication unit, wherein the multiplication unit is a compliant multiplication unit. After calculation, they are input into the phase accumulation unit for accumulation calculation, and then converted into the corresponding format by the data readout unit to obtain the multiplication and accumulation result (i.e., calculation result).

[0075] For example, the phase calibration unit is used to align the phase of the input time pulse and the frequency signal represented by the weight before the calculation starts, so as to ensure that the edge of the time pulse and the reference period of the frequency signal start at a fixed phase difference, effectively improving the accuracy and consistency of the phase accumulation and reducing the error. Figure 7 As shown, the phase calibration unit includes a D-type flip-flop and an inverter. The D-type flip-flop includes an input port D, an output port Q, and a clock input ClK. The time pulse signal T and the frequency signal F are input into the D-type flip-flop. The final output processed signal T_cal is input into the inverter together with T and F to finally obtain the superimposed signal T×F. Figure 8The waveforms corresponding to the signals F, T, T_cal and T×F are shown, where the dotted lines are the time points of phase alignment. In addition, to support the signed calculation requirements in the neural network, the embodiment of the present application also introduces automatic processing of the sign bits of the input data and weights. The sign of the product result is then determined by XOR logic, and the direction of phase accumulation (forward or reverse) is adjusted accordingly. This method can avoid traditional complement conversion or additional subtraction logic, and can achieve lightweight implementation of signed multiplication. Afterwards, the phase accumulation unit module receives a phase-aligned time pulse signal (representing input data) and a frequency signal (representing weight), and completes the multiplication and addition calculation by phase modulating each input pulse with the corresponding frequency and accumulating it in time. It should be noted that the phase accumulator also includes a storage unit, and the phase value is accumulated in the storage unit until the operation is completed. After completing a round of multiplication and addition operations, the calculation result is read out from the phase storage unit in a serial manner and can be further converted into the required output format through the digital time conversion interface. This module can support clock-controlled bit-by-bit output and adapt to data readout requirements under different precisions.

[0076] Furthermore, in a feasible embodiment, as Figure 6 As shown, the phase accumulator 300 also includes a signal jump protection unit 305, which is connected to the phase accumulator unit 303 and is used to latch the current multiplication and addition result when the time pulse signal or frequency signal jumps until the signal stabilizes and the operation is resumed.

[0077] Specifically, when the input time pulse signal or frequency signal jumps, to prevent phase erroneous updates caused by signal instability, the module can temporarily latch the current accumulated value and resume operation after the signal stabilizes. In the embodiment of the present application, the above mechanism can ensure the stability and robustness of the phase accumulation during the calculation process.

[0078] Specifically, if Figure 9 As shown, the jump protection unit includes a first delay device DELAY (VN and VP represent positive voltage input and negative voltage input respectively), a second delay device DELAY and a logic XOR gate connected in sequence; the first delay device DELAY is used to receive the input time pulse signal and frequency signal (i.e., SIGN), and output the intermediate processing signal SIGN_DE after delay processing, and input the intermediate processing signal to the second delay device DELAY, and then the second delay device delays the intermediate processing signal and outputs the delayed signal to the logic XOR gate; the logic XOR gate is used to perform XOR processing on the input time pulse signal and frequency signal with the corresponding delay signal, and output a jump control signal HOLD, which is used to represent the jump interval to latch the multiplication and addition operation result of the intermediate processing signal SIGN_DE in the jump interval.

[0079] Among them, the intermediate processing signal SIGN_DE is used as the protected signal output by the signal jump protection unit, and the jump control signal HOLD indicates the jump interval that needs to be protected through a high level. It can be understood that, if Figure 10 As shown, the intermediate processed signal SIGN_DE undergoes a first delay relative to the original signal SIGN. During the XOR process of the HOLD signal, the time point in the delayed SIGN_DE corresponding to the time point when the signal SIGN first rises to a high level does not change, so the XOR result is 1 (corresponding to a high level). The HOLD signal is obtained by XORing the twice-delayed signal with the original signal. When the twice-delayed signal first rises to a high level, the two levels are the same, and the HOLD signal changes to a low level. The high-level interval of the HOLD signal is the interval starting from the time point when the signal SIGN first rises to a high level, and the high-level duration is the sum of the two delays. The SIGN_DE signal jumps exactly in the middle of this jump interval. The control goal of the HOLD signal is to perform jump protection within a certain period before and after the SIGN_DE signal jumps, latching the multiplication and addition operation result, and maintaining computational stability and robustness. Figure 10 The dotted lines in the figure represent the time points at which the original signal changes. From left to right, the first dotted line represents the time point at which the signal changes from a low level to a high level, and the second dotted line represents the time point at which the signal changes from a high level to a low level.

[0080] In a feasible embodiment, the neural network computing system further includes a data segmentation module and a weight segmentation module, the data segmentation module is connected to the digital time conversion module, and the weight segmentation module is connected to the digital frequency conversion module;

[0081] The data segmentation module is used to split the input data into high-order data and low-order data based on a preset data splitting rule, and input the high-order data and low-order data into the digital time conversion module for independent encoding conversion to obtain a high-order time pulse signal and a low-order time pulse signal respectively;

[0082] The weight segmentation module is used to split the weight data into high-order weights and low-order weights based on the preset weight splitting rules, and input them into different digitally controlled frequency oscillators of the digital frequency conversion module for independent encoding conversion to obtain high-order frequency signals and low-order frequency signals.

[0083] To meet the high-precision requirements of neural network calculations, the present invention provides a method for segmenting data and then performing digital time conversion and digital frequency conversion. Specifically, data is processed using a high- and low-bit segmented calculation method. The input data and weight data are divided into high-order (MSB, Most Significant Bit) and low-order (LSB, Least Significant Bit) parts, respectively. These parts are encoded into corresponding time pulses and frequency signals through independent channels, and multiplication and accumulation are performed separately within the calculation unit, thereby supporting flexible mixed-precision operations.

[0084] To further adapt to the differentiated precision requirements of different computing layers in neural networks and improve the versatility and energy efficiency of hardware, this application proposes a bit width reconstruction mechanism that supports mixed-precision computing. This mechanism, through segmented processing and reconstruction of input data and weight data, flexibly supports multiplication and addition tasks of different bit widths (such as 4-bit, 8-bit, 9-bit, etc.) on a unified computing structure, thereby meeting the comprehensive performance requirements of model compression, precision tuning, and end-side deployment.

[0085] For example, the input data and weight data are first split into high and low bits before being sent to the multiplication and addition module. Taking a typical 9-bit × 8-bit multiplication operation as an example, the 9-bit input data can be split into the upper 5 bits (MSB) and the lower 4 bits (LSB), and the 8-bit weight can be split into the upper 4 bits (MSB) and the lower 4 bits (LSB). Each segment acts as an independent calculation unit and participates in the subsequent time-frequency multiplication calculation.

[0086] For example, Figure 11 As shown, after being split into high-order weights and low-order weights, they are input into independent digitally controlled frequency oscillators, which convert them into the form of frequency signals. The input data is input into the digital time conversion module after being split (not shown in the figure), and finally multiplication and addition calculations are performed respectively by the high-order calculation phase accumulator and the low-order calculation phase accumulator. It should be noted that when the digital time conversion module outputs a time pulse signal, the high-order time pulse signal and the low-order time pulse signal are alternately sent to the high-order calculation phase accumulator and the low-order calculation phase accumulator according to a preset period. Among them, the multiplication and addition results of each segment can be restored by the corresponding weighted synthesis logic to reconstruct the final high-precision multiplication output. Through this seed data-level "low-precision sub-calculation + weighted accumulation" mode, the present invention realizes the efficient implementation of high-bit width multiplication under limited hardware resource conditions.

[0087] In terms of hardware resource scheduling, the bit width configurations can share the same basic hardware units, including time encoders, frequency modulators, phase accumulators, etc., avoiding the problem of repeated deployment of independent calculation paths for different bit width configurations. This resource sharing mechanism not only improves the area utilization, but also only needs to adjust the bit width flag bit in the configuration register when switching the calculation precision between different layers in the network, without the need to reconfigure the hardware path, greatly improving the adaptive flexibility of the system.

[0088] The bit width reconfiguration mechanism provided by the embodiments of the present application achieves the following technical effects: flexible precision adaptation capability, supporting multi-precision calculation tasks between 4bit and 9bit; significantly reducing the hardware area and power consumption overhead caused by high bit width calculation; dynamically adjusting the precision configuration according to the neural network structure or runtime strategy, balancing inference accuracy and efficiency; effectively improving the practicality and expansibility of the end-side neural network processor.

[0089] Further, in a feasible embodiment, as shown in Figure 12 The phase accumulator includes a rectangular calculation array composed of a plurality of calculation units, each calculation unit is connected with a digital time conversion module and a digital control frequency oscillator in a digital frequency conversion module, each row of calculation units in the rectangular calculation array shares the same time pulse signal, and each column of calculation units in the rectangular calculation array shares the same frequency signal, and each calculation unit is used for multiplication and accumulation operation according to the received time pulse signal and frequency signal.

[0090] Specifically, the neural network calculation system can include a plurality of digital time conversion modules and a plurality of digital control frequency oscillators and a calculation array composed of a plurality of calculation units. At the calculation unit level, the data is split into high bits and low bits for calculation, the same input pulse is multiplied and accumulated with the frequency pulse encoded by the weight high bit and the weight low bit respectively, and the input multiplexing is realized at the calculation unit level, as shown in Figure 3As shown. This computing architecture adopts an array structure design, in which each row shares the same input time pulse signal, and each column shares the same frequency weight signal, realizing data path sharing and high-reuse input and output structure at the array level, greatly reducing the transmission resource overhead and power consumption on the chip, and improving the overall energy efficiency. During the calculation process of the computing array, the high-order time pulse signal, low-order time pulse signal, high-order frequency signal, and low-order frequency signal are calculated in pairs, and there are four combinations: high-order time pulse signal × high-order frequency signal, high-order time pulse signal × low-order frequency signal, low-order time pulse signal × high-order frequency signal, and low-order time pulse signal × low-order frequency signal. This array can solve the problems of strong sparsity, large differences in accuracy requirements, and the contradiction between power consumption and computing power in the end-side deployment of neural networks. Through designs such as data encoding, sparse processing, mixed precision support, and computing array multiplexing, it can effectively improve the system's computing efficiency and reduce power consumption.

[0091] The embodiment of the present application is equivalent to providing a data multiplexing mechanism based on a time-frequency orthogonal structure, which introduces a structural data sharing strategy at both the computing unit level and the computing array level to achieve efficient multiplexing of input pulses and frequency signals. Specifically, at the computing unit level, the present application adopts a unified time encoding method for input sparse data so that multiple computing paths can share the same time pulse signal; dense weight data is provided to multiple computing units in the form of frequency encoding, and multiple weights can be mapped to each computing path through different frequency signals to construct a reusable input-weight multiplication and addition path, which significantly reduces redundant calculations and input bandwidth load. At the computing array level, the embodiment of the present application further proposes a two-dimensional orthogonal mapping structure of "row time / column frequency", such as Figure 12 As shown, each array row shares the same input timing pulse, and each column shares the same frequency signal. This structure enables a single input timing pulse to drive the synchronous operation of the entire row of computing units, while each column completes the corresponding multiplication and addition calculations under the condition of receiving a unified frequency signal. This approach maximizes the reuse rate of input and weight signals while ensuring computational accuracy, achieving efficient parallel computing at the array level.

[0092] The above structural design effectively reduces the need for data movement in the chip architecture, reduces the pressure on input and output interfaces, and improves the throughput and resource utilization efficiency of the computing array, making it suitable for high-energy-efficiency execution scenarios of deep neural networks deployed on the edge. It also further improves data reuse efficiency at the array level. The structure of rows sharing time pulses and columns sharing frequency signals allows multiple computing units to reuse input signals and weight signals in parallel, significantly reducing energy consumption while maintaining calculation accuracy.

[0093] The neural network computing system provided in the embodiment of the present application includes an efficient neural network acceleration computing architecture, which is designed to optimize the computing performance and energy efficiency of the end-side neural network model. By combining the time pulse coding of sparse input data, a digitally controlled oscillator with precise frequency response, a phase-aware signed multiplication and accumulation computing unit, and a mixed-precision bit width reconstruction mechanism, efficient matrix multiplication operations are achieved. In addition, a time-frequency bidirectional data multiplexing mechanism based on orthogonal coding is adopted to significantly improve the data multiplexing rate and computing throughput. The above scheme not only solves the problems of sparsity and precision requirements in neural network calculations, but also provides a flexible solution in reducing power consumption and improving computing efficiency, which is suitable for the efficient deployment of end-side devices.

[0094] Furthermore, it should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the neural network computing system of the present application. More forms of simple transformations based on this technical concept are all within the scope of protection of the present application.

[0095] The embodiment of the present application also provides a neural network computing system device, which includes at least the neural network computing system in the above embodiment. The neural network computing system device can be a device with computing capabilities, such as a computer, server, chip or other computing processing device. The neural network computing system device includes the neural network computing system in the above embodiment, which can solve the technical problems of large resource overhead and high energy consumption in the current neural network computing strategy for sparse data. Compared with the prior art, the beneficial effects of the neural network computing system device provided in the embodiment of the present application are the same as the beneficial effects of the neural network computing system provided in the above embodiment, and the other technical features in the neural network computing system device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0096] It should be understood that the various parts of the embodiments of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any appropriate manner in any one or more embodiments or examples.

[0097] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the above claims.

[0098] The above is only an exemplary solution of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent processing scope of the present application.

Claims

1. A neural network computing system, characterized in that: The neural network computing system includes a digital time conversion module, a digital frequency conversion module and a phase accumulator, wherein the output ends of the digital time conversion module and the digital frequency conversion module are respectively connected to the input end of the phase accumulator; The digital time conversion module is used to compress the sparse data in the received input data, and convert it into a corresponding time pulse signal to output to the phase accumulator; The digital frequency conversion module is used to convert the received weight data into a corresponding frequency signal and output it to the phase accumulator; The phase accumulator is used to perform multiplication and accumulation calculations on the time pulse signal and the frequency signal based on the phase, and output a calculation result.

2. The neural network computing system according to claim 1, wherein: The digital time conversion module at least includes a synchronous counter, a clock source connected to the synchronous counter, and a sparse detection module; The synchronous counter and the clock source are used to provide a reference clock signal to the sparse detection module as a reference for time encoding; The sparse detection module is used to identify sparse data with an input of 0 in the input data and compress the sparse data in the time encoding stage to output a corresponding time pulse signal.

3. The neural network computing system according to claim 1, wherein: The digital frequency conversion module includes at least one digitally controlled oscillator, which is used to convert input weight data into corresponding frequency signals, wherein each weight data corresponds to an output frequency to maintain periodic oscillation.

4. The neural network computing system according to claim 3, wherein: The digitally controlled oscillator comprises at least a plurality of inverters connected in sequence, wherein the pull-up path and the pull-down path of the inverters are respectively connected in series with a controlled current source, and the digitally controlled oscillator is used to adjust the oscillation frequency by regulating voltage or digital control signal.

5. The neural network computing system according to claim 1, wherein: The phase accumulator at least comprises a phase calibration unit, a multiplication unit, a phase accumulation unit and a data readout unit connected in sequence; The phase calibration unit is used to align the phases of the time pulse signal and the frequency signal; The multiplication unit is used to multiply the phase-aligned time pulse signal and the frequency signal to obtain a product result, and process the phase-aligned time pulse signal and the frequency signal according to the exclusive OR logic to determine the accumulation direction of the product result; The phase accumulation unit is used to phase-modulate the phase-aligned time pulse signal and the frequency signal according to the accumulation direction and accumulate them in time to obtain a multiplication and addition operation result; The data readout unit is used to read the multiplication and addition operation result, and convert the multiplication and addition operation result into a calculation result in a target format through a preset digital time conversion interface.

6. The neural network computing system according to claim 5, wherein: The phase accumulator also includes a signal jump protection unit, which is connected to the phase accumulator unit and is used to latch the current multiplication and addition result when the time pulse signal or the frequency signal jumps until the signal stabilizes and the operation is resumed.

7. The neural network computing system according to claim 6, wherein: The jump protection unit includes a first delayer, a second delayer and a logic XOR gate connected in sequence; the first delayer is used to receive the input time pulse signal and frequency signal, and output an intermediate processing signal after delay processing, and input the intermediate processing signal to the second delayer, and then the second delayer delays the intermediate processing signal and outputs the delayed signal to the logic XOR gate; the logic XOR gate is used to perform XOR processing on the input time pulse signal and frequency signal with the corresponding delay signal, and output a jump control signal, and the jump control signal is used to characterize the jump interval, so as to latch the multiplication and addition operation result of the intermediate processing signal in the jump interval.

8. The neural network computing system according to any one of claims 1 to 7, wherein: The neural network computing system further includes a data segmentation module and a weight segmentation module, wherein the data segmentation module is connected to the digital time conversion module, and the weight segmentation module is connected to the digital frequency conversion module; The data segmentation module is used to split the input data into high-order data and low-order data based on a preset data splitting rule, and input the data into the digital time conversion module for independent coding conversion to obtain a high-order time pulse signal and a low-order time pulse signal; The weight segmentation module is used to split the weight data into high-order weights and low-order weights based on preset weight splitting rules, and input them into different digitally controlled frequency oscillators of the digital frequency conversion module for independent encoding conversion to obtain high-order frequency signals and low-order frequency signals.

9. The neural network computing system according to any one of claims 1 to 7, wherein: The phase accumulator includes a rectangular computing array consisting of multiple computing units, each computing unit is respectively connected to a digital time conversion module and a digitally controlled frequency oscillator in the digital frequency conversion module, each row of computing units in the rectangular computing array shares the same time pulse signal, and each column of computing units in the rectangular computing array shares the same frequency signal, and each of the computing units is used to perform multiplication and accumulation operations based on the received time pulse signal and frequency signal.

10. A neural network computing device, characterized in that: The neural network computing device comprises the neural network computing system according to any one of claims 1 to 9.