Bit serial computing transmission structure for high-speed data processing in chip interconnection system

By introducing a bit serial computing transmission structure and voltage mode logic VML interface in the traditional computing architecture, the bit serial processing unit is directly connected to the high-speed physical layer, which solves the problem of limited performance in high-speed data processing in traditional computing architecture, and realizes high throughput and low-latency data transmission.

CN120104539AActive Publication Date: 2025-06-06UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510170629.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

Traditional computing architectures face trade-offs between computing and data transmission when processing high-speed data, resulting in performance limitations, especially in bandwidth bottlenecks and hardware efficiency issues.

Method used

Design a bit serial computing transmission structure, through a customized voltage mode logic VML interface, the bit serial processing unit is directly connected to the high-speed physical layer, achieving tight integration of computing and communication. This structure combines digital circuits and analog circuits to form an efficient computing and data transmission architecture.

Benefits of technology

It realizes high throughput and high transmission frequency, reduces latency and energy consumption, improves the parallelism and efficiency of the computing process, and is suitable for high-speed data processing in chip interconnect systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104539A_ABST
    Figure CN120104539A_ABST
Patent Text Reader

Abstract

The invention discloses a bit serial computing transmission structure for high-speed data processing in a chip interconnection system, and belongs to the technical field of semiconductor integrated circuits. The architecture mainly integrates two bit serial multiply-accumulate modules, a bit serial adder and a high-speed voltage mode logic (VML) physical interface, wherein the VML interface mainly comprises a driver, an equalizer, a comparator, a bias current source and an RS trigger. The structure realizes direct connection of calculation and data transmission without a parallel-serial conversion circuit or a serial-parallel conversion circuit, and has the potential of reducing delay and energy consumption. The extensible chip interconnection component may be the basis of a modularized algorithm accelerator in the future, and different accelerator hardware structures are formed by expanding or reconstructing the basic component, so that different algorithms are adapted, and the reconfigurability of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a bit serial computing transmission structure, in particular to a bit serial computing transmission structure used for high-speed data processing in a chip interconnection system, and belongs to the technical field of semiconductor integrated circuits. Background Art

[0002] As the demand for data-intensive applications such as real-time image processing, artificial intelligence, and high-performance computing continues to grow, the limitations of traditional computing architectures in solving bandwidth bottlenecks and hardware efficiency issues are becoming increasingly apparent. Traditional systems often face a trade-off between computing and data transmission, where the high overhead of transmitting intermediate results between computing units severely limits overall system performance.

[0003] One promising solution is a compute-transfer architecture that integrates computation and data transfer functions into a unified hardware structure. This architecture has the potential to reduce latency and energy consumption by eliminating the intermediate buffering and data alignment steps typically required in traditional systems.

[0004] In this context, bit-serial processing offers unique advantages due to its simplicity, high-speed operation, and ability to process data in a compact format. Bit-serial data formats are mainly divided into single-precision data and double-precision data. In bit-serial computing, double-precision data usually consists of three parts: low-order bits, high-order bits, and header bits. Among them, the low-order bits and high-order bits correspond to the low-order and high-order parts of the data, respectively, while the header bit is used to identify the starting position of the data and control the operation of the bit-serial circuit.

[0005] Therefore, a bit-serial computation and transmission structure for high-speed data processing is proposed, which is suitable for chip interconnection systems and integrates computation and transmission functions in a unified framework. Summary of the invention

[0006] The technical problem to be solved by the present invention is to provide a bit serial computing transmission structure for high-speed data processing in a chip interconnection system, which directly connects the bit serial processing unit to the high-speed physical layer through a custom-designed voltage mode logic VML interface, such as Figure 1 As shown, computing and communication are more tightly integrated, thus achieving high throughput and high transmission frequency.

[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is: Figure 2 The bit serial multiply accumulator 2 and a Figure 8 The bit serial adder 8 is connected through a high-speed serial transmission interface VML module. The entire structural design combines digital circuits and analog circuits to form an efficient computing and data transmission architecture.

[0008] In this structure, 2 accepts input signals x1 and y1, or x2 and y2, representing two multipliers respectively. After these input data are processed, they are output by o1_l and o1_h, or o2_l and o2_h respectively, representing the low and high bits of the multiplication and accumulation results output from 2.

[0009] Taking the multiplication of two 8-bit fixed-point numbers as an example, the head signal is output after passing through seven D flip-flops 100, so that the output hout signal can be aligned with the output data o1_l and o1_h, or o2_l and o2_h.

[0010] The two 2 structures share the clock signal and the head signal in the structure, which means that their calculations are performed synchronously, effectively improving the parallelism and efficiency of the calculation process. The signal transmitted through the VML interface in the structure synchronizes the five data signals through the D flip-flop 101 and the clk_rx signal, and then inputs into 8.

[0011] In the structure, 8 is responsible for receiving the calculation results of the two 2 structures transmitted through the VML interface. 8 adds the results from the two 2 and outputs the final result through the low-order out_l and the high-order out_h. At the same time, the h signal is used as a header signal to identify the validity and alignment information of the output result.

[0012] For the eight-bit fixed-point signed number multiplication calculation 2, a base-4 Booth encoding circuit is used, as shown in Figure 3 The bit serial radix-4 Booth multiplier 3, such as Figure 7 The invention is composed of an adder 7 suitable for multiplication and accumulation calculation and a plurality of registers 200.

[0013] The main function of 3 is to perform multiplication operations. It can effectively process the multiplication of fixed-point numbers and generate bit-serial products, reducing the hardware resources required for parallel computing. In this architecture, the collaboration between 3 and 7 determines the computing performance of the entire 2.

[0014] 7 is a key component in the multiplication and accumulation calculation process. It receives the bit-serial product from 3 and accumulates it to the previous result, thereby achieving efficient calculation. 7 adopts serial bit-by-bit addition in this architecture, which can efficiently complete the addition operation with a small amount of hardware resources.

[0015] In order to ensure the correct alignment of each bit in the bit-serial calculation, 200 plays an important role as a cache module in the bit-serial calculation process. They store and cache the results of the previous calculation cycle and ensure that the new serial input is aligned with the previous calculation result. This design is crucial for bit-serial calculation because it can smoothly pass the bit-serial data stream to 7 for accumulation, avoiding the problem of data misalignment or misalignment during the calculation process.

[0016] The VML interface design uses Fig. 9 The VML driving circuit 9 is responsible for transmitting the serial data from the transmitting end of the VML interface to the receiving end of the VML interface in the form of differential signals.

[0017] At the receiving end of the VML interface, the structure includes two levels of cascaded Fig.10 The equalizer 10 shown is used to compensate for attenuation and distortion during signal transmission to ensure data integrity and stability. The function of 10 is to enhance the high frequency part of the signal and reduce the signal attenuation caused by factors such as transmission line bandwidth limitation, crosstalk or noise, thereby improving the data transmission quality.

[0018] 10 Then connect Fig.12 The differential signal comparator 12 shown can accurately identify the received signal and convert it into a digital signal. Since the transmission of high-speed signals may be affected by noise, clock offset and other problems, 12 receives the signal through the differential mode, which greatly improves the anti-interference ability and signal reliability.

[0019] The bias voltage or bias current is used to provide a stable operating point for the circuit, thereby ensuring that the amplifier operates in a suitable region. Although the bias current or voltage can be introduced externally, in order to reduce the dependence on peripheral circuits and simplify the system, the present invention uses an internal current source to provide the bias. Fig.11 As shown, this internal bias current source 11 not only provides a stable bias point for the circuit, but also enhances the integration and robustness of the structure. 10 and 12 both use 11 as a supporting module.

[0020] In addition, the receiving end of the VML interface is also equipped with an RS trigger module for signal shaping. The main function of the RS trigger is to shape the received signal, remove high-frequency noise and unnecessary waveform distortion, so as to ensure the stability and accuracy of the signal in the subsequent processing process. In this way, the RS trigger can effectively improve the signal processing capability of the structure and ensure high accuracy during data transmission.

[0021] In order to accurately simulate the loss effect of the transmission line, a differential transmission line simulated by a distributed model is used to connect the transmitting end 9 and the receiving end 10 of the VML interface. At high frequencies, the signal will experience a certain attenuation in the transmission line. The use of a distributed model can make a more accurate simulation of the transmission performance of the entire system.

[0022] The entire structure adopts a modular and scalable architecture. This design concept enables the structure to flexibly handle large-scale data. In traditional computing architectures, as the amount of data increases, bandwidth bottlenecks and power consumption issues often become the main limiting factors of structural performance. The design based on the combination of bit-serial architecture and VML interface greatly reduces the bandwidth requirements of the computing system. This serial transmission method can achieve higher transmission rates while maintaining lower power consumption, which has great advantages in long-distance transmission. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is a schematic diagram of a bit serial computing transmission structure for high-speed data processing in a chip interconnection system of the present invention; Figure 2 is a schematic diagram of a bit serial multiply-accumulator in the present invention; Figure 3 Schematic diagram of a bit-serial multiplier of a bit-serial multiply-accumulator in the present invention; Figure 4 It is a schematic diagram of the least significant bit circuit of the bit-serial multiplier of the present invention; Figure 5 It is a schematic diagram of the intermediate bit circuit of the bit serial multiplier of the present invention; Figure 6 It is a schematic diagram of the most significant bit circuit of the bit-serial multiplier of the present invention; Figure 7 is a schematic diagram of a bit-serial adder of a bit-serial multiply-accumulate device in the present invention; Figure 8 is a schematic diagram of a bit serial adder in the structure of the present invention; Fig. 9 It is a schematic diagram of the driver circuit of the VML interface in the present invention; Fig.10 It is a schematic diagram of an equalizer circuit of a VML interface in the present invention; Fig.11 It is a schematic diagram of a bias current source circuit of an equalizer and a differential comparator acting on a VML interface in the present invention; Fig.12 It is a schematic diagram of a differential signal comparator circuit of a VML interface in the present invention. DETAILED DESCRIPTION

[0023] In order to elaborate on the technical scheme adopted by the present invention to achieve the predetermined technical purpose, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only partial embodiments of the present invention, rather than all embodiments, and the technical means or technical features in the embodiments of the present invention can be replaced without paying creative labor. The present invention will be described in detail below with reference to the drawings.

[0024] like Figure 2 As shown, the rstn signal is the asynchronous reset signal of the D flip-flop. It should be noted that 3 itself does not rely on the rstn signal reset, but controls the reset through the head signal and input data. Specifically, when the head signal is set to 1, the input data is set to zero, and the output data of 3 is reset to 0 in a serial manner. In 2, the clock signal clk of the structure controls the synchronization of all operations to ensure that each calculation step is strictly coordinated in timing. The out signal represents the final bit-serial product output. Usually, after the multiplication calculation is completed, the accumulation operation is performed through 7 to finally obtain the calculation result required by the structure. out_d is an intermediate signal with an output delay of seven clock cycles. It plays a buffering role when the adder receives and processes the bit-serial calculation results, thereby avoiding data errors caused by timing asynchrony.

[0025] 3 in 2 has a radix-4 Booth encoding circuit. Radix-4 Booth encoding is a technique used to optimize the binary multiplication process. By reducing the number of partial products of multiplication, it can effectively reduce the hardware complexity and area of ​​the multiplier. Radix-4 Booth encoding reduces the number of these partial products by effectively grouping and encoding the input data, thereby reducing the hardware resources required for calculation. Specifically, Figure 4 , Figure 5 and Figure 6 They represent the basic components that make up 3, and the radix-4 Booth encoding circuit is divided into two parts: one is the part that determines whether addition or subtraction will be performed when the Booth code is not zero; the other part determines whether the Booth code needs to be right-shifted, that is, multiplying some partial products by 2, thereby reducing the complexity of the multiplication calculation.

[0026] Taking the multiplication of two eight-bit fixed-point signed numbers as an example, before the introduction of radix-4 Booth encoding, 3 usually requires eight units to implement bit-by-bit multiplication calculations, with each unit responsible for calculating a multiplication bit and the corresponding partial product. However, by introducing radix-4 Booth encoding, 3 only needs four units to complete the same calculation task. This optimization not only reduces the hardware resources required for calculation, but also significantly reduces the area and power consumption of 3. As the scale of calculation increases, this optimization can bring significant hardware savings and provide strong support for high-performance, low-power computing systems.

[0027] For bit-serial adders, 7 and 8 have slight differences, mainly in the control header. Since 3 outputs eight bits of data every ten cycles, there is a certain time difference and delay when the bit-serial output result of 7 is connected back to 7. Specifically, since 3 completes a complete multiplication operation every ten cycles, the output eight bits of data need to be transmitted through two clock cycles before they can be fully delivered to 7. This means that the result output by 7 will be delayed by two clock cycles before entering 7. In order to effectively handle this delay, the control header of 7 needs to be adjusted appropriately to ensure that 7 can correctly handle the carry and accumulation process of the bit-serial data. Therefore, the header of 7 needs to be delayed for a total of three clock cycles before controlling the carry logic of the high and low bits of 7. Since 8 does not have a back connection operation, its header only needs to be delayed for one clock cycle.

[0028] As the core module of the entire VML system, the main function of 9 is to convert the full-swing serial CMOS digital signal into a low-voltage differential analog signal LVDS. When designing 9, it is necessary not only to ensure that the structure can meet the required high transmission rate, but also to fully consider the integrity of the signal and the sufficiently strong driving capability to ensure that the data is not distorted during high-speed transmission and can operate stably under different working environments. The voltage divider circuit is composed of resistors R1-R9 and MOS tubes M31-M46, and is controlled by digital voltages Ctrl1-Ctrl4. When the Ctrl1-Ctrl4 signal is low, the corresponding resistor network path is activated and participates in the operation of the voltage divider circuit. By adjusting the state of these four control signals, up to 15 different voltage divider circuit configurations can be achieved, thereby providing 15 different drive signal swings to meet the requirements for signal strength under different transmission environments. In addition, the voltage divider circuit is also used in the startup circuit of the subsequent self-biased amplifier to play an auxiliary role.

[0029] Fig. 9 900 in FIG. 900, i.e., the self-biased differential amplifier A1 is composed of MOS tubes such as M10, M13, M15-M17, and forms a two-stage operational amplifier with M1. In order to reduce the capacitance value and ensure sufficient stability, the circuit uses MOS tube capacitor M7 as Miller compensation capacitor to improve the stability of the operational amplifier. In 900, by adjusting the upper node voltage of the differential drive tube and the node voltage in the voltage divider circuit, the two input terminals of the self-biased amplifier are formed, thereby realizing a negative feedback mechanism and converting V OH Clamped within the voltage range set by the voltage divider circuit. Fig. 9 The working principle of 901, that is, operational amplifier A2, is similar to this.

[0030] However, in actual operation, the M3-M6 tubes have a large area, high parasitic capacitance, and the operating frequency far exceeds the bandwidth of the operational amplifier, resulting in V OH and VOL Instability may occur. Therefore, in order to suppress these fluctuations, two larger MOS tube capacitors, M47 and M48, are added to the circuit to stabilize V OH and V OL This can effectively improve the stability of the signal and prevent signal distortion caused by parasitic effects.

[0031] In addition, an enable signal V is added to the circuit en , through the path formed by transistors such as M9, M14, M20, and M25, the switch of the driving circuit is controlled. en When it is at a high level, M9 and M20 will be turned on, pulling the lower node to a high level and pulling the upper node to a low level respectively, thereby causing M1 and M2 to stop working normally, thereby turning off the drive circuit.

[0032] The self-starting circuit of 900 is composed of M11-M13. en When it is low, the circuit enters the normal working state, M11 remains on, and M12 and M13 provide a current path to the ground through the diode connection to prevent the circuit from being in a zero current state. This design ensures that the circuit can start smoothly, and after the operational amplifier works normally, by adjusting the width-to-length ratio of M12, its source voltage meets V gs -V th ≤0, which eventually makes the startup circuit current zero. The startup circuit design on the 901 side is similar to this.

[0033] In order to ensure the accuracy of data transmission and reduce the bit error rate, it is usually necessary to introduce signal equalization technology into the structural design. Fig.10 The active continuous-time linear equalizer 10 shown in FIG. 1 becomes an important component for solving signal attenuation and distortion. 10 is a fully differential circuit structure, and its input end is connected to the source through a differential pair. The source is connected to resistor R2 and capacitor C2 respectively. For the signal, resistor R2 acts as an all-pass filter, which has no significant effect on the frequency response of the signal, while capacitor C2 provides a dedicated high-frequency signal channel. The combination of these two components is equivalent to forming a high-pass filter, which can effectively supplement the high-frequency component of the signal, thereby compensating for the signal distortion caused by the attenuation of the high-frequency component during the transmission process. In this way, 10 restores the output signal to a state close to the original state by compensating for the high-frequency part of the signal, thereby ensuring the identifiability of the signal at the receiving end.

[0034] Furthermore, when the functions of resistors and capacitors are realized by MOS tubes, the circuit can not only adjust the size of resistors and capacitors, but also adjust the current and signal gain and bandwidth by controlling the gate voltage. This adjustability enables 10 to flexibly cope with signal attenuation under different transmission conditions and adapt to a variety of channel environments. By adjusting VBP To control the current flowing through M1 and M5 tubes, thereby adjusting the transconductance g m , thereby controlling the gain and frequency response of the circuit. Specifically, adjusting V BP Voltage can change the position of zeros and poles in a circuit, thereby affecting the amplitude and phase characteristics of the signal.

[0035] Fig.11 The bias current source 11 shown is used to provide a stable operating point for the circuit, thereby ensuring that the amplifier operates in a suitable region. By selecting an appropriate number of bipolar transistors and a suitable resistance value, the target reference current can be accurately generated. The design of the reference current not only needs to meet the bias requirements of the circuit module, but also needs to consider the temperature compensation characteristics to ensure stable performance within the operating temperature range.

[0036] Regarding the startup problem of 11, at the moment the system is powered on, the initial current of all transistors may be zero. Since zero current is allowed to pass through the loop at this time, the entire circuit may fall into a "zero current lock" state, also known as a degenerate problem. Therefore, 11 introduces a self-starting circuit, which is composed of devices such as NM3, NM4 and PM3, and its function is to force the circuit out of the zero current state. Specifically, the self-starting circuit injects a startup current at the moment of power-on to push the main circuit into a normal working state. When the bias current is established, the self-starting circuit will automatically exit the working state to avoid affecting the normal operation of the circuit. In 11, NM3 and NM4, as part of the startup path, provide a short current path through the control signal, while PM3 assists in forming a startup loop. When the circuit enters steady-state operation, the current of the self-starting circuit will drop to zero and will no longer affect the main circuit.

[0037] like Fig.12The differential signal comparator 12 shown is one of the core components in the high-speed data receiving end circuit. Its main function is to process the analog differential signal received from the transmission channel and convert it into a full-swing CMOS digital signal for subsequent digital circuit processing. 12 adopts a two-stage operational amplifier structure. The first-stage operational amplifier is an amplifier with differential input and dual-end output, which is responsible for amplifying the input weak differential signal. The core components of this stage of operational amplifier are cross-coupled MOS tubes M3 and M4, which are coupled with the input differential signal as load resistors. This cross-coupling structure enables the difference between the two input signals to be effectively amplified. The MOS tubes M3 and M4 in the cross-coupling structure have complementary working characteristics, so that the positive and negative differences of the input signals can enhance each other, thereby increasing the gain and reducing the distortion of the circuit. In addition, the first-stage amplifier also includes diode-connected MOS tubes M5 and M6, which are combined with the current mirror component to form a current replicator. M5 and M6 are controlled by the current mirror to copy and transmit the current output from the first-stage operational amplifier to the two branches of the second-stage amplifier. In this way, the output current of the first-stage op amp is evenly distributed to the second stage, ensuring that the operation of the subsequent amplifier can remain consistent and run stably.

[0038] The second-stage op amp of 12 adopts a differential input single-ended output structure, which is responsible for converting the signal amplified in the first stage into a single-ended output, so as to adapt to the input requirements of the digital circuit. At this stage, the signal has been amplified enough to drive the CMOS logic circuit to work. The core components of the second-stage amplifier are the two pairs of MOS tubes M8 and M10, which, under the action of the differential input signal, influence each other through the current mirror to control the output swing. As the input signal changes, the current flowing through M5 and M6 will be significantly different. This current difference will eventually be transmitted to M8 and M10, causing the current difference between the two MOS tubes to gradually increase. When the current difference reaches a certain level, M8 will enter the deep linear region, that is, its working state will change from the normal linear amplification region to the strong nonlinear region, and the output signal will quickly reach a full-swing digital level. This process is the decision-making process of 12, in which the input signal is successfully converted into a clear logic level.

[0039] The present invention directly connects two MACs and an adder at the physical layer through a VML physical interface without parallel-to-serial conversion or serial-to-parallel conversion circuit, thereby realizing two-layer convolution calculation. At the same time, only one MAC can be retained to perform signal processing tasks similar to Gaussian filtering, or multiple MACs can form a higher-dimensional convolution calculation unit. This scalable chip interconnection component may be the basis of future modular algorithm accelerators. Different accelerator hardware structures can be formed by expanding or reconstructing the basic components to adapt to different algorithms and improve the reconfigurability of the system.

[0040] The above is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any professional and technical personnel familiar with the present invention can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement made to the above embodiments without departing from the content of the technical solution of the present invention, based on the technical essence of the present invention, within the spirit and principles of the present invention, still fall within the protection scope of the technical solution of the present invention.

Claims

1. A bit serial computing transmission structure for high-speed data processing in a chip interconnect system, characterized in that: At least one bit serial multiplication and accumulation device, each bit serial multiplication and accumulation device is used to receive an input signal and calculate a product result, the input signal includes a low-order, a high-order and a head-order part, wherein each multiplication and accumulation device has a bit serial multiplier of a base-4 Booth encoding circuit, an adder suitable for multiplication and accumulation calculation, and a plurality of registers, and the result output by the multiplication and accumulation device is in a bit serial format, including a low-order and a high-order output signal; At least one bit serial adder, used for receiving the output result of the bit serial multiplication and accumulation device and performing accumulation, the output result of the bit serial adder is a low-order and a high-order signal, and the accumulation process is performed by bit serial bit-by-bit addition, and the calculation result is output in a serial manner; A high-speed serial transmission interface for connection, connecting the bit-serial multiplier-accumulator and the bit-serial adder through a custom-designed voltage mode logic (VML) interface, wherein the VML interface includes a driver circuit for converting serial data from the transmitting end into a differential signal and transmitting it to the receiving end; an equalizer that uses a differential structure to compensate for attenuation and distortion during signal transmission and enhances the high-frequency portion to ensure signal integrity; A differential signal comparator is used to accurately identify the received differential signal and convert it into a digital signal to ensure the accuracy of high-speed data transmission; a bias current source is used to provide a stable operating point for the circuit, and an RS trigger is used to shape the signal, remove high-frequency noise and unnecessary waveform distortion, and ensure high-precision data transmission.

2. The bit serial computing transmission structure for high-speed data processing in a chip interconnection system according to claim 1, characterized in that: The bit-serial multiplier-accumulator shares the clock signal and the header bit signal in the structure to synchronize the calculation process.

3. The bit serial computing transmission structure for high-speed data processing in a chip interconnection system according to claim 1, characterized in that: The bit-serial adder in the bit-serial multiply-accumulator includes a head bit control signal with appropriate delay to ensure proper alignment of the bit-serial data when it is accumulated.

4. The bit serial computing transmission structure for high-speed data processing in a chip interconnection system according to claim 1, characterized in that: The driver circuit includes a voltage divider circuit that can provide different signal driving strengths to adapt to different transmission environments.

5. The bit serial computing transmission structure for high-speed data processing in a chip interconnection system according to claim 1, characterized in that: The equalizer adopts a differential structure, which effectively enhances the high-frequency part of the signal through the all-pass filter and the high-frequency signal channel, and reduces the impact of signal attenuation.

6. The bit serial computing transmission structure for high-speed data processing in a chip interconnection system according to claim 1, characterized in that: The differential signal comparator works through a two-stage operational amplifier structure, which can accurately identify and convert the received high-speed signal and enhance the anti-interference ability of the structure.

7. The bit serial computing transmission structure for high-speed data processing in a chip interconnection system according to claim 1, characterized in that: It adopts a modular design, supports combinations of different numbers of multipliers and adders, and can be flexibly expanded according to different algorithm requirements.

8. The bit serial computing transmission structure for high-speed data processing in a chip interconnection system according to claim 1, characterized in that: It combines digital circuits and analog circuits to form an efficient computing and data transmission architecture with lower power consumption and higher transmission rate.

Citation Information

Patent Citations

  • Approximate-computation-based binary weight convolution neural network hardware accelerator calculating module

    CN106909970A

  • Logarithmic approximation multiply-accumulator for convolutional neural network accelerator

    CN113360131A

  • Multi-input serial adder

    CN115509488A

  • Electroencephalogram signal hardware acceleration recognition system based on hybrid neural network

    CN116226717A

  • High-speed serial data receiving module

    CN117931712A