Multiply-accumulator with signed fixed-point number in superconductive rapid single magnetic flux sub-circuit and related products
By designing a signed fixed-point multiplication and accumulator in a superconducting fast single-flux quantum circuit and using logic gates and shift counters connected in series, simultaneous calculation of multiplication and addition is achieved, which solves the performance loss problem caused by the feedback loop in the RSFQ circuit and improves the operating frequency and efficiency.
Patent Information
- Application Number
- CN202410250099.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2025-09-05
AI Technical Summary
The multiplication and accumulation calculation circuit of traditional CMOS technology is difficult to transplant to RSFQ technology, resulting in the existence of the feedback loop in the RSFQ circuit greatly reducing the operating frequency and performance, and the multiplication and accumulation calculation circuit design of the existing RSFQ technology is insufficient.
A signed fixed-point multiplication and accumulation device for superconducting fast single-flux quantum circuits is designed. Logic gates, multiple shift counters, and data registers are connected in series to achieve simultaneous calculation of multiplication and addition operations and avoid the generation of feedback loops.
The operating frequency and performance of the RSFQ circuit are improved, the generation of feedback loops is avoided, and the computing efficiency is improved.
Smart Images

Figure CN120596060A_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to the field of multiplication and accumulation technology. More specifically, the present application relates to a multiplication and accumulation device, a multiplication and accumulation circuit, a multiplication and accumulation method, a multiplication and accumulation device, and a computer-readable storage medium for signed fixed-point numbers in a superconducting fast single-flux quantum circuit. Background Art
[0002] Multiplication and accumulation are fundamental operations in vector and matrix multiplication and are widely used in fields such as matrix computation and artificial intelligence. Furthermore, deep neural networks are state-of-the-art in a wide range of artificial intelligence tasks, including object recognition and speech recognition, which also involve a large number of multiplication and accumulation operations. To achieve high performance and high energy efficiency, significant research efforts have been devoted to exploring hardware acceleration methods for deep neural networks. Currently, most of these designs rely on complementary metal-oxide-semiconductor (CMOS) technology, which faces performance limitations as Moore's Law approaches its limits. To address this challenge, innovative technologies such as quantum computing, neuromorphic computing, approximate computing, and stochastic computing have emerged as potential solutions. Among these candidates, the superconducting rapid single-flux quantum (RSFQ) logic family based on Josephson junctions (JJs) is considered a highly promising solution.
[0003] However, conventional CMOS-based multiplication and accumulation circuits (including bit-parallel and bit-serial) usually have feedback loops (e.g. Figure 1 Directly migrating a CMOS-based multiplication-accumulation circuit structure to RSFQ technology would make it difficult to eliminate the feedback loop in the accumulation portion of the circuit, significantly reducing the operating frequency of the RSFQ circuit and resulting in significant performance loss. Furthermore, current research on RSFQ-based computing circuits has mostly focused on standalone addition or multiplication circuits, with little research on the design of multiplication-accumulation circuits that combine both multiplication and addition.
[0004] In view of this, there is an urgent need to provide a multiplication and accumulation scheme for signed fixed-point numbers in superconducting fast single-flux quantum circuits, so as to simultaneously calculate the multiplication and accumulation of signed fixed-point numbers, avoid feedback loops, and improve the operating frequency and performance of RSFQ circuits. Summary of the Invention
[0005] In order to at least solve one or more technical problems mentioned above, the present application proposes a multiplier-accumulator solution for signed fixed-point numbers in a superconducting fast single-flux quantum circuit in multiple aspects.
[0006] In a first aspect, the present application provides a multiplication and accumulation device for signed fixed-point numbers in a superconducting fast single-flux quantum circuit, comprising: a logic gate for performing a single-bit product operation on a target bit in the signed fixed-point number to obtain a single-bit partial product; a plurality of shift counters, wherein the plurality of shift counters are connected in series and are used to perform an accumulation operation on the single-bit partial product according to a shift count to obtain a multiplication-accumulation partial sum; a plurality of first data registers, which are correspondingly connected to the plurality of shift counters, are used to receive and transmit the multiplication-accumulation partial sum and perform a shift operation in the accumulation operation; and a plurality of second data registers, which are correspondingly connected to the plurality of first data registers, are used to receive and temporarily store the multiplication-accumulation partial sum.
[0007] In one embodiment, the single-bit partial product includes a first partial product and a second partial product, and the first partial product includes a partial product of a sign bit multiplied by a value bit multiplied by the signed fixed-point number; the second partial product includes a partial product of a sign bit multiplied by a value bit in the signed fixed-point number.
[0008] In another embodiment, in performing an accumulation operation on the single-bit partial product according to the shift count to obtain the multiplication-accumulation partial sum, the multiple shift counters are further used to: perform an accumulation operation on the partial product of the first part and the partial product of the second part under the same weight coefficient according to the shift count to obtain the first multiplication-accumulation partial sum and the second multiplication-accumulation partial sum respectively.
[0009] In another embodiment, the plurality of shift counters are further configured to: in response to completion of the target bit input for calculating the partial product of the first part or the partial product of the second part under the same weight coefficient, perform a pause operation and wait for ripple carry transfer.
[0010] In another embodiment, the multiple shift counters are further used to: in response to a first control signal, transfer the first multiplication-accumulation partial sum or the second multiplication-accumulation partial sum corresponding to the partial product of the first part or the partial product of the second part under the same weight coefficient to the multiple first data registers.
[0011] In another embodiment, the multiple first data registers are further used to: in response to a second control signal, reload the first multiplication-accumulation partial sum or the second multiplication-accumulation partial sum corresponding to the partial product of the first part or the partial product of the second part under the same weight coefficient into the multiple shift counters to realize the shift operation.
[0012] In another embodiment, the multiple first data registers are further used to: receive the first multiplication-accumulation partial sum and the second multiplication-accumulation partial sum corresponding to the partial product of the first part or the partial product of the second part under the same weight coefficient for the target times transmitted by the multiple shift counters under the first control signal of the target group number; and in response to the second control signal of the target group number, reload the first multiplication-accumulation partial sum or the second multiplication-accumulation partial sum corresponding to the partial product of the first part or the partial product of the second part under the same weight coefficient for the target times into the multiple shift counters to achieve the target number of bits moved.
[0013] In another embodiment, the plurality of first data registers are further configured to transfer the first multiplication-accumulation partial sum or the second multiplication-accumulation partial sum to the plurality of second data registers in response to a third control signal.
[0014] In another embodiment, the plurality of second data registers are further used to: transmit the first multiplication-accumulation partial sum and the second multiplication-accumulation partial sum to an external subtractor in response to a fourth control signal, so that the external subtractor performs a target subtraction operation based on the first multiplication-accumulation partial sum and the second multiplication-accumulation partial sum to obtain a final result of the multiplication-accumulation.
[0015] In a second aspect, the present application provides a multiplication-accumulation circuit for signed fixed-point numbers in a superconducting fast single-flux quantum circuit, comprising: a multiplication-accumulation circuit according to the aforementioned first aspect; and an external subtractor.
[0016] In a third aspect, the present application provides a multiplication and accumulation method for signed fixed-point numbers in a superconducting fast single-flux quantum circuit, comprising: performing a single-bit product operation on a target bit in the signed fixed-point number to obtain a single-bit partial product; performing an accumulation operation on the single-bit partial product according to a shift count to obtain a multiplication-accumulation partial sum; receiving and transmitting the multiplication-accumulation partial sum, and performing a shift operation in the accumulation operation; and receiving and temporarily storing the multiplication-accumulation partial sum.
[0017] In a fourth aspect, the present application provides a device for multiplying and accumulating signed fixed-point numbers in a superconducting fast single-flux quantum circuit, comprising: a processor; and a memory on which computer instructions for multiplying and accumulating signed fixed-point numbers in a superconducting fast single-flux quantum circuit are stored. When the computer instructions are executed by the processor, the multiplication and accumulation method in the aforementioned third aspect is implemented.
[0018] In the fifth aspect, the present application provides a computer-readable storage medium, characterized in that computer program instructions for multiplication and accumulation of signed fixed-point numbers in a superconducting fast single-flux quantum circuit are stored thereon, and when the computer program instructions are executed by one or more processors, the multiplication and accumulation method in the aforementioned third aspect is implemented.
[0019] The multiplication and accumulation device for signed fixed-point numbers in a superconducting fast single-flux quantum circuit, as provided above, implements simultaneous and direct multiplication and accumulation, as well as shift operations, through logic gates, multiple bit-serial shift counters, and multiple first and second data registers to obtain a multiplication-accumulation partial sum. This avoids the generation of feedback loops and significantly improves the operating frequency and performance of the RSFQ circuit. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other objects, features and advantages of the exemplary embodiments of the present application will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0021] Figure 1 is an exemplary schematic diagram showing a conventional multiplication-accumulation calculation circuit based on CMOS technology;
[0022] Figure 2 is an exemplary structural block diagram showing a multiplication and accumulation device for signed fixed-point numbers in a superconducting fast single-flux quantum circuit according to an embodiment of the present application;
[0023] Figure 3 is an exemplary schematic diagram illustrating a multiplication-accumulation calculation sequence executed by a multiplication-accumulation unit according to an embodiment of the present application;
[0024] Figure 4 is an exemplary schematic diagram showing a multiplication and accumulation device for signed fixed-point numbers in a superconducting fast single-flux quantum circuit according to an embodiment of the present application;
[0025] Figure 5 is an exemplary schematic diagram showing symbols of various components and a state machine in a multiplier-accumulator according to an embodiment of the present application;
[0026] Figure 6 is an exemplary structural block diagram showing a multiplication and accumulation circuit for signed fixed-point numbers in a superconducting fast single-flux quantum circuit according to an embodiment of the present application;
[0027] Figure 7 is an exemplary flow chart illustrating a method for multiplying and accumulating signed fixed-point numbers in a superconducting fast single-flux quantum circuit according to an embodiment of the present application;
[0028] Figure 8 1 is an exemplary structural block diagram showing a device for multiplying and accumulating signed fixed-point numbers in a superconducting fast single-flux quantum circuit according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0030] It should be understood that the terms "include" and "comprising" used in the description and claims of this application indicate the presence of described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0031] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this specification and claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" as used in this specification and claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.
[0032] As used in this specification and claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0033] As can be seen from the background description above, current multiplication and accumulation calculations are often designed based on traditional CMOS technology. However, as Moore's Law gradually approaches its limit, this technology has performance limitations. Currently, the JJ-based superconducting RSFQ logic family is considered to be a very promising solution. This is because superconducting RSFQ has extremely fast switching speed (~1ps) and extremely low switching energy (~10 -19Using this technology, the clock frequency of the device can be increased by 1 to 2 orders of magnitude (i.e., tens of gigahertz to hundreds of gigahertz (GHz)). In other words, under the same process, the logic gate delay of the RSFQ circuit is 1 to 2 orders of magnitude lower than that of the CMOS circuit. Studies have shown that the operating frequency of the T flip-flop (TFF) based on RSFQ technology can reach 770 GHz at liquid helium temperature, which makes RSFQ circuits different from CMOS circuits.
[0034] Specifically, CMOS circuits use electrical levels as the carrier for information storage and transmission, while RSFQ circuits use single magnetic flux quanta Φ0 and voltage pulses as the carriers for information storage and transmission, respectively. CMOS circuits use transistors as switching elements, while RSFQ circuits use Josephson junctions. The logic gates responsible for computation in CMOS circuits are mostly combinational, while the majority of logic gates in RSFQ circuits are sequential, requiring clock drive. In other words, each RSFQ logic gate has a latching function, enabling gate-level pipelining without additional registers. CMOS circuits use an equipotential clock scheme based on electrical levels, while RSFQ circuits use a stream clock scheme based on voltage pulses. In this mechanism, the presence of a feedback loop significantly reduces the operating frequency of RSFQ circuits. Furthermore, the sequential logic in CMOS circuits can retain electrical level signals for extended periods of time, while the logic gates in RSFQ circuits cannot. Under the control of a drive signal, pulse signals stored within the RSFQ gates can be destructively read out. Therefore, RSFQ circuits place even stricter demands on timing control.
[0035] Therefore, based on the differences between the RSFQ circuit and the CMOS circuit, if the multiplication and accumulation (including bit-parallel and bit-serial) calculation circuit structure based on CMOS technology is directly transplanted to RSFQ technology, it will be difficult to eliminate the feedback loop of the accumulation part in the circuit, which greatly reduces the operating frequency of the RSFQ circuit and causes a huge performance loss.
[0036] Figure 1 FIG. 1 is an exemplary schematic diagram showing a conventional multiplication-accumulation circuit based on CMOS technology. Figure 1 As shown in the figure, in a conventional CMOS-based multiplication-accumulation circuit (including bit-parallel and bit-serial circuits), multiplication-accumulation is performed by a processing unit ("PE") equipped with a multiplier and an adder. In this scenario, after completing the multiplication operation ("×") and the addition operation ("+"), the multiplication-accumulation part and the sum Psum are obtained. This sum needs to be fed back (for example, as shown by arrow A in the figure) to perform the next multiplication-accumulation operation, thereby forming a feedback loop.
[0037] However, in RSFQ circuits using a voltage-pulse-based streaming clock mechanism, the presence of a feedback loop significantly reduces the operating frequency of the RSFQ circuit, resulting in significant performance loss. Furthermore, current research on computational circuits based on RSFQ technology has mostly focused on standalone addition or multiplication circuits, with little research on the design of multiplication-accumulation circuits that combine multiplication and addition.
[0038] Based on this, an embodiment of the present application proposes a multiplier-accumulator for signed fixed-point numbers in a superconducting fast single-flux quantum circuit to avoid simultaneously calculating the multiplication and accumulation of signed fixed-point numbers, avoid feedback loops, and improve the operating frequency and performance of the RSFQ circuit.
[0039] For ease of understanding, before describing the embodiments of this application, the multiplication and accumulation calculation process of signed fixed-point numbers is first described in detail. It can be understood that the common form of multiplication and accumulation calculation is Z = X1*Y1+X2*Y2+…+X k *Y k +…, where X and Y represent signed fixed-point numbers, each of which can be expanded based on the n-bit signed fixed-point number represented by the following formula (1), which includes the binary form of the signed fixed-point number and the weight coefficient of each bit, where x n-1 is the sign bit. Therefore, the binary multiplication of two signed fixed-point numbers can be expressed by the following formula (2). From formula (2), it can be seen that the binary multiplication of signed fixed-point numbers mainly includes 4 terms, namely x n-1 *y n-1 Indicates the multiplication of two sign bits, 2 2n-2 is the weight coefficient of the partial product; Indicates that the values of x and y are multiplied, multiplied by the corresponding weight coefficients and then accumulated; Indicates that the sign bit of x is multiplied by each value bit of y, and then multiplied by the corresponding weight coefficient and accumulated; Indicates that the sign bit of y is multiplied by each numeric bit of x, multiplied by the corresponding weight coefficient and then accumulated.
[0040] Therefore, the binary multiplication of the above-mentioned signed fixed-point numbers can be divided into two parts. That is, the addition part ("positive") containing the partial product of the sign bit multiplication and the partial product of the value bit multiplication, which is also the first part in the context of this application, and the partial product of the first part contains the partial product of the sign bit multiplication and the partial product of the value bit multiplication. The subtraction part ("negative") containing the partial product of the sign bit and the value bit multiplication, which is also the second part in the context of this application, and the partial product of the second part contains the partial product of the sign bit and the value bit multiplication. In addition, it should be understood that the partial product in the context of this application refers to the result of multiplying two bits, and the partial sum refers to the intermediate result of accumulating the partial products according to the corresponding weight coefficients.
[0041] X=(x n-1 x n-2 …x0)2=-x n-1 *2 n-1 +x n-2 *2 n-2 +…+x0*2 0 (1)
[0042]
[0043] Taking 4-bit signed number multiplication as an example, its binary multiplication P = A * B can be expressed as formula (3). Taking 4-bit signed number multiplication and accumulation as an example, its binary multiplication S = A * B + C * D can be expressed as formula (4). By dividing the multiple terms in formula (4) into the addition part ("positive") and the subtraction part ("negative"), S = A * B + C * D = positive - negative. Specifically, the addition part ("positive") can be expressed as formula (5), and the subtraction part ("negative") can be expressed as formula (6).
[0044] P=A*B=(-a3*2 3 +a2*2 2 +a1*2+a0)*(-b3*2 3 +b2*2 2 +b1*2+b0)
[0045] =a3*b3*2 6 +a2*b2*2 4 +(a1*b2+a2*b1)*2 3 +(a0*b2+a1*b1+a2*b0)*2 2 +(a0*b1+a1*b0)*2+a0*b0-[(a2*b3+a3*b2)*2 5 +(a1*b3+a3*b1)*2 4 +(a0*b3+a3*b0)*2 3 ] (3)
[0046] S=A*B+C*D=(a3*b3+c3*d3)*2 6 +(a2*b2+c2*d2)*2 4 +(a1*b2+c1*d2+a2*b1+c2*d1)*2 3 +(a0*b2+c0*d2+a1*b1+c1*d1+a2*b0+c2*d0)*2 2+(a0*b1+c0*d1+a1*b0+c1*d0)*2+(a0*b0+c0*d0)-[(a2*b3+c2*d3+a3*b2+c3*d2)*2 5 +(a1*b3+c1*d3+a3*b1+c3*d1)*2 4 +(a0*b3+c0*d3+a3*b0+c3*d0)*2 3 ] (4)
[0047] positive=(a3*b3+c3*d3)*2 6 +(a2*b2+c2*d2)*2 4 +(a1*b2+c1*d2+a2*b1+c2*d1)*2 3 +(a0*b2+c0*d2+a1*b1+c1*d1+a2*b0+c2*d0)*2 2 +(a0*b1+c0*d1+a1*b0+c1*d0)*2+(a0*b0+c0*d0)(5)
[0048] negative=(a2*b3+c2*d3+a3*b2+c3*d2)*2 5 +(a1*b3+c1*d3+a3*b1+c3*d1)*2 4 +(a0*b3+c0*d3+a3*b0+c3*d0)*2 3 (6)
[0049] The above is the multiplication and accumulation calculation form. The specific implementation of this application is described in detail below with reference to the accompanying drawings.
[0050] Figure 2 FIG. 2 is an exemplary structural block diagram showing a multiplication and accumulation device 200 for a signed fixed-point number in a superconducting fast single-flux quantum circuit according to an embodiment of the present application. Figure 2 As shown in FIG, the multiplier-accumulator 200 may include a logic gate 201, a plurality of shift counters 202, a plurality of first data registers 203, and a plurality of second data registers 204.
[0051] In one embodiment, logic gate 201 can be used to perform a single-bit product operation on a target bit in a signed fixed-point number to obtain a single-bit partial product. The single-bit partial product includes a first partial product and a second partial product, and the first partial product includes a partial product of the multiplication of the sign bit and the value bit in the signed fixed-point number; the second partial product includes a partial product of the multiplication of the sign bit and the value bit in the signed fixed-point number. That is, the first part is the addition part ("positive") described above, and the second part is the subtraction part ("negative") described above.
[0052] Taking the above-mentioned 4-bit signed fixed-point number as an example, the target bits are, for example, a3, a2, a1, a0 and c3, c2, c1, c0, and the single-bit partial products are, for example, the partial products a3*b3, c3*d3 in the addition (first) part and the partial products a2*b3, c2*d3 in the subtraction (second) part. In one implementation scenario, the logic gate 201 is a logic AND gate (AND) in the RSFQ circuit, which implements single-bit multiplication through a logic AND operation. During the implementation process, a pulse is formed and stored inside the AND only when pulses are fed into both ends of the logic AND gate (for example, the A end and the B end) at the same time (that is, there are two bit inputs with a value of 1). When a pulse arrives at the clock signal end, the pulse stored in the AND is destructively read out from the output end (for example, the Q end).
[0053] In one embodiment, a plurality of shift counters 202 are connected in series and are used to perform an accumulation operation on the single-bit partial products according to the shift count to obtain a multiplication-accumulation partial sum. Specifically, the plurality of shift counters 202 are further used to perform an accumulation operation on the partial products of the first part and the partial products of the second part under the same weight coefficient according to the shift count to obtain a first multiplication-accumulation partial sum and a second multiplication-accumulation partial sum respectively. In some embodiments, the multiplication-accumulation partial sum of the partial products of the first part can be calculated first, and after the calculation of the first part is completed, the multiplication-accumulation partial sum of the partial products of the second part is calculated. Furthermore, when calculating the addition part and the subtraction part, the partial products under the same weight coefficient are first multiplied and accumulated.
[0054] In one implementation scenario, the aforementioned shift counter can be, for example, a T1 trigger, the operation of which is modulo 2 and carry 1, and has the function of a full adder, whose base sum is latched internally and driven to output by a control signal. A plurality of serially connected T1 triggers can be regarded as a counter, that is, an accumulator with a data input width of a single bit. Specifically, when a pulse is fed into the input end (e.g., the D end) of the T1 trigger and no pulse is stored inside the T1 trigger, a 1 pulse will be stored inside T1. When a pulse is fed into the input end (e.g., the D end) of the T1 trigger and a pulse is stored inside the T1 trigger, a pulse will be formed at its output end (e.g., the Q end). When the control signal arrives, the pulse stored in T1 will be destructively read out from the sum end (e.g., the sum end). In an embodiment of the present application, the base sum of the two bits added will be stored inside the T1 trigger, and its carry will be derived from, for example, the Q end.
[0055] As an example, assume that 32 serially connected shift counters are set up, which can form a 32-bit counter, and the accumulation operation of the partial products of the same weight coefficient in the first and second parts is achieved by counting, and the left side of the 32-bit counter is the low bit and the right side is the high bit. In one implementation scenario, the multiple shift counters 202 are further used to: in response to the completion of the input of the target bit of the partial product of the first part or the partial product of the second part under the same weight coefficient, perform a pause operation and wait for the ripple carry transfer. Taking the partial products a3*b3 and c3*d3 under the same weight coefficient in the addition part as an example, after the input of a3, b3, c3, and d3 is completed, the shift counter pauses data input and performs the accumulation operation. A carry may be generated during the accumulation process to wait for the completion of the ripple carry transfer operation.
[0056] In one implementation scenario, in response to a first control signal, a first multiplication-accumulation partial sum or a second multiplication-accumulation partial sum corresponding to the partial products of the first part or the partial products of the second part under the same weight coefficient is transferred to a plurality of first data registers. The first control signal is a read signal. That is, after the accumulation operation of the partial products under the same weight coefficient in the addition part or the subtraction part is completed, the control signal read arrives at the shift counter, and the multiplication-accumulation partial sum in the shift counter is transferred to the first data register.
[0057] In one embodiment, a plurality of first data registers 203 are connected to a plurality of shift counters 202 for receiving and transmitting the multiplication-accumulation partial sums and performing the shift operation in the accumulation operation. In some embodiments, the first data register 203 is a data register with dual clocks and dual output ports ("D2FF"). During implementation, the pulse fed into the D2FF storage terminal (e.g., the D terminal) is destructively read from one output terminal (e.g., the QA terminal) when a pulse arrives at one clock signal terminal (e.g., the clk1 terminal). When a pulse arrives at the other clock signal terminal (e.g., the clk2 terminal), the pulse stored in the D2FF is destructively read from the other output terminal (e.g., the QB terminal).
[0058] Specifically, in response to a second control signal, the first multiplication-accumulation partial sum or the second multiplication-accumulation partial sum corresponding to the partial products of the first part or the partial products of the second part under the same weight coefficient is reloaded into the plurality of shift counters to implement a shift operation. The second control signal is a shift signal. When the shift signal arrives at the clk1 port of the D2FF, the partial sum is reloaded into the shift counters via the QA port of the D2FF, with the value being one bit higher than the original value.
[0059] More specifically, the plurality of first data registers 203 are further configured to receive a first multiplication-accumulation partial sum or a second multiplication-accumulation partial sum corresponding to the partial products of the first part or the partial products of the second part under the same weight coefficient transferred a target number of times by the plurality of shift counters under the first control signal of the target group number; and, in response to the second control signal of the target group number, reload the first multiplication-accumulation partial sum or the second multiplication-accumulation partial sum corresponding to the partial products of the first part or the partial products of the second part under the same weight coefficient for the target number of times into the plurality of shift counters to achieve the target number of bits of movement. In other words, depending on the target number of bits to be moved, the first control signal and the second control signal of the target group number need to be input to transfer the target number of times between the plurality of shift counters and the plurality of first data registers.
[0060] For example, in one exemplary scenario, if a 2-bit left shift is required, another set of read and shift signals is input, causing the partial sums in the multiple shift counters to be transferred to the multiple first data registers again, and the multiple first data registers to reload the partial sums into the multiple shift counters again. Similarly, if a k-bit left shift is required, k-1 more sets of read and shift signals are input.
[0061] Furthermore, the plurality of first data registers 203 are further configured to transmit the first multiplication-accumulation partial sum or the second multiplication-accumulation partial sum to the plurality of second data registers in response to a third control signal. That is, after the multiplication-accumulation of the addition portion or the subtraction portion is completed, the multiplication-accumulation partial sum of the addition portion or the subtraction portion is transmitted to the plurality of second data registers based on the third control signal. The third control signal is a done signal.
[0062] In one embodiment, the plurality of second data registers 204 can be used to receive and temporarily store the multiplication-accumulation partial sums. Specifically, in response to a fourth control signal, the first multiplication-accumulation partial sum and the second multiplication-accumulation partial sum are transmitted to an external subtractor, so that the external subtractor performs a target subtraction operation based on the first multiplication-accumulation partial sum and the second multiplication-accumulation partial sum to obtain a final multiplication-accumulation result. The fourth control signal is a sample signal. That is, driven by the sample signal, the multiplication-accumulation partial sum of the addition part and the multiplication-accumulation partial sum of the subtraction part are input to the external subtractor, and the multiplication-accumulation partial sum of the addition part and the multiplication-accumulation partial sum of the subtraction part are subtracted in the external subtractor to obtain a final multiplication-accumulation result.
[0063] In some implementation scenarios, the plurality of second data registers may be, for example, DFFs, which are implemented by storing pulses fed into the input terminal (e.g., the D terminal). When a pulse arrives at the clock signal terminal, the pulse stored in the DFF is destructively read out from the output terminal (e.g., the Q terminal), thereby transmitting the multiplication-accumulation partial sum to the external subtractor. As can be seen from the foregoing, the multiplication-accumulation calculation of the addition part can be completed first, and then the multiplication-accumulation calculation of the subtraction part can be calculated. In some embodiments, when the multiplication-accumulation partial sum of the addition part is transmitted to the second data register, the multiplication-accumulation calculation of the subtraction part can be started.
[0064] As can be seen from the above description, the embodiments of the present application utilize specialized logic gates, multiple bit-serialized shift counters, and multiple first and second data registers within the RSFQ circuit to simultaneously and directly perform multiplication, accumulation, and shift operations to obtain the multiplication-accumulation partial sum. This avoids the generation of a feedback loop and significantly improves the operating frequency and performance of the RSFQ circuit.
[0065] Figure 3 1 is an exemplary diagram showing the multiplication and accumulation calculation sequence executed by the multiplication and accumulation unit according to an embodiment of the present application. Figure 3As shown in the figure, taking the multiplication and accumulation calculation of a 4-bit signed number as an example, the multiplication and accumulation device of the embodiment of the present application is used to first calculate the addition part ("positive"), then calculate the subtraction part ("negative"), and finally subtract the positive and negative in the external subtractor to obtain the final result of the multiplication and accumulation calculation. Among them, the multiplication and accumulation calculation order of the addition part ("positive") is: a3*b3+c3*d3; (a3*b3+c3*d3)*2 2 ;(a3*b3+c3*d3)*2 2 +a2*b2+c2*d2;[(a3*b3+c3*d3)*2 2 +a2*b2+c2*d2]*2;[(a3*b3+c3*d3)*2 2 +a2*b2+c2*d2]*2+a1*b2+c1*d2;[(a3*b3+c3*d3)*2 2 +a2*b2+c2*d2]*2+a1*b2+c1*d2+a2*b1+c2*d1;{[(a3*b3+c3*d3)*2 2 +a2*b2+c2*d2]*2+a1*b2+c1*d2+a2*b1+c2*d1}*2;{[(a3*b3+c3*d3)*2 2 +(a2*b2+c2*d2)]*2+a1*b2+c1*d2+a2*b1+c2*d1}*2+a0*b2+c0*d2;{[(a3*b3+c3*d3)*2 2 +(a2*b2+c2*d2)]*2+a1*b2+c1*d2+a2*b1+c2*d1}*2+a0*b2+c0*d2+a1*b1+c1*d1;{[(a3*b3+c3*d3)*2 2 +(a2*b2+c2*d2)]*2+a1*b2+c1*d2+a2*b1+c2*d1}*2+a0*b2+c0*d2+a1*b1+c1*d1+a2*b0+c2*d0;{{[(a3*b3+c3*d3)*2 2 +(a2*b2+c2*d2)]*2+a1*b2+c1*d2+a2*b1+c2*d1}*2+a0*b2+c0*d2+a1*b1+c1*d1+a2*b0+c2*d0}*2;{{[(a3*b3+c3*d3)*2 2+(a2*b2+c2*d2)]*2+a1*b2+c1*d2+a2*b1+c2*d1}*2+a0*b2+c0*d2+a1*b1+c1*d1+a2*b0+c2*d0}*2+a0*b1+c0*d1;{{[(a3*b3+c3*d3)*2 2 +(a2*b2+c2*d2)]*2+a1*b2+c1*d2+a2*b1+c2*d1}*2+a0*b2+c0*d2+a1*b1+ c1*d1+a2*b0+c2*d0}*2+a0*b1+c0*d1+a1*b0+c1*d0;{{{[(a3*b3+c3*d3)*2 2 +(a2*b2+c2*d2)]*2+a1*b2+c1*d2+a2*b1+c2*d1}*2+a0*b2+c0*d2+a1*b1+c1 *d1+a2*b0+c2*d0}*2+a0*b1+c0*d1+a1*b0+c1*d0}*2;{{{[(a3*b3+c3*d3)*2 2 +(a2*b2+c2*d2)]*2+a1*b2+c1*d2+a2*b1+c2*d1}*2+a0*b2+c0*d2+a1*b 1+c1*d1+a2*b0+c2*d0}*2+a0*b1+c0*d1+a1*b0+c1*d0}*2+a0*b0+c0*d0.
[0066] In some embodiments, the multiplication and accumulation calculation of the aforementioned addition part ("positive") can be expressed by the following pseudo code: positive = 0; positive + = a3*b3+c3*d3; positive < < = 2; positive + = a2*b2+c2*d2; positive < < = 1; positive + = a1*b2+c1*d2; positive + = a2*b1+c2*d1; positive < < = 1; positive + = a0*b2+c0*d2; positive + = a1*b1+c1*d1; positive + = a2*b0+c2*d0; positive < < = 1; positive + = a0*b1+c0*d1; positive + = a1*b0+c1*d0; positive < < = 1; positive + = a0*b0+c0*d0.
[0067] Furthermore, the order of multiplication and accumulation calculations for the subtraction part ("negative") is: a2*b3+c2*d3; a2*b3+c2*d3+a3*b2+c3*d2; (a2*b3+c2*d3+a3*b2+c3*d2)*2; (a2*b3+c2*d3+a3*b2+c3*d2)*2+a1*b3+c1*d3; (a2*b3+c2*d3+a3*b2+c3*d2)*2+a1*b3+c1*d3+a3*b1+c3*d1; [(a2*b3+c2*d3+a3*b2+c3*d2)*2+a1*b3+c1*d3+a3*b1+c3 *d1]*2;[(a2*b3+c2*d3+a3*b2+c3*d2)*2+a1*b3+c1*d3+a3*b1+c3*d1 ]*2+a0*b3+c0*d3;[(a2*b3+c2*d3+a3*b2+c3*d2)*2+a1*b3+c1*d3+a3* b1+c3*d1]*2+a0*b3+c0*d3+a3*b0+c3*d0;{[(a2*b3+c2*d3+a3*b2+c3 *d2)*2+a1*b3+c1*d3+a3*b1+c3*d1]*2+a0*b3+c0*d3+a3*b0+c3*d0}*2 3 .
[0068] Correspondingly, the multiplication and accumulation calculation of the subtraction part ("negative") can be expressed by the following pseudo code: negative = 0; negative + = a2*b3+c2*d3; negative + = a3*b2+c3*d2; negative < < = 1; negative + = a1*b3+c1*d3; negative + = a3*b1+c3*d1; negative < < = 1; negative + = a0*b3+c0*d3; negative + = a3*b0+c3*d0; negative < < = 3.
[0069] Depend on Figure 3It can be seen that the embodiment of the present application first calculates the single-bit partial products (such as a3b3, c3d3) under the same weight coefficient by using a bit-serial multiplier-accumulator, and accumulates the partial products under the same weight coefficient to obtain the multiplication-accumulation partial sum. Compared with the existing parallel algorithm, the embodiment of the present application also reduces the area overhead. Specifically, the existing parallel algorithm needs to calculate multiple partial products at a time and then accumulate them, so that multiple logic gates are designed in the circuit, which increases the area overhead, while the present application only designs one logic gate, reducing the area overhead in the multiplication-accumulator. Furthermore, the embodiment of the present application can calculate data of any bit width and any length, while the existing array-type processing unit limits the width and length of the processed data.
[0070] Figure 4 1 is an exemplary schematic diagram showing a multiplication and accumulation device for signed fixed-point numbers in a superconducting fast single-flux quantum circuit according to an embodiment of the present application. It should be understood that Figure 4 is the above Figure 2 A specific embodiment of the multiplication and accumulation device, so the above Figure 2 The description also applies to Figure 4 .
[0071] like Figure 4 As shown in FIG4 , the multiplier-accumulator may include a logic gate (AND 401), multiple shift counters (e.g., T1 flip-flop 402), multiple first data registers (e.g., D2FF 403), and multiple second data registers (e.g., DFF 404). As previously described, the multiple shift counters are connected in series and correspondingly connected to the multiple first data registers and the multiple second data registers.
[0072] In the multiplication-accumulation process, the addition part ("positive") can be calculated first. Specifically, in each clock cycle, two target bits are input via AND 401, which are respectively one bit of the two signed fixed-point numbers to be multiplied. In the addition part, the two target bits input are sign bits or value bits. Then, the multiplication of the two target bits is achieved through the logical AND operation of AND 401 to obtain a single-bit partial product, and the single-bit partial product is transmitted to, for example, a 32-bit counter for counting operation. In the addition part, the single-bit partial product includes the partial product of the sign bit multiplication and the partial product of the value bit multiplication.
[0073] Next, after the data required to calculate all partial products with the same weight coefficients has been input, AND 401 pauses data input and waits for the ripple carry propagation in 32-bit T1 flip-flop 402 to complete. This process requires a certain number of clock cycles. When the value in 32-bit T1 flip-flop 402 stabilizes, the read signal arrives at the rd port of T1 flip-flop 402, and the partial sums in 32-bit T1 flip-flop 402 are transmitted to 32 D2FFs 403.
[0074] In the next clock cycle, the shift signal arrives at the clk1 port of D2FF 403. The partial sum is reloaded into the 32-bit T1 flip-flop 402 via the QA port of D2FF 403, with the value one bit higher than the original value. In some implementations, if a 2-bit left shift is required, another set of read and shift signals is input. Similarly, if a k-bit left shift is required, k-1 more sets of read and shift signals are input.
[0075] Furthermore, AND 401 re-accepts the input of data, and the weight coefficient of the partial product obtained by multiplying the input data is the same as the weight coefficient of the partial sum in the 32-bit T1 flip-flop 402. By repeating the above operations, the multiplication and accumulation calculation of the addition part ("positive") is completed. Entering the next clock cycle, the read signal transmits the multiplication and accumulation partial sum of the addition part of the 32-bit T1 flip-flop 402 to 32 D2FFs 403, and in the next clock cycle, the done signal arrives at the clk2 port of D2FF 403, and the multiplication and accumulation partial sum of the addition part is temporarily stored in the 32-bit DFF 404 through the QB port of D2FF 403. It can be subsequently driven by the sample signal to be transmitted to an external subtractor (not shown in the figure).
[0076] Similar to the addition part mentioned above, the sum of the multiplication and accumulation parts of the subtraction part ("negative") can be calculated and transmitted to the external subtractor driven by the sample signal, so as to be subtracted in the external subtractor to obtain the final result of the multiplication and accumulation calculation.
[0077] In an exemplary scenario, the multiplication and accumulation operation of the multiplication and accumulation device of the embodiment of the present application is described in detail by taking the multiplication and accumulation calculation of a 4-bit signed number (eg, S=A*B+C*D) as an example.
[0078] Initialization: The read, done, and sample signals arrive at the shift counter in sequence, resetting the value in the shift counter to 0. This takes three clock cycles. Next, the shift counter first calculates the positive value.
[0079] First clock cycle: a3 and b3 are input into the shift counter and a logical AND operation is performed.
[0080] Second clock cycle: c3 and d3 are input to the shift counter for a logical AND operation. a3*b3 is transferred to the 32-bit counter for counting.
[0081] 3rd clock cycle: The shift counter pauses data input. c3*d3 is transferred to the 32-bit counter for counting.
[0082] 19th clock cycle: The read signal arrives at the shift counter, and the partial sum a3*b3+c3*d3 in the 32-bit counter is transferred to 32 D2FFs.
[0083] 20th clock cycle: The shift signal reaches the shift counter, the partial sum is shifted left by one bit, and the 32-bit counter is reloaded.
[0084] 21st clock cycle: The read signal arrives at the shift counter, and the partial sum (a3*b3+c3*d3)*2 in the 32-bit counter is transferred to 32 D2FFs.
[0085] 22nd clock cycle: The shift signal reaches the shift counter, the partial sum is shifted left one bit again, and reloaded into the 32-bit counter.
[0086] 23rd clock cycle: a2 and b2 are input into the shift counter and a logical AND operation is performed.
[0087] 24th clock cycle: c2 and d2 are input into the shift counter for a logical AND operation. a2*b2 is transferred to the 32-bit counter for counting.
[0088] 25th clock cycle: The shift counter pauses data input. c2*d2 is transferred to the 32-bit counter for counting.
[0089] 41st clock cycle: The read signal arrives at the shift counter.
[0090] 42nd clock cycle: The shift signal arrives at the shift counter.
[0091] 43rd clock cycle: a1 and b2 are input to the shift counter.
[0092] 44th clock cycle: c1 and d2 are input into the shift counter.
[0093] 45th clock cycle: a2 and b1 are input to the shift counter.
[0094] 46th clock cycle: C2 and D1 are input into the shift counter.
[0095] 47th clock cycle: The shift counter pauses data input.
[0096] 63rd clock cycle: The read signal arrives at the shift counter.
[0097] 64th clock cycle: The shift signal arrives at the shift counter.
[0098] 65th clock cycle: a0 and b2 are input to the shift counter.
[0099] 66th clock cycle: c0 and d2 are input to the shift counter.
[0100] 67th clock cycle: a1 and b1 are input to the shift counter.
[0101] 68th clock cycle: C1 and D1 are input into the shift counter.
[0102] 69th clock cycle: a2 and b0 are input to the shift counter.
[0103] 70th clock cycle: c2 and d0 are input to the shift counter.
[0104] 71st clock cycle: The shift counter pauses data input.
[0105] 87th clock cycle: The read signal arrives at the shift counter.
[0106] 88th clock cycle: The shift signal arrives at the shift counter.
[0107] 89th clock cycle: a0 and b1 are input to the shift counter.
[0108] 90th clock cycle: c0 and d1 are input into the shift counter.
[0109] 91st clock cycle: a1 and b0 are input to the shift counter.
[0110] 92nd clock cycle: c1 and d0 are input to the shift counter.
[0111] 93rd clock cycle: The shift counter pauses data input.
[0112] 109th clock cycle: The read signal arrives at the shift counter.
[0113] 110th clock cycle: The shift signal arrives at the shift counter.
[0114] 111th clock cycle: a0 and b0 are input to the shift counter.
[0115] 112th clock cycle: c0 and d0 are input into the shift counter.
[0116] 113th clock cycle: The shift counter pauses data input, and all data related to positive calculation is input.
[0117] 129th clock cycle: The read signal arrives at the shift counter and is positively transmitted to 32 D2FFs.
[0118] 130th clock cycle: The done signal reaches the shift counter and is positively transferred to the 32-bit data register.
[0119] 131st clock cycle: The sample signal reaches the shift counter, and the positive signal is transferred to the external memory. The shift counter begins counting the negative signal. a2 and b3 are input to the shift counter.
[0120] 132nd clock cycle: c2 and d3 are input into the shift counter.
[0121] 133rd clock cycle: a3 and b2 are input to the shift counter.
[0122] 134th clock cycle: C3 and D2 are input to the shift counter.
[0123] 135th clock cycle: The shift counter pauses data input.
[0124] 151st clock cycle: The read signal arrives at the shift counter.
[0125] 152nd clock cycle: The shift signal arrives at the shift counter.
[0126] 153rd clock cycle: a1 and b3 are input to the shift counter.
[0127] 154th clock cycle: c1 and d3 are input to the shift counter.
[0128] 155th clock cycle: a3 and b1 are input to the shift counter.
[0129] 156th clock cycle: C3 and D1 are input into the shift counter.
[0130] 157th clock cycle: The shift counter pauses data input.
[0131] 173rd clock cycle: The read signal arrives at the shift counter.
[0132] 174th clock cycle: The shift signal arrives at the shift counter.
[0133] 175th clock cycle: a0 and b3 are input to the shift counter.
[0134] 176th clock cycle: c0 and d3 are input to the shift counter.
[0135] 177th clock cycle: a3 and b0 are input to the shift counter.
[0136] 178th clock cycle: c3 and d0 are input to the shift counter.
[0137] 179th clock cycle: The shift counter suspends data input, and all data related to negative calculation is input.
[0138] 195th clock cycle: The read signal arrives at the shift counter.
[0139] 196th clock cycle: The shift signal arrives at the shift counter.
[0140] 197th clock cycle: The read signal arrives at the shift counter.
[0141] 198th clock cycle: The shift signal arrives at the shift counter.
[0142] 199th clock cycle: The read signal arrives at the shift counter.
[0143] 200th clock cycle: The shift signal reaches the shift counter, and the partial sum is shifted left by 3 bits to obtain negative.
[0144] 201st clock cycle: The read signal arrives at the shift counter and the negative signal is transferred to 32 D2FFs.
[0145] 202nd clock cycle: The done signal reaches the shift counter and the negative signal is transferred to the 32-bit data register.
[0146] Clock cycle 203: The sample signal reaches the shift counter, and the negative signal is passed to the external subtractor. The positive and negative signals are then subtracted in the external subtractor to obtain the final result of the multiplication-accumulation operation.
[0147] Figure 5 1 is an exemplary schematic diagram showing the symbols and state machines of the components in the multiplier-accumulator according to an embodiment of the present application. Figure 5 The symbols of the components of the multiplier-accumulator (e.g., DFF, AND, CB, D2FF, T1, and SPL) are shown in the upper part of Figures (a) and (b), and the corresponding state machines are shown in the lower part. Figure 5As shown in Figure (a), the DFF stores the pulses fed into the D terminal. When a pulse arrives at the clk terminal, the pulse stored in the DFF is destructively read out from the Q terminal. AND implements a logical AND function. A pulse is formed and stored inside the AND only when pulses are fed into both the A and B terminals simultaneously. When a pulse arrives at the clk terminal, the pulse stored in the AND is destructively read out from the Q terminal. The function of the concentrator ("CB") is to converge the pulses fed into the A and B terminals to the Q terminal, and the number of pulses remains unchanged. If there is a pulse at one of the input ports of the A and B terminals and no pulse at the other input port, one pulse will appear at the Q terminal. If one pulse is fed into both the A and B terminals simultaneously, two pulses will appear successively at the Q terminal.
[0148] like Figure 5 As shown in Figure (b), D2FF stores the pulse fed into the D terminal. When a pulse arrives at the clk1 terminal, the pulse stored in D2FF is destructively read out from the QA terminal. When a pulse arrives at the clk2 terminal, the pulse stored in D2FF is destructively read out from the QB terminal. T1 is an asynchronous modulo-2 counter. When a pulse is fed into the D terminal and there is no pulse stored inside T1, a 1 pulse will be stored inside T1. When a pulse is fed into the D terminal and a pulse is stored inside T1, a pulse will be formed at the Q terminal. When a pulse arrives at the rd terminal, the pulse stored in T1 will be destructively read out from the sum terminal. T1 can be regarded as a one-bit full adder with a latch function. The sum of the two bits will be stored inside T1, and the carry will be derived from the Q terminal. Furthermore, the splitter ("SPL", that is, the above-mentioned Figure 4 The function of the splitter shown in FIG is to copy one pulse fed into the D terminal into two pulses, which appear at the Q1 terminal and the Q2 terminal respectively.
[0149] Figure 6 FIG. 6 is an exemplary structural block diagram showing a multiplication and accumulation circuit 600 for a signed fixed-point number in a superconducting fast single-flux quantum circuit according to an embodiment of the present application. Figure 6As shown in , the multiplication and accumulation circuit 600 may include the multiplication and accumulation device 200 and the external subtractor 601 of the embodiment of the present application. According to the foregoing, the multiplication and accumulation device 200 of the embodiment of the present application may include a logic gate 201, a plurality of shift counters 202, a plurality of first data registers 203 and a plurality of second data registers 204. Among them, the logic gate (AND) 201 is responsible for realizing the multiplication of a single bit through a logical AND operation. A plurality of shift counters (such as T1) 202 are connected in series to form a 32-bit counter, and the accumulation operation of partial products with the same weight coefficient is realized by counting. The plurality of first data registers 203 can realize the multiplication by 2 operation in the bit serial multiplication and accumulation algorithm through logical left shift, and transfer the positive or negative in the algorithm to the plurality of second data registers 204 to be responsible for temporarily storing the positive or negative. Among them, for more details about the multiplication and accumulation device, please refer to the above Figure 2-Figure 5 The description of , this application will not be repeated here.
[0150] In one embodiment, the external subtractor 601 may be used to subtract the multiplication-accumulation partial sum of the addition part from the multiplication-accumulation partial sum of the subtraction part to obtain a final multiplication-accumulation result.
[0151] Figure 7 FIG. 7 is an exemplary flow chart showing a method 700 for multiplying and accumulating signed fixed-point numbers in a superconducting fast single-flux quantum circuit according to an embodiment of the present application. Figure 7 As shown in FIG, at step S701, a single-bit product operation is performed on the target bit of the signed fixed-point number to obtain a single-bit partial product. In one implementation scenario, a logic gate can be used to perform the single-bit product operation to obtain the single-bit partial product. The single-bit partial product may include a partial product of the sign bit multiplication and a partial product of the magnitude bit multiplication in the addition part, and a partial product of the sign bit multiplication and the magnitude bit multiplication in the subtraction part.
[0152] Next, at step S702, an accumulation operation is performed on the single-bit partial products according to the shift count to obtain a multiplication-accumulation partial sum. In some embodiments, the accumulation and counting operation can be performed using, for example, multiple serially connected T1 flip-flops. Specifically, the partial products with the same weight coefficient are accumulated to obtain the multiplication-accumulation partial sum corresponding to the addition part and the subtraction part.
[0153] Further, at step S703, the multiplication-accumulation partial sum is received and transferred, and a shift operation in the accumulation operation is performed. In one embodiment, a D2FF, for example, can be used to receive and transfer the multiplication-accumulation partial sum to implement the shift operation. Specifically, a control signal of a target number of groups can be set to transfer the target number of multiplication-accumulation partial sums to implement the shift operation. Finally, at step S704, the multiplication-accumulation partial sum is received and temporarily stored to subsequently calculate the final multiplication-accumulation result based on the multiplication-accumulation partial sums corresponding to the addition part and the subtraction part.
[0154] Figure 8 FIG. 8 is an exemplary structural block diagram showing a device 800 for multiplying and accumulating signed fixed-point numbers in a superconducting fast single-flux quantum circuit according to an embodiment of the present application. Figure 8 As shown in , the multiplication-accumulation device 800 of the present application may include a processor 801 and a memory 802, wherein the processor 801 and the memory 802 communicate with each other via a bus. The memory 802 stores program instructions for multiplication-accumulation of signed fixed-point numbers in a superconducting fast single-flux quantum circuit. When the program instructions are executed by the processor 801, the method steps described above in conjunction with the accompanying drawings are implemented: performing a single-bit product operation on the target bit of the signed fixed-point number to obtain a single-bit partial product; performing an accumulation operation on the single-bit partial product according to a shift count to obtain a multiplication-accumulation partial sum; receiving and transmitting the multiplication-accumulation partial sum and performing a shift operation in the accumulation operation; and receiving and temporarily storing the multiplication-accumulation partial sum.
[0155] According to the above description in conjunction with the accompanying drawings, those skilled in the art will also understand that the embodiments of the present application can also be implemented by software programs. Therefore, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores computer-readable instructions for multiplying and accumulating signed fixed-point numbers in a superconducting fast single-flux quantum circuit. When the computer-readable instructions are executed by one or more processors, the present application in conjunction with the accompanying drawings is implemented. Figure 2 The described method is used for the multiplication and accumulation of signed fixed-point numbers in superconducting fast single-flux quantum circuits.
[0156] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0157] It should be noted that although the operations of the present method are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in that particular order, or that all of the operations shown must be performed to achieve the desired results. Rather, the steps depicted in the flowcharts may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into a single step, and / or a single step may be broken down into multiple steps.
[0158] It should be understood that when the terms "first," "second," "third," and "fourth," etc., are used in the claims, specification, and drawings of this application, they are only used to distinguish different objects, rather than to describe a specific order. The terms "comprise" and "comprising" used in the specification and claims of this application indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0159] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this specification and claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" as used in this specification and claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.
[0160] Although the implementation methods of this application are as described above, the contents are only examples adopted to facilitate understanding of this application and are not intended to limit the scope and application scenarios of this application. Any technician in the technical field described in this application can make any modifications and changes in the form and details of implementation without departing from the spirit and scope disclosed in this application, but the scope of patent protection of this application shall still be based on the scope defined by the attached claims.
Claims
1. A signed fixed-point multiplier-accumulator for use in a superconducting fast single-flux quantum circuit, comprising: a logic gate for performing a single-bit product operation on a target bit in a signed fixed-point number to obtain a single-bit partial product; a plurality of shift counters, the plurality of shift counters being serially connected and configured to perform an accumulation operation on the single-bit partial products according to a shift count to obtain a multiply-accumulate partial sum; a plurality of first data registers, connected to the plurality of shift counters correspondingly, configured to receive and transmit the multiplication-accumulation partial sums and perform shift operations in the accumulation operations; as well as A plurality of second data registers are connected to the plurality of first data registers correspondingly, and are used to receive and temporarily store the multiplication-accumulation partial sums.
2. The multiplier-accumulator according to claim 1 , wherein the single-bit partial product comprises a first partial product and a second partial product, and the first partial product comprises a partial product of a sign bit multiplied by a value bit multiplied in the signed fixed-point number; and the second partial product comprises a partial product of a sign bit multiplied by a value bit in the signed fixed-point number.
3. The multiplication-accumulator according to claim 2 , wherein in performing an accumulation operation on the single-bit partial products according to a shift count to obtain a multiplication-accumulation partial sum, the plurality of shift counters are further configured to: Accumulation operations are performed on the partial products of the first part and the partial products of the second part under the same weight coefficient according to the shift count to obtain a first multiplication-accumulation partial sum and a second multiplication-accumulation partial sum respectively.
4. The multiplier-accumulator according to claim 3, wherein the plurality of shift counters are further configured to: In response to the completion of inputting the target bit of the calculation of the partial product of the first part or the partial product of the second part under the same weight coefficient, a pause operation is performed and ripple carry propagation is waited.
5. The multiplier-accumulator according to claim 4, wherein the plurality of shift counters are further configured to: In response to a first control signal, a first multiplication-accumulation partial sum or a second multiplication-accumulation partial sum corresponding to the partial product of the first part or the partial product of the second part under the same weight coefficient is transferred to the plurality of first data registers.
6. The multiplier-accumulator according to any one of claims 3 to 5, wherein the plurality of first data registers are further configured to: In response to the second control signal, the first multiplication-accumulation partial sum or the second multiplication-accumulation partial sum corresponding to the partial product of the first part or the partial product of the second part under the same weight coefficient is reloaded into the multiple shift counters to implement the shift operation.
7. The multiplier-accumulator according to claim 6, wherein the plurality of first data registers are further configured to: receiving a first multiplication-accumulation partial sum or a second multiplication-accumulation partial sum corresponding to the partial products of the first part or the partial products of the second part under the same weight coefficient transferred a target number of times by the plurality of shift counters under the first control signal of the target group number; and In response to the second control signal of the target group number, the first multiplication-accumulation partial sum or the second multiplication-accumulation partial sum corresponding to the partial product of the first part or the partial product of the second part under the same weight coefficient of the target times is reloaded into the multiple shift counters to achieve the target number of bits of movement.
8. The multiplier-accumulator according to claim 7, wherein the plurality of first data registers are further configured to: In response to a third control signal, the first multiply-accumulate partial sum or the second multiply-accumulate partial sum is transferred to the plurality of second data registers.
9. The multiplier-accumulator according to any one of claims 2 to 8, wherein the plurality of second data registers are further configured to: In response to a fourth control signal, the first multiplication-accumulation partial sum and the second multiplication-accumulation partial sum are transferred to an external subtractor, so that the external subtractor performs a target subtraction operation based on the first multiplication-accumulation partial sum and the second multiplication-accumulation partial sum to obtain a final result of multiplication and accumulation.
10. A multiplication and accumulation circuit for signed fixed-point numbers in a superconducting fast single-flux quantum circuit, comprising: The multiplier-accumulator according to any one of claims 1 to 9; as well as External subtractor.
11. A method for multiplying and accumulating signed fixed-point numbers in a superconducting fast single-flux quantum circuit, comprising: Perform a single-bit product operation on the target bit in the signed fixed-point number to obtain a single-bit partial product; performing an accumulation operation on the single-bit partial products according to the shift count to obtain a multiply-accumulate partial sum; receiving and transmitting the multiplication-accumulation partial sum, and performing a shift operation in the accumulation operation; as well as The multiplication-accumulation partial sum is received and temporarily stored.
12. A device for multiplying and accumulating signed fixed-point numbers in a superconducting fast single-flux quantum circuit, comprising: processor; as well as A memory having computer instructions for multiplying and accumulating signed fixed-point numbers in a superconducting fast single-flux quantum circuit stored therein, wherein when the computer instructions are executed by a processor, the multiplying and accumulating method according to claim 11 is implemented.
13. A computer-readable storage medium, characterized in that Computer program instructions for multiplication and accumulation of signed fixed-point numbers in a superconducting fast single-flux quantum circuit are stored thereon, and when the computer program instructions are executed by one or more processors, the multiplication and accumulation method according to claim 11 is implemented.