Multiply-accumulate operation method and device based on one-by-two-addition structure
By using a one-multiple two-multiple structure multiplication and accumulation operation method and device in the multiplication and accumulation operation circuit, the problem of difficulty in parallel degree optimization in traditional circuits is solved, and efficient multiplication and accumulation operation and flexible parallel degree support are realized.
Patent Information
- Application Number
- CN202510069511.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
The number of adders and multipliers in the traditional tree-like multiplication and accumulation operation circuit is large, which is not conducive to the optimization of parallelism in the reconstructible computing circuit, and cannot flexibly schedule parallelism, which limits its application in efficient computing.
The multiplication and accumulation operation method and device based on one multiplication and two addition structure is adopted, and the multiplication and accumulation operation is realized through a multiplication and two adders, and the control module generates a control signal based on the count warning signal to control the operation process.
The parallel operation of multiplier and adder is realized, the computing speed is improved, and it can flexibly support multi-parallel calculation tasks, with good scalability and flexibility.
Smart Images

Figure CN119987715A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital processing technology, and in particular to a multiplication-accumulation operation method and device based on a one-multiplication-two-addition structure. Background Art
[0002] In traditional multiplication-accumulation circuits, a tree structure is a commonly used design method. This structure realizes the gradual accumulation and multiplication of data through the cascade of multiple levels of adders and multipliers. However, the traditional tree structure requires that the number of adders and multipliers is a power of 2. This structural design method is not conducive to the optimization of parallelism in reconfigurable computing circuits. The multiplication-accumulation circuit with a tree structure can only support computing tasks with a single degree of parallelism and cannot effectively handle multiplication-accumulation calculations with multiple degrees of parallelism.
[0003] For example, a Chinese patent with authorization announcement number CN118519613B discloses a multiplication-accumulation operation cluster and a data processing method, including: multiple multiplication-accumulation units, each multiplication-accumulation unit is used to complete the multiplication-accumulation operation of at least one group of data to be processed, when the calculation dimension of the data to be processed is smaller than the total dimension of the multiplication-accumulation unit, the multiplication-accumulation operation of each group of data to be processed is completed by at least one multiplication-accumulation sub-unit in the multiplication-accumulation unit, and the multiplication-accumulation operation result of the data to be processed is output; and / or, when the data to be processed includes multiple sequentially arranged data blocks, the multiplication-accumulation operation of one data block is completed respectively by multiple multiplication-accumulation sub-clusters in the multiplication-accumulation operation cluster, and the multiplication-accumulation operation result of the data block is output to the next sequentially arranged multiplication-accumulation sub-cluster, until the multiplication-accumulation operation of all data blocks is completed, and the multiplication-accumulation operation result of the data to be processed is output.
[0004] The above patents have the problem raised by the background technology: the number of adders and multipliers in the traditional tree-structured multiplication-accumulation operation circuit is large, which is not conducive to the optimization of parallelism in the reconfigurable computing circuit. Therefore, the tree structure cannot flexibly schedule the parallelism according to actual needs, which limits its application in efficient computing. To solve the above problems, the present invention proposes a multiplication-accumulation operation method and device based on a one-multiplication-two-addition structure. Summary of the invention
[0005] In view of the shortcomings of the prior art, the main purpose of the present invention is to provide a multiplication-accumulation-addition method and device based on a one-multiplication-two-addition structure, which can effectively solve the problems in the background technology. The specific technical solution of the present invention is as follows:
[0006] A multiplication-accumulation-addition operation device based on a one-multiplication-two-addition structure, comprising:
[0007] Multiplication and accumulation module and control module; wherein,
[0008] The multiplication-accumulation module is used to obtain at least two sets of source data, and perform multiplication-accumulation operations on the source data through a multiplier and two adders to obtain result data;
[0009] The control module generates a control signal according to the counting warning signal to control the multiplication and accumulation operation process of the multiplication and accumulation module.
[0010] Specifically, the multiplication and accumulation module includes:
[0011] One multiplier and two adders; among them,
[0012] The multiplier performs multiplication calculation on the source data to obtain a product result;
[0013] The adder performs accumulation calculation on the product results to obtain a multiplication-accumulation result.
[0014] Specifically, the control module includes:
[0015] a plurality of multiplexers, a plurality of multiplexers, and a plurality of registers; wherein,
[0016] The multiplexer selects one signal to output from a plurality of input signals transmitted by the multiplexer or register according to the control signal;
[0017] The multiplexer distributes an input signal transmitted from the adder or the multiplexer to a plurality of registers or multiplexers according to the control signal;
[0018] The register receives and stores the signal transmitted by the multiplexer or the multiplexer according to the control signal.
[0019] A multiplication-accumulation operation method based on a one-multiplication-two-addition structure, comprising:
[0020] Obtain N sets of source data, where N ≥ 2;
[0021] Inputting the source data into a multiplier to perform product calculation to obtain a product result;
[0022] Calculate the number of clock cycles T and the counting warning signal according to a preset adder to generate a control signal;
[0023] Inputting the multiplication result into a first adder, performing accumulation calculation according to a control signal, and obtaining a first accumulation result;
[0024] The first accumulation result is input into the second adder, and accumulation calculation is performed according to the control signal to obtain a multiplication-accumulation result.
[0025] Specifically, the step of calculating the number of clock cycles T and the counting warning signal according to a preset adder to generate a control signal includes:
[0026] According to the multiply-accumulate calculation process and the preset adder clock cycle number T, initialize the counting warning signal;
[0027] Generate multiple control signals according to the counting warning signal, where each control signal controls a multiplexer, a multiplexer or a register.
[0028] Specifically, the step of initializing the counting warning signal according to the multiply-accumulate calculation process and the preset adder clock cycle number T includes:
[0029] When n ≤ N - T + 1, when each group of source data is subjected to multiply-accumulate calculation, the counting warning signal is set to 0, where n is the number of groups of input source data;
[0030] When n = N - T + 2, at this time, the reciprocal (T - 1) groups of source data are input for multiply-accumulate calculation, and the counting warning signal is set to 4;
[0031] When n = N - T + 3, at this time, the reciprocal (T - 2) groups of source data are input for multiply-accumulate calculation, and the counting warning signal is set to 3;
[0032] When n = N, the last group of source data is input for multiply-accumulate calculation, and the counting warning signal is set to 2.
[0033] Specifically, the step of inputting the product result into the first adder and performing accumulation calculation according to the control signal to obtain the first accumulation result includes:
[0034] When 0 < n ≤ T, input the product results calculated from the first T groups of source data into the first adder, and each group of product results is added to 0 respectively to obtain a single addition result;
[0035] When T + 1 ≤ n ≤ N, input the product results calculated from the (T + 1)th to Nth groups of source data into the first adder, and accumulate them with the single addition result according to the clock cycle number to obtain the first accumulation result.
[0036] Specifically, the step of, when T + 1 ≤ n ≤ N, inputting the product results calculated from the (T + 1)th to Nth groups of source data into the first adder and accumulating them with the single addition result according to the clock cycle number to obtain the first accumulation result includes:
[0037] When T + 1 ≤ n ≤ 2T, when inputting the product results calculated from the (T + 1)th to 2Tth groups of source data into the first adder, the product result of the nth group is added to the result of the (n - T)th group in the single addition result respectively to obtain a double addition result;
[0038] When 2T+1≤n≤3T, when the product results calculated from the (2T+1)th to 3Tth groups of source data are input into the first adder, the product results of the nth group are respectively added to the results of the (nTth)th group in the secondary addition results to obtain a third addition result;
[0039] Until the accumulation calculation process of N groups of source data in the first adder is completed, a first accumulation result is obtained.
[0040] Specifically, the step of inputting the first accumulation result into a second adder, performing accumulation calculation according to a control signal, and obtaining a multiplication-accumulation result includes:
[0041] Each group of results corresponding to the first accumulated result is input into a plurality of registers for temporary storage;
[0042] The results stored in the register are sequentially input into the second adder according to the control signal, and accumulation calculation is performed to obtain the multiplication and accumulation result.
[0043] Specifically, the step of sequentially inputting the results stored in the register into the second adder according to the control signal, performing accumulation calculation, and obtaining the multiplication and accumulation result includes:
[0044] Inputting the (N-T+1)th group of results and the (N-T+2th group of results in the first cumulative result into a second adder, performing addition calculation, and obtaining a primary cumulative result;
[0045] Inputting the (N-T+3)th group of results and the (N-T+4th group of results in the first cumulative result into a second adder for adding and calculating, thereby obtaining a secondary cumulative result;
[0046] The first accumulation result and the second accumulation result are input into the second adder for addition calculation to obtain a first multiplication and accumulation result;
[0047] The accumulation calculation process of the Nth group of results in the second adder is completed to obtain the multiplication and accumulation result.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] The present invention realizes the multiplication and accumulation operation of data through a multiplication and accumulation module with a one-multiplication-two-addition structure, realizes the parallel operation of the multiplier and the adder, so that one multiplication operation and two addition operations can be completed in each clock cycle, thereby improving the operation speed and being able to more flexibly support multi-parallel computing tasks; the multiplication and accumulation calculation process is controlled by a control module, and different control logics and operation processes can be configured according to different application scenarios; the present invention realizes efficient multiplication and accumulation operation through an efficient multiplication and accumulation operation structure and a flexible control module, and has good scalability and flexibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 Schematic diagram of the structure of a multiplication-accumulation-addition operation device based on a one-multiplication-two-addition structure in Embodiment 1 of the present invention;
[0051] Figure 2 This is a flowchart of a multiplication-accumulation-addition operation method based on a one-multiplication-two-addition structure in Embodiment 2 of the present invention;
[0052] Figure 3 The req waveform diagram of two 8-point multiplication and accumulation seamless input in Embodiment 3 of the present invention;
[0053] Figure 4 t1 to t5 in Example 3 of the present invention represent interval schematic diagram;
[0054] Figure 5 This is a req waveform diagram of two 7-point multiplication and accumulation operations and t5 being 1 in Embodiment 3 of the present invention. DETAILED DESCRIPTION
[0055] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0056] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0057] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0058] Example 1
[0059] This embodiment provides a multiplication-accumulation-addition operation device based on a one-multiplication-two-addition structure, such as Figure 1 As shown, the multiplication-accumulation operation device based on a one-multiplication-two-addition structure comprises:
[0060] Multiplication and accumulation module and control module; wherein,
[0061] The multiplication-accumulation module is used to obtain at least two sets of source data, and perform multiplication-accumulation operations on the source data through a multiplier and two adders to obtain result data;
[0062] The control module generates a control signal according to the counting warning signal to control the multiplication and accumulation operation process of the multiplication and accumulation module.
[0063] In this embodiment, a multiplication-accumulation operation device based on a one-multiplication-two-addition structure includes a multiplication-accumulation module and a control module. Source data to be multiplied and accumulated are input into the multiplication-accumulation module. The multiplication-accumulation module adopts a one-multiplication-two-addition structure, i.e., a multiplier and two adders. The acquired source data is sent to the multiplier. The multiplier performs a multiplication operation on two sets of source data to generate a product result. The product result is sent to the first adder. The first adder adds the product results in sequence according to the number of clock cycles to obtain the result of the first addition operation. The result of the first addition operation is sent to the second adder. The second adder adds the results of the first addition operation in sequence to obtain the final multiplication-accumulation result. Through the one-multiplication-two-addition structure, efficient multiplication-accumulation operation is realized. The parallel operation of the multiplier and the adder enables one multiplication operation and two addition operations to be completed in each clock cycle, thereby improving the operation speed.
[0064] Specifically, the control module generates a control signal according to the counting warning signal to control the multiplication and accumulation process of the multiplication and accumulation module. The control signal controls the data transmission process. The multiplication and accumulation module adjusts its working state according to these control signals, such as starting a new multiplication and accumulation operation, pausing the current operation, resetting the internal state, etc. The existence of the control module makes the multiplication and accumulation operation device more flexible. The operation process can be adjusted as needed, such as pausing the operation, resetting the state, etc. At the same time, different control logics and operation processes can be configured according to different application scenarios; the multiplication and accumulation operation device based on the one-multiplication-two-addition structure realizes efficient multiplication and accumulation operations through an efficient operation structure and a flexible control module, and has good scalability and flexibility.
[0065] The present invention realizes the multiplication and accumulation operation of data through a multiplication and accumulation module with a one-multiplication-two-addition structure, realizes the parallel operation of the multiplier and the adder, so that one multiplication operation and two addition operations can be completed in each clock cycle, thereby improving the operation speed and being able to more flexibly support multi-parallel computing tasks; the multiplication and accumulation calculation process is controlled by a control module, and different control logics and operation processes can be configured according to different application scenarios; the present invention realizes efficient multiplication and accumulation operation through an efficient multiplication and accumulation operation structure and a flexible control module, and has good scalability and flexibility.
[0066] Furthermore, the multiplication and accumulation module includes:
[0067] One multiplier and two adders; among them,
[0068] The multiplier performs multiplication calculation on the source data to obtain a product result;
[0069] The adder performs accumulation calculation on the product results to obtain a multiplication-accumulation result.
[0070] In this embodiment, the multiplication-accumulation module adopts a one-multiplication-two-addition structure, including a multiplier and two adders, and the multiplier receives at least two sets of source data as input, and the two sets of data can be any type of numerical values, such as integers, floating-point numbers, etc. The multiplier performs a multiplication operation, multiplies the input source data, and generates a product result. The multiplier can efficiently perform the multiplication operation and provide an accurate product result for the subsequent accumulation operation.
[0071] Specifically, two adders perform cumulative calculation on the product results, and input the product results into the first adder. The first adder sequentially adds the product results according to the number of clock cycles to obtain the result of the first addition operation. The result of the first addition operation is sent to the second adder. The second adder sequentially adds the results of the first addition operation to obtain the final multiplication and accumulation result. The accumulation of products is achieved through the addition operation of the two adders. Through the accumulation operation, the sum of a series of products can be calculated. The one-multiplication-two-addition structure enables the multiplier and the adder to work in parallel, thereby improving the operation efficiency. While the multiplier is calculating the product, the adder can prepare for the accumulation operation, thereby reducing the waiting time of the operation. This structural design allows flexible configuration according to specific needs. For example, the number of adders can be adjusted to meet different accumulation requirements. Since the multiplier and the adder work in parallel, the hardware resources can be fully utilized to reduce unnecessary waiting time and power consumption. This structural design helps to reduce overall power consumption and has the advantages of high efficiency, flexibility, scalability and low power consumption. It is suitable for various application scenarios that require efficient multiplication and accumulation operations.
[0072] Furthermore, the control module includes:
[0073] a plurality of multiplexers, a plurality of multiplexers, and a plurality of registers; wherein,
[0074] The multiplexer selects one signal to output from a plurality of input signals transmitted by the multiplexer or register according to the control signal;
[0075] The multiplexer distributes an input signal transmitted from the adder or the multiplexer to a plurality of registers or multiplexers according to the control signal;
[0076] The register receives and stores the signal transmitted by the multiplexer or the multiplexer according to the control signal.
[0077] In this embodiment, the control module includes a plurality of multiplexers, a plurality of multiplexers and a plurality of registers, such as Figure 1As shown in the figure, Demux is a multiplexer with 1 input and multiple outputs. It distributes an input signal to multiple output ports according to the control signal, which can be distributed to registers or multiplexers. The multiplexer can flexibly distribute input signals according to the control signal, which improves the flexibility and scalability of the system and improves resource utilization. Through the multiplexer, system resources can be used more effectively, thereby improving the overall performance of the system. Mux is a multiplexer with multiple inputs and 1 output. It selects one signal from multiple input signals for output according to the control signal. The basic structure of the multiplexer includes multiple input terminals, one output terminal and a control terminal. Multiple signals are input from the data input terminal of the selector, and the selector determines which data bit to output according to the control signal. The multiplexer can select one of the multiple input signals for output, thereby improving the efficiency of signal transmission and reducing the complexity of the system. Through the multiplexer, the connection lines and complexity of the system can be reduced.
[0078] Specifically, Reg is a register. The register is determined by a control signal, a clock signal, and a reset signal to determine whether to accept new input data or clear the stored data. In the control module, the register receives and stores data transmitted by a multiplexer or a multiplexer according to the control signal. By storing data and instructions and quickly accessing these data when needed, the processing speed of the system is improved and the system stability is improved. The control module separates the functional modules such as the multiplexer, the multiplexer, and the register to facilitate the maintenance and upgrade of the system; through the combined use of the multiplexer and the multiplexer, the selection, allocation, and routing of signals can be flexibly realized, which improves the flexibility and scalability of the system; the register can quickly store and access data, which improves the processing speed and stability of the system. At the same time, the use of the multiplexer and the multiplexer also improves the efficiency of signal transmission and processing. The control module realizes efficient selection, allocation, and storage of signals through the combined use of multiple multiplexers, multiplexers, and registers, which not only improves the flexibility and scalability of the system, but also improves the processing speed and stability of the system.
[0079] Example 2
[0080] This embodiment provides a multiplication-accumulation method based on a one-multiplication-two-addition structure. Figure 2 As shown, the multiplication-accumulation operation method based on a one-multiplication-two-addition structure includes:
[0081] S101, obtaining N groups of source data, where N≥2;
[0082] S102, inputting the source data into a multiplier to perform product calculation to obtain a product result;
[0083] S103, calculating the number of clock cycles T and the counting warning signal according to a preset adder, and generating a control signal;
[0084] S104, inputting the multiplication result into a first adder, performing accumulation calculation according to a control signal, and obtaining a first accumulation result;
[0085] S105 , input the first accumulation result into a second adder, perform accumulation calculation according to a control signal, and obtain a multiplication-accumulation result.
[0086] In this embodiment, at least two groups of source data are first obtained, each group of source data includes source data 1 and source data 2, and each group of source data is input into the multiplier in turn for multiplication operation, and the input source data 1 and source data 2 are multiplied to obtain the product result; the preset adder calculation clock cycle number T defines the timing of the accumulation operation, and the counting warning signal is used to trigger the generation of the control signal when a specific condition is met, so as to realize precise control of the accumulation operation process and ensure the accuracy and stability of the operation. The timing and conditions of the accumulation operation can be flexibly adjusted through the preset clock cycle number and the counting warning signal.
[0087] Specifically, the product result calculated by the multiplier is input into the first adder, and the first adder performs an accumulation operation according to the control signal, and adds the product results of each group respectively to obtain a first accumulation result, thereby realizing the accumulation of the product results. Through the accumulation operation of the first adder, the pressure of data storage and transmission can be reduced, and the operation efficiency can be improved; the first accumulation result is input into the second adder, and the second adder performs an accumulation operation again according to the control signal, and adds the first accumulation results in sequence to obtain a final multiplication-accumulation result, thereby realizing the final accumulation operation and obtaining the result of the multiplication-accumulation operation. Through further accumulation of the second adder, it can be ensured that all the product results are correctly accumulated together, thereby obtaining an accurate multiplication-accumulation result. The multiplication-accumulation operation method based on the one-multiplication-two-addition structure realizes the accuracy and efficiency of the multiplication-accumulation operation through precise step control and efficient accumulation operation.
[0088] Further, the step of calculating the number of clock cycles T and the counting warning signal according to a preset adder to generate a control signal includes:
[0089] S201, initializing a counting warning signal according to a multiplication-accumulation calculation process and a preset adder to calculate the number of clock cycles T;
[0090] S202 . Generate a plurality of control signals according to the counting warning signal, wherein each control signal controls a multiplexer, a multiplexer or a register.
[0091] In this embodiment, the overall process of multiplication and accumulation calculation needs to be considered first, including multiplication operation, multiple accumulation operations and the timing relationship between different operation steps. The preset adder calculation clock cycle number T is the basis for the generation of the control signal, and the time interval of each operation in the entire multiplication and accumulation calculation process is determined. The initialization of the count warning signal is to trigger a warning when a specific number of clock cycles is reached, thereby generating a corresponding control signal. When the counter reaches the preset value, a count warning signal will be output; by initializing the count warning signal, it can be ensured that the multiplication and accumulation calculation process can be carried out according to the predetermined timing, making the entire calculation process more stable and controllable.
[0092] Specifically, a corresponding control signal is generated according to the technical warning signal, and each control signal is used to control a corresponding circuit element, such as a multiplexer, a multiplexer and a register. By controlling the working process of the element through the control signal, precise control of different circuit elements can be achieved, thereby ensuring the correctness and efficiency of the multiplication and accumulation calculation process, so that the circuit elements can operate in a predetermined order and conditions, avoiding calculation errors or performance degradation caused by operational confusion. By initializing the counting warning signal and generating multiple control signals, precise control of different circuit elements can be achieved, thereby ensuring the stability and controllability of the entire calculation process.
[0093] Further, the method of calculating the number of clock cycles T according to the multiplication-accumulation calculation process and the preset adder to initialize the counting warning signal includes:
[0094] S301, when n≤N-T+1, when each set of source data is multiplied and accumulated, the counting warning signal is set to 0, where n is the number of input source data sets;
[0095] S302, when n=N-T+2, the inverse (T-1) group of source data is input for multiplication and accumulation calculation, and the counting warning signal is set to 4;
[0096] S303, when n=N-T+3, the inverse (T-2) group of source data is input to perform multiplication and accumulation calculation, and the counting warning signal is set to 3;
[0097] S304, when n=N, input the last set of source data to perform multiplication and accumulation calculation, and set the counting warning signal to 2.
[0098] In this embodiment, when n groups of source data are input, when n ≤ N - T + 1, the counting warning signal is set to 0 when each group of source data is subjected to multiply-accumulate calculation. When the penultimate (T - 1) group of source data is input for multiply-accumulate calculation, the counting warning signal is set to 4; when the penultimate (T - 2) group of source data is input for multiply-accumulate calculation, the counting warning signal is set to 3; when the last group of source data is input for multiply-accumulate calculation, the counting warning signal is set to 2. By setting the counting warning signal, the multiply-accumulate module is informed of how many more groups of source data need to be calculated.
[0099] Adders and multipliers generally require multiple clock cycles to complete one calculation. The number of clock cycles required for the multiplier and adder to complete one calculation is called the pipeline stage number. We assume that the pipeline establishment times of both the multiplier and adder are 4 clock cycles; the counting warning signal is denoted as the last4 signal. The following is an explanation of the last4 signal:
[0100] Suppose it is necessary to perform multiply-accumulate calculation on 10 groups of source data (i.e., 10 src1s and 10 src2s). Then, when the 1st to 7th groups of data are input, the last4 input signal remains 0; when the 8th group of data is input, the last4 input signal is 4; when the 9th group of data is input, the last4 input signal is 3; when the 10th group of data is input, the last4 input signal is 2 (that is, when the penultimate 3rd group of data is input, the last4 signal is 4; when the penultimate 2nd group of data is input, the last4 signal is 3; when the last 1st group of data is input, the last4 signal is 2; in other cases, the last4 signal is 0). Control signals 1 to 12 are used as control signals for multiplexers, demultiplexers, and registers, and their values are all controlled by the last4 signal.
[0101] Further, the step of inputting the product result into the first adder and performing accumulative calculation according to the control signal to obtain the first accumulative result includes:
[0102] S401. When 0 < n ≤ T, input the product results calculated from the first T groups of source data into the first adder, and each group of product results is added to 0 respectively to obtain a single addition result;
[0103] S402. When T + 1 ≤ n ≤ N, input the product results calculated from the (T + 1)-th to N-th groups of source data into the first adder, and accumulate them with the single addition result according to the number of clock cycles respectively to obtain the first accumulative result.
[0104] Further, the step of when T + 1 ≤ n ≤ N, inputting the product results calculated from the (T + 1)-th to N-th groups of source data into the first adder, and accumulating them with the single addition result according to the number of clock cycles respectively to obtain the first accumulative result includes:
[0105] S501, when T+1≤n≤2T, when the product results calculated from the (T+1)th to 2Tth groups of source data are input into the first adder, the product results of the nth group are respectively added to the results of the (nTth)th group in the first addition result to obtain a second addition result;
[0106] S502, when 2T+1≤n≤3T, when the product results calculated from the (2T+1)th to 3Tth groups of source data are input into the first adder, the product results of the nth group are respectively added to the results of the (nTth)th group in the secondary addition results to obtain a cubic addition result;
[0107] S503, until the accumulation calculation process of N groups of source data in the first adder is completed to obtain a first accumulation result.
[0108] In this embodiment, the product results are accumulated according to the number of clock cycles T. When the product results calculated by the first T groups of source data are input into the first adder, the product results of each group are added to 0 respectively, which is equivalent to not performing the accumulation operation, but retaining the original product results to obtain a single addition result; when the product results calculated by the (T+1)th to 2Tth groups of source data are input into the first adder, the product results of the nth group are added to the results of the (nT)th group in the single addition result respectively, and these product results are input into the first adder, and are sequentially accumulated with the previous product results (i.e., the first T groups) to obtain a secondary addition result, thereby realizing the step-by-step accumulation of the source data product results. Addition ensures that each product result is correctly added to its corresponding previous result; when the product results calculated from the (2T+1)th to 3Tth groups of source data are input into the first adder, the product results of the nth group are respectively added to the results of the (nT)th group in the secondary addition results, and these product results are added to the previous (i.e., the first 2T groups) accumulated results to obtain three addition results, and so on, each T group of product results is input into the first adder and accumulated with the previous addition results until all N groups of source data are processed, and by sequentially adding the addition result of each clock cycle to the addition result of the previous cycle, it is ensured that all source data are correctly processed and accumulated.
[0109] For example, the multiplication result calculated by the multiplier will be transmitted to the adder 1 for accumulation operation. When the multiplication results of the first four groups of numbers are input into the adder 1, they will be accumulated with 0; when the multiplication result of the fifth group of numbers is input into the adder 1, it will be accumulated with the addition result of the first group of numbers; when the multiplication result of the sixth group of numbers is input into the adder 1, it will be accumulated with the addition result of the second group of numbers; when the multiplication result of the seventh group of numbers is input into the adder 1, it will be accumulated with the addition result of the third group of numbers; when the multiplication result of the eighth group of numbers is input into the adder 1, it will be accumulated with the addition result of the fourth group of numbers; when the multiplication result of the ninth group of numbers is input into the adder 1, it will be accumulated with the addition results of the first and fifth groups of numbers, and so on, until the accumulation calculation process of N groups of source data in the first adder is completed to obtain the first accumulation result.
[0110] Further, the step of inputting the first accumulation result into a second adder and performing accumulation calculation according to a control signal to obtain a multiplication-accumulation result includes:
[0111] S601, each group of results corresponding to the first accumulated result is input into a plurality of registers for temporary storage;
[0112] S602: Input the results stored in the register into the second adder in sequence according to the control signal, perform accumulation calculation, and obtain a multiplication-accumulation result.
[0113] Further, the step of sequentially inputting the results stored in the register into the second adder according to the control signal to perform accumulation calculation to obtain the multiplication and accumulation result includes:
[0114] S701, input the (N-T+1)th group of results and the (N-T+2th group of results in the first cumulative result into a second adder, perform addition calculation, and obtain a primary cumulative result;
[0115] S702, input the (N-T+3)th group of results and the (N-T+4th group of results in the first cumulative result into a second adder, perform addition calculation, and obtain a secondary cumulative result;
[0116] S703, inputting the first accumulation result and the second accumulation result into a second adder, performing addition calculation, and obtaining a first multiplication and accumulation result;
[0117] S704, until the accumulation calculation process of the Nth group of results in the second adder is completed to obtain a multiplication-accumulation result.
[0118] In this embodiment, first, each group of results in the first accumulation result is stored in a plurality of registers respectively. These registers are used as temporary storage units to store intermediate results for use in subsequent accumulation calculations. The registers provide the ability to quickly access and store data, which helps to speed up the calculation process. The results stored in the registers are read in sequence according to the control signal, and the read data is input into the second adder for accumulation operation. Through the guidance of the control signal, it can be ensured that the accumulation calculation is performed in a predetermined order, thereby ensuring the correctness of the result.
[0119] Specifically, the (N-T+1)th group of results and the (N-T+2)th group of results in the first accumulation result are input into the second adder for addition calculation to obtain a first accumulation result, the (N-T+3)th group of results and the (N-T+4)th group of results in the first accumulation result are input into the second adder for addition calculation to obtain a second accumulation result, the first accumulation result and the second accumulation result are input into the second adder for addition calculation to obtain a multiplication-accumulation result. By merging the results of multiple accumulations, we can gradually approach the final multiplication-accumulation result; continue to input the remaining first accumulation results (i.e., the (N-T+5)th group to the Nth group of results) into the second adder in sequence for accumulation calculation, and this process continues until all the first accumulation results are processed to finally obtain a multiplication-accumulation result; by ensuring that all results are correctly accumulated, an accurate multiplication-accumulation result can be obtained.
[0120] Assuming that there are N groups of source data to be input into the multiplication and accumulation calculation module, the N-3 group of numbers and the previous accumulation results will be input into Reg_2 for temporary storage, the N-2 group of numbers and the previous accumulation results will be input into Reg_3, and then input into adder 2 for addition in the next clock cycle, while the N-1 group of numbers will be input into Reg_2 for temporary storage, and the N group of numbers will be input into Reg_3 for temporary storage, and will also be input into adder 2 for addition in the next clock cycle. The addition results of the N-3 group of numbers and the N-2 group of numbers will be temporarily stored in Reg_2 again, and the addition results of the N-1 group of numbers and the N group of numbers will be sent to Reg_3 for temporary storage. In the next clock cycle, the results in the two registers will be sent to adder 2 for the last addition operation, and the addition result is the final output result of the multiplication and accumulation.
[0121] Specifically, in all the multiplication and accumulation functions, an output is instantiated at the top level of the module: last4_o. The definition of a multiplication and accumulation is: a row of matrix A is multiplied by a column of matrix B and then added together. This is called a multiplication and accumulation. The number of rows or columns is called the number of multiplication and accumulation points. When a multiplication and accumulation reaches the last four numbers, last4_o needs to be set to the corresponding value to prompt the computing resource controller how to calculate. Specific rules:
[0122] When the number read from BANK is not the last four, last4_o is set to 0;
[0123] When the memory access reads the fourth last number from the BANK and outputs it to the external data buffer, last4_o is set to 1 at the same time and remains at that value until the third last number is read.
[0124] The memory access reads the third last number from the BANK and outputs it to the external data buffer, and at the same time sets last4_o to 2 and keeps it until the second last number is read;
[0125] When the memory access reads the second-to-last number from the BANK and outputs it to the external data buffer, last4_o is set to 3 and remains so until the first-to-last number is read.
[0126] The memory access reads the last number from the BANK and outputs it to the external data buffer, and at the same time sets last4_o to 4 and keeps it until the first number of the next multiplication and accumulation calculation is read;
[0127] For the case where the number of points in a multiplication and accumulation calculation is less than 4:
[0128] The multiplication and accumulation point is 1: set last4_o to 0 at the beginning, directly set last4_o to 4 when sending numbers, and set it to 0 when idle. (The actual change of last4_o should be as follows: 0→4→0→4→0→4…);
[0129] The multiplication and accumulation point is 2: set last4_o to 0 at the beginning, set last4_o to 3 when sending the second to last number of each multiplication and accumulation, and set it to 4 when sending the first to last number. (The actual change of last4_o should be as follows: 0→3→4→0→3→4→0→3→4…);
[0130] The multiplication and accumulation point is 3: set last4_o to 0 at the beginning, set last4_o to 2 when sending the third last number of each multiplication and accumulation, set it to 3 when sending the second last number, and set it to 4 when sending the first last number. (The actual change of last4_o should be as follows: 0→2→3→4→0→2→3→4→0…).
[0131] Specifically, within the same memory access module, all parallel computing paths share the same last4_o. For example, there may be a time interval between reading the fourth last number and reading the third last number (such as a large ping-pong), and at this time, last4_o needs to be kept at 1.
[0132] Specifically, all memory accesses that use multiplication and accumulation, and the rules for transferring source data from memory to PE are as follows:
[0133] (1) When the number of points in a multiplication and accumulation is less than 8, there must be at least 8 beats between each two multiplication and accumulation calculations in the same path (referring to the interval between the last number sent by the previous multiplication and accumulation memory access and the first number sent by the next multiplication and accumulation memory access), otherwise PE will make calculation errors.
[0134] (2) When the number of points in a multiplication and accumulation is greater than or equal to 8, the memory access can transmit source data at will, and every two multiplication and accumulation calculations in the same path can output numbers seamlessly (that is, there can be no gap between the last number sent out by the previous multiplication and accumulation memory access and the first number sent out by the next multiplication and accumulation memory access).
[0135] Specifically, for all memory accesses that use multiplication and accumulation, the PE sends the result data to the memory access. Each multiplication and accumulation calculation only outputs one result. For each multiplication and accumulation, the time it takes for the PE to output the result = the time it takes for the memory access to output the last source data of each multiplication and accumulation + delay. The delay is what the computing resource controller sends to the memory access. The delay is different for different points, and can be roughly divided into the following situations:
[0136] (1) When the number of points in a multiplication and accumulation = 1, delay = 19.
[0137] (2) When the number of points added in a multiplication is 2, delay is 19.
[0138] (3) When the number of points of a multiplication and accumulation = 3, delay = 19.
[0139] (4) When the number of points added in a multiplication operation is greater than or equal to 4, delay = 19.
[0140] Example 3
[0141] In this embodiment, two simulations of 8-point multiplication and accumulation are given. The calculation delay of the adder and the multiplier is 4 clock cycles. req is the input valid signal. The interval of each multiplication and accumulation is 0clk. The waveform in VCS is as follows Figure 3 As shown in the figure, the hand-drawn blue line is the boundary between the first and second multiplication and accumulation. It can be found that if the multiplication and accumulation of eight points are carried out seamlessly, the ADD1 of the first multiplication and accumulation calculation has just completed the third calculation (summing up the first two results of ADD1), and the ADD1 of the second multiplication and accumulation calculation has immediately started the first calculation (calculating the fourth to last and the third to last results of ADD0).
[0142] Therefore, it can be expected that, due to the reuse of the second adder (ADD1), if the next multiplication and accumulation is started without gaps for the case of a relatively small number of points (less than 8), the following situation may occur: before the third calculation of ADD1 of the first multiplication and accumulation calculation is completed (summing the first two results of ADD1), the first calculation of ADD1 of the second multiplication and accumulation calculation has already started (calculating the addition of the fourth-to-last and third-to-last results of ADD0).
[0143] We set Skew = t1-t2, where t1 is the time when the first multiplication and accumulation calculation pulls up the req signal of the third calculation of ADD1, and t2 is the time when the second multiplication and accumulation calculation pulls up the req signal of the first calculation of ADD1. There is no guarantee that there will be no overlap, and Skew ≥ 0 is required.
[0144] After analysis, it is known that the most unfavorable situation is when the number is increased too quickly continuously. Therefore, the subsequent analysis is based on the situation where the number is increased continuously in each multiplication and accumulation.
[0145] like Figure 4 As shown, t1~t5 is a schematic diagram of the interval, wherein t3 is the time interval between the first time req is pulled high for the second multiplication and accumulation ADD1 and the first time req is pulled high for the first multiplication and accumulation ADD1; t4 is the time interval between the third time req is pulled high for the first multiplication and accumulation ADD1 and the first time req is pulled high for the first multiplication and accumulation ADD1.
[0146] Therefore, there is an equation: Skew = t1-t2 = t3-t4, and t3 is obviously equal to N+t5, where N means the number of points in a multiplication and accumulation, and t5 means the time interval between the last point of the first multiplication and accumulation and the first point of the second multiplication and accumulation. And t4 is a fixed value (because of continuous carry), and its value is found to be 8clk through simulation.
[0147] Therefore, to satisfy Skew ≥ 0, we have the following inequality:
[0148] Skew=t1-t2=t3-t4=N+t5-8≥0
[0149] That is: t5 ≥ 8-N. Since t5 must be greater than or equal to 0, there is a requirement for t5 when N < 8. Since ADD1 does not need to be called when N = 1 or 2, it does not need to be considered. Only the cases when N = 3, 4, 5, 6, and 7 are considered.
[0150] as follows Figure 5As shown, it is a waveform diagram when N=7 and t5=1. It can be found that the condition of Skew≥0 is just met, and after verification, it is found that the result is calculated correctly. The case of N=3~7 is similar and will not be repeated here. For the case of N≥8, there is no requirement for t5.
[0151] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A multiplication-accumulation-addition operation device based on a one-multiplication-two-addition structure, characterized in that: Comprising: A multiply-accumulate module and a control module; wherein, The multiply-accumulate module is used to obtain at least two sets of source data, and perform multiply-accumulate operations on the source data through a multiplier and two adders to obtain result data; The control module generates a control signal according to a count warning signal to control the multiply-accumulate operation process of the multiply-accumulate module.
2. The multiplication-accumulation-addition operation device based on a one-multiplication-two-addition structure according to claim 1, characterized in that: The multiply-accumulate module includes: A multiplier and two adders; wherein, The multiplier performs multiplication calculation on the source data to obtain a product result; The adder accumulates the product results to obtain a multiply-accumulate result.
3. The multiplication-accumulation-addition operation device based on a one-multiplication-two-addition structure according to claim 1, characterized in that: The control module includes: Multiple multiplexers, multiple demultiplexers and multiple registers; wherein, The multiplexer selects one signal from multiple input signals transmitted by the demultiplexer or register according to a control signal and outputs it; The demultiplexer distributes an input signal transmitted by the adder or demultiplexer to multiple registers or multiplexers according to a control signal; The register receives and stores the signal transmitted by the multiplexer or demultiplexer according to a control signal.
4. A multiplication-accumulation method based on a one-multiplication-two-addition structure, characterized in that: Comprising: Obtain N sets of source data, where N≥2; Input the source data into the multiplier for multiplication calculation to obtain a product result; Generate a control signal according to a preset adder calculation clock cycle number T and a count warning signal; Input the product result into the first adder, and perform accumulation calculation according to the control signal to obtain a first accumulation result; Input the first accumulation result into the second adder, and perform accumulation calculation according to the control signal to obtain a multiply-accumulate result.
5. The multiplication-accumulation operation method based on a one-multiplication-two-addition structure according to claim 4, characterized in that: The generating a control signal according to a preset adder calculation clock cycle number T and a count warning signal includes: Initializing the count warning signal according to the multiply-accumulate calculation process and the preset adder calculation clock cycle number T; Generate multiple control signals according to the count warning signal, wherein each control signal controls a demultiplexer, a multiplexer or a register.
6. The multiplication-accumulation operation method based on a one-multiplication-two-addition structure according to claim 5, characterized in that: The initializing the count warning signal according to the multiply-accumulate calculation process and the preset adder calculation clock cycle number T includes: When n≤N-T+1, when each set of source data is subjected to multiply-accumulate calculation, the count warning signal is set to 0, where n is the number of input source data sets; When n=N-T+2, at this time, the reciprocal (T-1) groups of source data are input for multiply-accumulate calculation, and the count warning signal is set to 4; When n=N-T+3, at this time, the reciprocal (T-2) groups of source data are input for multiply-accumulate calculation, and the count warning signal is set to 3; When n=N, input the last set of source data for multiply-accumulate calculation, and the count warning signal is set to 2.
7. The multiplication-accumulation operation method based on a one-multiplication-two-addition structure according to claim 4, characterized in that: The inputting the product result into the first adder and performing accumulation calculation according to the control signal to obtain a first accumulation result includes: When 0<n≤T, input the product results calculated from the first T groups of source data into the first adder, and each product result is added to 0 respectively to obtain a single addition result; When T+1≤n≤N, input the product results calculated from the (T+1)th to Nth groups of source data into the first adder, and accumulate them with the single addition result according to the clock cycle number respectively to obtain a first accumulation result.
8. The multiplication-accumulation operation method based on a one-multiplication-two-addition structure according to claim 7, characterized in that: When T+1≤n≤N, the product results calculated from the (T+1)th to Nth groups of source data are input into the first adder, and the product results are accumulated with the first addition results according to the number of clock cycles to obtain the first accumulation result, including: When T+1≤n≤2T, when the product results calculated from the (T+1)th to 2Tth groups of source data are input into the first adder, the product results of the nth group are respectively added to the results of the (nTth)th group in the first addition result to obtain a second addition result; When 2T+1≤n≤3T, when the product results calculated from the (2T+1)th to 3Tth groups of source data are input into the first adder, the product results of the nth group are respectively added to the results of the (nTth)th group in the secondary addition results to obtain a third addition result; Until the accumulation calculation process of N groups of source data in the first adder is completed, a first accumulation result is obtained.
9. The multiplication-accumulation operation method based on a one-multiplication-two-addition structure according to claim 4, characterized in that: The step of inputting the first accumulated result into a second adder and performing an accumulated calculation according to a control signal to obtain a multiplication-accumulation result comprises: Each group of results corresponding to the first accumulated result is input into a plurality of registers for temporary storage; The results stored in the register are sequentially input into the second adder according to the control signal, and accumulation calculation is performed to obtain the multiplication and accumulation result.
10. The multiplication-accumulation operation method based on a one-multiplication-two-addition structure according to claim 9, characterized in that: The step of sequentially inputting the results stored in the register into the second adder according to the control signal, performing accumulation calculation, and obtaining the multiplication and accumulation result comprises: Inputting the (N-T+1)th group of results and the (N-T+2th group of results in the first cumulative result into a second adder, performing addition calculation, and obtaining a primary cumulative result; Inputting the (N-T+3)th group of results and the (N-T+4th group of results in the first cumulative result into a second adder for adding and calculating, thereby obtaining a secondary cumulative result; The first accumulation result and the second accumulation result are input into the second adder for addition calculation to obtain a first multiplication and accumulation result; The accumulation calculation process of the Nth group of results in the second adder is completed to obtain the multiplication and accumulation result.
Citation Information
Patent Citations
Multiply-accumulate operation cluster and data processing method
CN118519613B