Pure digital memory multiplication circuit architecture and its operation method
Patent Information
- Application Number
- CN202210686720.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2042-06-17
AI Technical Summary
[0006]鉴于上述问题,本发明的目的是提供一种纯数字存内乘法电路架构及其运算方法,以解决现有的纯数字计算方法运算效率低的问题
[0032]The pure digital in-memory multiplication circuit architecture and its operation method provided by this invention, by setting the preprocessing of the judgment module, can realize that reading and calculation are only performed when the bit of A is 1. In this way, if A is a random number, the entire operation cycle will be reduced by half, and the storage and calculation efficiency will be doubled. In addition, the preprocessing of the judgment module can be piped (the preprocessing of the next input is performed while the current calculation is performed), so the preprocessing is almost time-consuming. Therefore, the overall efficiency can be doubled compared with the original scheme. By setting two non-interfering read clock domains and digital clock domains, the read operation and calculation operation of the value B can be separated. Since the value B has been pre-stored in the shift register of the digital clock domain during the pure digital storage and calculation process, it is not necessary to read the value B during the operation and memory process. Therefore, the pure digital in-memory multiplication circuit architecture and its operation method provided by this invention are no longer affected by the read speed of SA (read device)/memory device, thereby greatly improving the overall operation performance.
Smart Images

Figure CN115167810B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of arithmetic circuit design technology, and more specifically, to a pure digital in-memory multiplication circuit architecture and its operation method. Background Technology
[0002] In the traditional von Neumann architecture, due to the separation of computation and storage, storage bandwidth and energy consumption become the main bottlenecks of the entire system as computing speeds increase. Therefore, in-memory computing has been proposed in the field of computing. In-memory computing is a computing method proposed in recent years. Its biggest feature is that it combines storage and computation, thereby solving the main bottlenecks caused by storage bandwidth and energy consumption in the entire computing system.
[0003] In actual calculations, in-memory computation has two directions: analog computation and pure digital computation. For pure digital multiplication, the computing circuit architecture currently mainly used in the industry is as follows: Figure 1 As shown, the specific storage and computing structure is as follows: Figure 2 As shown, by Figure 1 and Figure 2 As can be seen, to calculate A*B (assuming A and B are two numbers with arbitrary bit widths), A is first stored in a shift register, and B is stored in the corresponding storage device. The storage device is equipped with a corresponding readout device SA (the storage device and the corresponding readout device SA constitute a readout unit). Then, each bit of A output by the shift register is ANDed with the readout data of the readout unit (1-bit multiplication) to complete the corresponding multiplication operation. Finally, the results of each readout are accumulated through the shift register to obtain the final calculation result.
[0004] However, the above-mentioned pure numerical calculation method has the following drawbacks: during the operation, it is necessary to read all the bits of A, but only the result of A being 1 will affect the final result. If an int32 operation is to be performed, regardless of whether it is 0 or 1, A needs to be read 32 times, which requires 32 cycles, thus reducing the overall storage and calculation efficiency.
[0005] Based on the above technical requirements, there is an urgent need for a method that can effectively improve the efficiency of pure digital in-memory operations. Summary of the Invention
[0006] In view of the above problems, the purpose of this invention is to provide a pure digital in-memory multiplication circuit architecture and its operation method to solve the problem of low operation efficiency of existing pure digital calculation methods.
[0007] The present invention provides a pure digital in-memory multiplication circuit architecture for calculating the formula A*B, wherein both the numerical values A and B are binary numbers; characterized in that it includes a judgment module and an arithmetic module; wherein,
[0008] The judgment module is used to perform pre-operation on the value A to extract the bits in the value A that are equal to 1;
[0009] The calculation module is used to perform calculations based on the formula A*B, which calculates the bit pairs in the value A that are equal to 1.
[0010] Furthermore, a preferred embodiment is that the judgment module is equipped with a register, the value of bit 0 of the register is preset to 1, and the subsequent bits of the register are used to store the values of each corresponding bit of the value A in sequence; wherein, when storing data into the register, only the bits of the value A with a value of 1 are stored.
[0011] The judgment module outputs the 0 bit of the register and the bits in the value A that are equal to 1 to the operation module from the least significant bit to the most significant bit through the register.
[0012] Furthermore, in a preferred embodiment, the arithmetic module includes a read clock domain and a digital clock domain, wherein the digital clock domain includes an offset register and an accumulator; wherein,
[0013] The clock reading section is used to read the value B in the calculation formula and store it in the displacement register;
[0014] The shift register is used to output the shifted value of the value B, which corresponds to the bit in the value A that is equal to 1.
[0015] The accumulator is used to accumulate the shift values output by the shift register for each shift.
[0016] Furthermore, in a preferred embodiment, the digital clock domain further includes a result register, which is used to store the accumulation results of each accumulator.
[0017] Furthermore, in a preferred embodiment, the result register is also used to store the value B in the calculation formula read from the clock domain.
[0018] Furthermore, in a preferred embodiment, the judgment module outputs the value of bit 0 of the register to the read clock domain section, so that the read clock domain section can read the value B in the calculation formula and store it in the shift register;
[0019] The judgment module outputs the bits equal to 1 in the value A to the digital clock field from the least significant bit to the most significant bit through the register.
[0020] Furthermore, a preferred embodiment is that after all the values stored in the register have been output, the value of bit 0 in the register is reset to 0; and,
[0021] The judgment module is also used to judge the value of bit 0 of the register. When the value of bit 0 of the register is 0, the operation of the pure digital in-memory multiplication circuit architecture is stopped.
[0022] Furthermore, in a preferred embodiment, the read clock domain includes a storage device and a readout device, wherein the read clock domain reads the value B in the formula to be calculated through the storage device in conjunction with the readout device.
[0023] On the other hand, the present invention also provides a pure digital in-memory arithmetic method, wherein the method performs an operation on the formula A*B using the aforementioned pure digital in-memory multiplication circuit architecture; the method includes:
[0024] The judgment module performs a pre-operation on the value A to extract the bits in the value A that are equal to 1;
[0025] The calculation module is used to perform calculations based on the formula A*B, which calculates the bit pairs in the value A that are equal to 1.
[0026] Furthermore, in a preferred embodiment, the judgment module includes a register, and the arithmetic module includes a read clock domain and a digital clock domain, the digital clock domain including a shift register and an accumulator; and the pre-operation of the value A by the judgment module to extract the bits equal to 1 in the value A, and the calculation of the formula A*B by the arithmetic module based on the bits equal to 1 in the value A, include:
[0027] The judgment module outputs the 0 bit of the register and the bits of value A that are equal to 1 to the operation module from the least significant bit to the most significant bit.
[0028] The clock read section reads the value B in the calculation formula based on the 0 bit value of the register and stores it in the shift register;
[0029] The shift register outputs the shifted value of value B, which corresponds to the bit in value A that is equal to 1.
[0030] The accumulator accumulates the shift values output by the shift register to calculate the formula A*B.
[0031] Compared with the prior art, the pure digital in-memory multiplication circuit architecture and its operation method according to the present invention have the following advantages:
[0032] The pure digital in-memory multiplication circuit architecture and its operation method provided by this invention, by setting the preprocessing of the judgment module, can realize that reading and calculation are only performed when the bit of A is 1. In this way, if A is a random number, the entire operation cycle will be reduced by half, and the storage and calculation efficiency will be doubled. In addition, the preprocessing of the judgment module can be piped (the preprocessing of the next input is performed while the current calculation is performed), so the preprocessing is almost time-consuming. Therefore, the overall efficiency can be doubled compared with the original scheme. By setting two non-interfering read clock domains and digital clock domains, the read operation and calculation operation of the value B can be separated. Since the value B has been pre-stored in the shift register of the digital clock domain during the pure digital storage and calculation process, it is not necessary to read the value B during the operation and memory process. Therefore, the pure digital in-memory multiplication circuit architecture and its operation method provided by this invention are no longer affected by the read speed of SA (read device) / memory device, thereby greatly improving the overall operation performance.
[0033] To achieve the foregoing and related objectives, one or more aspects of the invention include the features that will be described in detail below and particularly pointed out in the claims. The following description and accompanying drawings illustrate certain exemplary aspects of the invention. However, these aspects indicate only a few of the various ways in which the principles of the invention can be used. Furthermore, the invention is intended to include all such aspects and their equivalents. Attached Figure Description
[0034] Other objects and results of the invention will become more apparent and readily understood with reference to the following description taken in conjunction with the accompanying drawings and the contents of the claims, and with a more complete understanding of the invention. In the drawings:
[0035] Figure 1 For existing pure digital computing circuit architecture;
[0036] Figure 2 for Figure 1 The in-memory computing structure of the pure digital computing circuit architecture shown;
[0037] Figure 3 The pure digital in-memory multiplication circuit architecture provided by this invention;
[0038] Figure 4 A flowchart of the pure digital in-memory computation method provided by the present invention;
[0039] In all the accompanying drawings, the same reference numerals indicate similar or corresponding features or functions. Detailed Implementation
[0040] In the following description, numerous specific details are set forth for illustrative purposes and to provide a thorough understanding of one or more embodiments. However, it will be apparent that these embodiments may also be implemented without these specific details. In other instances, well-known structures and devices are shown in block diagram form for ease of description of one or more embodiments.
[0041] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. The terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, unless otherwise explicitly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0042] Figure 3 The structure of the pure digital in-memory multiplication circuit architecture provided by the present invention is shown.
[0043] Depend on Figure 3 As shown, the pure digital in-memory multiplication circuit architecture provided by the present invention is used to calculate the formula: A*B, where both the value A and the value B are binary numbers; and the pure digital in-memory multiplication circuit architecture provided by the present invention includes a judgment module and an operation module; wherein, the judgment module is used to perform a pre-operation on the value A to extract the bits in the value A that are equal to 1; the operation module is used to calculate the formula: A*B based on the bits in the value A that are equal to 1.
[0044] The pure digital in-memory multiplication circuit architecture provided by this invention can extract the bits that are equal to 1 in the value A by setting the preprocessing of the judgment module, so that reading and calculation are only performed when the bit of A is 1, thereby reducing the entire operation cycle by half and doubling the storage and calculation efficiency.
[0045] Specifically, to implement the preprocessing of the judgment module, a register is set up in the judgment module. The value of bit 0 of the register is preset to 1. The subsequent bits of the register are used to store the values of each corresponding bit of the value A in sequence. When storing data into the register, only the bits of the value A that are 1 are stored. In the subsequent calculation process, the judgment module outputs bit 0 of the register and the bits of the value A that are equal to 1 to the operation module from the least significant bit to the most significant bit through the register.
[0046] Specifically, to realize the storage and calculation function of the arithmetic module, the arithmetic module includes two separate, unaffected read clock domains and digital clock domains; specifically, to realize the calculation function of the digital clock domain, the digital clock domain includes a shift register and an accumulator; wherein, the read clock domain is used to read the value B in the calculation formula and store it in the shift register; the shift register is used to output the shifted value corresponding to the bit in the value A that is equal to 1 after shifting the value B; the accumulator is used to accumulate the shifted values output by the shift register.
[0047] In the arithmetic module, the shift register and accumulator can be used to accumulate all the shifted bit values corresponding to the bits in the value A that are equal to 1 after the value B is shifted, thereby realizing the calculation of the entire formula.
[0048] It should be noted that in traditional storage and accounting (such as...) Figure 1 and Figure 2 As shown, each calculation requires a read operation on value B, followed by a bitwise AND operation with each bit of value A. This process is cumbersome and limited by the read frequency, resulting in low computational efficiency. However, the pure digital in-memory multiplication circuit architecture provided by this invention performs only one read operation on value B through the read clock domain, places it in the shift register of the digital clock domain, then directly shifts B in the shift register according to each bit of value A, and finally performs an accumulation operation, thus realizing the operation between values A and B. Clearly, the pure digital in-memory multiplication circuit architecture provided by this invention allows all operations to be performed in the traditional digital domain, without being limited by the read frequency, significantly accelerating the overall computational speed.
[0049] It should be noted that the judgment module outputs the 0th bit of the register and the value of the bit equal to 1 in the value A to the operation module from the least significant bit to the most significant bit through the register. The shift value output by the shift register corresponds to the first bit of the value A output by the judgment module from the least significant bit to the most significant bit through the register. The shift register outputs the value of the nth bit of the value A. The shift register outputs the value of the value B shifted n bits to the right.
[0050] Furthermore, to enable data storage during each step, the digital clock domain also includes a result register, which stores the accumulation results of each accumulator step; and the result register also stores the value B in the formula to be calculated read by the read clock domain.
[0051] It should be noted that this judgment module is also used to calculate the bit width of the value A in the formula to be calculated.
[0052] Furthermore, in order to read the value B in the formula to be calculated by the read clock domain, the read clock domain may include a storage device and a readout device. The read clock domain reads the value B in the formula to be calculated by the storage device in conjunction with the readout device.
[0053] Specifically, the judgment module outputs the value of bit 0 of the register to the read clock field, so that the read clock field can read the value B in the calculation formula and store it in the shift register; the judgment module outputs the bits of value A that are equal to 1 to the digital clock field from the least significant bit to the most significant bit through the register.
[0054] In actual operation, after all the values stored in the register are output, the value of bit 0 of the register is returned to 0; and the judgment module is also used to judge the value of bit 0 of the register. When the value of bit 0 of the register is 0, the operation of the pure digital in-memory multiplication circuit architecture is stopped.
[0055] On the other hand, to explain in detail the operation process of the pure digital in-memory multiplication circuit architecture provided by the present invention, Figure 4 The flowchart of the pure digital in-memory computation method provided by the present invention is shown, by... Figure 4 It is understood that the present invention also provides a pure digital in-memory arithmetic method, which performs operations using the aforementioned pure digital in-memory multiplication circuit architecture; the method includes:
[0056] The judgment module performs a pre-operation on the value A to extract the bits in value A that are equal to 1;
[0057] This calculation module is used to perform calculations based on the formula A*B, which calculates the bit pairs in the value A that are equal to 1.
[0058] Furthermore, in a preferred embodiment, the judgment module includes a register, and the arithmetic module includes a read clock domain and a digital clock domain, the digital clock domain including a shift register and an accumulator; and the pre-operation of the value A by the judgment module to extract the bits equal to 1 in value A, and the calculation of the formula A*B by the arithmetic module based on the bits equal to 1 in value A, include:
[0059] The judgment module outputs the 0 bit of the register and the bits in the value A that are equal to 1 from the least significant bit to the most significant bit of the register to the operation module.
[0060] The clock field reads the value B in the calculation formula based on the 0 bit value of the register and stores it in the shift register.
[0061] The shift register outputs the shift value corresponding to the bit in value A that is equal to 1 after shifting the value B.
[0062] The accumulator accumulates the shift values output by the shift register to calculate the formula A*B.
[0063] Specifically, based on the 0-bit value of the register, the value B in the formula to be calculated is read through the read clock field and stored in the shift register;
[0064] The judgment module calculates the bit width n of the value A in the formula to be calculated, and outputs the value An of the nth bit of the value A in sequence based on the clock number N generated by the preset clock control unit (only the case where An = 1 is output); where the clock number N is an integer and increments from 0, and n corresponds to the clock number N;
[0065] Determine whether the value An of the nth bit of the value A currently output by the judgment module is equal to 1. If it is equal to 1, then generate the shifted value of value B after shifting n bits through the shift register.
[0066] The shift value generated by the shift register is accumulated into the result register by the accumulator;
[0067] Based on the increment of the clock number N, repeat the above steps until the calculation of the highest bit of the value A is completed.
[0068] In addition, it should be noted that if the value An of the nth bit of the current output value A of the bit width statistics output unit is not equal to 1, then the calculation of that bit of the value A is stopped, and the next bit of the value A is calculated directly based on the increment of the clock number N.
[0069] As can be seen from the above specific embodiments, the pure digital in-memory multiplication circuit architecture and its operation method provided by the present invention have at least the following advantages:
[0070] 1. By setting the preprocessing of the judgment module, it is possible to read and calculate only when the bit of A is 1. In this way, if A is a random number, the entire calculation cycle will be reduced by half and the storage and calculation efficiency will be doubled.
[0071] 2. The preprocessing of the judgment module can be pipelined (preprocessing the next input while the current calculation is being performed), so preprocessing takes almost no time. Therefore, the overall efficiency can be doubled compared to the original scheme. By setting up two non-interfering read clock domains and digital clock domains, the read operation of value B can be separated from the calculation operation. Since value B has been pre-stored in the shift register of the digital clock domain during the pure digital storage and calculation process, it is not necessary to read value B during the operation and storage process. Therefore, the pure digital in-memory multiplication circuit architecture and its operation method provided by this invention are no longer affected by the read speed of SA (read device) / memory device, thereby greatly improving the overall operation performance.
[0072] As per the above reference Figure 3 and Figure 4 The pure digital in-memory multiplication circuit architecture and its operation method according to the present invention are described by way of example. However, those skilled in the art should understand that various modifications can be made to the pure digital in-memory multiplication circuit architecture and its operation method proposed in the present invention without departing from the scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the content of the appended claims.
Claims
1. A purely digital in-memory multiplication circuit architecture for calculating formula A B, of which Both numerical values A and B are binary numbers; the feature is that it includes a judgment module and an arithmetic module; wherein, The judgment module is used to perform pre-operation on the value A to extract the bits in the value A that are equal to 1; The calculation module is used to calculate the formula A based on the bit pairs equal to 1 in the value A. B performs the calculation; wherein, a register is set in the judgment module, the value of bit 0 of the register is preset to 1, and the subsequent bits of the register are used to store the values of each corresponding bit of the value A in sequence; wherein, when storing in the register, only the bits of the value A with a value of 1 are stored.
2. The pure digital in-memory multiplication circuit architecture as described in claim 1, characterized in that, The judgment module outputs the 0 bit of the register and the bits in the value A that are equal to 1 to the operation module from the least significant bit to the most significant bit through the register.
3. The pure digital in-memory multiplication circuit architecture as described in claim 2, characterized in that, The arithmetic module includes a read clock domain and a digital clock domain, wherein the digital clock domain includes an offset register and an accumulator; wherein... The clock reading section is used to read the value B in the formula and store it in the shift register; The shift register is used to output the shifted value of the value B, which corresponds to the bit in the value A that is equal to 1. The accumulator is used to accumulate the shift values output by the shift register for each shift.
4. The pure digital in-memory multiplication circuit architecture as described in claim 3, characterized in that, The digital clock domain also includes a result register, which stores the accumulation results of the accumulator.
5. The pure digital in-memory multiplication circuit architecture as described in claim 4, characterized in that, The result register is also used to store the value B in the formula read from the clock field.
6. The pure digital in-memory multiplication circuit architecture as described in claim 3, characterized in that, The judgment module outputs the value of bit 0 of the register to the read clock field, so that the read clock field can read the value B in the formula and store it in the shift register; The judgment module outputs the bits equal to 1 in the value A to the digital clock field from the least significant bit to the most significant bit through the register.
7. The pure digital in-memory multiplication circuit architecture as described in claim 3, characterized in that, After all the values stored in the register have been output, the value of bit 0 in the register is reset to 0; and, The judgment module is also used to judge the value of bit 0 of the register. When the value of bit 0 of the register is 0, the operation of the pure digital in-memory multiplication circuit architecture is stopped.
8. The pure digital in-memory multiplication circuit architecture as described in any one of claims 3 to 7, characterized in that, The read clock domain includes a storage device and a readout device. The read clock domain reads the value B in the formula through the storage device and the readout device.
9. A pure digital in-memory arithmetic method, characterized in that, The method utilizes the pure digital in-memory multiplication circuit architecture as described in any one of claims 1 to 8 to perform formula A. B performs the calculation; the method includes: The judgment module performs a pre-operation on the value A to extract the bits in the value A that are equal to 1; The calculation module uses a formula based on the bit pairs equal to 1 in the value A: A B performs the calculation; wherein, a register is set in the judgment module, the value of bit 0 of the register is preset to 1, and the subsequent bits of the register are used to store the values of each corresponding bit of the value A in sequence; wherein, when storing in the register, only the bits of the value A with a value of 1 are stored.
10. The pure digital in-memory arithmetic method as described in claim 9, characterized in that, The arithmetic module includes a read clock domain and a digital clock domain, the digital clock domain including a shift register and an accumulator; and the judgment module performs a pre-operation on the value A to extract the bits equal to 1 in the value A, and the arithmetic module calculates a formula based on the bit pairs equal to 1 in the value A: A B's calculations include: The judgment module outputs the 0 bit of the register and the bits of value A that are equal to 1 to the operation module from the least significant bit to the most significant bit. The clock read section reads the value B in the calculation formula based on the 0 bit value of the register and stores it in the shift register; The shift register outputs the shifted value of value B, which corresponds to the bit in value A that is equal to 1. The accumulator accumulates the shift values output by the shift register to achieve formula A. Calculation of B.
Citation Information
Patent Citations
Multi-bit multiplexing multiply-add operation device, neural network operation system, and electronic apparatus
CN111694544A