Floating point in-memory calculation method and device for simultaneously processing index and mantissa, and storage medium

By pre-sorting and parallel computing of the exponents and mantissa in CIM, the problem of large overhead and low computing efficiency in traditional CIM circuits is solved, and efficient and low-power floating-point in-memory calculation is realized.

CN120447866APending Publication Date: 2025-08-08NANJING INST OF INTELLIGENT TECH INST OF MICROELECTRONICS OF THE CHINESE ACAD OF
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510514330.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

When performing floating-point data calculations, traditional CIM circuits have problems such as large communication overhead between mantissa and exponent, low computing efficiency, and large area overhead for preprocessing units.

Method used

By pre-sorting the exponential and mantissa data, parallel calculations are performed to reduce shift processing, parallel alignment of exponents and mantissa in CIM is achieved, pre-processing units are avoided, and communication overhead is reduced.

Benefits of technology

The calculation efficiency and accuracy of floating-point in-memory calculations are improved, power consumption and area overhead are reduced, and accuracy losses are avoided due to excessive difference between the index and mantissa.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120447866A_ABST
    Figure CN120447866A_ABST
Patent Text Reader

Abstract

The invention discloses a floating-point in-memory calculation method and device for simultaneous processing of indexes and mantissas and a storage medium, and belongs to the technical field of in-memory calculation, and the method comprises the steps: obtaining a shared index of a floating-point number in an index calculation array according to sorted index data; according to the sorted mantissa data, mantissa products of the floating-point number in each row of the mantissa calculation array in each calculation period are obtained, and the calculation period number is consistent with the digit of the input mantissa of the floating-point number; respectively shifting each row of mantissa product in each calculation period based on the sharing index; the sum of multiple rows of shift mantissa products under each shift period is obtained, and the shift period number is consistent with the calculation period number; the product accumulation operation result of the floating-point number is obtained according to the sum of multiple rows of shift mantissa products in each shift period, a preprocessing unit for aligning mantissa in advance is reduced, and high efficiency and low communication overhead are realized in a floating-point number MAC task through parallel index addition and mantissa multiplication steps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of in-memory calculation technology, and in particular relates to a floating-point in-memory calculation method, device and storage medium for simultaneous processing of exponent and mantissa. Background Art

[0002] In recent years, with breakthroughs in key technologies such as big data and artificial intelligence (AI), emerging intelligent applications, such as edge computing and smart living, have emerged rapidly. These emerging intelligent applications often require frequent memory access when processing events. However, the von Neumann architecture, the most commonly used data processing architecture, is implemented by separating memory and computing units. Initially, computing units and storage units developed in parallel, meaning that the storage speed of memory and the computing speed of the arithmetic unit were consistent. However, as semiconductor technology continued to advance in line with Moore's Law, computing units advanced in speed while storage increased in integration. This resulted in a lag between the storage speed of memory and the computing speed of computing units, commonly known as the "memory wall."

[0003] Research shows that the time and power consumption required to access data are far greater than those required for computation, and this difference will become more pronounced as technology continues to advance. The future direction of development is to achieve efficient and low-overhead computing on devices with limited resources. To address this, technicians have proposed a memory-based CIM (Computing in Memory) architecture. This new computer architecture directly utilizes memory to perform logical operations, eliminating the need to move data between memory and the processor. This significantly improves data processing efficiency and reduces device operating power consumption.

[0004] With the development of deep learning algorithms, the scale and complexity of neural networks continue to increase. These large-scale neural network models are now widely used in computer vision, natural language processing, and recommendation systems. Their diverse application scenarios demonstrate their unique importance. Traditional integer neural networks are inherently disadvantaged in large-scale model calculations due to their limited data representation range and low computational accuracy. Floating-point neural networks effectively address this issue. With their wider numerical representation range and computational accuracy, floating-point neural networks are more suitable for use in large-scale neural network models.

[0005] Data processing based on deep learning algorithms involves a large number of MAC (Multiply Accumulate) tasks. Existing CIM circuits with MAC functionality, implemented using SRAM (Static Random-Access Memory), are highly efficient for integer operations. However, traditional CIM circuits require a preprocessing unit and a storage unit to perform floating-point operations. In the preprocessing unit, the exponents of the input data are compared to obtain the largest reference exponent. The difference between the remaining exponents and the largest exponent is then calculated, and the mantissas of each group are right-shifted by the difference. Only after the right-shifted mantissas are input into the storage unit for calculation. This involves significant communication overhead between the mantissa and the exponent, and the mantissa can only participate in calculations after the exponents of both groups have been calculated. This results in low computational efficiency and a large area overhead for the preprocessing unit. Summary of the Invention

[0006] The purpose of the present invention is to provide a floating-point in-memory calculation method, device and storage medium for simultaneous processing of exponents and mantissas. By first sorting the exponent data and mantissa data and then calculating and aligning them in parallel in the CIM, the shift processing in the floating-point calculation process can be reduced, the calculation efficiency and calculation accuracy of the in-memory calculation are greatly improved, and the communication overhead is reduced.

[0007] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0008] In a first aspect, the present invention provides a floating-point in-memory calculation method for simultaneously processing an exponent and a mantissa, comprising: Obtain the shared exponent of the floating-point number in the exponent calculation array according to the sorted exponent data; According to the sorted mantissa data, the mantissa products of the floating-point numbers in each row of the mantissa calculation array are obtained in each calculation cycle, wherein the number of calculation cycles is consistent with the number of digits of the floating-point input mantissa; Shifting the mantissa products of each row in each calculation cycle based on the shared index; Obtain the sum of the products of the shifted mantissas of multiple rows in each shift cycle, where the number of shift cycles is consistent with the number of calculation cycles; The multiplication and accumulation result of the floating-point number is obtained according to the sum of the products of the multiple-row shift mantissas in each shift cycle.

[0009] Optionally, obtaining the shared exponent of the floating-point number in the exponent calculation array according to the sorted exponent data includes: Inputting multiple sets of sorted input indices into corresponding rows of the weight index array in parallel; Add the input index in the input weight index array to the weight index corresponding to the pre-stored sort order in the corresponding row to obtain the sum of the indexes of each row; Get the maximum exponent sum based on the sum of the exponents of each row, and use the maximum exponent sum as a floating-point number for in-memory calculation of the shared exponent.

[0010] Optionally, obtaining the mantissa products of the floating-point numbers in each row of the mantissa calculation array in each calculation cycle according to the sorted mantissa data includes: Inputting the multiple groups of sorted input mantissas serially into corresponding rows of the weight mantissa array according to the bit order of the input mantissas; In the input cycle of each bit data of the input mantissa, the input mantissa in the input weight mantissa array is multiplied with the weight mantissa corresponding to the sort pre-stored in the corresponding row to obtain the mantissa product of each row.

[0011] Optionally, the shifting the mantissa products of each row in each calculation cycle based on the shared index includes: Get the index sum difference between each row index sum and the maximum index sum; During the input cycle of the corresponding bit data, the mantissa products of the corresponding rows are shifted according to the index and difference of each row.

[0012] Optionally, obtaining the sum of the shifted mantissa products of multiple rows in each shift period includes: in each shift period, summing the shifted mantissa products of all rows in the mantissa calculation array through a multi-level addition tree to obtain the sum of the shifted mantissa products in the shift period.

[0013] Optionally, obtaining the floating-point multiplication and accumulation result according to the sum of the products of the multiple-row shift mantissas includes: summing the sum of the products of the multiple-row shift mantissas in all shift cycles to obtain the floating-point multiplication and accumulation result.

[0014] Optionally, when obtaining the shared exponents and the products of the mantissas of each row of floating-point numbers, the floating-point numbers to be calculated are pre-divided into multiple groups of floating-point number arrays according to the sorting results, and the calculation processes of the shared exponents and the products of the mantissas of each row of each group of floating-point number arrays are all parallel.

[0015] In a second aspect, the present invention provides a floating-point in-memory computing device for simultaneously processing an exponent and a mantissa, comprising: Shared index acquisition module: used to obtain the shared index of the floating-point number in the index calculation array according to the sorted index data; Mantissa product acquisition module: used to obtain the mantissa product of the floating-point number in each row of the mantissa calculation array in each calculation cycle according to the sorted mantissa data, wherein the number of calculation cycles is consistent with the number of digits of the floating-point input mantissa; A mantissa product shift module is configured to shift the mantissa products of each row in each calculation cycle based on the shared exponent; Adder module: used to obtain the sum of the products of the shifted mantissas of multiple rows in each shift cycle, where the number of shift cycles is consistent with the number of calculation cycles; Shift accumulation module: used to obtain the multiplication and accumulation operation result of floating-point numbers according to the sum of the products of multiple row shift mantissas in each shift cycle.

[0016] Optionally, the sharing index acquisition module includes: Index addition unit: used to perform addition operations on the input index in the input weight index array and the weight index corresponding to the pre-stored sort order in the corresponding row; Comparator unit: used to obtain the maximum index sum based on multiple groups of index sums; The mantissa product acquisition module includes: Mantissa multiplication unit: used to multiply the input mantissa in the input weight mantissa array with the weight mantissa corresponding to the sorting pre-stored in the corresponding row; The mantissa product shift module comprises: Subtractor unit: used to obtain the exponential sum difference between each group of exponential sums and the maximum exponential sum; Shifter unit: used to shift the product of the mantissa of the corresponding row according to the index and difference of each row.

[0017] In a third aspect, the present invention provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the floating-point in-memory calculation method for simultaneously processing the exponent and mantissa as described in any one of the first aspects.

[0018] Compared with the prior art, the present invention achieves the following beneficial effects: by proposing a method in which the exponent is calculated in a CIM, and the mantissa is calculated in a CIM and aligned in parallel, the mantissa does not need to be aligned in advance, the weighted exponent and the input exponent are directly added in the array, the exponent sum is compared to obtain the maximum value, and the difference between each group of exponent sums and the maximum exponent sum is calculated, the input mantissa does not need to be preprocessed in a peripheral circuit, the input mantissa and the weighted mantissa are directly multiplied in the array, and then shifted according to the difference, thereby minimizing the communication overhead between the exponent and the mantissa, making floating-point in-memory calculations more efficient, eliminating the need for a preprocessing unit, reducing area overhead, and achieving low power consumption and low area in floating-point MAC tasks. In addition, by pre-sorting the exponential data before calculation, the problem of large precision loss caused by excessive differences between the sums of each group of exponents and the shared exponent and excessive mantissa product shift operations can be avoided, thereby improving the accuracy of the calculation results. The present invention also groups floating-point numbers for parallel calculations, further improving calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1FIG2 is a flow chart of a floating-point in-memory calculation method for simultaneously processing the exponent and mantissa in one embodiment of the present invention;

[0020] Figure 2 FIG2 is a schematic diagram of a calculation process of a floating-point exponent calculation unit in an embodiment of the present invention;

[0021] Figure 3 FIG. 1 is a schematic diagram of the calculation process of a floating-point mantissa calculation unit in an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0023] Example 1

[0024] This embodiment provides a floating-point in-memory calculation method for simultaneously processing the exponent and mantissa, including: Obtain the shared exponent of the floating-point number in the exponent calculation array according to the sorted exponent data; According to the sorted mantissa data, the mantissa products of the floating-point numbers in each row of the mantissa calculation array are obtained in each calculation cycle, wherein the number of calculation cycles is consistent with the number of digits of the floating-point input mantissa; Shifting the mantissa products of each row in each calculation cycle based on the shared index; Obtain the sum of the products of the shifted mantissas of multiple rows in each shift cycle, where the number of shift cycles is consistent with the number of calculation cycles; The multiplication and accumulation result of the floating-point number is obtained according to the sum of the products of the multiple-row shift mantissas in each shift cycle.

[0025] To minimize the communication overhead between the exponent and mantissa and improve computational accuracy, a method based on pre-sorting the exponents is proposed, where the exponent is calculated in the CIM and the mantissa is calculated and aligned in parallel. This method eliminates the need for a preprocessing unit, reducing area overhead because exponent addition and mantissa multiplication are performed in parallel. A pipeline is designed to handle these computational steps, improving operational efficiency.

[0026] Example 2

[0027] Based on Example 1, this example also makes the following design.

[0028] The BF16 floating-point format consists of a 7-bit mantissa, an 8-bit exponent, a 1-bit sign bit, and a 1-bit hidden bit. In the floating-point MAC, exponents are added and mantissas are multiplied. The largest exponent sum is used as the final exponent. The mantissa product is shifted based on the difference between the largest exponent sum and the sums of the other exponents. The resulting floating-point number must be converted to the BF16 format. Typically, the sign bit, hidden bit, and mantissa bits are combined into 9 bits and converted to two's complement for multiplication. By default, the weighted mantissa is always positive, meaning 8 bits.

[0029] In this embodiment, the floating-point number to be calculated is 64 groups of BF16. The floating-point in-memory calculation method for simultaneously processing the exponent and mantissa by the floating-point exponent calculation unit and the floating-point mantissa calculation unit is divided into the following steps:

[0030] Step 1: If Figure 1 As shown, through the floating-point exponent calculation unit, 64 groups of 8-bit weight exponents are arranged in descending order of size and pre-stored in 64 rows of the weight index array. The 64 groups of 8-bit input exponents are also arranged in the corresponding order and input in parallel to the corresponding rows of the weight index array. According to the sorting result, the 64 groups of floating-point exponent data in the weight index array are evenly divided into 8 exponent calculation units, and each exponent calculation unit performs the following operations in parallel: the input exponent of each row is added to the corresponding weight exponent to generate a 9-bit exponent sum, and the maximum exponent sum MAX<8:0> is compared among the 8 groups of exponent sums. The difference between the 8 groups of exponent sums and MAX<8:0> is then calculated to generate 8 groups of 9-bit difference values, and each group of difference values is used as the number of bits that need to be shifted for the subsequent mantissa product of the corresponding group.

[0031] Step 2: If Figure 2 As shown, through the floating-point mantissa calculation unit, 64 groups of 8-bit weight mantissas are pre-stored in 64 rows of the weight mantissa array in the order of corresponding exponents, and 64 groups of 9-bit input mantissas (the highest bit is the sign bit) are also arranged in the corresponding order and serially input from MSB (Most Significant Bit) to LSB (Least Significant Bit) into the corresponding row of the weight mantissa array (9 cycles are required). According to the sorting result, the 64 groups of floating-point mantissa data in the weight mantissa array are evenly divided into 8 mantissa calculation units, and each mantissa calculation unit performs the following operations in parallel: the input mantissa of each row is multiplied by the corresponding weight mantissa to generate an 8-bit product, and the 8 groups of 8-bit products are arithmetically right-shifted according to the 8 groups of 9-bit differences calculated in the first step. The shifted data passes through a multi-level addition tree to add each group of mantissa products, and the serial output data is shifted and accumulated through 9 cycles to output the final circuit.

[0032] In the first cycle, the MSBs of 64 data sets are multiplied by 64 sets of 8-bit weighted mantissas to produce 64 sets of 8-bit mantissa products. These mantissa products are then right-shifted based on the corresponding exponents and differences. These 64 sets of right-shifted mantissa products are then passed through a 7-level adder tree to produce a 14-bit sum. The same process applies to cycles 2 through 9. During the shift-accumulation process, the data in the first cycle requires inversion and addition, while the data in the remaining cycles does not. The data from the nine cycles is shifted and accumulated, with the highest bit sign extended and the lowest bit padded with 0 to produce the final result.

[0033] In this embodiment, due to the pre-sorting of the exponential data, it is possible to avoid the difference between each group of exponents and the calculated value and the maximum exponent being too large, avoid too much data being discarded due to excessive shifting of the mantissa product, and reduce the loss of precision.

[0034] In this embodiment, the floating-point in-memory calculation method also supports the configuration of integer INT8, floating-point BF16 and FP32. The overall floating-point in-memory calculation array is 64×64 bits, and the floating-point mantissa calculation unit and floating-point exponent calculation unit are both 64×32 bits. The two large modules are each divided into 4 sub-modules. For BF16 floating-point numbers, 4 different groups of tasks can be processed simultaneously because the weight exponent and mantissa of BF16 are both 8 bits; for INT8 data types, four groups of tasks can also be processed simultaneously; for FP32 data types, the sign bit is 1 bit, the hidden bit is 1 bit, the exponent bit is 8 bits, and the mantissa bit is 23 bits. Compared with BF16, only the mantissa bit is different. When calculating with FP32, the hidden bit and mantissa bits are also combined together, which is also 24 bits. Therefore, during calculation, the mantissa calculation module needs to use three sub-modules to store the weight mantissa, while the exponent bit only needs one sub-module. When calculating the mantissa, 11 cycles are required. The FP32 mantissa input data plus the sign bit is 25 bits. The three sub-modules serially input MSB—MSB-8 (9 bits), MSB-9—MSB-16 (8 bits), and MSB-17—LSB (8 bits) in sequence. The calculation method of the serial input MSB—MSB-8 (9 bits) sub-module is the same as the BF16 mantissa calculation method. The data calculated by the mantissa input MSB of the other two sub-modules does not need to be inverted and added. After the calculation is completed, the three groups of data are shifted and accumulated. The second sub-module and the third sub-module are shifted 8 bits and 16 bits respectively, the high-order sign bit is extended, and the low-order bit is filled with 0. Then the three groups of data are accumulated to calculate the result of the FP32 data type.

[0035] Example 3

[0036] This embodiment provides a floating-point in-memory computing device for simultaneously processing an exponent and a mantissa, comprising: Shared index acquisition module: used to obtain the shared index of the floating-point number in the index calculation array according to the sorted index data; Mantissa product acquisition module: used to obtain the mantissa product of the floating-point number in each row of the mantissa calculation array in each calculation cycle according to the sorted mantissa data, wherein the number of calculation cycles is consistent with the number of digits of the floating-point input mantissa; A mantissa product shift module is configured to shift the mantissa products of each row in each calculation cycle based on the shared exponent; Adder module: used to obtain the sum of the products of the shifted mantissas of multiple rows in each shift cycle, where the number of shift cycles is consistent with the number of calculation cycles; Shift accumulation module: used to obtain the multiplication and accumulation operation result of floating-point numbers according to the sum of the products of the shift mantissas of multiple rows in each shift cycle; The sharing index acquisition module includes: Index addition unit: used to perform addition operations on the input index in the input weight index array and the weight index corresponding to the pre-stored sort order in the corresponding row; Comparator unit: used to obtain the maximum index sum based on multiple groups of index sums; The mantissa product acquisition module includes: Mantissa multiplication unit: used to multiply the input mantissa in the input weight mantissa array with the weight mantissa corresponding to the sorting pre-stored in the corresponding row; The mantissa product shift module includes: Subtractor unit: used to obtain the exponential sum difference between each group of exponential sums and the maximum exponential sum; Shifter unit: used to shift the product of the mantissa of the corresponding row according to the index and difference of each row.

[0037] Example 4

[0038] This embodiment provides a computer storage medium having a computer program stored thereon. When the computer program is executed by a processor, the floating-point in-memory calculation method for simultaneously processing the exponent and mantissa as described in any step of Embodiment 2 is implemented.

[0039] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0040] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0041] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0042] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0043] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which are all protected by the present invention.

Claims

1. A floating-point in-memory calculation method for simultaneously processing exponent and mantissa, characterized in that: include: Obtain the shared exponent of the floating-point number in the exponent calculation array according to the sorted exponent data; According to the sorted mantissa data, the mantissa products of the floating-point numbers in each row of the mantissa calculation array are obtained in each calculation cycle, wherein the number of calculation cycles is consistent with the number of digits of the floating-point input mantissa; Shifting the mantissa products of each row in each calculation cycle based on the shared index; Obtain the sum of the products of the shifted mantissas of multiple rows in each shift cycle, where the number of shift cycles is consistent with the number of calculation cycles; The multiplication and accumulation result of the floating-point number is obtained according to the sum of the products of the multiple-row shift mantissas in each shift cycle.

2. The floating-point in-memory calculation method for simultaneously processing the exponent and mantissa according to claim 1, characterized in that: The step of obtaining the shared exponent of the floating-point number in the exponent calculation array according to the sorted exponent data includes: Inputting multiple sets of sorted input indices into corresponding rows of the weight index array in parallel; Add the input index in the input weight index array to the weight index corresponding to the pre-stored sort order in the corresponding row to obtain the sum of the indexes of each row; Get the maximum exponent sum based on the sum of the exponents of each row, and use the maximum exponent sum as a floating-point number for in-memory calculation of the shared exponent.

3. The floating-point in-memory calculation method for simultaneously processing the exponent and mantissa according to claim 2, characterized in that: The step of obtaining the mantissa products of the floating-point numbers in each row of the mantissa calculation array in each calculation cycle according to the sorted mantissa data comprises: Inputting the multiple groups of sorted input mantissas serially into corresponding rows of the weight mantissa array according to the bit order of the input mantissas; In the input cycle of each bit data of the input mantissa, the input mantissa in the input weight mantissa array is multiplied with the weight mantissa corresponding to the sort pre-stored in the corresponding row to obtain the mantissa product of each row.

4. The floating-point in-memory calculation method for simultaneously processing the exponent and mantissa according to claim 3, characterized in that: The shifting of the mantissa products of each row in each calculation cycle based on the shared index includes: Get the index sum difference between each row index sum and the maximum index sum; During the input cycle of the corresponding bit data, the mantissa products of the corresponding rows are shifted according to the index and difference of each row.

5. The floating-point in-memory calculation method for simultaneously processing the exponent and mantissa according to claim 1, characterized in that: The method of obtaining the sum of the shifted mantissa products of multiple rows in each shift period includes: in each shift period, summing the shifted mantissa products of all rows in the mantissa calculation array through a multi-level addition tree to obtain the sum of the shifted mantissa products in the shift period.

6. The floating-point in-memory calculation method for simultaneously processing the exponent and mantissa according to claim 1, characterized in that: The method of obtaining the floating-point multiplication and accumulation result according to the sum of the products of the multiple-row shift mantissas includes summing the sum of the products of the multiple-row shift mantissas in all shift cycles to obtain the floating-point multiplication and accumulation result.

7. The floating-point in-memory calculation method for simultaneously processing the exponent and mantissa according to claim 1, characterized in that: When obtaining the shared exponents of floating-point numbers and the products of the mantissas of each row, the floating-point numbers to be calculated are pre-divided into multiple groups of floating-point arrays according to the sorting results, and the calculation processes of the shared exponents and the products of the mantissas of each row of each group of floating-point arrays are all parallel.

8. A floating-point in-memory computing device for processing exponents and mantissas simultaneously, characterized in that: include: Shared index acquisition module: used to obtain the shared index of the floating-point number in the index calculation array according to the sorted index data; Mantissa product acquisition module: used to obtain the mantissa product of the floating-point number in each row of the mantissa calculation array in each calculation cycle according to the sorted mantissa data, wherein the number of calculation cycles is consistent with the number of digits of the floating-point input mantissa; A mantissa product shift module is configured to shift the mantissa products of each row in each calculation cycle based on the shared exponent; Adder module: used to obtain the sum of the products of the shifted mantissas of multiple rows in each shift cycle, where the number of shift cycles is consistent with the number of calculation cycles; Shift accumulation module: used to obtain the multiplication and accumulation operation result of floating-point numbers according to the sum of the products of multiple row shift mantissas in each shift cycle.

9. The floating-point in-memory computing device for simultaneously processing exponents and mantissas according to claim 8, wherein: The sharing index acquisition module includes: Index addition unit: used to perform addition operations on the input index in the input weight index array and the weight index corresponding to the pre-stored sort order in the corresponding row; Comparator unit: used to obtain the maximum index sum based on multiple groups of index sums; The mantissa product acquisition module includes: Mantissa multiplication unit: used to multiply the input mantissa in the input weight mantissa array with the weight mantissa corresponding to the sorting pre-stored in the corresponding row; The mantissa product shift module includes: Subtractor unit: used to obtain the exponential sum difference between each group of exponential sums and the maximum exponential sum; Shifter unit: used to shift the product of the mantissa of the corresponding row according to the index and difference of each row.

10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the floating-point in-memory calculation method for simultaneously processing the exponent and mantissa according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • SRAM (Static Random Access Memory) floating point memory internal calculation architecture and calculation method

    CN121349406A

  • Sram floating point in-memory computing architecture and computing method

    CN121349406B