Multi-core-particle-based multifunctional storage and calculation integrated accelerator

By adopting a multi-core particle design in the multi-functional in-memory computing accelerator, in-memory computing core particles share resources with near-memory peripheral circuit core particles, the problems of area limitation, function fixation and resource waste caused by single-core particle design are solved, and more efficient hardware resource utilization and function expansion are achieved.

CN120067032APending Publication Date: 2025-05-30INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510076905.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing multifunctional in-memory computing accelerator uses a single core, resulting in limited area, fixed functions and difficulty in scaling, and there are problems of wasted hardware resources.

Method used

A multi-functional memory and computing integrated accelerator based on multi-core particles is designed. Through the combination of several in-memory computing core particles and near-memory peripheral circuit core particles, some in-memory computing core particles share the resources of near-memory peripheral circuit core particles, and realize flexible function calls and expansion.

Benefits of technology

It reduces the overhead of peripheral circuit area, improves the utilization rate of hardware resources, and enhances the functional flexibility and scalability of the accelerator.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067032A_ABST
    Figure CN120067032A_ABST
Patent Text Reader

Abstract

The invention provides a multi-core-particle-based multifunctional storage and calculation integrated accelerator, which comprises a plurality of in-storage calculation core particles and a plurality of near-storage peripheral circuit core particles, and is characterized in that at least part of the in-storage calculation core particles share resources of one or more near-storage peripheral circuit core particles, the accelerator is configured as follows: when logic calculation or four arithmetic calculation of integer data is executed, resources of calculation core particles in a memory are scheduled to perform in-situ calculation, and an in-situ calculation result is obtained; when calculation of the floating-point type data is executed, resources of the in-memory calculation core particles and resources of the near-memory peripheral circuit core particles are scheduled for calculation together, and the calculation result of the floating-point number is obtained.According to the technical scheme, the in-memory calculation core particles and the peripheral circuit core particles are separated and mutually independent, the expandability is good, and the calculation accuracy is high. The plurality of in-memory computing core particles can share the resources of the peripheral circuit core particles, so that the problem of waste of hardware resources is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of chiplet integration and computing-in-memory, specifically to the field of computing-in-memory accelerators, and more specifically, to a multi-functional computing-in-memory accelerator based on multi-chiplets. Background Art

[0002] Computing-in-memory combines part of the computing logic with the memory, reducing the data transfer volume between the processor and the memory, thereby better meeting the requirements of applications such as neural networks and scientific computing. Computing-in-memory technology can be divided into near-memory computing and in-memory computing technologies. Among them, in-memory computing has advantages in accelerating matrix-vector multiplication operators, while near-memory computing can support vector operations and floating-point calculations. Current research focuses on designing accelerators based on in-memory computing methods or near-memory computing methods to accelerate different application scenarios, and these accelerators often use a single chiplet.

[0003] Some researchers have proposed multi-functional in-memory computing accelerators. However, existing multi-functional in-memory computing accelerators use a single chiplet. Due to the limited area of a single chiplet and the fixed functional modules inside the chiplet, the current multi-functional computing-in-memory accelerators can only accelerate preset limited functions, and have the disadvantages of small scale and difficulty in flexible function expansion.

[0004] In addition, in current research, since a large number of peripheral circuits of multi-functional modules only for the use of this chiplet are integrated inside a single chiplet of a single accelerator, and its in-memory computing array is strictly bound to the peripheral circuits in one chiplet, and there is a situation where it is only configured as one function during actual use, the peripheral circuits of other remaining functional modules will be idle, which will cause waste of area and hardware resources.

[0005] Therefore, in existing multi-functional in-memory computing accelerators using a single chiplet, on the one hand, a large number of multi-functional modules only for the use of this chiplet are integrated in a single chiplet, and there is a situation where only some functions are used while other functional modules are idle, which easily leads to the problem of waste of the area and hardware resources of the accelerator; on the other hand, the area of a single chiplet is limited and the functional modules inside a single chiplet are fixed, resulting in the problems of fixed functions and poor scalability of these accelerators.

[0006] It should be noted that: This background art is only used to introduce relevant information of the present invention to help understand the technical solution of the present invention, but it does not mean that the relevant information is necessarily prior art. The relevant information is submitted and disclosed together with the solution of the present invention. Without evidence indicating that the relevant information has been publicly available before the filing date of the present invention, the relevant information should not be regarded as prior art. Summary of the Invention

[0007] Therefore, the object of the present invention is to overcome the defects of the above-mentioned prior art and provide a multi-die-based multi-functional memory-computation integrated accelerator.

[0008] The object of the present invention is achieved by the following technical solutions:

[0009] According to a first aspect of the present invention, there is provided a multi-die-based multi-functional memory-computation integrated accelerator, the accelerator comprising a plurality of in-memory computing dies and a plurality of near-memory peripheral circuit dies, at least some of the in-memory computing dies sharing the resources of one or more near-memory peripheral circuit dies, wherein the accelerator is configured to: when performing logical calculations or four arithmetic operations on integer data, schedule the resources of the in-memory computing dies for in-situ calculation to obtain an in-situ calculation result; when performing calculations on floating-point data, schedule the resources of the in-memory computing dies and the near-memory peripheral circuit dies to jointly perform calculations to obtain a calculation result of the floating-point number.

[0010] In some embodiments of the present invention, the in-memory computing die comprises a plurality of crossbar switch arrays, each crossbar switch array comprising a plurality of memory cells, wherein the plurality of crossbar switch arrays are configured to: when performing logical calculations or four arithmetic operations on integer data, read and store the integer data participating in the calculation through the memory cells, and perform logical calculations or four arithmetic operations on the integer data in-situ to obtain an in-situ calculation result; when performing calculations on floating-point data, read and store the floating-point data participating in the calculation through the memory cells; wherein the in-memory computing die represents the floating-point data and / or the in-situ calculation result as a digital signal and transmits it to the near-memory peripheral circuit die that jointly calculates with it to complete the calculation.

[0011] In some embodiments of the present invention, when performing multiply-accumulate calculations on integer data, the accelerator is further configured to: schedule the crossbar switch array of the in-memory computing die to perform multiply-accumulate calculations to obtain an intermediate multiply-accumulate calculation result, and transmit the intermediate multiply-accumulate calculation result to the near-memory peripheral circuit die that jointly calculates with it; schedule the resources of the near-memory peripheral circuit die that jointly calculates with it, and perform shift-accumulate calculations according to the intermediate multiply-accumulate calculation result to obtain a multiply-accumulate calculation result of the integer data.

[0012] In some embodiments of the present invention, the in-memory computing die further includes an input / output register, a data input driver, and a sense amplifier. The data input driver is configured to input integer data and / or floating-point data participating in the calculation into a plurality of crossbar arrays. The sense amplifier is configured to control the crossbar arrays to perform logical calculations or four arithmetic calculations on integer data, and read floating-point data represented by digital signals and / or in-situ calculation results represented by digital signals from the crossbar arrays. The input / output register is configured to temporarily store integer data and / or floating-point data participating in the calculation, as well as temporarily store the in-situ calculation results represented by digital signals and the floating-point data represented by digital signals read by the sense amplifier.

[0013] In some embodiments of the present invention, the in-memory computing die further includes a sample-and-hold circuit. When performing multiply-accumulate calculations on integer data, the sample-and-hold circuit is configured to sample and hold the intermediate multiply-accumulate calculation results represented by analog signals obtained by the plurality of crossbar arrays performing multiply-accumulate calculations, and transmit them to the near-memory peripheral circuit die for joint calculation.

[0014] In some embodiments of the present invention, the plurality of crossbar arrays are further configured to store floating-point data by storing the exponent and mantissa of the floating-point data, where the exponent and mantissa of the floating-point data are stored in different crossbar arrays respectively.

[0015] In some embodiments of the present invention, the near-memory peripheral circuit die includes: a digital-to-analog converter, configured to convert integer data participating in the calculation into integer data represented by analog signals and transmit them to the in-memory computing die; a register file, configured to store data, including storing the in-situ calculation results; a digital circuit module, configured to align each bit value of two floating-point data participating in the calculation; and a floating-point calculation unit, configured to perform calculations between the two aligned floating-point data.

[0016] In some embodiments of the present invention, the near-memory peripheral circuit die includes an analog-to-digital converter and a shift-accumulation circuit. When performing multiply-accumulate calculations on integer data, the near-memory peripheral circuit die is configured to: use the analog-to-digital converter to perform analog-to-digital conversion on the intermediate multiply-accumulate calculation results to obtain the intermediate multiply-accumulate calculation results represented by digital signals; and use the shift-accumulation circuit to perform shift-accumulation operations based on the intermediate multiply-accumulate calculation results to obtain the multiply-accumulate calculation results of the integer data.

[0017] In some embodiments of the present invention, when performing calculations on floating-point data, the near-memory peripheral circuit die is configured to: obtain the floating-point data participating in the calculation, use the digital circuit module to align each bit value of the two floating-point data; and input the two aligned floating-point data into the floating-point calculation unit to obtain the calculation results of the floating-point numbers.

[0018] In some embodiments of the present invention, the digital circuit module is further configured to perform the function of an activation function, perform the function of a trigonometric function, and / or perform the function of a pooling operation.

[0019] Compared with the prior art, the advantages of the present invention are as follows:

[0020] First, the accelerator of the present invention is provided with a plurality of in-memory computing dies and a plurality of near-memory peripheral circuit dies, and some of the in-memory computing dies share the resources of the near-memory peripheral circuit dies, which reduces the area overhead of the peripheral circuit as a whole and improves the utilization rate of area and hardware resources. Second, the in-memory computing is separated from the peripheral circuit, and independent in-memory computing dies and near-memory peripheral circuit dies are provided, so that the accelerator can call the resources of the in-memory computing dies and / or near-memory peripheral circuit dies according to the computing function requirements, improving the flexibility of the hardware function application of the accelerator. Finally, according to the computing requirements, the numbers of the in-memory computing dies and near-memory peripheral circuit dies can be expanded respectively, and the scalability is good. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The following further describes the embodiments of the present invention with reference to the drawings, where:

[0022] Figure 1 FIG. is a schematic structural diagram of a multi-die-based multi-functional memory-computation integrated accelerator according to an embodiment of the present invention;

[0023] Figure 2 FIG. is a schematic structural diagram of an in-memory computing die according to an embodiment of the present invention;

[0024] Figure 3 FIG. is a schematic structural diagram of a near-memory peripheral circuit die according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below through specific embodiments with reference to the drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0026] As mentioned in the background art section, in the existing multi-functional in-memory computing accelerators using a single die, on the one hand, a large number of multi-functional modules only for this die are integrated in a single die, and there is a situation where only some functions are used while other function modules are idle, which easily leads to the problem of waste of the area and hardware resources of the accelerator;; on the other hand, the area of a single die is limited and the function modules inside a single die are fixed, resulting in the problems of fixed functions and poor scalability of these accelerators.

[0027] To solve the above problems, the inventors have found through research that the reasons for the above problems include: existing accelerators integrate the functions of in-memory computing and peripheral circuits in a single die, forming a fixed functional module. Therefore, according to an embodiment of the present invention, a multi-die-based multi-functional in-memory computing accelerator is proposed. For the problems in the first aspect, the present invention is provided with a number of in-memory computing dies and a number of near-memory peripheral circuit dies, and enables some in-memory computing dies to share the resources of the near-memory peripheral circuit dies, avoiding the need to separately integrate near-memory peripheral circuit dies around each in-memory computing die, reducing the area overhead of the peripheral circuits as a whole, and improving the utilization rate of area and hardware resources; for the problems in the second aspect, the present invention separates in-memory computing from the peripheral circuits, and sets up independent in-memory computing dies and near-memory peripheral circuit dies, enabling the accelerator to call the resources of the in-memory computing dies and / or near-memory peripheral circuit dies according to the computing function requirements, improving the flexibility of hardware function applications. At the same time, according to the computing requirements, the number ratio of the in-memory computing dies and the near-memory peripheral circuit dies can be adjusted, or a near-memory peripheral circuit die that meets the computing requirements can be extended and connected, with good scalability.

[0028] To better understand the present invention, the overall structural principle of the accelerator, the structural principle of the in-memory computing die, and the structural principle of the near-memory peripheral circuit die will be described below with reference to the accompanying drawings.

[0029] I. Overall Structural Principle of the Accelerator

[0030] According to an embodiment of the present invention, the accelerator can be used for matrix-vector calculations, such as logical calculations, four arithmetic calculations, and multiply-accumulate calculations of matrix vectors. Matrix-vector calculations are divided into low-precision data calculations and high-precision data calculations. Among them, high-precision data calculations refer to operations in which the range of numbers involved in the operation (such as addends and / or subtrahends) far exceeds the range that can be represented by the standard data type (integer or real). Among them, calculations involving integer data are called low-precision calculations; since floating-point data has a longer number of bits, generally 32 or 64 bits, calculations involving floating-point data are called high-precision data calculations. For example, addition and multiplication of floating-point data belong to high-precision data calculations.

[0031] According to an embodiment of the present invention, the accelerator includes a number of in-memory computing dies and a number of near-memory peripheral circuit dies, and at least some in-memory computing dies share the resources of one or more near-memory peripheral circuit dies. The accelerator is configured to: when performing logical calculations or four arithmetic calculations of integer data, schedule the resources of the in-memory computing dies for in-situ calculations to obtain in-situ calculation results; when performing calculations of floating-point data, schedule the resources of the in-memory computing dies and the near-memory peripheral circuit dies to perform calculations together to obtain the calculation results of floating-point numbers. Calculations of floating-point data include four arithmetic calculations and multiply-accumulate calculations of floating-point data.

[0032] According to an embodiment of the present invention, the accelerator is further configured to: when performing multiply-accumulate calculations on integer data, schedule the in-memory computing dielets to perform multiply-accumulate calculations to obtain intermediate multiply-accumulate calculation results, and transmit the intermediate multiply-accumulate calculation results to the near-memory peripheral circuit dielets that jointly calculate with it through the in-memory computing dielets; schedule the resources of the near-memory peripheral circuit dielets that jointly calculate with the in-memory computing dielets, and perform shift accumulation calculations based on the intermediate multiply-accumulate calculation results to obtain the multiply-accumulate calculation results of the integer data.

[0033] According to an embodiment of the present invention, refer to Figure 1 , which is a schematic diagram of the multi-dielet multi-functional memory-computing integrated accelerator structure. It includes 6 in-memory computing dielets and 3 near-memory peripheral circuit dielets. The two types of dielets are connected in a grid topology and data is transmitted through the network on the substrate. Among them, the two dielets that do not perform data transmission are connected in a straight line, and the two dielets that need to perform data transmission are connected by a straight line with an arrow. Therefore, data is not transmitted between the in-memory computing dielets, nor between the near-memory peripheral circuit dielets. Data can only be transmitted between the in-memory computing dielets and the near-memory peripheral circuit dielets. The same in-memory computing dielet can transmit data to multiple near-memory peripheral circuit dielets, and the in-memory computing dielet can receive data from multiple near-memory peripheral circuit dielets. In the figure, each near-memory peripheral circuit dielet is connected to 2 in-memory computing dielets and can perform data transmission. Therefore, these 2 in-memory computing dielets share one near-memory peripheral circuit dielet.

[0034] According to an embodiment of the present invention, when performing layout design on the in-memory computing dielets and the near-memory peripheral circuit dielets on the substrate, a heuristic algorithm can be used for functional layout. First, an initial layout state is randomly set; then, the heuristic algorithm is used for iterative optimization, and the layout is adjusted according to the rules of the algorithm; finally, after a certain number of rounds, the optimized final layout is obtained as the layout result.

[0035] According to an embodiment of the present invention, the heuristic algorithm is a type of existing algorithm. Taking the simulated annealing algorithm as an example, by setting an optimization function, the value of this optimization function is used as the optimization goal. The heuristic algorithm will optimize the layout so that the value of the optimization function is adjusted in the increasing or decreasing direction.

[0036] Schematically, the optimization function is to calculate the data exchange volume between all dielets, and the dielet layout is optimized by optimizing the value of this function in the direction with less data exchange volume. Initially, the value of the optimization function is calculated for a random layout. Then, each adjustment method is set to randomly swap the layout positions of two dielets, recalculate the value of this optimization function, and determine whether the new solution is better or worse based on the difference between the new value and the old value of the optimization function. If it is better, the better solution will be selected; if it is worse, the worse solution will be selected with a preset probability. Since the simulated annealing algorithm has an initial temperature, which decreases as the number of iterations increases, this temperature determines the preset probability of accepting a worse solution in each iteration, thus giving the algorithm the ability to jump out of the local optimal solution. After the number of iterations reaches the preset maximum number, the last layout is taken as the final result.

[0037] II. In-Memory Computing Dielet Structure Principle

[0038] According to an embodiment of the present invention, refer to Figure 2 , which is a schematic diagram of the structure of an in-memory computing dielet. The in-memory computing dielet includes a data input driver, multiple crossbar switch arrays, sense amplifiers, input / output registers, and a sample and hold circuit. Among them, the data input driver, sense amplifiers, and sample and hold circuit are respectively connected to the crossbar switch arrays. Each crossbar switch array includes multiple memory cells. The devices used for the memory cells can be volatile memory devices, such as static random-access memory (SRAM for short), etc., or non-volatile memory devices, such as resistive random-access memory (ReRAM for short), etc. The internal structure of the memory cell may consist of multiple transistors and include necessary metal-oxide-semiconductor field-effect transistors (MOSFET for short) for controlling the reading and writing of the memory cell.

[0039] Among them, the data input driver is used to input data. The crossbar switch array is used to store the data participating in the calculation and perform logical calculations on integer data, four arithmetic operations on integer data, or multiply-accumulate operations on integer data in situ. The sense amplifier is used to read and amplify the data. The input / output register is used to temporarily store the input data and output data of the in-memory computing die. The sample-and-hold circuit is used to temporarily store the analog signal output by the crossbar switch array. The analog signal output by the crossbar switch array: refers to the accumulated current value in the crossbar switch array. After activating multiple rows of data in the crossbar switch array and completing the calculation, the calculation result obtained by the crossbar switch array will be accumulated in the form of an analog signal of current. Before converting the analog signal output by the crossbar switch array into a digital signal, the sample-and-hold circuit is used to hold the analog signal to wait for conversion into a digital signal.

[0040] The structural principles of each part of the in-memory computing die in the above embodiments will be introduced in detail below:

[0041] 1) Data input driver

[0042] According to an embodiment of the present invention, the data input driver is used to input integer data and / or floating-point data participating in the calculation into a plurality of crossbar switch arrays.

[0043] 2) A plurality of crossbar switch arrays

[0044] According to an embodiment of the present invention, a plurality of crossbar switch arrays are configured to: when performing logical calculations or four arithmetic operations on integer data, read and store the integer data participating in the calculation through its storage unit, and perform logical calculations or four arithmetic operations on integer data in situ to obtain an in-situ calculation result; when performing calculations on floating-point data, read and store the floating-point data participating in the calculation through its storage unit. Among them, in-situ calculation means directly performing calculations at the position where the data is stored in the storage unit (such as calculating in the crossbar switch array); in other words, the data participating in the calculation does not leave the in-memory computing die, and the calculation process occurs inside the in-memory computing die.

[0045] According to an embodiment of the present invention, a plurality of crossbar switch arrays are further configured to: perform multiply-accumulate calculations on integer data to obtain an intermediate result of the multiply-accumulate calculation represented by an analog signal, and transmit the intermediate result of the multiply-accumulate calculation to the near-memory peripheral circuit die for joint calculation through the sample-and-hold circuit.

[0046] According to an embodiment of the present invention, multiple crossbar switch arrays are further configured to store floating-point data by storing the exponent and mantissa of the floating-point data, wherein the exponent and mantissa of the floating-point data are stored in different crossbar switch arrays respectively. Among them, the floating-point data in negative form will be converted into positive form by adding an offset value to the whole to achieve its storage. The technical solution of this embodiment can at least achieve the following beneficial technical effects: Since the number of bits of floating-point data is long, the present invention stores it in a separated form of exponent and mantissa, which can reduce the amount of data stored and transmitted and is also beneficial to calculation.

[0047] 3) Sense amplifier

[0048] According to an embodiment of the present invention, the sense amplifier is used to control the crossbar switch array to perform logical calculations or four arithmetic calculations of integer data. Among them, when performing logical calculations of integer data: the sense amplifier can cooperate with other logic gate circuits in the crossbar switch array to achieve logical calculations of integer data, such as AND, OR, NOT, etc. By controlling the on and off states of each switch in the crossbar switch array and the detection and amplification of signals by the sense amplifier, logical calculations of integer data can be achieved. When performing four arithmetic calculations of integer data: the sense amplifier is used to detect and amplify the integer data signals participating in the calculation, and realize the four arithmetic calculations of integer data, such as addition, subtraction, multiplication and division, through the corresponding circuit structure in the crossbar switch array.

[0049] According to an embodiment of the present invention, the sense amplifier is also used to read floating-point data represented by digital signals and in-situ calculation results represented by digital signals from multiple crossbar switch arrays.

[0050] According to an embodiment of the present invention, after the sense amplifier of the in-memory computing die represents the floating-point data and / or the in-situ calculation result as a digital signal, it is necessary to transmit the floating-point data and / or the in-situ calculation result represented by the digital signal to the near-memory peripheral circuit die that jointly calculates with it. If the digital signal transmitted to the near-memory peripheral circuit die is floating-point data, the near-memory peripheral circuit die performs the calculation of the floating-point data. If the digital signal transmitted to the near-memory peripheral circuit die is the in-situ calculation result, the result is directly stored in the near-memory peripheral circuit die.

[0051] 4) Input / output register

[0052] According to an embodiment of the present invention, the input / output register is used to temporarily store the integer data and / or floating-point data participating in the calculation, as well as the in-situ calculation result represented by the digital signal and the floating-point data represented by the digital signal read by the sense amplifier.

[0053] 5) Sample and hold circuit

[0054] According to an embodiment of the present invention, a sample-and-hold circuit is used to sample and hold the intermediate results of multiply-accumulate calculations obtained by multiple crossbar arrays during the execution of integer data multiply-accumulate calculations, and transmit them to the near-memory peripheral circuit die for co-computation, waiting to be converted into digital signals.

[0055] According to an embodiment of the present invention, during the execution of integer data multiply-accumulate calculations, the accelerator is further configured to: schedule the crossbar arrays of the in-memory computing die to perform multiply-accumulate calculations to obtain intermediate results of multiply-accumulate calculations represented by analog signals, and transmit the intermediate results of multiply-accumulate calculations to the near-memory peripheral circuit die for co-computation through a sample-and-hold unit; schedule the resources of the near-memory peripheral circuit die for co-computation, and perform shift-accumulate calculations based on the intermediate results of multiply-accumulate calculations to obtain the results of integer data multiply-accumulate calculations.

[0056] III. Structure and Principle of Near-Memory Peripheral Circuit Die

[0057] According to an embodiment of the present invention, refer to Figure 3 , which is a schematic diagram of the structure of the near-memory peripheral circuit die. The near-memory peripheral circuit die includes a digital-to-analog converter (DAC), an analog-to-digital converter (ADC), a shift-accumulate circuit, a controller, a register file, a digital circuit module, and a floating-point calculation unit or a combination thereof.

[0058] Among them, the DAC is used to implement the conversion from digital signals to analog signals, the ADC is used to implement the conversion from analog signals to digital signals, the shift-accumulate circuit is used to implement shift-accumulate operations, the controller is used to generate control signals to control the operation of each part of the near-memory peripheral circuit die, the register file is used to store the input data and output data of the near-memory peripheral circuit die, the floating-point calculation unit is used to perform calculations on floating-point data, and the digital circuit module is used for function expansion.

[0059] According to an embodiment of the present invention, the digital-to-analog converter (DAC) is used to convert integer data into integer data represented by analog signals during the execution of relevant calculations on integer data, so as to transmit the integer data represented by analog signals to the in-memory computing die. At this time, the in-memory computing die inputs the transmitted integer data into the crossbar array through its data input driver.

[0060] According to an embodiment of the present invention, a register file is used to store data, including storing the in-situ calculation results obtained when the in-memory computing die performs logical calculations or four arithmetic calculations on integer data. For example, the calculation of integer data is a + b + c. First, the in-memory computing die calculates a + b. Then, the in-situ calculation result obtained from a + b is transmitted to the register file of the near-memory peripheral circuit die that jointly calculates with it for storage. Finally, the digital-to-analog converter transmits the result to the in-memory computing die to add the in-situ calculation result obtained from a + b to c to obtain the final result, which is also written into the register file for storage. The technical solution of this embodiment can at least achieve the following beneficial technical effects: Since directly using a crossbar array to store in-situ calculation results will face problems such as slow writing speed and limited device life of the array, it is impossible to achieve efficient storage. By writing the in-situ calculation results into the register file of the near-memory peripheral circuit die, the present invention reduces the burden on the array, thereby increasing the service life of the array, and can quickly store the in-situ calculation results, improving the calculation efficiency.

[0061] According to an embodiment of the present invention, when performing multiply-accumulate calculations on integer data, after the in-memory computing die transmits the intermediate result of the analog signal multiply-accumulate calculation to the near-memory peripheral circuit die that jointly calculates with it through a sample-and-hold unit, the near-memory peripheral circuit die is configured to: use an analog-to-digital converter ADC to perform analog-to-digital conversion on the intermediate result of the multiply-accumulate calculation to obtain the intermediate result of the multiply-accumulate calculation represented by a digital signal; use a shift-accumulation circuit to perform a shift-accumulation operation based on the intermediate result of the multiply-accumulate calculation represented by the digital signal to obtain the multiply-accumulate calculation result of the integer data.

[0062] According to an embodiment of the present invention, the digital circuit module is further configured to align each digit value of the two floating-point data participating in the calculation; the floating-point calculation unit is configured to perform calculations between the two aligned floating-point data. That is, when performing calculations on floating-point data, the near-memory peripheral circuit die is configured to: obtain the floating-point data participating in the calculation, and use the digital circuit module to align each digit value of the two floating-point data; input the two aligned floating-point data into the floating-point calculation unit, and the floating-point calculation unit completes the calculation of the floating-point data to obtain the calculation result of the floating-point number. The calculations of floating-point data include operations such as addition, subtraction, multiplication, and division, and the alignment method includes calculating the difference between the exponents of the two floating-point data participating in the calculation and aligning the two floating-point data according to the difference.

[0063] According to an embodiment of the present invention, the digital circuit module is further configured to perform the function of an activation function, perform the function of a trigonometric function, and / or perform the function of a pooling operation. The functions of the digital circuit module can be user-defined. For example, if specific functions such as performing the activation function or the trigonometric function are required, a non-linear processing unit can be integrated into the digital circuit module. If the function of performing the pooling operation is required, a pooling unit can be integrated. Structures such as the non-linear processing unit and the pooling unit are defined and provided by the user. It should be understood that this is only for illustration, and other computing units related to neural networks can also be integrated, such as the weight calculation of backpropagation. The present invention is not limited thereto.

[0064] The technical solution of the above embodiment of the digital circuit module can at least achieve the following beneficial technical effects: In existing research, a large number of multi-functional modules are integrated inside a single accelerator. For example, if a single chip supports functions 1, 2, and 3, then the function modules of functions 1, 2, and 3 need to be integrated inside it. However, in actual use, it will only be configured to use one function. For example, if it is configured to function 1, then the function modules that support functions 2 and 3 will not be used, resulting in waste. Therefore, compared with the original multi-functional in-memory computing accelerator, the present invention integrates more complex multi-functional modules in the near-memory peripheral circuit die, and the function can be extended by only modifying the digital circuit module of the near-memory peripheral circuit die, which is beneficial to the scalability of the function. Moreover, multiple in-memory computing dies can share the resources of the near-memory peripheral circuit die. When the in-memory computing die does not need to use this function, it can be utilized by other in-memory computing dies, which not only reduces the area of the peripheral circuit but also increases the utilization rate of the peripheral circuit, avoiding resource waste.

[0065] According to an embodiment of the present invention, the working principle of the accelerator of the present invention will be fully described below:

[0066] When performing integer data calculation, the data transmitted through the DAC in the near-memory peripheral circuit die reaches the in-memory computing die. After receiving the data, the in-memory computing die drives the input through the data input and inputs it to the crossbar array. The data received on each storage unit of the crossbar array is calculated in situ with the data stored therein, and the in-situ calculation result represented by an analog signal is obtained.

[0067] Among them, for integer logical or four arithmetic calculations, the in-situ calculation result (such as the result obtained by logical NOT, multiplication result, or division result) is read out from the crossbar array through a sense amplifier and stored in the input / output register of the in-memory computing die, so as to be written into the register file of the near-memory peripheral circuit die.

[0068] For integer multiply-accumulate calculations, the in-situ calculation results in the crossbar array (i.e., the intermediate results of the multiply-accumulate calculations obtained from the multiply-accumulate calculations) are transmitted to the near-memory peripheral circuit die through a sample-and-hold circuit. The near-memory peripheral circuit die uses the ADC in the latter to convert the intermediate results of the multiply-accumulate calculations into digital-signal-represented intermediate results of the multiply-accumulate calculations, and then uses a shift-accumulation circuit to perform shift-accumulation operations to obtain the final multiply-accumulate calculation results. Subsequently, the multiply-accumulate calculation results are stored in the register file.

[0069] When performing calculations on floating-point data, the in-memory computing die reads out the floating-point data stored in the crossbar array through a sense amplifier and places it in the input / output register for transmission to the near-memory peripheral circuit die. The near-memory peripheral circuit die completes the calculations based on the transmitted floating-point data using digital circuit modules and floating-point calculation units and writes the floating-point calculation results into the register file.

[0070] If the in-situ calculation results, multiply-accumulate calculation results, or floating-point calculation results stored in the register file are required for neural network inference, further calculations can continue to be completed using the non-linear processing units and / or pooling units pre-integrated in the digital circuit modules.

[0071] Generally speaking, the technical solutions of the embodiments of the present invention can at least achieve the following beneficial technical effects: The present invention separates multifunctional modules (such as floating-point data calculations, activation function calculations, pooling operations, etc.) and designs them as independent near-memory peripheral circuit dies. By means of the interaction between the in-memory computing die and the near-memory peripheral circuit die, each in-memory computing die no longer needs to integrate multifunctional modules internally, but multiple in-memory computing dies share the multifunctional modules in the near-memory peripheral circuit die, avoiding resource duplication and waste. Each in-memory computing die performs in-memory calculations internally and conducts near-memory calculations through communication with the near-memory peripheral circuit die, that is, integrating the basic integer calculation and storage functions in the in-memory computing die and performing more complex floating-point calculations in the near-memory peripheral circuit die. Thus, overall, it realizes the support for multifunctions such as efficient writing of in-situ calculation results, integer data calculations, and near-memory floating-point data calculations.

[0072] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.

[0073] The present invention can be a system, method, and / or computer program product. The computer program product can include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present invention.

[0074] A computer-readable storage medium can be a tangible device that retains and stores instructions for use by an instruction execution device. A computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing.

[0075] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the technical field to understand the embodiments disclosed herein.

Claims

1. A multi-core-based multi-functional storage and computing accelerator, characterized in that: The accelerator includes a plurality of in-memory computing cores and a plurality of near-memory peripheral circuit cores, at least some of the in-memory computing cores share resources of one or more near-memory peripheral circuit cores, wherein: The accelerator is configured to: When performing logical calculations or arithmetic calculations on integer data, the resources of the in-memory computing core are scheduled to perform in-situ calculations to obtain in-situ calculation results. When performing calculations on floating-point data, the resources of the in-memory computing core particles and the near-memory peripheral circuit core particles are scheduled to perform the calculations together to obtain the floating-point calculation results.

2. The accelerator according to claim 1, characterized in that The in-memory computing core comprises a plurality of crossbar switch arrays, each crossbar switch array comprises a plurality of storage units, wherein the plurality of crossbar switch arrays are configured as follows: When performing a logic calculation or four arithmetic calculation of integer data, the integer data involved in the calculation is read and stored through a storage unit, and the logic calculation or four arithmetic calculation of the integer data is performed in situ to obtain an in-situ calculation result; When performing calculation of floating-point data, the floating-point data involved in the calculation is read and stored through the storage unit; The in-memory computing core represents the floating-point data and / or the in-situ computing result as a digital signal, and transmits it to the near-memory peripheral circuit core that performs the computing together with it to complete the computing.

3. The accelerator according to claim 2, characterized in that When performing multiplication and addition calculations on integer data, the accelerator is further configured to: The crossbar switch array of the in-memory computing core particle is scheduled to perform multiplication and addition calculations, obtain the intermediate results of the multiplication and addition calculations, and transmit the intermediate results of the multiplication and addition calculations to the near-memory peripheral circuit core particles that perform the calculations together with the core particles; The resources of the near-memory peripheral circuit core particles that perform the calculation together are scheduled, and the shift accumulation calculation is performed according to the intermediate results of the multiplication and addition calculation to obtain the multiplication and addition calculation results of the integer data.

4. The accelerator according to claim 2, characterized in that: The in-memory computing core also includes input and output registers, data input drivers and sense amplifiers, wherein: A data input driver, used for inputting integer data and / or floating point data involved in calculation into a plurality of crossbar switch arrays; A sensitive amplifier is used to assist the crossbar switch array in performing logic calculations or four arithmetic calculations of integer data, and to read floating-point data represented by digital signals and / or in-situ calculation results represented by digital signals from the crossbar switch array; The input and output registers are used to temporarily store integer data and / or floating-point data involved in the calculation, and to temporarily store in-situ calculation results represented by digital signals and floating-point data represented by digital signals read by the sense amplifier.

5. The accelerator according to claim 2, characterized in that: The in-memory computing core also includes a sample-and-hold circuit, wherein: When performing multiplication and addition calculations on integer data, the sampling and holding circuit is used to sample and hold the intermediate results of the multiplication and addition calculations represented by analog signals obtained by performing multiplication and addition calculations on multiple crossbar switch arrays, so as to transmit them to the near-memory peripheral circuit core particles for the joint calculations.

6. The accelerator according to claim 2, characterized in that: The plurality of crossbar switch arrays are further configured to: The floating point data is stored by storing the exponent and the mantissa of the floating point data, wherein the exponent and the mantissa of the floating point data are respectively stored in different crossbar switch arrays.

7. The accelerator according to claim 2, characterized in that: The near storage peripheral circuit core particle includes: A digital-to-analog converter, used to convert integer data involved in the calculation into integer data represented by analog signals for transmission to the in-memory computing core particles; Register files, used to store data, including storing in-situ computation results; A digital circuit module, used to align each bit value of two floating point data involved in the calculation; The floating-point calculation unit is used to perform calculations between two aligned floating-point data.

8. The accelerator according to claim 5, characterized in that The near-memory peripheral circuit core comprises an analog-to-digital converter and a shift-accumulate circuit, wherein when performing multiplication and addition calculations of integer data, the near-memory peripheral circuit core is configured as follows: Performing analog-to-digital conversion on the intermediate result of the multiplication and addition calculation using an analog-to-digital converter to obtain the intermediate result of the multiplication and addition calculation represented by a digital signal; The shift-accumulate circuit is used to perform shift-accumulate operations based on the intermediate results of multiplication-addition calculations to obtain the multiplication-addition calculation results of the integer data.

9. The accelerator according to claim 7, characterized in that: When performing calculations on floating-point data, the near memory peripheral chip is configured as follows: Obtain floating-point data involved in the calculation, and use a digital circuit module to align the values ​​of each bit of two floating-point data; The two aligned floating-point data are input into the floating-point calculation unit to obtain the calculation result of the floating-point number.

10. The accelerator according to claim 7, characterized in that: The digital circuit module is also used to execute the function of an activation function, the function of a trigonometric function and / or the function of a pooling operation.