Near storage computing system based on heterogeneous storage and control method thereof
By adopting heterogeneous memory and multiplexed design in the near-memory computing system, the problems of low efficiency and high power consumption of the von Neumann architecture processor system are solved, and a near-memory computing system with high efficiency and low power consumption are realized.
Patent Information
- Application Number
- CN202510047350.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-13
AI Technical Summary
In the prior art, the processor system of the von Neumann architecture is inefficient and has high power consumption when processing data, making it difficult to meet the rapidly growing computing power demand. In addition, near-access computing systems have problems such as single computing types and difficult to generalize memory structures, and the prior art is mostly circuit-level and difficult to generalize.
A near-access computing system based on heterogeneous storage is proposed, which uses static random access memory (SRAM) and magnetic random access memory (MRAM) to form heterogeneous memory. Combining the computing cluster module and the control unit, the combination of multiplexing design and heterogeneous storage can reduce data transfer, improve computing efficiency and reduce power consumption.
Through heterogeneous storage and multiplexing design, computing efficiency is significantly improved, power consumption is reduced, and circuit resources is saved, solving the problems of low efficiency and high power consumption of traditional processor systems, and at the same time achieving the universality of near-storage computing systems.
Smart Images

Figure CN120029966A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer architecture, and in particular to a near-memory computing system based on heterogeneous storage and a control method thereof. Background Art
[0002] With the rapid development of information processing technology and artificial intelligence technology, the demand for computing power is also growing rapidly. When the von Neumann architecture calculates data, it needs to continuously move the data in the memory to the arithmetic logic unit, and then move the results back to the memory. This computing paradigm is not only inefficient, but also accompanied by huge power consumption requirements. Therefore, the processor system based on the von Neumann architecture cannot meet the growing computing power requirements. Near-memory computing is a new computing paradigm in recent years. It places the memory and computing unit together at the chip layout and architecture level, thereby greatly reducing the movement of data, reducing power consumption and greatly improving computing efficiency. However, near-memory computing is also facing the problems of single computing type and difficult universal memory structure at this stage. At this stage, near-memory computing technology is mostly at the circuit level, and secondary optimization is required between different process nodes, making it difficult to universalize. At the same time, FFT (fast Fourier transform) and CNN (convolutional neural network) are the core computing in the fields of signal processing and artificial intelligence, respectively, and are widely used in scenarios such as radar signal processing, target recognition and text analysis. However, traditional scholars usually design and optimize the two types of calculations separately, and their dedicated FFT or CNN accelerators usually occupy a large amount of circuit resources in the chip, thereby increasing power consumption.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the invention
[0004] The main purpose of the embodiments of the present application is to propose a near-memory computing system based on heterogeneous storage and a control method thereof, which can avoid unnecessary data movement operations, thereby improving the computing efficiency of the system and reducing power consumption, thereby saving circuit resources.
[0005] To achieve the above-mentioned purpose, an embodiment of the present application provides a near-memory computing system based on heterogeneous storage on one hand, the system includes a static random access memory, a magnetic random access memory, a control unit and a computing cluster module, the static random access memory and the magnetic random access memory constitute a heterogeneous memory, the computing cluster module includes a plurality of processing units, and the plurality of processing units adopt a circuit multiplexing design, the output end of the static random access memory is connected to the first input end of the computing cluster module, the first output end of the computing cluster module is connected to the input end of the static random access memory, the second output end of the computing cluster module is connected to the first input end of the control unit, the first output end of the control unit is connected to the second input end of the computing cluster module, the second output end of the control unit is connected to the input end of the magnetic random access memory, and the output end of the magnetic random access memory is connected to the second input end of the control unit, wherein:
[0006] The static random access memory is used to store input data of the computing cluster module;
[0007] The magnetic random access memory is used to store weight data of the computing cluster module;
[0008] The control unit is used to control data transmission between the magnetic random access memory and the computing cluster module;
[0009] The computing cluster module is used to perform three-level operations according to the input data and the weight data to obtain a near-memory computing result, thereby realizing the acceleration function of the accelerator.
[0010] In some embodiments, the equivalent circuit model of the computing cluster module includes a first-level multiplication operation unit, a second-level addition operation unit, a third-level butterfly operation unit, a first multiplexer, and a second multiplexer, wherein:
[0011] The first-level multiplication unit is used to perform multiplication operations;
[0012] The second-stage addition unit is used to perform addition operations;
[0013] The third-level butterfly operation unit is used to perform butterfly operation;
[0014] The first multiplexer and the second multiplexer are used to select corresponding data according to different computing modes.
[0015] In some embodiments, the first-stage multiplication operation unit includes a first multiplier, a second multiplier, a third multiplier and a fourth multiplier, the second-stage addition operation unit includes a first adder and a second adder, and the third-stage butterfly operation unit includes a third adder, a fourth adder, a fifth adder and a sixth adder.
[0016] In some embodiments, the output end of the first multiplier and the output end of the second multiplier are both connected to the input end of the first adder, the output end of the third multiplier and the output end of the fourth multiplier are both connected to the input end of the second adder, the output end of the first adder is respectively connected to the first input end of the third adder and the first input end of the fourth adder, the output end of the first multiplexer is respectively connected to the second input end of the third adder and the second input end of the fourth adder, the output end of the fourth adder is connected to the input end of the first multiplexer, the output end of the second adder is respectively connected to the first input end of the fifth adder and the first input end of the sixth adder, the output end of the first multiplexer is respectively connected to the second input end of the fifth adder and the second input end of the sixth adder, and the output end of the fifth adder is connected to the input end of the second multiplexer.
[0017] To achieve the above object, another aspect of an embodiment of the present application provides a control method for a near-memory computing system based on heterogeneous storage, the control method comprising the following steps:
[0018] Determine the working mode of the near-memory computing system and obtain input data and weight data from heterogeneous memory;
[0019] transmitting the input data and the weight data to a computing cluster according to the working mode of the near-storage computing system;
[0020] Based on the computing cluster, a near-memory operation is performed on the input data and the weight data to obtain a near-memory computing result.
[0021] In some embodiments, the operating modes of the near-memory computing system include a fast Fourier transform operating mode and a convolutional neural network operating mode.
[0022] In some embodiments, performing near-memory operations on the input data and the weight data based on the computing cluster to obtain near-memory computing results includes:
[0023] Based on the computing cluster, performing a first-level multiplication operation on the input data and the weight data to obtain a product operation result;
[0024] Performing a second-level addition operation on the product operation result to obtain an addition operation result;
[0025] A third-level butterfly operation is performed on the addition and operation results to obtain the near-memory calculation result.
[0026] In some embodiments, the first stage multiplication operation comprises:
[0027] If the working mode of the near-memory computing system is a fast Fourier transform working mode, the input data is represented by first sampled signal data and second sampled signal data, and the weight data is represented by a rotation factor;
[0028] Performing multiplication calculation on the real part of the second sampling signal data and the real part of the rotation factor by a first multiplier to obtain a first multiplication calculation result;
[0029] Multiplying the imaginary part of the second sampling signal data and the imaginary part of the rotation factor by a second multiplier to obtain a second multiplication result;
[0030] Multiplying the imaginary part of the second sampling signal data and the real part of the rotation factor by a third multiplier to obtain a third multiplication result;
[0031] Performing multiplication calculation on the real part of the second sampling signal data and the imaginary part of the rotation factor by a fourth multiplier to obtain a fourth multiplication calculation result;
[0032] Combining the first multiplication calculation result, the second multiplication calculation result, the third multiplication calculation result and the fourth multiplication calculation result to obtain a product operation result of the fast Fourier transform working mode;
[0033] If the working mode of the near-memory computing system is a convolutional neural network working mode, the input data is represented by a first operand, a second operand, a third operand, and a fourth operand, and the weight data is represented by a fifth operand, a sixth operand, a seventh operand, and an eighth operand;
[0034] multiplying the first operand and the fifth operand by a first multiplier to obtain a fifth multiplication result;
[0035] multiplying the second operand and the sixth operand by a second multiplier to obtain a sixth multiplication result;
[0036] multiplying the third operand and the seventh operand by a third multiplier to obtain a seventh multiplication result;
[0037] multiplying the fourth operand and the eighth operand by a fourth multiplier to obtain an eighth multiplication result;
[0038] The fifth multiplication calculation result, the sixth multiplication calculation result, the seventh multiplication calculation result and the eighth multiplication calculation result are combined to obtain the product operation result of the convolutional neural network working mode.
[0039] In some embodiments, the second stage addition operation comprises:
[0040] If the working mode of the near-memory computing system is a fast Fourier transform working mode;
[0041] Subtracting the first multiplication calculation result from the second multiplication calculation result by a first adder to obtain a first addition operation result;
[0042] Adding the third multiplication calculation result and the fourth multiplication calculation result by a second adder to obtain a second addition operation result;
[0043] Combining the first addition and operation results with the second addition and operation results to obtain an addition and operation result of the fast Fourier transform working mode;
[0044] If the working mode of the near-memory computing system is a convolutional neural network working mode;
[0045] Adding the fifth multiplication result and the sixth multiplication result by a first adder to obtain a third addition result;
[0046] Adding the seventh multiplication result and the eighth multiplication result by a second adder to obtain a fourth addition result;
[0047] Combining the third addition and operation result with the fourth addition and operation result, the addition and operation result of the convolutional neural network working mode is obtained.
[0048] In some embodiments, the third stage butterfly operation includes:
[0049] If the working mode of the near-memory computing system is a fast Fourier transform working mode, the first sampling signal data is obtained through a first multiplexer and a second multiplexer;
[0050] Subtracting the first addition operation result from the real part of the first sampling signal data through a third adder to obtain the real part of the subtraction operation result;
[0051] Adding the first addition operation result and the real part of the first sampling signal data by a fourth adder to obtain the real part of the addition operation result;
[0052] Adding the second addition operation result and the imaginary part of the first sampling signal data by a fifth adder to obtain the imaginary part of the subtraction operation result;
[0053] subtracting the second addition operation result from the imaginary part of the first sampling signal data through a sixth adder to obtain the imaginary part of the addition operation result;
[0054] Combining the real part of the subtraction operation result, the imaginary part of the subtraction operation result, the real part of the addition operation result and the imaginary part of the addition operation result to obtain the calculation result of the accelerator in the fast Fourier transform working mode;
[0055] If the working mode of the near-memory computing system is a convolutional neural network working mode, the output data of the fourth adder is selected through the first multiplexer, and the output data of the fifth adder is selected through the second multiplexer;
[0056] Accumulating the first addition operation result and the output data of the fourth adder through a fourth adder to obtain a first accumulation operation result;
[0057] Accumulating the second addition operation result and the output data of the fifth adder through a fifth adder to obtain a second accumulation operation result;
[0058] Combine the first cumulative operation result with the second cumulative operation result to obtain the calculation result of the accelerator in the convolutional neural network working mode.
[0059] The embodiments of the present application include at least the following beneficial effects: The present application provides a near-memory computing system based on heterogeneous storage and a control method thereof. The scheme forms a heterogeneous memory through a static random access memory and a magnetic random access memory. Through heterogeneous storage, not only can SRAM be used to meet the demand for fast update of input data in FFT / CNN calculations, but MRAM can also be used to significantly reduce the power consumption and area occupancy of the circuit, obtain input data and weight data and determine the working mode of the near-memory computing system, map the input data and weight data to the computing cluster, and the processing units in the computing cluster are connected through multiplexing. The data in the SRAM and MRAM are directly transferred from the memory to the PE for calculation, without the need to write to the register, eliminating a large number of data moving operations, thereby improving computing efficiency and reducing power consumption, thereby saving circuit resources, and finally performing accumulation operations on the input data and weight data, thereby improving the operating efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a structural diagram of a near-memory computing system based on heterogeneous storage provided in an embodiment of the present application;
[0061] Figure 2 It is a schematic diagram of the steps of a control method of a near-memory computing system based on heterogeneous storage provided by an embodiment of the present application;
[0062] Figure 3 It is a schematic diagram of the circuit model structure of the PE in the computing cluster module provided in the embodiment of the present application;
[0063] Figure 4 It is a schematic diagram of a model of three-level operation in the fast Fourier transform working mode provided in an embodiment of the present application;
[0064] Figure 5 It is a schematic diagram of a model of three-level operations in a convolutional neural network working mode provided in an embodiment of the present application;
[0065] Figure 6 It is a schematic diagram of a pipeline of multiplexed PEs provided in an embodiment of the present application;
[0066] Figure 7 It is a schematic diagram of the step framework of the three-level operation of the processing unit provided in the embodiment of the present application. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of systems and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0068] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0069] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0071] Reference Figure 1 , Figure 1 A schematic diagram of a near memory computing system based on heterogeneous storage provided by an embodiment of the present invention, referring to Figure 1 The system includes a static random access memory, a magnetic random access memory, a control unit and a computing cluster module, wherein the static random access memory and the magnetic random access memory constitute a heterogeneous memory, the computing cluster module includes a plurality of processing units, the plurality of processing units are connected by multiplexing, the output end of the static random access memory is connected to the first input end of the computing cluster module, the first output end of the computing cluster module is connected to the input end of the static random access memory, the second output end of the computing cluster module is connected to the first input end of the control unit, the first output end of the control unit is connected to the second input end of the computing cluster module, the second output end of the control unit is connected to the input end of the magnetic random access memory, and the output end of the magnetic random access memory is connected to the second input end of the control unit, wherein:
[0072] The static random access memory is used to store input data of the computing cluster module;
[0073] The magnetic random access memory is used to store the weight data of the computing cluster module;
[0074] The control unit is used for controlling data transmission between the magnetic random access memory and the computing cluster module;
[0075] The computing cluster module is used to perform three-level operations based on input data and weight data to obtain near-memory computing results and realize the acceleration function of the accelerator.
[0076] In some specific embodiments, it should be noted that in order to improve the computing efficiency, a near-memory computing system based on heterogeneous storage is first designed, such as Figure 1As shown in the figure. In this system, static random access memory (SRAM) and magnetic random access memory (MRAM) are used to form heterogeneous memory. SRAM has higher read and write flexibility and faster read and write speed, while MRAM has the characteristics of non-volatility, high storage density and low static power consumption. At the same time, the input data in FFT / CNN calculation needs to be updated continuously, and the weight is not updated or the update frequency is low. Therefore, SRAM is used to store the input in FFT / CNN calculation, and MRAM is used to store the weight in FFT / CNN calculation. Through heterogeneous storage, not only can SRAM be used to meet the fast update requirements of input data in FFT / CNN calculation, but MRAM can also be used to significantly reduce the power consumption and area occupied by the circuit. The near-memory computing system includes a total of 4 computing clusters, each cluster includes 32 processing units (PE). The data in SRAM and MRAM are directly transferred from the memory to PE for calculation, without writing to the register, eliminating a large number of data movement operations, thereby improving computing efficiency and reducing power consumption. Each computing cluster has a private data SRAM corresponding to it, while the MRAM is shared between clusters.
[0077] Furthermore, it should be noted that the equivalent circuit model of the computing cluster module includes a first-level multiplication operation unit, a second-level addition operation unit, a third-level butterfly operation unit, a first multiplexer and a second multiplexer, wherein the first-level multiplication operation unit is used to perform multiplication operations; the second-level addition operation unit is used to perform addition operations; the third-level butterfly operation unit is used to perform butterfly operations; the first multiplexer and the second multiplexer are used to select corresponding data according to different computing modes.
[0078] Furthermore, it should be noted that the first-stage multiplication operation unit includes a first multiplier, a second multiplier, a third multiplier and a fourth multiplier, the second-stage addition operation unit includes a first adder and a second adder, and the third-stage butterfly operation unit includes a third adder, a fourth adder, a fifth adder and a sixth adder, wherein the output end of the first multiplier and the output end of the second multiplier are both connected to the input end of the first adder, the output end of the third multiplier and the output end of the fourth multiplier are both connected to the input end of the second adder, and the output ends of the first adder and the third adder are respectively connected to the input end of the third multiplier. The first input end of the first adder is connected to the first input end of the fourth adder, the output end of the first multiplexer is connected to the second input end of the third adder and the second input end of the fourth adder respectively, the output end of the fourth adder is connected to the input end of the first multiplexer, the output end of the second adder is connected to the first input end of the fifth adder and the first input end of the sixth adder respectively, the output end of the first multiplexer is connected to the second input end of the fifth adder and the second input end of the sixth adder respectively, and the output end of the fifth adder is connected to the input end of the second multiplexer.
[0079] Among them, the near-memory computing system includes the design of PE, that is, the design of FFT / CNN multiplexing PE. FFT and CNN have great commonality in two kinds of calculations, and most of their circuit structures can be reused through design, thus saving circuit resources. The designed FFT / CNN multiplexing processing unit is as follows Figure 3 As shown, the circuit is divided into three levels of operations and two multiplexers (Multiplexer, MUX). The three levels of operations are multiplication, addition and butterfly operations. In the first level of multiplication, a total of 4 signed multipliers are included. Each multiplier can perform multiplication operations on two multipliers. The two multipliers are input data and weight data. The input comes from SRAM and the weight comes from MRAM. Through the first level of multiplication, the 4 multipliers output 4 products and transmit them to the second level of operation. In the second level of addition, a total of 2 adders are included, and the operation results of the first level are added in pairs. Through the second level of addition, the two adders output two addition sums and transmit them to the third level of butterfly operation. In the third level of butterfly operation, a total of 4 adders are included, and the operation results of the second level and the output of MUX are added for the second time. Depending on the operation mode, MUX selects different data. In FFT mode, MUX selects the operand of external input. In CNN mode, MUX selects the result of butterfly operation. With the support of pipeline, accumulation operation can be realized. Through the third-level butterfly operation, the calculation result of FFT / CNN is finally output. In this design, negative numbers are represented by complement code, so that the subtraction operation is completed by adder.
[0080] See also Figure 2 and Figure 7The embodiment of the present application also provides a control method for a near-memory computing system based on heterogeneous storage, which can implement the above-mentioned near-memory computing system based on heterogeneous storage, and the system includes:
[0081] S100, determining the working mode of the near-memory computing system and obtaining input data and weight data from the heterogeneous memory;
[0082] It should be noted that, in some embodiments, the working modes of the near-memory computing system include a fast Fourier transform working mode and a convolutional neural network working mode.
[0083] In some specific embodiments, the instructions in the instruction SRAM are read out, and the input data and weight data are taken out from the data SRAM and MRAM respectively according to the instructions, and sent to the PE of the computing cluster to further determine the working mode of the near-memory computing system, which is divided into FFT mode and CNN mode.
[0084] S200, transmitting input data and weight data to the computing cluster according to the working mode of the near-storage computing system;
[0085] It should be noted that, in some embodiments, data transmitted to the PE is mapped according to different working modes.
[0086] S300 , based on the computing cluster, performing near-memory operations on input data and weight data to obtain near-memory computing results.
[0087] It should be noted that, in some embodiments, step S300 may include steps S310 to S330;
[0088] S310, based on the calculation cluster, performing a first-level multiplication operation on the input data and the weight data to obtain a product operation result;
[0089] Specifically, Figure 4 As shown, if the working mode of the near-memory computing system is the fast Fourier transform working mode, the input data is represented by the first sampling signal data and the second sampling signal data, and the weight data is represented by the rotation factor; the real part of the second sampling signal data and the real part of the rotation factor are multiplied by the first multiplier to obtain a first multiplication calculation result; the imaginary part of the second sampling signal data and the imaginary part of the rotation factor are multiplied by the second multiplier to obtain a second multiplication calculation result; the imaginary part of the second sampling signal data and the real part of the rotation factor are multiplied by the third multiplier to obtain a third multiplication calculation result; the real part of the second sampling signal data and the imaginary part of the rotation factor are multiplied by the fourth multiplier to obtain a fourth multiplication calculation result; the first multiplication calculation result, the second multiplication calculation result, the third multiplication calculation result and the fourth multiplication calculation result are combined to obtain the product operation result of the fast Fourier transform working mode.
[0090] In this embodiment, in FFT mode, the input data is a sampled signal, which is divided into a real part and an imaginary part; the weight data is a rotation factor, which is also divided into a real part and an imaginary part. Figure 4 In , X1 and X2 are sampled data, wherein X1 is the first sampled signal data, X2 is the second sampled signal data, and W is the rotation factor. In the multiplication operation, the real and imaginary parts of X2 and the real and imaginary parts of W are multiplied in pairs by multipliers to obtain 4 products.
[0091] like Figure 5 As shown, if the working mode of the near-memory computing system is the convolutional neural network working mode, the input data is represented by the first operand, the second operand, the third operand and the fourth operand, and the weight data is represented by the fifth operand, the sixth operand, the seventh operand and the eighth operand; the first operand and the fifth operand are multiplied by the first multiplier to obtain the fifth multiplication result; the second operand and the sixth operand are multiplied by the second multiplier to obtain the sixth multiplication result; the third operand and the seventh operand are multiplied by the third multiplier to obtain the seventh multiplication result; the fourth operand and the eighth operand are multiplied by the fourth multiplier to obtain the eighth multiplication result; the fifth multiplication result, the sixth multiplication result, the seventh multiplication result and the eighth multiplication result are combined to obtain the multiplication result of the convolutional neural network working mode.
[0092] In this embodiment, in the CNN mode, the input data is four operands of A1-A4, and the weight data is four operands of W1-W4, where A1 represents the first operand, A2 represents the second operand, A3 represents the third operand, A4 represents the fourth operand, W1 represents the fifth operand, W2 represents the sixth operand, W3 represents the seventh operand, and W4 represents the eighth operand. In the multiplication operation, A1 and W1, A2 and W2, A3 and W3, A4 and W4 are multiplied in pairs to obtain 4 products.
[0093] S320, performing a second-level addition operation on the product operation result to obtain an addition and operation result;
[0094] Specifically, Figure 4 As shown, if the working mode of the near-memory computing system is the fast Fourier transform working mode; the first multiplication calculation result is subtracted from the second multiplication calculation result through the first adder to obtain a first addition and operation result; the third multiplication calculation result is added to the fourth multiplication calculation result through the second adder to obtain a second addition and operation result; the first addition and operation result is combined with the second addition and operation result to obtain the addition and operation result of the fast Fourier transform working mode.
[0095] In this embodiment, in FFT mode, the first adder is used as a subtractor to subtract the second product from the first product; and the second adder is used to add the third and fourth products.
[0096] like Figure 5 As shown, if the working mode of the near-memory computing system is the convolutional neural network working mode; the fifth multiplication calculation result and the sixth multiplication calculation result are added by the first adder to obtain a third addition and operation result; the seventh multiplication calculation result and the eighth multiplication calculation result are added by the second adder to obtain a fourth addition and operation result; the third addition and operation result and the fourth addition and operation result are combined to obtain the addition and operation result of the convolutional neural network working mode.
[0097] In this embodiment, in CNN mode, the four products are added in pairs.
[0098] S330, performing a third-level butterfly operation on the addition and operation results to obtain a near-memory calculation result.
[0099] Specifically, Figure 4 As shown, if the working mode of the near-memory computing system is the fast Fourier transform working mode, the first sampling signal data is obtained through the first multiplexer and the second multiplexer; the first addition and operation result is subtracted from the real part of the first sampling signal data through the third adder to obtain the real part of the subtraction operation result; the first addition and operation result is added to the real part of the first sampling signal data through the fourth adder to obtain the real part of the addition operation result; the second addition and operation result is added to the imaginary part of the first sampling signal data through the fifth adder to obtain the imaginary part of the subtraction operation result; the second addition and operation result is subtracted from the imaginary part of the first sampling signal data through the sixth adder to obtain the imaginary part of the addition operation result; the real part of the subtraction operation result, the imaginary part of the subtraction operation result, the real part of the addition operation result and the imaginary part of the addition operation result are combined to obtain the calculation result of the accelerator in the fast Fourier transform working mode.
[0100] In this embodiment, in FFT mode, MUX is used to select X1, and it is sent to the butterfly operation together with the result of the addition operation, and the result of the addition operation is cross-added (subtracted) with X1 to obtain the final results Y1 and Y2, where Y1 represents the real part of the addition operation result and the imaginary part of the addition operation result, and Y2 represents the real part of the subtraction operation result and the imaginary part of the subtraction operation result.
[0101] like Figure 5As shown, if the working mode of the near-memory computing system is the convolutional neural network working mode, the output data of the fourth adder is selected through the first multiplexer, and the output data of the fifth adder is selected through the second multiplexer; the first addition and operation result and the output data of the fourth adder are accumulated through the fourth adder to obtain a first accumulation result; the second addition and operation result and the output data of the fifth adder are accumulated through the fifth adder to obtain a second accumulation result; the first accumulation result and the second accumulation result are combined to obtain the calculation result of the accelerator in the convolutional neural network working mode.
[0102] In this embodiment, in the CNN mode, the MUX no longer selects the external input, but selects the output result of the butterfly operation back and retransmits it to the input of the butterfly operation. The current butterfly operation result is continuously added to the result of the next level addition operation, such as Figure 6 As shown, with the support of the pipeline, the accumulation function is finally realized, and the accumulation results B1 and B2 are obtained, where B1 represents the first accumulation operation result and B2 represents the second accumulation operation result.
[0103] In summary, the computing system constructed by the embodiment of the present invention can accelerate FFT and CNN calculations and be applied to the field of radar signal processing. In the clutter suppression and coherent accumulation of radar signal processing, a large number of FFT calculations are required. After obtaining the range-Doppler map, CNN can be used for target recognition and tracking. Therefore, the computing system constructed by the present invention can accelerate the core calculations in radar signal processing and improve the operating efficiency of the system.
[0104] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0105] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A near-memory computing system based on heterogeneous storage, characterized in that: The system includes a static random access memory, a magnetic random access memory, a control unit and a computing cluster module, wherein the static random access memory and the magnetic random access memory constitute a heterogeneous memory, the computing cluster module includes a plurality of processing units, and the plurality of processing units adopt a circuit multiplexing design, the output end of the static random access memory is connected to the first input end of the computing cluster module, the first output end of the computing cluster module is connected to the input end of the static random access memory, the second output end of the computing cluster module is connected to the first input end of the control unit, the first output end of the control unit is connected to the second input end of the computing cluster module, the second output end of the control unit is connected to the input end of the magnetic random access memory, and the output end of the magnetic random access memory is connected to the second input end of the control unit, wherein: The static random access memory is used to store input data of the computing cluster module; The magnetic random access memory is used to store weight data of the computing cluster module; The control unit is used to control data transmission between the magnetic random access memory and the computing cluster module; The computing cluster module is used to perform three-level operations according to the input data and the weight data to obtain a near-memory computing result, thereby realizing the acceleration function of the accelerator.
2. The system according to claim 1, characterized in that The equivalent circuit model of the computing cluster module includes a first-level multiplication operation unit, a second-level addition operation unit, a third-level butterfly operation unit, a first multiplexer and a second multiplexer, wherein: The first-level multiplication unit is used to perform multiplication operations; The second-stage addition unit is used to perform addition operations; The third-level butterfly operation unit is used to perform butterfly operation; The first multiplexer and the second multiplexer are used to select corresponding data according to different computing modes.
3. The system according to claim 2, characterized in that The first-stage multiplication operation unit includes a first multiplier, a second multiplier, a third multiplier and a fourth multiplier, the second-stage addition operation unit includes a first adder and a second adder, and the third-stage butterfly operation unit includes a third adder, a fourth adder, a fifth adder and a sixth adder.
4. The system according to claim 3, characterized in that The output end of the first multiplier and the output end of the second multiplier are both connected to the input end of the first adder, the output end of the third multiplier and the output end of the fourth multiplier are both connected to the input end of the second adder, the output end of the first adder is respectively connected to the first input end of the third adder and the first input end of the fourth adder, the output end of the first multiplexer is respectively connected to the second input end of the third adder and the second input end of the fourth adder, the output end of the fourth adder is connected to the input end of the first multiplexer, the output end of the second adder is respectively connected to the first input end of the fifth adder and the first input end of the sixth adder, the output end of the first multiplexer is respectively connected to the second input end of the fifth adder and the second input end of the sixth adder, and the output end of the fifth adder is connected to the input end of the second multiplexer.
5. A control method for a near-memory computing system based on heterogeneous storage, characterized in that: The control method comprises the following steps: Determine the working mode of the near-memory computing system and obtain input data and weight data from heterogeneous memory; transmitting the input data and the weight data to a computing cluster according to the working mode of the near-storage computing system; Based on the computing cluster, a near-memory operation is performed on the input data and the weight data to obtain a near-memory computing result.
6. The method according to claim 5, characterized in that The working modes of the near-memory computing system include a fast Fourier transform working mode and a convolutional neural network working mode.
7. The method according to claim 6, characterized in that The performing near-memory operation on the input data and the weight data based on the computing cluster to obtain near-memory calculation results includes: Based on the computing cluster, performing a first-level multiplication operation on the input data and the weight data to obtain a product operation result; Performing a second-level addition operation on the product operation result to obtain an addition and operation result; A third-level butterfly operation is performed on the addition and operation results to obtain the near-memory calculation result.
8. The method according to claim 7, characterized in that The first-level multiplication operation includes: If the working mode of the near-memory computing system is a fast Fourier transform working mode, the input data is represented by first sampled signal data and second sampled signal data, and the weight data is represented by a rotation factor; Performing multiplication calculation on the real part of the second sampling signal data and the real part of the rotation factor by a first multiplier to obtain a first multiplication calculation result; Multiplying the imaginary part of the second sampling signal data and the imaginary part of the rotation factor by a second multiplier to obtain a second multiplication result; Multiplying the imaginary part of the second sampling signal data and the real part of the rotation factor by a third multiplier to obtain a third multiplication result; Multiplying the real part of the second sampling signal data and the imaginary part of the rotation factor by a fourth multiplier to obtain a fourth multiplication result; Combining the first multiplication calculation result, the second multiplication calculation result, the third multiplication calculation result and the fourth multiplication calculation result to obtain a product operation result of the fast Fourier transform working mode; If the working mode of the near-memory computing system is a convolutional neural network working mode, the input data is represented by a first operand, a second operand, a third operand, and a fourth operand, and the weight data is represented by a fifth operand, a sixth operand, a seventh operand, and an eighth operand; multiplying the first operand and the fifth operand by a first multiplier to obtain a fifth multiplication result; multiplying the second operand and the sixth operand by a second multiplier to obtain a sixth multiplication result; multiplying the third operand and the seventh operand by a third multiplier to obtain a seventh multiplication result; multiplying the fourth operand and the eighth operand by a fourth multiplier to obtain an eighth multiplication result; The fifth multiplication calculation result, the sixth multiplication calculation result, the seventh multiplication calculation result and the eighth multiplication calculation result are combined to obtain the product operation result of the convolutional neural network working mode.
9. The method according to claim 8, characterized in that The second-level addition operation includes: If the working mode of the near-memory computing system is a fast Fourier transform working mode; Subtracting the first multiplication calculation result from the second multiplication calculation result by a first adder to obtain a first addition operation result; Adding the third multiplication calculation result and the fourth multiplication calculation result by a second adder to obtain a second addition operation result; Combining the first addition and operation results with the second addition and operation results to obtain an addition and operation result of the fast Fourier transform working mode; If the working mode of the near-memory computing system is a convolutional neural network working mode; Adding the fifth multiplication result and the sixth multiplication result by a first adder to obtain a third addition result; Adding the seventh multiplication result and the eighth multiplication result by a second adder to obtain a fourth addition result; Combining the third addition and operation result with the fourth addition and operation result, the addition and operation result of the convolutional neural network working mode is obtained.
10. The method according to claim 9, characterized in that The third-level butterfly operation includes: If the working mode of the near-memory computing system is a fast Fourier transform working mode, the first sampling signal data is obtained through a first multiplexer and a second multiplexer; Subtracting the first addition operation result from the real part of the first sampling signal data through a third adder to obtain the real part of the subtraction operation result; Adding the first addition operation result and the real part of the first sampling signal data by a fourth adder to obtain the real part of the addition operation result; Adding the second addition operation result and the imaginary part of the first sampling signal data by a fifth adder to obtain the imaginary part of the subtraction operation result; subtracting the second addition operation result from the imaginary part of the first sampling signal data through a sixth adder to obtain the imaginary part of the addition operation result; Combining the real part of the subtraction operation result, the imaginary part of the subtraction operation result, the real part of the addition operation result and the imaginary part of the addition operation result to obtain the calculation result of the accelerator in the fast Fourier transform working mode; If the working mode of the near-memory computing system is a convolutional neural network working mode, the output data of the fourth adder is selected through the first multiplexer, and the output data of the fifth adder is selected through the second multiplexer; Accumulating the first addition operation result and the output data of the fourth adder through a fourth adder to obtain a first accumulation operation result; Accumulating the second addition operation result and the output data of the fifth adder through a fifth adder to obtain a second accumulation operation result; Combine the first cumulative operation result with the second cumulative operation result to obtain the calculation result of the accelerator in the convolutional neural network working mode.
Citation Information
Patent Citations
Energy efficient compute near memory binary neural network circuits
CN112862061A
SoC system with in-memory / near-memory computing module
CN114356840A
Three-dimensional convolutional neural network accelerator on complex field and method
CN116596034A
Core group memory processsing chip design
US20240272821A1
Method and system for realizing vector operations
WO2012145986A1