In-memory computing system, method and device with dynamically adjustable precision
By designing the dynamic division module of feature value bits, the configurable in-memory calculation module of dynamic division of weight bits and the multiplication and accumulation result shift addition module in the in-memory computing system, dynamic adjustment of accuracy is achieved, solving the problems of limited performance and insufficient accuracy adjustment capabilities of traditional computing systems, and improving satellite data processing capabilities and computing efficiency.
Patent Information
- Application Number
- CN202510542216.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The bottleneck of data transmission speed between the processor and memory in traditional computing systems leads to performance limitations, and in satellite applications, different tasks have different requirements for computing accuracy and lack dynamic adjustment capabilities.
Design an in-memory computing system with dynamic accuracy adjustable accuracy, including a dynamic division module for eigenvalue bits, a configurable in-memory computing module for dynamic division of weight bits, and a multiplication and accumulation result shift addition module. Through the bit copying and slicing function and dynamic mapping of in-memory computing core units, dynamic adjustment of accuracy is achieved.
It realizes flexible adjustment of calculation accuracy according to task requirements, improves data processing capabilities and calculation efficiency, and is suitable for satellite tasks with different accuracy requirements, extending the life of satellite tasks.
Smart Images

Figure CN120067043A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of in-memory computing, and in particular, to an in-memory computing system, method, and device with dynamically adjustable precision. Background Art
[0002] With the increasing complexity of satellite missions, the requirements for data processing capabilities and computing efficiency are constantly rising. However, traditional computing systems usually rely on data transmission between the central processing unit (CPU) and the memory. The main bottleneck of this architecture lies in the "memory wall" problem between the processor and the memory, that is, the data transmission speed is much lower than the processing speed, resulting in limited overall system performance. To solve this problem, in-memory computing technology has emerged. In-memory computing embeds processing units in the memory, reducing the need for data transmission and significantly improving data processing speed and energy efficiency. This architecture is suitable for application scenarios that require a large amount of data processing and low latency, such as satellite real-time image processing, data compression, and complex signal processing.
[0003] However, in satellite applications, due to different tasks having different requirements for computing precision, the system needs to have the ability to dynamically adjust the computing precision. For example, image processing tasks require high-precision computing to ensure the accuracy of processing results, while low-power tasks only need relatively low precision to save computing resources. In this case, an in-memory computing system and method with dynamically adjustable precision can flexibly adjust the computing precision according to specific task requirements, optimize computing efficiency, and extend the mission life of the satellite.
[0004] Therefore, designing an in-memory computing system and method that can dynamically adjust computing precision can not only improve the satellite's data processing capabilities but also flexibly respond to different precision requirements in various complex tasks, thus having important technical significance and broad application prospects in satellite applications. Summary of the Invention
[0005] Aiming at the deficiencies in the prior art, the purpose of the present invention is to provide an in-memory computing system, method, and device with dynamically adjustable precision.
[0006] An in-memory computing system with dynamically adjustable precision according to the present invention includes: an eigenvalue bit dynamic partitioning module, a configurable in-memory computing module with weight bit dynamic partitioning, and a multiply-accumulate result shift-and-add module; The eigenvalue bit dynamic partitioning module is used to dynamically and adjustably split the eigenvalue into different precisions and send it to the configurable in-memory computing module. The configurable in-memory computing module contains multiple in-memory computing core units, and each in-memory computing core unit stores different bits of the weight. The multiply-accumulate result shift-and-add module is used to perform shift-and-add operations on the multiply-accumulate results of the in-memory computing core units with different weights to obtain a complete output eigenvalue; In the eigenvalue bit dynamic partitioning module, a bit replication and bit splitting function are set. The bit replication is to replicate a 1-bit value twice and send it to two branches. The bit splitting is to split a 2-bit value into two 1-bit values and send them to two branches.
[0007] Preferably, the bit replication and bit splitting functions include: If the input eigenvalue is 8-bit abcdefgh, the allocation relationship is a:#0, b:#1, c:#2, d:#3, e:#4, f:#5, g:#6, h:#7; if the input eigenvalue is 4-bit abcd, the allocation relationship is a:#0, a:#1, b:#2, b:#3, c:#4, c:#5, d:#6, d:#7; if the input eigenvalue is 2-bit ab, the allocation relationship is a:#0, a:#1, a:#2, a:#3, b:#4, b:#5, b:#6, b:#7; if the input eigenvalue is 1-bit a, the allocation relationship is a:#0, a:#1, a:#2, a:#3, a:#4, a:#5, a:#6, a:#7.
[0008] Preferably, the configurable in-memory computing module contains eight identical in-memory computing core units. Each in-memory computing core unit stores different bits of weights and can dynamically map 1 bit of weights.
[0009] Preferably, an addition and splicing function are set in the multiply-accumulate result shift addition module.
[0010] Preferably, the addition function is to add the two-way outputs to obtain a result. The splicing function is to directly splice the two-way outputs to obtain a result.
[0011] Preferably, the criterion for selecting addition or splicing is whether the two-way outputs belong to the same output channel of the neural network layer. If so, the addition function is selected; if not, the splicing function is selected.
[0012] According to an in-memory computing method with dynamically adjustable precision provided by the present invention, it includes: Step S1, map the neural network weight parameters to the configurable in-memory computing module according to the precision. Step S2, send the eigenvalue to the eigenvalue bit dynamic partitioning module for replication and partitioning according to the precision. Step S3, send the multiply-accumulate result after in-memory computing to the multiply-accumulate result shift addition module for addition and splicing to obtain the output eigenvalue.
[0013] A memory - in - computing device with dynamically adjustable precision provided by the present invention, the device includes a processor and a memory, and executable program instructions are stored in the memory. When the processor calls the program instructions in the memory, the processor is used to execute the steps of the precision - dynamically adjustable memory - in - computing method.
[0014] Compared with the prior art, the present invention has the following beneficial effects: A memory - in - computing system and method with dynamically adjustable precision provided by the present invention support the inference acceleration of a neural network model with dynamically adjustable precision and improve the efficiency of neural network inference by designing an eigenvalue bit dynamic division module, a configurable memory - in - computing module with weight bit dynamic division, and a multiply - accumulate result shift - add module. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] By reading the following detailed description of non - restrictive embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent: Figure 1 It is a schematic structural diagram of a memory - in - computing system with dynamically adjustable precision provided by the present invention; Figure 2 It is a schematic structural diagram of the eigenvalue bit dynamic division module provided by the present invention; Figure 3 It is a schematic structural diagram of the configurable memory - in - computing module with weight bit dynamic division provided by the present invention; Figure 4 It is a schematic structural diagram of the multiply - accumulate result shift - add module provided by the present invention; Figure 5 It is a schematic flowchart of a memory - in - computing method with dynamically adjustable precision provided by the present invention.
[0016] DESCRIPTION OF THE REFERENCE NUMERALS: DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0018] Figure 1 It is a schematic structural diagram of a memory - in - computing system with dynamically adjustable precision provided by the present invention, as Figure 1As shown in the figure, the present invention provides an in-memory computing system with dynamically adjustable precision, including an eigenvalue bit dynamic partitioning module 101, a configurable in-memory computing module 102 with dynamically partitioned weight bits, and a multiply-accumulate result shift-and-add module 103, where: The eigenvalue bit dynamic partitioning module 101 is used to dynamically and adjustably split the eigenvalue into different precisions and send it to the configurable in-memory computing module 102.
[0019] The high or low of the bits determines the calculation precision. For example, INT8 calculation precision means that fixed-point 8-bit eigenvalues and fixed-point 8-bit weights are calculated, and INT4 calculation precision means that fixed-point 4-bit eigenvalues and fixed-point 4-bit weights are calculated, and so on. Generally, INT8 and higher bits are considered high-precision calculations, and INT4 and lower bits are low-precision calculations. The higher the precision, the more bits of the weights and eigenvalues, and the greater the required hardware storage and calculation amount; the lower the precision, the fewer bits of the weights and eigenvalues, and the smaller the required hardware storage and calculation amount.
[0020] In the present invention, a bit replication and splitting function is set in the eigenvalue bit dynamic partitioning module 101.
[0021] The configurable in-memory computing module 102 contains eight identical in-memory computing core units 1021, and each in-memory computing core unit 1021 stores different bits of the weights and can dynamically map 1 bit of the weights.
[0022] In the present invention, an in-memory computing function is set in the configurable in-memory computing module 102.
[0023] The multiply-accumulate result shift-and-add module 103 is used to shift and add the multiply-accumulate results of the in-memory computing core units 1021 with different weights to obtain a complete output eigenvalue.
[0024] In the present invention, an addition and splicing function is set in the multiply-accumulate result shift-and-add module 103.
[0025] An in-memory computing system and method with dynamically adjustable precision provided by the present invention support the inference acceleration of a neural network model with dynamically adjustable precision by designing an eigenvalue bit dynamic partitioning module, a configurable in-memory computing module with dynamically partitioned weight bits, and a multiply-accumulate result shift-and-add module, and improve the efficiency of neural network inference.
[0026] In the present invention, Figure 2 is a schematic structural diagram of the eigenvalue bit dynamic partitioning module provided by the present invention. As Figure 2 shown, the present invention provides an eigenvalue bit replication and splitting function. Figure 2The black dots in it indicate the selection of the copy or split function. If it is the copy function, two copies of the 1-bit value are copied and sent to two branches; if it is the split function, the 2-bit value is split into two 1-bit values and sent to two branches. If the input eigenvalue is 8-bit abcdefgh, its final allocation relationship is a:#0, b:#1, c:#2, d:#3, e:#4, f:#5, g:#6, h:#7; if the input eigenvalue is 4-bit abcd, its final allocation relationship is a:#0, a:#1, b:#2, b:#3, c:#4, c:#5, d:#6, d:#7; if the input eigenvalue is 2-bit ab, its final allocation relationship is a:#0, a:#1, a:#2, a:#3, b:#4, b:#5, b:#6, b:#7; if the input eigenvalue is 1-bit a, its final allocation relationship is a:#0, a:#1, a:#2, a:#3, a:#4, a:#5, a:#6, a:#7.
[0027] Specifically, the implementation method of the eigenvalue bit copy and split function is as follows: copy and split operations are performed according to the number of consecutive 0s in the input eigenvalue, with the aim of removing consecutive 0s in the 8-bit input eigenvalue and supplementing it to 8 bits through copying. For example, for the 8-bit eigenvalue "01001000", the consecutive 0s in the 0th and 1st bits are removed, and the consecutive 0s in the 4th and 5th bits are removed. The remaining eigenvalue at this time is 4 bits, and then it is copied to become "00111100", and then it is split into 8 bits "0", "0", "1", "1", "1", "1", "0", "0", and transmitted to the in-memory computing core units 1021 from #0 to #7. Another example is the 8-bit eigenvalue "00100000", where the consecutive 0s in the 0th, 1st, 2nd, and 3rd bits are removed, and the consecutive 0s in the 6th and 7th bits are removed. The remaining eigenvalue at this time is 2 bits, and then it is copied to become "11110000", and then it is split into 8 bits "1", "1", "1", "1", "0", "0", "0", "0", and transmitted to the in-memory computing core units 1021 from #0 to #7. This automated copy and split function can automatically determine the position of consecutive 0s and reduce the computational overhead of this part, and automatically perform bit copy and split operations.
[0028] In the present invention, Figure 3 is a schematic structural diagram of the configurable in-memory computing module for dynamic partitioning of weight bits provided by the present invention, as Figure 3As shown in the figure, the present invention provides an in-memory computing function. Different bits of the weights are dynamically divided into different in-memory computing core units 1021. If the weight is 8 bits ijklmnop, its final allocation relationship is p:#0, o:#1, n:#2, m:#3, l:#4, k:#5, j:#6, i:#7; if the input feature value is 4 bits ijkl, its final allocation relationship is l:#0, k:#1, j:#2, i:#3, l:#4, k:#5, j:#6, i:#7; if the input feature value is 2 bits ij, its final allocation relationship is j:#0, i:#1, j:#2, i:#3, j:#4, i:#5, j:#6, i:#7; if the input feature value is 1 bit i, its final allocation relationship is i:#0, i:#1, i:#2, i:#3, i:#4, i:#5, i:#6, i:#7.
[0029] The structural schematic diagrams described above are only illustrative. The in-memory computing module can be composed of an analog in-memory computing module or a digital in-memory computing module, and the selected memory can be composed of SRAM, RRAM, Flash, etc. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0030] In the present invention, Figure 4 is the structural schematic diagram of the multiplication-accumulation result shift-and-add module provided by the present invention. As Figure 4 shown, the present invention provides an addition and splicing function. Figure 4 The black dots in indicate the selection of addition or splicing function. If it is the addition function, the two-way outputs are added to obtain the result; if it is the splicing function, the two-way outputs are directly spliced to obtain the result. The criterion for selecting addition or splicing is based on whether the two-way outputs belong to the same output channel of the neural network layer. If so, the addition function is selected; if not, the splicing function is selected.
[0031] Figure 5 is the flowchart of an in-memory computing method with dynamically adjustable precision provided by the present invention. As Figure 5 shown, the present invention provides an in-memory computing method with dynamically adjustable precision, including: Step S1, mapping the neural network weight parameters to the configurable in-memory computing module according to the precision. For example, an 8-bit weight is mapped to the in-memory computing core units 1021 from #0 to #7, a 4-bit weight is mapped to two groups of in-memory computing core units 1021 from #0 to #3 and from #4 to #7, and so on; Step S2: Send the eigenvalue to the eigenvalue bit dynamic partitioning module for replication and partitioning according to the precision. For example, the 8-bit eigenvalue is sliced and sent to the in-memory computing core units #0 to #7 of the memory. The 4-bit eigenvalue is replicated and sliced, and the 0th bit is sent to #0 to #1, the 1st bit is sent to #2 to #3, the 2nd bit is sent to #4 to #5, the 3rd bit is sent to #6 to #7, and so on. Step S3: Send the multiplication and accumulation result after in-memory computing to the multiplication and accumulation result shift addition module for addition and splicing. For example, the 8-bit * 8-bit multiplication and accumulation result of #0 to #7 is shifted and added to obtain the output eigenvalue. The 4-bit * 4-bit multiplication and accumulation results of #0 to #3 and #4 to #7 are shifted and added and spliced to obtain the output eigenvalue.
[0032] The present invention also provides an in-memory computing device with dynamically adjustable precision, including: a processor and a memory. The memory stores executable program instructions. When the processor calls the program instructions in the memory, the processor is used to execute the steps of the in-memory computing method with dynamically adjustable precision as described above.
[0033] The present invention also provides a computer-readable storage medium for storing a program, and when the program is executed, it implements the steps of the in-memory computing method with dynamically adjustable precision as described above.
[0034] It should be noted that those skilled in the art can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "platform" here.
[0035] In addition, the embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When at least one processor of the user device executes the computer-executable instructions, the user device executes the various possible methods described above. Among them, the computer-readable medium includes a computer storage medium and a communication medium, and the communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer. An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in the user device. Of course, the processor and the storage medium can also exist as discrete components in the communication device.
[0036] The present application also provides a program product, which includes a computer program stored in a readable storage medium. At least one processor of the server can read the computer program from the readable storage medium, and the execution of the computer program by the at least one processor enables the server to implement the method according to any one of the above embodiments of the present invention.
[0037] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0038] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. An in-memory computing system with dynamically adjustable precision, characterized in that: include: A module for dynamically dividing eigenvalue bits, a module for dynamically dividing weight bits and a module for shifting and adding multiplication and accumulation results; The eigenvalue bit dynamic division module is used to dynamically and adjustably split the eigenvalue into different precisions and send them to the configurable in-memory computing module, the configurable in-memory computing module includes a plurality of in-memory computing core units, each of which stores different bits of weights, and the multiplication and accumulation result shift addition module is used to shift and add the multiplication and accumulation results of the in-memory computing core units with different weights to obtain a complete output eigenvalue; The bit replication and bit splitting functions are set in the characteristic value bit dynamic division module. The bit replication is: copying 1 bit value twice and sending it to two forks; the bit splitting is: splitting 2 bit value into two 1 bit values and sending them to two forks.
2. The in-memory computing system with dynamically adjustable precision according to claim 1, characterized in that: The bit copying and bit splitting functions include: If the input characteristic value is 8 bits abcdefgh, the allocation relationship is a:#0, b:#1, c:#2, d:#3, e:#4, f:#5, g:#6, h:#7; if the input characteristic value is 4 bits abcd, the allocation relationship is a:#0, a:#1, b:#2, b:#3, c:#4, c:#5, d:#6, d:#7; if the input characteristic value is 2 bits ab, the allocation relationship is a:#0, a:#1, a:#2, a:#3, b:#4, b:#5, b:#6, b:#7; if the input characteristic value is 1 bit a, the allocation relationship is a:#0, a:#1, a:#2, a:#3, a:#4, a:#5, a:#6, a:#7.
3. The in-memory computing system with dynamically adjustable precision according to claim 1, characterized in that: The configurable in-memory computing module includes eight identical in-memory computing core units, each of which stores different bits of weight and can dynamically map 1 bit of the weight.
4. The in-memory computing system with dynamically adjustable precision according to claim 1, characterized in that: The multiplication and accumulation result shift addition module is provided with addition and splicing functions.
5. The in-memory computing system with dynamically adjustable precision according to claim 4, characterized in that: The adding function is to add the two outputs to obtain the result; the splicing function is to directly splice the two outputs to obtain the result.
6. The in-memory computing system with dynamically adjustable precision according to claim 5, characterized in that: The criterion for selecting addition or splicing is whether the two outputs belong to the same output channel of the neural network layer. If so, the addition function is selected; if not, the splicing function is selected.
7. An in-memory computing method with dynamically adjustable precision, based on the in-memory computing system with dynamically adjustable precision as claimed in any one of claims 1 to 6, characterized in that: include: Step S1, mapping the neural network weight parameters to the configurable in-memory computing module according to the accuracy; Step S2, sending the eigenvalues to the eigenvalue bit dynamic division module for replication and division according to the precision; Step S3, sending the multiplication and accumulation results calculated in the memory to the multiplication and accumulation result shift addition module for addition and splicing to obtain the output eigenvalue.
8. An in-memory computing device with dynamically adjustable precision, characterized in that: The device includes a processor and a memory, wherein the memory stores executable program instructions. When the processor calls the program instructions in the memory, the processor is used to execute the steps of the in-memory computing method with dynamically adjustable precision as described in claim 7.
Citation Information
Patent Citations
Hardware fingerprint information generation method and system based on national cryptographic algorithm
CN111709044A
Static random access memory in-memory computing system and method
CN119513032A
In-memory computing neural network accelerator assisted by digital multiply-accumulate core and acceleration method
CN119721148A
Neural network acceleration apparatus and method, and device and computer storage medium
WO2023116314A1
Computing-in-memory circuit based on BNN algorithm acceleration
WO2025035581A1
Cited By
Lightweight deployment method and system of large model on edge computing device
CN120821479A
A method and system for lightweight deployment of a large model on an edge computing device
CN120821479B