On-chip off-chip collaborative triple modular redundancy in-memory computing system, chip and satellite

Through on-chip, off-chip collaborative three-mode redundant design and multi-level error correction technology, the problem of multi-bit error in in-memory computing systems in high radiation environments is solved, and high-reliability and high-precision neural network computing is realized, which improves the robustness and energy efficiency of the system.

CN120256194AInactive Publication Date: 2025-07-04SHANGHAI JIAOTONG UNIV +1

Patent Information

Application Number
CN202510733569.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing in-memory computing systems face multi-bit error problems in high-radiation environments. Existing error correction methods such as Hamming code cannot be effectively processed, resulting in a decrease in the reliability and accuracy of the calculation results. Traditional redundant storage methods also find it difficult to recover data when multi-bit errors.

Method used

The on-chip off-chip collaborative three-mode redundancy design is adopted, combined with the Hamming code decoding circuit and the three-mode redundant voting machine. Through multi-level error correction and redundant storage technology, it ensures accuracy and reliability in multi-bit errors, including using the Hamming code decoding circuit in the on-chip unit to correct 1-bit errors, and determining the correct data through the three-mode redundant voting machine in the off-chip unit.

Benefits of technology

It significantly improves the robustness of the in-memory computing system, ensures the accuracy of neural network computing and the reliability of the system, meets the needs of high energy efficiency and low power consumption, and can effectively handle multi-bit errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256194A_ABST
    Figure CN120256194A_ABST
Patent Text Reader

Abstract

The invention provides an on-chip and off-chip collaborative triple modular redundancy in-memory computing system, a chip and a satellite. The on-chip and off-chip collaborative triple modular redundancy in-memory computing system comprises an on-chip unit and an off-chip unit, the on-chip unit comprises an SRAM (Static Random Access Memory), a Hamming code decoding circuit and an in-memory calculation core; the Hamming code decoding circuit is used for checking and correcting 1-bit errors of weights and characteristic values in the SRAM and the in-memory calculation core; the off-chip unit comprises a FLASH and a triple modular redundancy voter; the FLASH stores original weights and characteristic values, multi-bit error checking is carried out through the Hamming code decoding circuit, correct data are read from the Flash through the triple modular redundancy voter, and the correct data are transmitted to the SRAM. Through the innovative multi-stage error correction design and the triple modular redundancy technology, the accuracy of neural network calculation and the reliability of the system can still be ensured in the face of multi-bit errors, the robustness of the in-memory calculation system is remarkably improved, and the calculation requirements of high energy efficiency, low power consumption and high reliability are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of chip technology, and specifically, to an in-memory computing system, a chip, and a satellite with on-chip and off-chip collaborative triple modular redundancy. Background Art

[0002] With the continuous increase in satellite communication and data processing requirements, on-board intelligent computing systems are facing increasingly complex computing tasks, especially the issues of computing reliability and accuracy in high-radiation environments. The traditional von Neumann architecture can no longer meet these requirements of high real-time performance, high energy efficiency, and high reliability. Therefore, in-memory computing technology has emerged.

[0003] By tightly integrating data storage and computing, and directly performing computations inside the storage unit, the data transmission time is significantly reduced, thereby improving the computing efficiency and energy efficiency. In-memory computing is suitable for data-intensive applications such as neural networks and graph computing, especially in on-board computing and intelligent devices, where there are broad application prospects. These applications usually require a large number of multiply-accumulate operations and have very high requirements for real-time performance and high energy efficiency. Compared with the traditional von Neumann architecture, in-memory computing technology provides a more superior solution for efficient data processing.

[0004] RRAM (Resistive Random Access Memory) is widely used in in-memory computing architectures. The non-volatility, radiation resistance, and high-density storage capabilities of RRAM make it an ideal storage medium. RRAM stores information by changing its resistance state and directly reads the stored data during the computing process, achieving the tight integration of storage and computing. Although RRAM technology has the above advantages, it also faces some challenges, especially in terms of high reliability and high-precision computing. In practical applications, especially in high-radiation environments such as satellites and spacecraft, in-memory computing systems face bit errors, especially single-event upsets and multi-bit errors caused by space radiation. This may lead to incorrect computing results in in-memory computing systems, especially during high-precision computations, such as the calculation of weights and eigenvalues during neural network inference. Without effective error detection and correction, these errors will cause a decrease in the accuracy of the computing results, seriously affecting the performance and reliability of the system.

[0005] Existing technologies usually adopt simple error detection and correction methods, and the most commonly used one is the Hamming code, which is a coding scheme that can detect and correct single-bit errors. The Hamming code detects and corrects errors by introducing redundant bits and has good performance, especially in environments where single-bit errors are relatively common. However, the limitation of the Hamming code is that it can only handle single-bit errors and cannot cope with the situation of multi-bit errors. In high-radiation environments, multiple bit errors may occur simultaneously, which will cause the existing systems to be unable to effectively recover the original data and affect the reliability of the computing results.

[0006] 1) Error correction technology based on Hamming code: Hamming code is a classic error detection and correction coding scheme that can detect and correct single-bit errors through redundant bits. This method is widely adopted in current RRAM in-memory computing systems. For example, some in-memory computing systems use Hamming code to encode the weights of neural networks and detect and correct single-bit errors in data through corresponding decoding circuits. This technology is effective in low-radiation environments and is widely used because its encoding and decoding processes are relatively simple and the computational overhead is small. However, Hamming code performs poorly in the face of multi-bit errors. Especially in the space environment, due to the high occurrence frequency of single-event upsets or multi-bit errors, the error recovery ability of Hamming code seems inadequate.

[0007] 2) Traditional redundant storage technology: Another method to solve the errors in in-memory computing is to use redundant storage technology, that is, to store multiple copies of data (such as dual-mode redundancy or triple-mode redundancy) to increase the fault tolerance of the system. For example, in an in-memory computing system, some critical data is stored in multiple copies in different storage units. When an error occurs, the system determines the correct data by comparing the data of different copies. However, although this method can improve the reliability of the system to a certain extent, it also has certain limitations. First, redundant storage increases the hardware overhead, especially in computing systems that need to process a large amount of data, and the cost of redundant storage is very high. Second, when multi-bit errors occur, the redundant storage method still cannot effectively recover the data. Especially when bit errors occur simultaneously in multiple copies, the system may not be able to correctly determine the correct data copy.

[0008] In view of the deficiencies in the error correction ability, poor reliability of calculation results, and insufficient redundancy design in the existing in-memory computing systems, a highly reliable and high-precision error correction in-memory computing system is needed to solve the above problems. Summary of the Invention

[0009] Aiming at the deficiencies in the prior art, the purpose of the present invention is to provide an in-memory computing system, a chip, and a satellite with on-chip and off-chip collaborative triple-mode redundancy.

[0010] An in-memory computing system with on-chip and off-chip collaborative triple-mode redundancy provided by the present invention includes: an on-chip unit and an off-chip unit; The on-chip unit includes SRAM, a Hamming code decoding circuit, and an in-memory computing core; the Hamming code decoding circuit checks and corrects 1-bit errors in the weights and eigenvalues in SRAM and the in-memory computing core; The off-chip unit includes a FLASH and a triple modular redundancy voter; the FLASH stores the original weights and eigenvalues, performs multi-bit error checking through a Hamming code decoding circuit, and reads the correct data from the Flash through the triple modular redundancy voter and transfers it to the SRAM.

[0011] Preferably, the Hamming code decoding circuit performs data checking and error correction on the SRAM and the in-memory computing core, including: - Error correction step for writing weights to the RRAM; - Error correction step for reading weights from the RRAM; - Error correction step for reading the eigenvalue SRAM and in-memory computing.

[0012] Preferably, the error correction step for writing weights to the RRAM includes: Read the data code and check code of the weights from the SRAM, and determine whether there is an error through the Hamming code decoding circuit. If a 1-bit error occurs, refresh and correct the SRAM; if there is no error, write the data code and check code of the weights to the RRAM.

[0013] Preferably, the error correction step for reading weights from the RRAM includes: Read the data code and check code of the weights from the RRAM, and determine whether there is an error through the Hamming code decoding circuit. If a 1-bit error occurs, refresh and correct the RRAM.

[0014] Preferably, the error correction step for reading the eigenvalue SRAM and in-memory computing includes: Read the data code and check code of the eigenvalues from the SRAM, and determine whether there is an error through the Hamming code decoding circuit. If a 1-bit error occurs, refresh and correct the SRAM. If there is no error, transfer the data code of the eigenvalues to the in-memory computing core and perform in-memory computing with the weights in the RRAM.

[0015] Preferably, the off-chip unit further includes a DDR. When a multi-bit error is detected, the correct data is read from the Flash through the triple modular redundancy voter, read into the DDR, and then transferred to the SRAM.

[0016] Preferably, the reading method of the triple modular redundancy voter includes: The triple modular redundancy voter determines the final output by comparing the same piece of data stored in three Flashes and following the principle of the minority obeying the majority; If the bits stored in the three Flashes of the same piece of data are "0", "1", "1", the output result is "1".

[0017] A chip according to the present invention includes the on-chip and off-chip cooperative triple modular redundancy in-memory computing system described above.

[0018] A satellite provided by the present invention includes the chip described above.

[0019] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention proposes a highly reliable and high-precision error-correcting in-memory computing system. Through innovative multi-level error-correcting design and triple modular redundancy technology, it can ensure the accuracy of neural network computing and the reliability of the system even in the face of multi-bit errors, significantly improving the robustness of the in-memory computing system and meeting the computing requirements of high energy efficiency, low power consumption, and strong reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent: Figure 1 It is a schematic structural diagram of a reconfigurable Hamming code EDAC in-memory computing error-correcting scheme provided by an embodiment of the present invention.

[0021] Figure 2 It is a schematic structural diagram of an in-memory computing system with on-chip and off-chip collaborative triple modular redundancy provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0023] As Figure 1 shown, the reconfigurable Hamming code EDAC in-memory computing error-correcting scheme designs multi-level Hamming code error-correcting coding for the weights and eigenvalues of the neural network, performs Hamming code error-correcting coding on the weights in the RRAM array, and performs Hamming code error-correcting coding on the eigenvalues in the SRAM. Through the multi-level Hamming code error-correcting coding design, the mapping of the neural network weights and eigenvalues is correct. Based on this, a high-precision in-memory computing error-correcting design based on Hamming code EDAC is realized, making the relative error of the neural network mapping less than 10 -6 .

[0024] One implementation solution is as follows: Bit flipping in the memory of the in-memory computing system will cause errors in the in-memory computing results. Therefore, this patent proposes a reconfigurable Hamming code EDAC in-memory computing error-correcting scheme. This error-correcting design includes error-correcting design for writing weights to RRAM, error-correcting design for reading weights from RRAM, and error-correcting design for reading eigenvalues from SRAM and in-memory computing, as Figure 1 shown.

[0025] 1) In the error correction design of writing weights into RRAM, the data code and parity code of the weights are read out from SRAM, and the Hamming code decoding circuit is used to determine whether there is an error. If a 1-bit error occurs, SRAM is refreshed and corrected. If there is no error, the data code and parity code of the weights are written into the RRAM of the in-memory computing core; 2) In the error correction design of reading weights from RRAM, the data code and parity code of the weights are read out from RRAM, and the Hamming code decoding circuit is used to determine whether there is an error. If a 1-bit error occurs, RRAM is refreshed and corrected; 3) In the error correction design of reading eigenvalue SRAM and in-memory computing, the data code and parity code of the eigenvalues are read out from SRAM, and the Hamming code decoding circuit is used to determine whether there is an error. If a 1-bit error occurs, SRAM is refreshed and corrected. If there is no error, the data code of the eigenvalues is transmitted to the in-memory computing core and in-memory computing is performed with the weights in RRAM. This method corrects the eigenvalues and weights through Hamming code, realizes high-precision eigenvalue and weight mapping, and obtains accurate in-memory computing multiply-accumulate results.

[0026] As Figure 2 shown, an in-memory computing system with on-chip and off-chip collaborative triple modular redundancy. For multi-bit errors that cannot be corrected by Hamming code, the original weight and eigenvalue data are stored in Flash in three copies, and a triple modular redundancy voter is designed. When a multi-bit error occurs, the Hamming code parity bit can only detect the existence of the error. After 1-bit error correction, if the Hamming code parity bit still shows the existence of the error, it means that a multi-bit error has occurred. At this time, the correct data is read from Flash through the triple modular redundancy voter to DDR and then transmitted to SRAM for the next calculation. Based on this, a highly reliable in-memory computing system design based on triple modular redundancy is realized, further ensuring the mapping correctness of neural network weights and eigenvalues.

[0027] In a specific implementation, the reconfigurable Hamming code EDAC in-memory computing error correction scheme can only solve 1-bit errors. To solve the errors caused by multi-bit flips, this patent proposes an in-memory computing system with on-chip and off-chip collaborative triple modular redundancy, as Figure 2As shown. When multiple-bit errors are detected in the on-chip SRAM, such as 2-bit errors, data needs to be read from the off-chip Flash back into the SRAM. The data is stored in three copies in the Flash, and the correct data is read out by a triple modular redundancy voter. The triple modular redundancy voter determines the final output by comparing the same piece of data stored in three Flash memories and following the principle of the minority obeying the majority. For example, when the bits stored in the three Flash memories of the same piece of data are "0", "1", and "1", according to the principle of the minority obeying the majority, the output result at this time is "1". Since the occurrence frequency of single-event upsets is not high, it is difficult to have the situation where two of the three Flash memories of the same piece of data have bit flips at the same time. Therefore, this method can effectively solve the single-event upset problem of on-chip memories.

[0028] Although the present invention proposes a technical solution for a reconfigurable Hamming code EDAC in-memory computing error correction design and an on-chip and off-chip collaborative triple modular redundancy in-memory computing system, there are still other alternative solutions available during the implementation process. For example, other types of error correction codes can be considered to replace the Hamming code for data error correction. For the redundant storage part, different types of redundancy methods, such as dual modular redundancy or quadruple modular redundancy, can also be considered to further improve the reliability of the system.

[0029] The present invention also provides a chip that employs the above-mentioned on-chip and off-chip collaborative triple modular redundancy in-memory computing system. The present invention also provides a satellite that employs the above-mentioned chip.

[0030] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific implementation manners, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.

Claims

1. An in-memory computing system with on-chip and off-chip collaborative triple modular redundancy, characterized in that, Comprising: On-chip units and off-chip units; The on-chip units include SRAM, a Hamming code decoding circuit, and an in-memory computing core; The Hamming code decoding circuit checks and corrects 1-bit errors in the weights and eigenvalues in the SRAM and the in-memory computing core; The off-chip units include FLASH and a triple modular redundancy voter; the FLASH stores the original weights and eigenvalues, performs multi-bit error checking through the Hamming code decoding circuit, and reads the correct data from the Flash through the triple modular redundancy voter and transfers it to the SRAM.

2. The in-memory computing system with on-chip and off-chip collaborative triple modular redundancy according to claim 1, wherein The Hamming code decoding circuit performing data checking and error correction on the SRAM and the in-memory computing core includes: - Weight writing RRAM error correction step; - Weight reading RRAM error correction step; - Eigenvalue reading SRAM and in-memory computing error correction step.

3. The in-memory computing system with on-chip and off-chip collaborative triple modular redundancy according to claim 2, wherein The weight writing RRAM error correction step includes: Reading the data code and parity code of the weight from the SRAM, determining whether there is an error through the Hamming code decoding circuit. If a 1-bit error occurs, refreshing and correcting the SRAM; if there is no error, writing the data code and parity code of the weight into the RRAM.

4. The in-memory computing system with on-chip and off-chip cooperative triple modular redundancy according to claim 2, wherein The weight reading RRAM error correction step includes: Reading the data code and parity code of the weight from the RRAM, determining whether there is an error through the Hamming code decoding circuit. If a 1-bit error occurs, refreshing and correcting the RRAM.

5. The in-memory computing system with on-chip and off-chip collaborative triple modular redundancy according to claim 2, characterized in that, The eigenvalue reading SRAM and in-memory computing error correction step includes: Reading the data code and parity code of the eigenvalue from the SRAM, determining whether there is an error through the Hamming code decoding circuit. If a 1-bit error occurs, refreshing and correcting the SRAM; if there is no error, transferring the data code of the eigenvalue to the in-memory computing core and performing in-memory computing with the weights in the RRAM.

6. The in-memory computing system with on-chip and off-chip collaborative triple modular redundancy according to claim 1, wherein The off-chip unit further includes DDR. When a multi-bit error is detected, the correct data is read from the Flash through the triple modular redundancy voter and transferred to the DDR and then to the SRAM.

7. The in-memory computing system with on-chip and off-chip collaborative triple modular redundancy according to claim 6, characterized in that The reading method of the triple modular redundancy voter includes: The triple modular redundancy voter determines the final output by comparing the same data stored in three Flashes and following the principle of the minority obeying the majority; If the bits stored in the three Flashes of the same data are "0", "1", "1", the output result is "1".

8. A chip, characterized in that, An in-memory computing system with on-chip and off-chip collaborative triple modular redundancy as described in any one of claims 1 to 7.

9. A satellite, characterized in that, A chip as described in claim 8.

Citation Information

Patent Citations

  • Fault-tolerant method of storage module of picosatellite based on FPGA

    CN101615147A

  • High-reliability lightweight satellite-borne intelligent calculation acceleration system and method

    CN116383133A

  • Dynamic instantiation-based error correction in-memory computing system, method and apparatus

    CN120011133A

  • Radiation hardening memory

    CN211124024U

Cited By

  • Error correction protection method and system for integrated circuit calibration

    CN121858356A

  • Neural network hardware reasoning fault tolerance method and system for BNCT radiation environment and medium

    CN122309231A

  • Neural network hardware inference fault-tolerant methods, systems, and media for bnct radiation environments

    CN122309231B