Triple Modular Redundancy (TMR) Radiation Hardened Memory System
The TMR memory system addresses radiation-induced errors in space-based applications by interfacing with existing memory systems to detect and correct errors using triple redundant memory and associated circuitry, ensuring reliable operation and cost-effective error recovery.
Patent Information
- Application Number
- JP2025507229
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-10
- Filing Date
- 2023-07-20
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-07-20
AI Technical Summary
Memory systems in space-based applications suffer from high radiation levels causing single-bit, multi-bit, and single-event functional interrupt errors, rendering them non-functional, and shielding with dense materials is impractical due to weight constraints.
A TMR memory system interfaces between a processor and a memory controller/PHY, using triple redundant memory and associated circuitry for error detection and correction through scrubbing, scrubber circuits, and error recovery, without modifying the memory controller/PHY, to ensure reliable operation in high-radiation environments.
The TMR system provides improved error correction and recovery capabilities in terms of cost, reliability, and speed, suitable for space-based applications, including radar and communication systems, by detecting and correcting errors using a majority vote among redundant memory modules.
Smart Images

Figure 2025528793000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE DISCLOSURE
[0001] This disclosure relates to memory systems, and more particularly to the use of triple modular redundancy (TMR) memory systems to provide radiation-hardened memory operation. [Background technology]
[0002]
[0002] Memory systems deployed in space-based applications are exposed to relatively high radiation levels, which can cause a significant increase in single-bit, multi-bit, and single-event functional interrupt (SEFI) errors. In some cases, these errors can render the memory non-functional. While shielding can reduce radiation exposure, this approach is impractical due to the added weight (e.g., requiring lead or similarly dense materials), especially in space-based applications where stringent weight constraints may be imposed. [Brief explanation of the drawings]
[0003] [Figure 1] 1 illustrates an implementation of a triple modular redundancy (TMR) memory system in accordance with certain embodiments of the present disclosure. [Figure 2]
[0004] 2 is a block diagram of the TMR memory system of FIG. 1 configured in accordance with certain embodiments of the present disclosure. [Figure 3]
[0005] 3 is a block diagram of the TMR reliability controller of FIG. 2 configured in accordance with certain embodiments of the present disclosure. [Figure 4]
[0006] FIG. 3 is a block diagram of the microcontroller of FIG. 2 configured in accordance with certain embodiments of the present disclosure. [Figure 5]
[0007] 1 is a flowchart illustrating a methodology for providing a TMR radiation-hardened memory according to an embodiment of the present disclosure. [Figure 6]
[0008] FIG. 1 is a block diagram of a processing platform configured to provide TMR radiation-hardened memory, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0004]
[0009] The following detailed description proceeds with reference to exemplary embodiments, many of which alternatives, modifications, and variations will be apparent in light of this disclosure.
[0005]
[0010] Techniques are provided herein for triple modular redundancy (TMR) memory systems configured to provide radiation-hardened memory operation. As described above, memory systems deployed in space-based applications are exposed to relatively high radiation levels, which can cause a significant increase in single-bit, multi-bit, and single-event functional interrupt (SEFI) errors. In some cases, these errors can render the memory at least temporarily non-functional. While shielding can reduce radiation exposure, this approach is impractical due to the added weight (e.g., requiring lead or similarly dense materials), especially in space-based applications where stringent weight constraints may be imposed.
[0006]
[0011] One approach to solving this problem is to modify the design of commercially available memory controllers and / or associated physical interfaces (PHYs) to commercially available memory chips (e.g., integrated circuits or ICs) to include circuitry for correcting bit errors. However, this approach is generally not capable of correcting longer sequences of bit errors (e.g., Single Event Functional Interrupt, or SEFI, conditions). This approach also tends to slow down memory operation, making it incompatible with increasingly faster memory chips, such as faster Double Data Rate 4 (DDR4) and DDR5 memory chips. Additionally, PHY development is costly and time-consuming. For example, it can take more than a year to develop a PHY.
[0007]
[0012] To this end, according to embodiments of the present disclosure, a TMR memory system is disclosed that interfaces between a processor and a commercially available memory controller / PHY for memory chips without requiring modification of the memory controller / PHY. The TMR memory system provides improved reliability for memory operations in higher radiation environments, such as space-based applications. The disclosed system uses triple redundant memory and associated circuitry to perform memory scrubbing and error recovery balanced with the processor's execution of mission software from those memories, as described in more detail below. The disclosed TMR memory system can be used with electronic systems in a wide variety of applications, including, for example, radar systems and communication systems, which may be deployed in space-based applications (e.g., satellite-based platforms) or other high-radiation environments, although other applications will be apparent.
[0008]
[0013] According to an embodiment, a TMR memory system interfacing between a processor and a memory controller / PHY includes a redundancy comparator configured to detect differences between data redundantly stored in a first memory, a second memory, and a third memory. The redundancy comparator is further configured to identify a memory error based on the detected differences. The memory system also includes an error collection buffer configured to store memory addresses associated with the memory errors. The memory system further includes a memory scrubber circuit configured to overwrite erroneous data with corrected data at the memory addresses associated with the memory errors. The corrected data is based on a majority vote among the three memories. The memory system further includes a priority arbiter configured to arbitrate between the overwrite operation of the memory scrubber and functional memory accesses associated with software execution executed by a processor configured to utilize the memory system (e.g., a mission application).
[0009]
[0014] It will be appreciated that the techniques described herein may provide improved error correction and recovery capabilities in terms of cost, reliability, and speed of operation compared to systems that provide physical shielding or require memory controller / PHY modifications. Numerous embodiments and applications will become apparent in light of this disclosure. System Architecture
[0010]
[0015] FIG. 1 illustrates a TMR memory system implementation 100 according to certain embodiments of the present disclosure. Implementation 100 is shown to include a processor 110, a TMR system 120, and three redundant memory systems 160a, 160b, and 160c. Each of redundant memory systems 160 is shown to include a memory controller 130, a memory PHY 140, and a memory chip 150. In some embodiments, these components 130, 140, and 150 can be any suitable chips, including commercially available chips. Redundant memory systems 160 are each configured to store and provide read / write access to copies of software and data used by processor 110, for example, in executing a mission application (e.g., signal processing, radar processing, communications, etc.). In the absence of any bit errors, the copies of software and data stored in the three memory chips 150a, 150b, and 150c are identical, and thus redundancy serves to identify such errors based on detectable differences. The use of three memory chips allows for the use of a majority voting scheme, as described below, although in some embodiments additional memory chips may be used to increase redundancy.
[0011]
[0016] The operation of the TMR system 120 is described in more detail below, but broadly speaking, the TMR system is configured to interface between the processor 110 and triple redundant memory and to use redundancy to detect and correct errors (e.g., errors caused by radiation effects or from other sources).
[0012]
[0017] 2 is a block diagram of the TMR memory system 120 of FIG. 1 configured in accordance with certain embodiments of the present disclosure. The TMR system is shown to include a TMR reliability controller 210 and a microcontroller 230. The operation of the TMR reliability controller 210 and the microcontroller 230 is described in more detail below. However, broadly speaking, the TMR reliability controller 210 is configured to enable the processor 110, through the memory controller 130 and the memory PHY 140, to read and write data to the memory chips 150 over the data bus 200 and provide increased protection against memory errors. In some embodiments, the microcontroller 230 communicates with the processor 110 over the configuration bus 220 and is configured to adjust various parameters associated with the operation of the TMR reliability controller 210 and provide overall control of the TMR system 120, as described in more detail below.
[0013]
[0018] Figure 3 is a block diagram of the TMR reliability controller 210 of Figure 2, configured in accordance with certain embodiments of the present disclosure. The TMR reliability controller 210 is shown to include an error collection buffer 300, a redundancy comparator 310, a priority arbiter 320, and a memory scrubber 330. The components of the TMR reliability controller operate under the control of a microcontroller 230.
[0014]
[0019] The redundancy comparator 310 is configured to detect differences between the data redundantly stored in the three memories 160 and identify memory errors based on the detected differences. The redundancy comparator 310 accesses the redundant memory chips 150 through the memory controller 130 and PHY 140 using any suitable technique associated with the particular memory controller / PHY selected for use in the application. Discovered errors are stored in an error collection buffer 300 configured to store a memory address associated with each memory error. In some embodiments, the error collection buffer 300 may be implemented as a circular buffer that may be monitored 340 by the microcontroller 230. In some embodiments, the error collection buffer may be configured to generate an interrupt to the microcontroller 230 to signal the presence of a new error.
[0015]
[0020] The memory scrubber circuit 330 is configured to overwrite erroneous data with corrected data at memory addresses associated with the memory errors. The corrected data is generated based on a majority vote performed among the first memory 160a, the second memory 160b, and the third memory 160c. Random bit errors, such as those caused by radiation, are relatively unlikely to occur at the same address and bit location in two or more of the memories, so a majority vote can be used to correct an error in one of the memories based on a consensus value obtained from the other two memories. Because only the memory addresses associated with the memory errors are stored, the memory overwrite operation is a read-modify-write memory operation, allowing the erroneous data to be reread, repaired, and written back to the memory at that address.
[0016]
[0021] In some embodiments, the rate at which memory scrubbing is performed may be set by the microcontroller 230 through scrub rate signaling 370. For example, the scrub rate may be increased when a SEFI occurs, as described below.
[0017]
[0022] The memory scrubber circuitry is also configured to monitor traffic on data bus 200 to determine whether a functional memory write is being performed by processor 110 to a memory address that is in the process of being scrubbed (e.g., about to be overwritten for correction). If so, the memory scrubber cancels the overwrite because it is not necessary and could potentially corrupt memory if performed after the processor has completed the functional write.
[0018]
[0023] The priority arbiter 320 is configured to arbitrate between overwrites performed by the memory scrubber and functional memory accesses associated with software execution performed by the processor 110. The priority arbiter throttles the scrubbing rate based on guidance provided by the microcontroller 230 (priority signaling 360), as described below.
[0019]
[0024] In some embodiments, the TMR reliability controller 210 may also be configured to use any suitable type of error correction coding (ECC) technique to detect and repair errors as an additional mechanism to the memory scrubbing process. This functionality may be switched on or off based on a mode setting 350 provided by the microcontroller 230.
[0020]
[0025] Figure 4 is a block diagram of the microcontroller 230 of Figure 2, configured in accordance with certain embodiments of the present disclosure. The microcontroller 230 is shown to include a microprocessor 400, a local memory 410, and an ECC system 430.
[0021]
[0026] The microprocessor 400 is configured to monitor 340 the error collection buffer 300 of the TMR reliability controller 210 and trigger operation of the memory scrubber circuit 330 in response to memory errors stored in the buffer. In some embodiments, the monitoring may be performed by reading the buffer. In some embodiments, the monitoring may be achieved through an interrupt generated by the buffer when an error is stored. The microprocessor may clear errors from the buffer after detection.
[0022]
[0027] Microprocessor 400 is also configured to determine whether the error retrieved from the buffer is a relatively simple single-bit error, or whether the error is a multi-bit error or is associated with a more serious SEFI condition that requires reinitialization of one or more of memory controllers 130 in addition to memory scrubbing (e.g., operating in SEFI recovery mode).
[0023]
[0028] Local memory 410 is configured to store copies 420 of the contents of configuration registers of memory controller 130 so that these copies can be used to restore or refresh the memory controller in response to detecting a SEFI condition. In some embodiments, the microprocessor may increase the scrub rate 370 during SEFI recovery and / or turn off error correction for the memory undergoing SEFI recovery to allow faster recovery from this more severe error condition.
[0024]
[0029] In some embodiments, the microprocessor may set a priority 360 for the priority arbiter 320 based on a trade-off between overwrites performed by the memory scrubber and functional memory accesses by the processor 110. The priority may be determined based on guidance from the processor 110 provided on the configuration bus 220 and may be related to mission parameters or other considerations. For example, in some cases, correcting errors may be paramount to mission success, so overwrites by the memory scrubber may be set at a higher priority. However, in other cases, allowing the processor to execute mission software with minimal interruptions due to error correction may be paramount to mission success, so functional memory accesses by the processor may be set at a higher priority.
[0025]
[0030] In some embodiments, the microcontroller 230 may power cycle the external memory 150a, 150b, or 150c to clear any invalid states that are not recoverable by a reset or command.
[0026]
[0031] In some embodiments, the microprocessor may control mode setting 350 to cause the TMR reliability controller to include or exclude ECC functionality as an additional operation to the scrubbing function. The mode setting decision may be made, in part, based on guidance from processor 110 that is also provided on configuration bus 220.
[0027]
[0032] In some embodiments, the microprocessor may be configured to monitor the rate at which errors are being detected through the error collection buffer 300 and detect an increase or decrease in those error rates. In response to detecting such a change in the error rate, the microprocessor may increase or decrease the scrub rate 370 accordingly.
[0028]
[0033] In some embodiments, ECC system 430 is configured to maintain the integrity of local memory 410 by using any suitable ECC technique to detect and correct errors that may occur in local memory 410.
[0029]
[0034] In some embodiments, microprocessor 400 may be configured to detect stuck bit errors, where the bit remains stuck at a 1 or 0 state despite repeated scrubbing attempts. The stuck bit may be reported back to processor 110 over configuration bus 220 so that the processor may attempt to avoid using the memory location containing the stuck bit. method system
[0030]
[0035] FIG. 5 is a flowchart illustrating a methodology 500 for providing a TMR radiation-hardened memory according to an embodiment of the present disclosure. As can be seen, the illustrative method 500 includes several phases and subprocesses, the sequence of which may vary from embodiment to embodiment. However, taken as a whole, these phases and subprocesses form a process for providing a TMR radiation-hardened memory in accordance with certain embodiments disclosed herein, as described above and illustrated in, for example, FIGS. 1-4. However, as will become apparent in light of this disclosure, other system architectures may be used in other embodiments. To this end, the correlation of the various functions shown in FIG. 5 with the specific components illustrated therein is not intended to imply any architectural and / or usage limitations. Rather, other embodiments may include, for example, varying degrees of integration, where multiple functionalities are effectively performed by one system. Numerous variations and alternative configurations will become apparent in light of the present disclosure.
[0031]
[0036] In one embodiment, method 500 begins at operation 510 by detecting differences between data redundantly stored in a first memory, a second memory, and a third memory.
[0032]
[0037] In operation 520, memory errors are identified based on the detected differences, and in operation 530, memory addresses associated with the memory errors are stored in an error collection memory.
[0033]
[0038] In operation 540, a memory scrub is performed in which the erroneous data is overwritten with corrected data at the memory address associated with the memory error. In some embodiments, the memory overwrite operation is a read-modify-write memory operation. In some embodiments, the corrected data is generated based on a majority vote performed among the first memory, the second memory, and the third memory.
[0034]
[0039] In operation 550, arbitration is performed between memory scrubbing operations and functional memory accesses associated with mission software execution (e.g., performed by a processor configured to utilize the memory system), the arbitration being based on a priority associated with a trade-off between the memory scrubber's error correction activities to increase data reliability and the timely execution of the mission software.
[0035]
[0040] Of course, in some embodiments, additional actions may be performed as previously described in connection with the system, for example, the operation rate of the memory scrubber circuit may be increased if the memory error is associated with a single event functional interrupt (SEFI) condition, as opposed to a single bit error.
[0036]
[0041] In some embodiments, copies of configuration registers of the controllers associated with the first, second, and third memories may be stored in local memory, and the configuration registers may be restored from the local memory copies in response to determining that a memory error is associated with a SEFI condition.
[0037]
[0042] In some embodiments, the overwrite may be canceled in response to detecting that a functional write is being performed at a memory address associated with a memory error. Illustrative System
[0038]
[0043] 6 is a block diagram of a processing platform 600 configured to provide TMR radiation-hardened memory in accordance with an embodiment of the present disclosure. In some embodiments, platform 600, or portions thereof, may be hosted on or otherwise incorporated within an electronic system of a space-based platform where radiation hardening is particularly useful, including a data communications system, a radar system, a computing system, or any type of embedded system. The disclosed techniques may also be used to improve memory reliability in other platforms, including data communications devices, personal computers, workstations, laptop computers, tablets, touchpads, portable computers, handheld computers, cellular telephones, smartphones, or messaging devices. In certain embodiments, any combination of different devices may be used.
[0039]
[0044] In some embodiments, platform 600 may include any combination of processor 110, memory 160a, 160b, 160c, TMR system 120, network interface 640, input / output (I / O) system 650, user interface 660, display element 664, and storage system 670. As further seen in the figure, a bus and / or interconnect 690 may also be provided to enable communication between the various components listed above and / or other components not shown. Platform 600 may be coupled to a network 694 through network interface 640 to enable communication with other computing devices, platforms, devices to be controlled, or other resources. Other components and functionality not reflected in the block diagram of FIG. 6 will become apparent in light of this disclosure, and it will be appreciated that other embodiments are not limited to any particular hardware configuration.
[0040]
[0045] Processor 110 may be any suitable processor and may include one or more coprocessors or controllers, such as an audio processor, a graphics processing unit, or a hardware accelerator, to assist in executing mission software and / or any control and processing operations associated with platform 600. In some embodiments, processor 110 may be implemented as any number of processor cores. A processor (or processor core) may be any type of processor, such as a microprocessor, embedded processor, digital signal processor (DSP), graphics processor (GPU), tensor processing unit (TPU), network processor, field programmable gate array, or other device configured to execute code. A processor may be a multithreaded core in that it may include more than one hardware thread context (or “logical processor”) per core. Processor 110 may be implemented as a complex instruction set computer (CISC) or reduced instruction set computer (RISC) processor. In some embodiments, processor 110 may be configured as an x86 instruction set compatible processor.
[0041]
[0046] Memory 160 comprises three redundant memories 160a, 160b, and 160c as previously described and may be implemented using any suitable type of digital storage including, for example, DDR3, DDR4, and / or DDR5 SDRAM.
[0042]
[0047] Storage system 670 may be implemented as a non-volatile storage device such as, but not limited to, one or more of a hard disk drive (HDD), a solid state drive (SSD), a universal serial bus (USB) drive, an optical disk drive, a tape drive, an internal storage device, an external storage device, flash memory, battery-backed synchronous DRAM (SDRAM), and / or a network-accessible storage device. In some embodiments, storage 670 may comprise technology to increase storage performance-enhancing protection for valuable digital media when multiple hard drives are included.
[0043]
[0048] Processor 110 may be configured to execute operating system (OS) 680, which may comprise any suitable operating system, such as Google Android (Google Inc., Mountain View, CA), Microsoft Windows® (Microsoft Corp., Redmond, WA), Apple OS X (Apple Inc., Cupertino, CA), Linux®, or a real-time operating system (RTOS). As will be appreciated in light of this disclosure, the techniques provided herein may be implemented regardless of the particular operating system provided with platform 600, and thus may be implemented using any suitable existing or later-developed platform.
[0044]
[0049] The network interface circuitry 640 may be any suitable network chip or chipset that enables wired and / or wireless connections between the platform 600 and / or other components of the network 694, thereby enabling the platform 600 to communicate with other local and / or remote computing systems and / or other resources. Wired communications may conform to existing (or future developed) standards, such as, for example, Ethernet. Wireless communications may conform to existing (or future developed) standards, such as, for example, cellular communications, including LTE (Long Term Evolution) and 5G, Wireless Fidelity (Wi-Fi), Bluetooth, and / or Near Field Communication (NFC). Exemplary wireless networks include, but are not limited to, wireless local area networks, wireless personal area networks, wireless metropolitan area networks, cellular networks, and satellite networks.
[0045]
[0050] I / O system 650 may be configured to interface between various I / O devices and other components of platform 600. The I / O devices may include, but are not limited to, a user interface 660 and a display element 664. User interface 660 may include devices (not shown), such as a touchpad, keyboard, mouse, and the like, to enable a user to control the system. Display element 664 may be configured to display information to a user. I / O system 650 may include a graphics subsystem configured to perform processing of images for rendering on display element 664. The graphics subsystem may be, for example, a graphics processing unit or a visual processing unit (VPU). An analog or digital interface may be used to communicatively couple the graphics subsystem and the display element. For example, the interface may be any of High-Definition Multimedia Interface (HDMI), DisplayPort, Wireless HDMI, and / or any other suitable interface using wireless high-definition compliant technology. In some embodiments, the graphics subsystem may be integrated into processor 110 or any chipset of platform 600.
[0046]
[0051] It will be appreciated that in some embodiments, the various components of platform 600 may be combined or integrated into a system-on-chip (SoC) architecture. In some embodiments, the components may be hardware components, firmware components, software components, or any suitable combination of hardware, firmware, or software.
[0047]
[0052] TMR system 120 is configured to provide radiation-hardened reliability for memory 160, as previously described. TMR system 120 may include any or all of the circuits / components illustrated in FIGS. 2-4, as described above. These components may be implemented or otherwise used with a variety of suitable software and / or hardware coupled to or otherwise forming part of platform 600. These components may additionally or alternatively be implemented or otherwise used with user I / O devices capable of providing information to and receiving information and commands from a user.
[0048]
[0053] In various embodiments, platform 600 may be implemented as a wireless system, a wired system, or a combination of both. When implemented as a wireless system, platform 600 may include components and interfaces suitable for communicating over a wireless shared medium, such as one or more antennas, transmitters, receivers, transceivers, amplifiers, filters, control logic, etc. Examples of wireless shared media may include portions of a wireless spectrum, such as the radio frequency spectrum, etc. When implemented as a wired system, platform 600 may include components and interfaces suitable for communicating over a wired communication medium, such as input / output adapters, physical connectors for connecting the input / output adapters to corresponding wired communication media, network interface cards (NICs), disk controllers, video controllers, audio controllers, etc. Examples of wired communication media may include wires, cable metal leads, printed circuit boards (PCBs), backplanes, switch fabrics, semiconductor materials, twisted pair wires, coaxial cables, optical fibers, etc.
[0049]
[0054] Various embodiments may be implemented using hardware elements, software elements, or a combination of both. Examples of hardware elements may include processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, ASICs, programmable logic devices, digital signal processors, FPGAs, logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application program interfaces, instruction sets, computational code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. The decision whether to implement an embodiment using hardware and / or software elements may vary according to any number of factors, such as desired computation rate, power level, thermal tolerance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed, and other design or performance constraints.
[0050]
[0055] Some embodiments may be described using the terms "coupled" and "connected," along with their derivatives. These terms are not intended as synonyms for each other. For example, some embodiments may be described using the terms "connected" and / or "coupled" to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but yet still cooperate or interact with each other.
[0051]
[0056] The various embodiments disclosed herein may be implemented in various forms of hardware, software, firmware, and / or special-purpose processors. For example, in one embodiment, at least one non-transitory computer-readable storage medium has encoded thereon instructions that, when executed by one or more processors, cause one or more of the methodologies disclosed herein to be performed. The instructions may be encoded using a suitable programming language, such as C, C++, Object-Oriented C, Java, JavaScript, Visual Basic .NET, Beginner's All-Purpose Symbolic Instruction Code (BASIC), or alternatively, using a custom or proprietary instruction set. The instructions may be provided in the form of one or more computer software applications and / or applets tangibly embodied on a memory device and executable by a computer having any suitable architecture. In one embodiment, the system may be hosted on a given website and implemented, for example, using JavaScript or another suitable browser-based technology. For example, in certain embodiments, the system may leverage processing resources provided by a remote computer system accessible via network 694. The computer software applications disclosed herein may include any number of different modules, sub-modules, or other components of different functionality, and may provide information to or receive information from other components. These modules may be used to communicate with input and / or output devices, such as a display screen, a touch-sensitive surface, a printer, and / or any other suitable device. It will be appreciated that other components and functionality not reflected in the figures will be apparent in light of this disclosure, and that other embodiments are not limited to any particular hardware or software configuration. Accordingly, in other embodiments, platform 600 may include additional, fewer, or alternative sub-components compared to those included in the illustrative embodiment of FIG. 6 .
[0052]
[0057] The non-transitory computer-readable medium described above may be any suitable medium for storing digital information, such as a hard drive, a server, flash memory, and / or random access memory (RAM), or a combination of memories. In alternative embodiments, the components and / or modules disclosed herein may be implemented in hardware, including gate-level logic such as a field programmable gate array (FPGA), or alternatively, special-purpose semiconductors such as an application-specific integrated circuit (ASIC). Still other embodiments may be implemented in a microcontroller having several input / output ports for receiving and outputting data and several built-in routines for performing the various functionality disclosed herein. It will be apparent that any suitable combination of hardware, software, and firmware may be used, and that other embodiments are not limited to any particular system architecture.
[0053]
[0058] Some embodiments may be implemented using, for example, a machine-readable medium or article that may store instructions or sets of instructions that, when executed by a machine, cause the machine to perform methods, processes, and / or operations in accordance with the embodiments. Such a machine may include, for example, any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, process, or the like, and may be implemented using any suitable combination of hardware and / or software. A machine-readable medium or article may include any suitable type of memory unit, memory device, memory article, memory medium, storage device, storage article, storage medium, and / or storage unit, such as, for example, memory, removable or non-removable media, erasable or non-erasable media, writable or rewritable media, digital or analog media, hard disk, floppy disk, compact disk read-only memory (CD-ROM), compact disk recordable (CD-R) memory, compact disk rewriteable (CD-RW) memory, optical disk, magnetic medium, magneto-optical medium, removable memory card or disk, various types of digital versatile disks (DVDs), tape, cassette, or the like. The instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, encrypted code, and the like, implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.
[0054]
[0059] Unless otherwise specified, it may be recognized that terms such as "processing," "calculating," "computing," "determining," or the like refer to the actions and / or processes of a computer or computing system, or similar electronic computing device, that manipulate and / or transform data represented as physical quantities (e.g., electrons) in the registers and / or memory units of the computer system into other data similarly represented as physical quantities in the registers, memory units, or other such information storage transmission or representation of the computer system. Embodiments are not limited in this context.
[0055]
[0060] The terms “circuit” or “circuitry,” as used in any embodiment herein, are functional and may comprise, for example, alone or in any combination, hardwired circuitry, programmable circuitry such as a computer processor with one or more individual instruction processing cores, state machine circuitry, and / or firmware that stores instructions executed by the programmable circuitry. Circuitry may include a processor and / or controller configured to execute one or more instructions to perform one or more operations described herein. Instructions may be embodied, for example, as an application, software, firmware, etc., configured to cause circuitry to perform any of the foregoing operations. Software may be embodied as a software package, code, instructions, instruction set, and / or data recorded on a computer-readable storage device. Software may be embodied or implemented to include any number of processes, which in turn may be embodied or implemented to include any number of threads, etc., in a hierarchical manner. Firmware may be embodied as hard-coded (e.g., non-volatile) code, instructions or instruction sets, and / or data in a memory device. Circuitry may be embodied collectively or individually as circuitry forming part of a larger system, e.g., an integrated circuit (IC), an application-specific integrated circuit (ASIC), a system-on-a-chip (SoC), a desktop computer, a laptop computer, a tablet computer, a server, a smartphone, etc. Other embodiments may be implemented as software executed by a programmable control device. In such cases, the terms "circuit" or "circuitry" are intended to include combinations of software and hardware, such as a programmable control device or processor capable of executing software. As described herein, various embodiments may be implemented using hardware elements, software elements, or any combination thereof.Examples of hardware elements may include a processor, a microprocessor, a circuit, a circuit element (e.g., a transistor, a resistor, a capacitor, an inductor, etc.), an integrated circuit, an application specific integrated circuit (ASIC), a programmable logic device (PLD), a digital signal processor (DSP), a field programmable gate array (FPGA), a logic gate, a register, a semiconductor device, a chip, a microchip, a chipset, etc.
[0056]
[0061] Numerous specific details have been set forth herein to provide a thorough understanding of the embodiments. However, it will be understood that other embodiments may be practiced without these specific details or with an otherwise different set of details. It will be further appreciated that specific structural and functional details disclosed herein may be representative of example embodiments and are not necessarily intended to limit the scope of the present disclosure. In addition, while subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described herein. Rather, the specific features and acts described herein are disclosed as example forms of implementing the claims. Further illustrative embodiments
[0057]
[0062] The following examples relate to further embodiments, from which numerous permutations and configurations will be apparent.
[0058]
[0063] Example 1 is a memory system comprising: a redundancy comparator configured to detect differences between data redundantly stored in a first memory, a second memory, and a third memory; wherein the redundancy comparator is further configured to identify a memory error based on the detected differences; an error collection buffer configured to store memory addresses associated with the memory errors; a memory scrubber circuit configured to overwrite erroneous data with corrected data at the memory addresses associated with the memory errors; wherein the corrected data is generated based on a majority vote performed among the first memory, the second memory, and the third memory; and a priority arbiter configured to arbitrate between the overwrite performed by the memory scrubber and functional memory accesses associated with software execution performed by a processor configured to utilize the memory system.
[0059]
[0064] Example 2 includes the memory system of Example 1, further comprising a microcontroller configured to monitor the error collection buffer and trigger operation of the memory scrubber circuit in response to a memory error.
[0060]
[0065] Example 3 includes the memory system of Example 2, wherein the microcontroller is configured to set priorities for the priority arbiter based on a trade-off between overwrites performed by the memory scrubber and functional memory accesses.
[0061]
[0066] Example 4 includes the memory system of Example 2 or 3, wherein the microcontroller is configured to determine that the memory error is associated with a single event functional interrupt (SEFI) condition and, in response to the determination, increase an operation rate of the memory scrubber circuit.
[0062]
[0067] Example 5 includes the memory system of Example 4, wherein the microcontroller comprises a local memory, the microcontroller is configured to store copies of configuration registers of the controllers associated with the first memory, the second memory, and the third memory in the local memory, and to restore the configuration registers based on the stored copies in response to determining that the memory error is associated with a SEFI condition.
[0063]
[0068] Example 6 includes the memory system of Example 4 or 5, wherein the microcontroller is configured to power cycle one or more of the first memory, the second memory, and the third memory in response to determining that the memory error is associated with a SEFI condition.
[0064]
[0069] Example 7 includes the memory system of any one of Examples 1-6, wherein the memory scrubber circuit is configured to cancel the overwrite in response to detecting that a functional write is being performed at a memory address associated with a memory error.
[0065]
[0070] Example 8 is a space-based processing system comprising: a processor configured to execute mission software; and a memory system, the memory system comprising: a redundancy comparator configured to detect differences between data redundantly stored in a first memory, a second memory, and a third memory, wherein the redundancy comparator is further configured to identify a memory error based on the detected differences; an error collection buffer configured to store memory addresses associated with the memory errors; a memory scrubber circuit configured to overwrite erroneous data with corrected data at the memory addresses associated with the memory errors, wherein the corrected data is generated based on a majority vote performed among the first memory, the second memory, and the third memory; and a priority arbiter configured to arbitrate between the overwrite performed by the memory scrubber and functional memory accesses associated with execution of the mission software.
[0066]
[0071] Example 9 includes the space-based processing system of Example 8, wherein the memory system comprises a microcontroller configured to monitor the error collection buffer and trigger operation of the memory scrubber circuit in response to a memory error.
[0067]
[0072] Example 10 includes the space-based processing system of Example 9, wherein the microcontroller is configured to set priorities for the priority arbiter based on a trade-off between overwrites performed by the memory scrubber and functional memory accesses.
[0068]
[0073] Example 11 includes the space-based processing system of Example 9 or 10, wherein the microcontroller is configured to determine that the memory error is associated with a single event functional interrupt (SEFI) condition and, in response to the determination, increase an operation rate of the memory scrubber circuit.
[0069]
[0074] Example 12 includes the space-based processing system of Example 11, wherein the microcontroller comprises a local memory, the microcontroller stores copies of configuration registers of the controllers associated with the first memory, the second memory, and the third memory in the local memory, and is configured to restore the configuration registers based on the stored copies in response to determining that a memory error is associated with a SEFI condition.
[0070]
[0075] Example 13 includes the space-based processing system of Example 11 or 12, wherein the microcontroller is configured to power cycle one or more of the first memory, the second memory, and the third memory in response to determining that the memory error is associated with a SEFI condition.
[0071]
[0076] Example 14 includes the space-based processing system of Examples 8-13, wherein the memory scrubber circuit is configured to cancel the overwrite in response to detecting that a functional write is being performed at a memory address associated with a memory error.
[0072]
[0077] Example 15 is a method for providing a radiation-hardened memory, the method comprising: detecting, by a redundancy comparator, differences between data redundantly stored in a first memory, a second memory, and a third memory; identifying a memory error based on the detected differences by the redundancy comparator; storing, by an error collection buffer, a memory address associated with the memory error; generating, by a memory scrubber circuit, corrected data based on a majority vote performed among the first memory, the second memory, and the third memory; overwriting, by the memory scrubber circuit, the erroneous data with the corrected data at the memory address associated with the memory error; and arbitrating, by a priority arbiter, between the overwrite performed by the memory scrubber and functional memory accesses associated with software execution performed by a processor configured to utilize the memory system.
[0073]
[0078] Example 16 includes the method of example 15, further comprising setting a priority for the priority arbiter based on a trade-off between overwrites performed by the memory scrubber and functional memory accesses.
[0074]
[0079] Example 17 includes the method of example 15 or 16, further comprising determining that the memory error is associated with a single event functional interrupt (SEFI) condition and, in response to the determination, increasing an operation rate of the memory scrubber circuit.
[0075]
[0080] Example 18 includes the method of Example 17, further comprising storing copies of configuration registers of the controllers associated with the first memory, the second memory, and the third memory in a local memory, and in response to determining that the memory error is associated with a SEFI condition, restoring the configuration registers based on the stored copies.
[0076]
[0081] Example 19 includes the method of Example 17 or 18, further comprising power cycling one or more of the first memory, the second memory, and the third memory in response to determining that the memory error is associated with a SEFI condition.
[0077]
[0082] Example 20 includes the method of any one of Examples 15-19, further comprising canceling the overwrite in response to detecting that a functional write is being performed at a memory address associated with a memory error.
[0078]
[0083] The terms and expressions used herein are used as terms of description and not of limitation, and the use of such terms and expressions is not intended to exclude any equivalents of the shown and described features (or portions thereof), recognizing that various modifications are possible within the scope of the claims. Accordingly, the claims are intended to cover all such equivalents. Various features, aspects, and embodiments have been described herein. The features, aspects, and embodiments are susceptible to combination with one another, as well as variations and modifications, as will be appreciated in light of this disclosure. The present disclosure should therefore be considered to embrace such combinations, variations, and modifications. It is intended that the scope of the present disclosure be limited not by this Detailed Description, but rather by the claims appended hereto. Future applications claiming priority to this application may claim the disclosed subject matter differently and may generally include any set of one or more elements as variously disclosed or otherwise illustrated herein.
Claims
1. 1. A memory system comprising: a redundancy comparator configured to detect differences between data redundantly stored in the first memory, the second memory, and the third memory, wherein the redundancy comparator is further configured to identify a memory error based on the detected differences; an error collection buffer configured to store memory addresses associated with the memory errors; a memory scrubber circuit configured to overwrite erroneous data with corrected data at the memory address associated with the memory error, wherein the corrected data is generated based on a majority vote performed among the first memory, the second memory, and the third memory; a priority arbiter configured to arbitrate between the overwrites performed by the memory scrubber and functional memory accesses associated with software execution performed by a processor configured to utilize the memory system; A memory system comprising:
2. 10. The memory system of claim 1, further comprising a microcontroller configured to monitor the error collection buffer and to trigger operation of the memory scrubber circuit in response to the memory error.
3. 3. The memory system of claim 2, wherein the microcontroller is configured to set a priority for the priority arbiter based on a trade-off between the overwrites performed by the memory scrubber and the functional memory accesses.
4. 3. The memory system of claim 2, wherein the microcontroller is configured to determine that the memory error is associated with a single event functional interrupt (SEFI) condition and to increase an operation rate of the memory scrubber circuit in response to the determination.
5. 5. The memory system of claim 4, wherein the microcontroller comprises a local memory, the microcontroller configured to store copies of configuration registers of controllers associated with the first memory, the second memory, and the third memory in the local memory, and to restore the configuration registers based on the stored copies in response to the determination that the memory error is associated with a SEFI condition.
6. 5. The memory system of claim 4, wherein the microcontroller is configured to power cycle one or more of the first memory, the second memory, and the third memory in response to the determining that the memory error is associated with a SEFI condition.
7. 2. The memory system of claim 1, wherein the memory scrubber circuitry is configured to cancel the overwrite in response to detecting that a functional write is being performed at the memory address associated with the memory error.
8. 1. A space-based processing system comprising: a processor configured to execute mission software; Memory system and The memory system comprises: a redundancy comparator configured to detect differences between data redundantly stored in the first memory, the second memory, and the third memory, wherein the redundancy comparator is further configured to identify a memory error based on the detected differences; an error collection buffer configured to store memory addresses associated with the memory errors; a memory scrubber circuit configured to overwrite erroneous data with corrected data at the memory address associated with the memory error, wherein the corrected data is generated based on a majority vote performed among the first memory, the second memory, and the third memory; a priority arbiter configured to arbitrate between the overwrites performed by the memory scrubber and functional memory accesses associated with execution of the mission software; A space-based processing system comprising:
9. 9. The space-based processing system of claim 8, wherein the memory system comprises a microcontroller configured to monitor the error collection buffer and trigger operation of the memory scrubber circuit in response to the memory error.
10. 10. The space-based processing system of claim 9, wherein the microcontroller is configured to set a priority for the priority arbiter based on a trade-off between the overwrites performed by the memory scrubber and the functional memory accesses.
11. 10. The space-based processing system of claim 9, wherein the microcontroller is configured to determine that the memory error is associated with a single event functional interrupt (SEFI) condition and, in response to the determination, increase an operation rate of the memory scrubber circuit.
12. 12. The space-based processing system of claim 11, wherein the microcontroller comprises a local memory, the microcontroller configured to store copies of configuration registers of controllers associated with the first memory, the second memory, and the third memory in the local memory, and to restore the configuration registers based on the stored copies in response to the determination that the memory error is associated with a SEFI condition.
13. 12. The space-based processing system of claim 11, wherein the microcontroller is configured to power cycle one or more of the first memory, the second memory, and the third memory in response to the determination that the memory error is associated with a SEFI condition.
14. 9. The space-based processing system of claim 8, wherein the memory scrubber circuit is configured to cancel the overwrite in response to detecting that a functional write is being performed at the memory address associated with the memory error.
15. 1. A method for providing a radiation-hardened memory, the method comprising: detecting, by a redundancy comparator, a difference between the data redundantly stored in the first memory, the second memory, and the third memory; identifying, by the redundancy comparator, a memory error based on the detected difference; storing, by an error collection buffer, a memory address associated with the memory error; generating corrected data based on a majority vote performed among the first memory, the second memory, and the third memory by a memory scrubber circuit; overwriting, by the memory scrubber circuit, the erroneous data at the memory address associated with the memory error with the corrected data; arbitrating, by a priority arbiter, between the overwrites performed by the memory scrubber and functional memory accesses associated with software execution performed by a processor configured to utilize the memory system; A method comprising:
16. 16. The method of claim 15, further comprising setting a priority for the priority arbiter based on a trade-off between the overwrites performed by the memory scrubber and the functional memory accesses.
17. 16. The method of claim 15, further comprising: determining that the memory error is associated with a single event functional interrupt (SEFI) condition; and increasing an operation rate of the memory scrubber circuit in response to the determination.
18. 18. The method of claim 17, further comprising: storing copies of configuration registers of controllers associated with the first memory, the second memory, and the third memory in a local memory; and, in response to determining that the memory error is associated with a SEFI condition, restoring the configuration registers based on the stored copies.
19. 20. The method of claim 17, further comprising power cycling one or more of the first memory, the second memory, and the third memory in response to the determining that the memory error is associated with a SEFI condition.
20. 16. The method of claim 15, further comprising canceling the overwrite in response to detecting that a functional write is being performed at the memory address associated with the memory error.
Citation Information
Patent Citations
Memory error correction system
JP1992119442A
Satellite line connector
JP1998143445A
Refresh scheme in a memory controller
US10593391B2
Adaptive memory scrub rate
US20120284575A1
Method and apparatus for scrubbing memory
US7913147B2