Error management for memory devices
By introducing timer circuitry and independent error management components for logic gates into the memory subsystem, the problem of memory subsystem failure due to errors was solved, ensuring the stability and safety of autonomous devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MICRON TECHNOLOGY INC
- Filing Date
- 2025-08-11
- Publication Date
- 2026-04-21
AI Technical Summary
The memory subsystem is prone to failure due to errors in autonomous devices, causing the host to receive potentially erroneous data, which affects the accuracy and security of decision-making.
An independent error management component, including a timer circuit system and logic gates, is used to trigger a signal interrupt when a specific type of error is detected, preventing erroneous data from being transmitted to the host and ensuring the stability and security of the memory subsystem.
It effectively prevents the host from receiving erroneous data due to errors in the memory subsystem, improves the decision-making accuracy and security of autonomous devices, and reduces adverse results caused by errors.
Smart Images

Figure CN121901005A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure generally relate to memory systems and subsystems, and more specifically, to error management for memory devices. Background Technology
[0002] The memory subsystem may include one or more memory devices for storing data. For example, the memory devices may be non-volatile memory devices and volatile memory devices. Generally, the host system can utilize the memory subsystem to store data at the memory devices and retrieve data from the memory devices.
[0003] Vehicles are increasingly reliant on memory subsystems to provide storage for previously mechanical, stand-alone, or non-existent components. Vehicles may contain computing systems, which can serve as the host for the memory subsystems. The computing systems can run applications that provide the functionality of the components. Vehicles can be driver-operated, driverless (autonomous), and / or partially autonomous. Memory devices can be extensively used by the computing systems within vehicles. Summary of the Invention
[0004] One aspect of the invention relates to a device comprising: a controller configured to independently manage a first type of error indication and a second type of error indication, wherein each error indication indicates an error in data received from a corresponding data source or a failure of the corresponding data source, or any combination thereof; wherein the controller includes a timer circuitry and is further configured to: receive the second type of error indication; and selectively route the second type of error indication to the timer circuitry such that the timer circuitry triggers a signal interruption from the device if the timer circuitry does not receive a reset signal at the timer circuitry for a specific time period, wherein a timeout signal causes the controller to enter a reduced power state to prevent one or more errors associated with the second type of error indication from being transmitted outside the device.
[0005] Another aspect of the invention relates to a method comprising: receiving an error indication of a first type or a second type, wherein the error indication is used to indicate an error in data received from a corresponding data source or a failure of the corresponding data source, or any combination thereof; in response to the error indication being of the first type, routing the error indication to a processing resource of a memory subsystem such that the processing resource manages the error indication; and in response to the error indication being of the second type, routing the error indication to a timer circuitry of the memory subsystem such that the timer circuitry replaces the processing resource and triggers a signal interruption of the memory subsystem in the event that no reset signal is received at the timer circuitry, wherein the signal interruption of the memory subsystem prevents data associated with the error indication from being transmitted outside the memory subsystem.
[0006] Another aspect of the invention relates to a device comprising: a logic gate; and a timer circuit system coupled to the logic gate, the timer circuit system being configured to: receive a signal indicating a specific type of error indication among a plurality of types, wherein each of the plurality of types of the error indication is for an error in data from a corresponding data source, or a malfunction in the corresponding data source, or any combination thereof; and, if no reset signal is received at the timer circuit system within a specific time period, provide a timeout signal to the logic gate, wherein the reset signal resets the timer circuit system to prevent the timeout period of the timer circuit system from expiring; and the logic gate is configured to: in response to receiving the timeout signal from the timer circuit system, output a trigger signal to trigger a signal interruption of the device to prevent data associated with the error indication from being transmitted from the device. Attached Figure Description
[0007] This disclosure will be more fully understood from the detailed description given below and from the accompanying drawings of various embodiments thereof.
[0008] Figure 1 The description includes examples of computing systems that include memory subsystems operating according to some embodiments of the present disclosure.
[0009] Figure 2 Examples of error management components that manage errors in association with an operating computing system, according to some embodiments of this disclosure, are described.
[0010] Figure 3 The description includes examples of computing systems that include a memory subsystem controller having an error management component that operates according to some embodiments of the present disclosure.
[0011] Figure 4This is a flowchart of an example method for managing errors in association with an operating computing system, according to some embodiments of the present disclosure.
[0012] Figure 5 Examples of systems including a computing system in a vehicle are described according to some embodiments of the present disclosure. Detailed Implementation
[0013] This disclosure relates to error management for memory devices, such as those in automotive scenarios (e.g., autonomous vehicles). The memory subsystem can be a storage system, a storage device, a memory module, or a combination thereof. Examples of memory subsystems are storage systems, such as solid-state drives (SSDs), universal flash memory (UFS) drives, etc. The following is combined with… Figure 1 Describe examples of storage devices and memory modules. Generally, a host system may utilize a memory subsystem that includes one or more components, such as a memory device for storing data. The host system can provide data to be stored in the memory subsystem and can request data to be retrieved from the memory subsystem. As an example, a vehicle may include a memory subsystem, such as an SSD, UFS, etc. The memory subsystem may be used for data storage by various components of the vehicle, such as applications running on the vehicle's host system.
[0014] Autonomous devices (autonomous vehicles, drones, vacuum cleaners, industrial robots, medical robots, etc.) can autonomously make decisions and perform specific tasks based on various inputs. These inputs can be obtained from various sources, such as various sensors, data networks, user input, pre-loaded data, external inputs, etc. These inputs collectively enable the autonomous device to analyze its environment, make decisions, and operate as intended with minimal human intervention. Due to the nature of the domains where autonomous devices are used, the accuracy and safety of certain types of inputs are crucial. For example, in situations where autonomous decision-making is safety-related, erroneous or incorrect inputs (e.g., data) can lead to adverse outcomes that could endanger individuals. Memory devices may contain components (e.g., CPU, firmware, etc.) that detect errors and report them to the host to alert it. However, due to errors, these components themselves can often be prone to failure, which could interrupt the host's ability to act as a decision-making entity.
[0015] This disclosure addresses the aforementioned and other problems by providing a means for independently managing errors that may be particularly prone to occur in memory devices (e.g., components primarily handling errors). For example, various embodiments of this disclosure provide a hardware component (e.g., a circuit system) capable of independently handling errors and terminating (alternatively referred to as "dropping") communication with the host upon detecting such an error. This ensures that the memory device's ability to manage errors is not interrupted by the aforementioned errors, thereby preventing the host from receiving potentially erroneous data and, consequently, from making decisions based on unreliable information.
[0016] The figures in this document follow a numbering convention, where the first one or a few digits correspond to the figure number, and the remaining digits identify the elements or components within the figure. Similar elements or components between different figures can be identified by using similar digits. For example, 112 could refer to... Figure 1 Component "12" in the text, and similar components in Figure 2 This can be referred to as 212. Similar elements within the diagram can be referenced using hyphens and additional numbers or letters. Such similar elements can be generally referenced without hyphens and additional numbers or letters. For example, Figure 2 Elements 222-1, 222-2, ..., 222-N in the figures can be collectively referred to as 222. As used herein, the indicators “N,” “M,” or “X,” especially with respect to reference numerals in the figures, indicate that several specific features may be included. It should be understood that elements shown in the various embodiments herein may be added, interchanged, and / or eliminated to provide several additional embodiments of this disclosure. Furthermore, it should be understood that the scale and relative dimensions of the elements provided in the figures are intended to illustrate a particular embodiment of the invention and should not be considered as intended to be limiting.
[0017] Figure 1 This description describes an example computing system 100 that includes a memory subsystem 104 (alternatively referred to as memory device 104) and operates according to some embodiments of this disclosure. The computing system 100 may be a computing device, such as a desktop computer, laptop computer, web server, mobile device, vehicle (e.g., an airplane, drone, train, car, or other means of transport), device with Internet of Things (IoT) capabilities, embedded computer (e.g., an embedded computer included in a vehicle, industrial equipment, or networked business device), or such a computing device including memory and processing means.
[0018] The computing system 100 includes a host system 102 coupled to one or more memory subsystems 104. For example, the host system 102 may be a computing system included in a vehicle, and the computing system may run applications that provide component functionality for the vehicle. In some embodiments, the host system 102 is coupled to different types of memory subsystems 104. Figure 1This describes an example of a host system 102 coupled to a memory subsystem 104. As used herein, “coupled to” or “coupled with” generally refers to a connection between components, which can be an indirect or direct communication connection (e.g., without an intermediary component), whether wired or wireless, including, for example, electrical, optical, magnetic, and similar connections.
[0019] Host system 102 includes or is coupled to processing resources, memory resources, and network resources. As used herein, a “resource” is a physical or virtual component within computing system 100 with limited availability. For example, processing resources include processing devices, memory resources include a memory subsystem 104 for secondary storage and a main memory device (not specifically described) for primary storage, and network resources include network interfaces (not specifically described). A processing device may be one or more processor chipsets that can execute software stacks. A processing device may include one or more cores, one or more caches, a memory controller (e.g., an NVDIMM controller), and a storage protocol controller (e.g., a PCIe controller, a SATA controller). For example, host system 102 uses memory subsystem 104 to write data to and read data from memory subsystem 104.
[0020] The host system 102 may run one or more applications. For example, the applications may run on an operating system (not specifically described) executed by the host system 102. An operating system is system software that manages computer hardware and software resources and provides public services to applications. An application is a set of instructions that can be executed to perform a specific task. For example, an application may be a black box application for a vehicle, but the embodiments are not limited thereto.
[0021] Host system 102 can be coupled to memory subsystem 104 via a physical host interface. Examples of physical host interfaces include, but are not limited to, Serial Advanced Technology Attachment (SATA) interfaces, PCIe interfaces, Universal Serial Bus (USB) interfaces, Fibre Channel, Serial Attached SCSI (SAS), Small Computer System Interface (SCSI), Double Data Rate (DDR) memory bus, Dual In-line Memory Module (DIMM) interfaces (e.g., DIMM slot interfaces supporting Double Data Rate (DDR)), Open NAND Flash Interface (ONFI), Double Data Rate (DDR), Low Power Double Data Rate (LPDDR), or any other interface. The physical host interface can be used to transfer data between host system 102 and memory subsystem 104. Host system 102 may further utilize an NVM Fast (NVMe) interface to access non-volatile memory device 116 when memory subsystem 104 is coupled to host system 102 via a PCIe interface. The physical host interface provides an interface for transmitting control, address, data and other signals between the memory subsystem 104 and the host system 102. Figure 1 The memory subsystem 104 is illustrated as an example. Generally, the host system 102 can access multiple memory subsystems via the same communication connection, multiple individual communication connections, and / or combinations of communication connections.
[0022] For example, host system 102 can control memory subsystem 104 and / or send requests (e.g., commands) to memory subsystem 104 to store data in or read data from memory subsystem 104. For example, host system 102 can use memory subsystem 104 to provide storage for black-box applications. Data to be written or read, as specified by a host request, is referred to as "host data." The host request may contain logical address information. Logical address information may be a logical block address (LBA), which may include or be accompanied by a partition number. Logical address information is the location associated between the host system and the host data. Logical address information may be part of the metadata of the host data. The LBA may correspond to (e.g., dynamically mapped to) a physical address, such as a physical block address (PBA), which indicates the physical location in memory where the host data is stored.
[0023] The memory subsystem 104 may include media, such as one or more volatile memory devices 115, one or more non-volatile memory devices 116, or a combination thereof. The volatile memory device 115 may be, but is not limited to, random access memory (RAM), such as dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), and resistive DRAM (RDRAM).
[0024] The memory subsystem 104 may be a storage device, a memory module, or a hybrid of a storage device and a memory module. Examples of storage devices include SSDs, flash drives, Universal Serial Bus (USB) flash drives, embedded multimedia controller (eMMC) drives, universal flash memory (UFS) drives, secure digital cards (SD cards), and hard disk drives (HDDs). Examples of memory modules include dual in-line memory modules (DIMMs), small form factor DIMMs (SO-DIMMs), and various types of non-volatile dual in-line memory modules (NVDIMMs).
[0025] Examples of non-volatile memory devices 116 include NAND flash memory. NAND flash memory includes, for example, two-dimensional NAND (2D NAND) and three-dimensional NAND (3D NAND). Non-volatile memory devices 116 can be other types of non-volatile memory, such as read-only memory (ROM), phase-change memory (PCM), selectable memory, other chalcogenide-based memories, ferroelectric transistor random access memory (FeTRAM), ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), spin-transfer torque (STT)-MRAM, conductive bridged RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), NOR flash memory, electrically erasable programmable read-only memory (EEPROM), and three-dimensional cross-point memory. Cross-point non-volatile memory arrays can perform bit storage based on volume resistance variations combined with stacked cross-network format data access arrays. In addition, compared to many flash-based memories, cross-point non-volatile memories can perform in-situ write operations, where non-volatile memory cells can be programmed without first erasing the non-volatile memory cells.
[0026] The memory subsystem controller 106 (or, for simplicity, controller 106) can communicate with memory devices 115, 116 to perform operations such as reading data, writing data, erasing data, and other such operations at memory devices 115, 116. The memory subsystem controller 106 may include hardware such as one or more integrated circuits and / or discrete components or combinations thereof. The hardware may include a digital circuit system having dedicated (i.e., hard-coded) logic for performing the operations described herein. The memory subsystem controller 106 may be a microcontroller, a dedicated logic circuit system (e.g., a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc.), or other suitable circuit system.
[0027] The memory subsystem controller 106 may include a processing means 108 (e.g., a processor, which may be a central processing unit (CPU)) configured to execute instructions stored in local memory 110. Local memory 110 may be, for example, static random access memory (SRAM). In the illustrated example, the local memory 110 of the memory subsystem controller 106 is an embedded memory configured to store instructions for performing various processes, operations, logical flows, and routines for controlling the operations of the memory subsystem 104 (including handling communication between the memory subsystem 104 and the host system 102). For example, local memory 110 may store instructions executable by processor 108 and / or operating components 114, as will be further described herein. As used herein, "processor" may alternatively be referred to as "processing resource".
[0028] In some embodiments, local memory 110 may include memory registers for storing memory pointers, fetch data, etc. For example, local memory 110 may also include ROM for storing microcode. Although Figure 1 The instance memory subsystem 104 has been described as including a memory subsystem controller 106, but in another embodiment of this disclosure, the memory subsystem 104 may not include a memory subsystem controller 106, but may instead rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem 104). In some embodiments, the memory subsystem 104 may be a managed NAND (MNAND) device, wherein an external controller (e.g., controller 106) is packaged together with one or more NAND dies (e.g., non-volatile memory device 116).
[0029] Generally, the memory subsystem controller 106 can receive information or operations from the host system 102 and can translate the information or operations into instructions or appropriate information to achieve the desired access to the non-volatile memory device 116 and / or the volatile memory device 115. The memory subsystem controller 106 may be responsible for other operations, such as wear leveling operations, error detection and / or correction operations, encryption operations, caching operations, and address translation between logical addresses (e.g., logical block addresses) and physical addresses (e.g., physical block addresses) associated with the non-volatile memory device 116. The memory subsystem controller 106 may further include a host interface circuitry for communicating with the host system 102 via a physical host interface. The host interface circuitry can translate queries received from the host system 102 into commands to access the non-volatile memory device 116 and / or the volatile memory device 115, and translate responses associated with the non-volatile memory device 116 and / or the volatile memory device 115 into information for the host system 102.
[0030] like Figure 1As shown, the memory subsystem 104 may include an error management component 112 and an operation component 114. Although Figure 1 While not shown in the diagrams to avoid obscuring the schematics, the error management component 112 may include various circuitry systems to facilitate the aspects of this disclosure described herein. In some embodiments, the error management component 112 and / or the operation component 114 may include firmware, dedicated circuitry systems in the form of ASICs, FPGAs, state machines, hardware processing devices, and / or other logic circuitry systems that allow the error management component 112 and / or the operation component 114 to orchestrate and / or perform the operations described herein.
[0031] Operation component 114 (which may be firmware or hardware or any combination thereof) manages and / or controls the operation of computing system 100 (e.g., autonomous device). In some embodiments, operation component 114 may be a portion (e.g., an integrated portion) of processor 108 (e.g., CPU).
[0032] Operation component 114 allows host 102 to operate autonomously or in autonomous mode (alternatively referred to as "task mode") to analyze the environment based on various inputs, make decisions, and operate as intended with minimal human intervention. Furthermore, operation component 114 ensures that the operation of computing system 100 meets safety standards, such as those defined by safety standard ISO 26262. These functionalities provided by operation component 114 to meet the requirements of autonomous mode and / or safety standards may include error detection, error correction, or error notification to host 102, and others.
[0033] Operation component 114 may be the “primary” error management entity for handling or managing errors in computing system 100 (or at least memory subsystem 104). However, some errors that may be particularly prone to occur in operation component 114 (and / or processor 108) (causing the error management / notification capabilities of operation component 114 itself to become ineffective due to errors) may be managed independently at error management component 112 (which may be a “secondary” error management entity).
[0034] Therefore, errors in various components of the computing system 100 that could adversely affect the capabilities of the operational component 114 can instead be managed (e.g., configured) at the error management component 112 rather than the operational component 114. For example, the controller 106 can route error indications of those types of errors to the error management component 112 instead of the operational component 114. The error management component 112 can prevent received error indications or erroneous / incorrect data associated with these indications from being further provided (e.g., reported) to the host 102. This provides an auxiliary and independent way to handle errors (without the error management component 112) that would otherwise disrupt the functionality of the memory subsystem 104, which would further allow the memory subsystem 104 to report potentially incorrect or erroneous data to the host 102. Figures 2 to 4 Further details of this process will be provided.
[0035] Figure 2 This section describes an example of an error management component 212 that manages errors (e.g., error messages) in association with an operating computing system, according to some embodiments of this disclosure. The error management component 212 may be similar to... Figure 1 Error management component 112 as described in the document.
[0036] like Figure 2 As described herein, error management component 212 may include timer circuitry 224 (referred to simply as "timer"), which may be called timer circuitry 224. In some embodiments, timer circuitry 224 may be a watchdog timer circuitry (WDT), but embodiments are not limited thereto. As used herein, the term WDT refers to a dedicated timer circuitry configured to reset the system (e.g., memory subsystem 100) or enable the system to reset if it is not reset upon the expiration of a timeout period (e.g., timeout interval).
[0037] The timer circuit system 224 can receive error indications 222-1, 222-2, and 222-N (collectively referred to as error indications 222, or simply indications 222) from various "sources". The timer circuit system 224 may include a memory (e.g., RAM, flash memory, EEPROM, non-volatile RAM (NVRAM), SD card) capable of storing (e.g., at least temporarily) the error indications 222. The error indications 222 (alternatively referred to as "error messages", "error notifications", or the like) may be error indications of the corresponding source, such as errors in data received from the corresponding source and / or malfunctions of the corresponding source.
[0038] Various operational problems (such as errors) can be reported in error indication 222. Operational problems may include, but are not limited to, errors such as: data corruption, memory exhaustion, performance degradation, hardware failure, power loss, temperature-related problems (such as temperature sensor malfunction and / or incorrect / erroneous readings), clock failure, computing system (e.g., Figure 1 Voltage faults in the computing system 100 and / or memory subsystem (e.g., memory subsystem 104) described herein (e.g., incorrect or faulty voltage levels of one or more voltage regulators (e.g., low dropout regulators (LDOs), AC / DC converters, DC / DC buck converters, switched capacitors, etc.)), CPU reset faults (e.g., in the memory subsystem (e.g., Figure 1 During the initialization process of the memory subsystem 104 described herein, issues such as bad memory blocks, firmware errors, connectivity problems, and security vulnerabilities may occur.
[0039] Error indications (e.g., indication 222) associated with these operational problems may be generated at corresponding sources, detectors, etc., such as temperature sensors configured to monitor the temperature of components of computing system 100, voltage sensors configured to monitor the output voltage level of a voltage regulator, clock monitors configured to monitor the integrity of clock signals, current sensors configured to monitor the current flowing through various components or circuits of computing system 100, and error correction code (ECC) components configured to detect errors in data using CRC, parity check data, etc. In some embodiments, error indication 222 may be managed as an "asynchronous event" independent of the timing requirements of data signals, clock signals, etc., which allows the signal indicating error indication 222 to be used without the need for a timing synchronization mechanism.
[0040] Error indications associated with some of those operational problems can be configured (e.g., pre-configured during the initialization phase (e.g., boot phase) of memory subsystem 104) to be routed to timer circuitry system 224, rather than to a different entity, e.g. Figure 1 The operating component 114 is described herein. For example, these instructions may be associated with incorrect or erroneous data that would interrupt the operation of the operating component 114.
[0041] More specifically, the error indication 222 that can be directly routed to the timer circuitry 224 may include an error indication of a clock failure (e.g., sent from a clock signal generator and / or detector), an error indication from a temperature sensor, an error indication of a voltage failure (e.g., detected by a detector configured to monitor the voltage level of a voltage regulator), and an error indication of a CPU reset failure, but embodiments are not limited thereto. As further illustrated herein, these error indications 222, when received by the timer circuitry 224, typically cause a trigger signal interruption (alternatively referred to as a “link interruption”) to prevent erroneous and / or incorrect data (that triggered the error indication 222) from being transmitted to the host 102.
[0042] like Figure 2 As described, the timer circuit system 224 can output an "ERROR_OUT" signal (alternatively referred to as a "timeout signal") in response to the expiration of a "timeout period". This is in addition to the output from the controller 106 (e.g., ...). Figure 1 In the event of a reset signal provided by the operating component 114 and / or processor 108 as described herein, the "timeout period" of the timer circuitry 224 may expire. For example, during operation of the computing system 100 and / or memory subsystem 104, the controller 106 (e.g., operating component 114) may reset the timer circuitry 224 before the timeout period expires (e.g., periodically). This may continue unless a failure of the memory subsystem 104 (and / or operating component 114) prevents a reset signal from being provided to the timer circuitry 224. While embodiments are not limited thereto, the timeout value may be configured by the operating component 114, for example, during the initialization phase (e.g., boot phase) of the computing system 100 and / or memory subsystem 104.
[0043] An "ERROR_OUT" signal can be provided to logic gate 228. Although the embodiment is not limited to this, logic gate 228 can be an OR gate. Logic gate 228 receives two input signals: one input signal is the "ERROR_OUT" signal from timer circuitry 224, and the other input signal is a "RESET" signal, such as... Figure 2 As shown in the image.
[0044] The output signal from logic gate 228 (alternatively referred to as the "trigger signal") can trigger a signal interrupt based on its value (e.g., the logic value of the signal). For example, assuming that input signal 226-2 has been driven "high" regardless of whether the timeout period (e.g., that of timer circuitry 224) has expired, then when the timeout period of timer circuitry 224 has expired, the "ERROR_OUT" signal 226-1 can be asserted. This further asserts the output signal from logic gate 228, which can trigger a signal interrupt.
[0045] As used herein, the term "signal interruption" refers to the loss (e.g., intentional loss) of communication between two entities (e.g., between host 102 and memory subsystem 104). For example, a "signal interruption" can be achieved by resetting memory subsystem 104 or placing memory subsystem 104 into a low-power state (e.g., inactive, power-sleep, or power-off). Figure 2 As described, the operation of error management component 212 primarily involves simplified hardware components that are less prone to errors (e.g., timer circuitry 224, logic gates 228, etc.). These error indications can be routed to error management component 212. In contrast, operating component 114, which relies on more complex firmware and CPU-based operations, is more susceptible to such errors. Therefore, managing these errors independently at error management component 212, which handles simplified hardware components, prevents firmware or CPU failures from causing controller 106 to fail to filter errors before they are provided to host 102.
[0046] A signal interrupt triggered by the timer circuitry 224 prevents the host 102 from participating in the decision-making process based on erroneous or incorrect data that would otherwise be provided by the controller 106 (e.g., attributed to a malfunction of the operating component 114). This helps avoid security risks, especially when the host 102 relies on real-time data obtained from the memory subsystem 104 for its decision-making process. By ensuring that only accurate data is used, particularly within short time frames, the overall security and reliability of the computing system 100 are improved.
[0047] The timer circuit system 224 may further provide information associated with error indications (e.g., one of the indications 222 received at the timer circuit system 224) via a communication channel 223 (e.g., a sideband channel). The communication channel 223 may include one or more pins, such as general purpose input / output (GPIO) pins. In some embodiments, in addition to the primary communication channel, the communication channel may also be an auxiliary communication channel (e.g., a sideband channel). The communication channel as a sideband channel can operate in parallel with the primary communication channel, which improves the control of the host 102 and / or the vehicle control system (e.g., Figure 3 The response time of the vehicle control system 318 described herein.
[0048] Communication channel 223 can be used as a means to convey additional details / information about error indication 222 (which is stored in the timer circuitry 224) outside of timer 224 (e.g., to a host computer). This additional details / information may include the source of the error (e.g., from a temperature sensor), the type of error detected, the severity of the error, the timestamp of the error occurrence, etc. For example, host computer 102 may detect a “signal interruption” after detecting a lack of communication from controller 106 for a specific time period. In this case, host computer 102 (e.g., an automotive system application) that is constantly monitoring controller 106 can request or poll details of indication 222 from controller 106.
[0049] Figure 3 The description includes an example of a computing system 300 having a memory subsystem controller 306 with an error management component 312 operating according to some embodiments of the present disclosure. The computing system 300, the memory subsystem controller 306 (referred to simply as controller 306), the error management component 312, and the hosts 302-1, ..., 302-M (collectively referred to as hosts 302) may be respectively similar to... Figure 1 The computing system 100, memory subsystem controller 106, error management component 112 and host 102 described herein.
[0050] like Figure 3 As described, the host 302 and controller 306 can be further coupled to the vehicle control component 318. For example, when operating the vehicle autonomously or partially autonomously, the vehicle control component 318 can manage the physical controls of one or more vehicles (based on requests, commands, etc. received from the host 302 or input data received from the controller 306). For example, the physical controls of the vehicle that can be managed by the vehicle control component 318 may include start control for switching ignition or controlling the starting of the vehicle; steering or steering mechanism for turning the steering wheel or controlling the steering of the vehicle; the vehicle's route or direction; throttle control for increasing or decreasing the throttle or accelerator or controlling the speed of the vehicle, thus changing the speed of the vehicle; pressing or releasing the brakes; turning the turn signals on / off; controlling the lights on the vehicle (e.g., by turning the headlights, parking brake, fog lights, etc. on / off); activating warning signals (e.g., horn; hazard lights); locking or unlocking doors; activating windshield wipers; parking sensors or controls; and / or changing the gears of the vehicle, etc.
[0051] The host 302 may be based on a slave memory subsystem (e.g. Figure 1 The memory subsystem 104 described herein) and / or various sensors (e.g. Figure 5The vehicle control system 318 operates based on inputs provided by the sensors 544 described herein. The vehicle control system 318 can also operate to provide the functionality described herein based on requests, commands, etc., provided from the host 302 and / or inputs provided from the memory subsystem 104 and / or the various sensors 544. This includes inputs related to errors (e.g., error type), error sources, and / or their indications (e.g., errors). Figure 2 The information associated with the instruction 222 described herein may be transmitted from the controller 306 to at least one of the host 302 and / or the vehicle control system 318, for example, via one or more pins (e.g., GPIO pins).
[0052] Figure 4 It is based on some embodiments of this disclosure for managing and operating a computing system (e.g. Figure 1 The flowchart illustrates an example method 430 for resolving errors associated with a computing system 100. The method can be executed by processing logic, which may include hardware (e.g., processing device, circuitry, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions that run or execute on the processing device), or a combination thereof. In some embodiments, the error is caused by or uses... Figure 1 The illustrated memory subsystem controller 106 executes the method. Although shown in a specific sequence or order, the order of processes may be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes may be executed in different orders, and some processes may be executed in parallel. In addition, one or more processes may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.
[0053] At 432, method 430 may include receiving an error indication of a first type or a second type, wherein the error indication is used to indicate an error in data received from a corresponding data source or a failure of the corresponding data source, or any combination thereof. At 434, method 430 may further include routing the error indication to a memory subsystem (e.g., [missing information]) in response to the error indication being of the first type. Figure 1 The processing resources (e.g., memory subsystem 104) described herein Figure 1 The processing resource 108 described herein is used to indicate a processing resource management error.
[0054] At 436, method 430 may further include responding to an error indication of a second type (e.g., Figure 2 The error indication is routed to the timer circuitry of the memory subsystem 104 (e.g., the indication 222 described herein) instead of the error indication. Figure 2The timer 224 described herein causes the timer circuitry 224 to replace the processing resource 108 and triggers a signal interrupt of the memory subsystem 104 if no reset signal is received at the timer circuitry 224. The signal interrupt of the memory subsystem 104 prevents data associated with error indication 222 from being transferred outside the memory subsystem 104. The signal interrupt of the memory subsystem 104 may include placing the memory subsystem 104 in a low-power state, resetting the memory subsystem, or any combination thereof.
[0055] In some embodiments, information associated with indication 222 may be transmitted (e.g., transmitted to a host) in response to a request for said information. Figure 1 , 3 (and hosts 102, 302, and 502 described in sections 5 respectively). The information may be information about the source of the error, the error type associated with indicator 222, or any combination thereof. Furthermore, the information may be transmitted via general purpose input / output (GPIO) pins.
[0056] Figure 5 This section describes an example of a system 546 comprising a computing system 500 in a vehicle, according to some embodiments of the present disclosure. The computing system 500 may include a memory subsystem 504, which, for simplicity, is described as including a controller 506 and a non-volatile memory device 516, but is similar to... Figure 1 The memory subsystem 104 is described herein. The computing system 500 and therefore the host 502 may be directly coupled to several sensors 544, as described with respect to sensor 544-4, or coupled via transceiver 552 to several sensors 544, as described with respect to sensors 544-1, 544-2, 544-3, 544-5, 544-6, 544-7, 544-8, ..., 544-X (collectively referred to as sensors 544). The transceiver 552 is capable of receiving data from the sensors 544 wirelessly (e.g., via radio frequency communication). In at least one embodiment, each of the sensors 544 may wirelessly communicate with the computing system 500 via transceiver 552. In at least one embodiment, each of the sensors 544 is directly connected to the computing system 500 (e.g., via wires or optical fibers).
[0057] Vehicle 550 may be a car (e.g., a sedan, van, truck, etc.), a connected vehicle (e.g., a vehicle with computing capabilities to communicate with an external server), an autonomous vehicle (e.g., a vehicle with automated capabilities such as self-driving), a drone, an aircraft, a ship, and / or any item used to transport people and / or goods. Sensor 544 in Figure 5The description includes instance attributes. For example, sensors 544-1, 544-2, and 544-3 are cameras that collect data from the front of vehicle 550. Sensors 544-4, 544-5, and 544-6 are microphone sensors that collect data from the front, middle, and rear of vehicle 550. Sensors 544-7, 544-8, and 544-X are cameras that collect data from the rear of vehicle 550. As another example, sensors 544-5 and 544-6 are tire pressure sensors. As another example, sensor 544-4 is a navigation sensor, such as a Global Positioning System (GPS) receiver. As another example, sensor 544-6 is a speedometer. As another example, sensor 544-4 represents several engine sensors, such as a temperature sensor, pressure sensor, voltmeter, ammeter, tachometer, fuel gauge, etc. As another example, sensor 544-4 represents a camera. Video data can be received from any of the sensors 544, including the camera, associated with vehicle 550. In at least one embodiment, the video data may be compressed by the host 502 before being provided to the memory subsystem 504.
[0058] The host computer 502 can execute instructions to provide an overall control system and / or operating system for the vehicle 550. The host computer 502 may be a controller designed to assist in the automated operation of the vehicle 550. For example, the host computer 502 may be an Advanced Driver Assistance System (ADAS) controller. ADAS can monitor data to prevent accidents and provide warnings of potential unsafe situations. For example, ADAS can monitor sensors in the vehicle 550 and gain control over the operation of the vehicle 550 to avoid accidents or injuries (e.g., to avoid accidents in the event of an incapacitated user of the vehicle). The host computer 502 may need to act quickly and make decisions to avoid accidents. The memory subsystem 504 can store reference data in a non-volatile memory device 516, allowing the host computer 502 to compare data from the sensor 544 with the reference data to make rapid decisions.
[0059] Sensor 544 can be composed of one or more detectors ( Figure 5(Not specified) Monitoring, the detector can indicate errors associated with the sensor (e.g., errors in data obtained at sensor 544, sensor 544 malfunction, etc., and others). Although embodiments are not limited thereto, the detector may be embedded in sensor 544 and / or integrated as part of sensor 544. When an error is detected, the detector can generate an error indication and route the error indication to host 502 or controller 506. When the error indication is routed to controller 506 (e.g., to error management components 112, 312), controller 506 typically prevents sensor 544 from providing its measurement / sensing data to host 502 and / or prevents the received error indication itself from being provided (further routed) to host 502. This may occur during timers (e.g. Figure 2 When the "timeout period" of timer 224 described herein has expired, as described herein.
[0060] Some parts of the foregoing detailed description have been presented based on the algorithms and symbolic representations of operations on data bits within computer memory. These algorithmic descriptions and representations are the means by which those skilled in the art of data processing most effectively communicate the essence of their work to others skilled in the art. Algorithms are, and generally are, conceived herein as self-consistent sequences of operations that lead to desired results. Operations are those that require the physical manipulation of physical quantities. Typically, but not necessarily, these quantities take the form of electrical or magnetic signals that can be stored, combined, compared, and otherwise manipulated. It has proven convenient, sometimes primarily for general reasons, to refer to these signals as bits, values, elements, symbols, characters, items, numbers, or the like.
[0061] However, it should be remembered that all these and similar terms should be associated with appropriate physical quantities and are merely convenient labels for application to those quantities. This disclosure may relate to the operation and processes of a computer system or similar electronic computing device that manipulate and transform data representing physical (electronic) quantities in the registers and memories of the computer system into other data similarly represented in the memory or registers of the computer system or other such information storage systems.
[0062] This disclosure also relates to apparatus for performing the operations described herein. Such apparatus may be specifically constructed for its intended purpose, or may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in a computer. This computer program may be stored in a machine-readable storage medium, such as, but not limited to, various types of disks, semiconductor-based memories, magnetic cards or optical cards, or other types of media suitable for storing electronic instructions.
[0063] This disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, the instructions being used to program a computer system (or other electronic device) to perform processes according to this disclosure. The machine-readable medium includes mechanisms for storing information in a form readable by a machine (e.g., a computer).
[0064] In the foregoing description, embodiments of the present disclosure have been described with reference to specific exemplary embodiments. It will be apparent that various modifications can be made to the embodiments of the present disclosure without departing from the broader spirit and scope set forth in the appended claims. Therefore, the description and drawings should be considered illustrative rather than limiting.
Claims
1. A device (100; 500) for error management, comprising: A controller (106; 306; 506) configured to independently manage first-type error indications and second-type error indications (222-1, ..., 222-N), wherein each error indication indicates an error in data received from a corresponding data source or a failure of the corresponding data source, or any combination thereof; The controller includes a timer circuitry (224) and is further configured to: Receive the error indication of the second type; and The error indication of the second type is selectively routed to the timer circuitry so that if the timer circuitry does not receive a reset signal at the timer circuitry for a specific time period, a signal interruption from the device is triggered, wherein the timeout signal causes the controller to enter a reduced power state to prevent one or more errors associated with the error indication of the second type from being transmitted to the outside of the device.
2. The device of claim 1, wherein the controller further includes processing resources (108), wherein the controller is configured to: Receive the error indication of the first type; and The error indication of the first type is selectively routed to the processing resource of the controller, such that the error indication of the first type is managed by the processing source.
3. The device of claim 2, wherein the processing resources are further configured to periodically reset the timer circuitry of the controller to prevent the timeout period of the timer circuitry from expiring.
4. The device of claim 2, wherein the processing resources are further configured to selectively alert the first type of error indication to the host (102; 302-1, 302-2, 302-3, 302-M; 502).
5. The device according to any one of claims 1 to 4, wherein the information associated with the error indication of the first type is accessed by the host (102; 302-1, 302-2, 302-3, 302-M; 502) via a communication channel (223) including general purpose input / output (GPIO) pins.
6. The device according to any one of claims 1 to 4, wherein the error indication of the second type is used to indicate: Clock signal malfunction; Temperature sensor malfunction; The voltage regulator is malfunctioning; or A failure during the CPU reset process in the initialization phase; or any combination thereof.
7. The device according to any one of claims 1 to 4, wherein the signal interruption from said device further results in: The device is in a reduced power state; or The device reset; or any combination thereof.
8. A method for error management, comprising: Receive a first type or a second type of error indication (222-1, ..., 222-N), wherein the error indication is used to indicate an error in the data received from the corresponding data source or a failure of the corresponding data source, or any combination thereof; In response to the error indication being of the first type, the error indication is routed to the memory subsystem (104); 504) processing resources (108) are used to manage the error indication; and In response to the error indication being of the second type, the error indication is routed to the timer circuitry (224) of the memory subsystem to cause the timer circuitry to replace the processing resource and trigger a signal interruption of the memory subsystem if no reset signal is received at the timer circuitry, wherein the signal interruption of the memory subsystem prevents data associated with the error indication from being transmitted outside the memory subsystem.
9. The method of claim 8, further comprising, in response to the reset signal and to prevent data associated with the error indication from being transmitted outside the memory subsystem: Place the memory subsystem in a low-power state; or Reset the memory subsystem.
10. The method according to any one of claims 8 to 9, further comprising: In response to receiving a request for information associated with the error indication, the information is transmitted from the timer circuit system; and The information mentioned includes information associated with the source of the error, the type of the error, or any combination thereof.
11. The method of claim 10, further comprising transmitting the information via a communication channel (223) including general purpose input / output (GPIO) pins.
12. A device (100; 500) for error management, comprising: Logic gates (228); and A timer circuit system (224) coupled to the logic gate, the timer circuit system being configured to: Receive a signal indicating a specific type of error indication (222-1, ..., 222-N) among a plurality of types, wherein each of the plurality of types indicates an error in data from a corresponding data source, or a malfunction in the corresponding data source, or any combination thereof; and If no reset signal is received at the timer circuit system within a specific time period, a timeout signal is provided to the logic gate, wherein the reset signal resets the timer circuit system to prevent the timeout period of the timer circuit system from expiring; and The logic gate is configured as follows: In response to receiving the timeout signal from the timer circuitry, a trigger signal is output to trigger a signal interruption of the device to prevent data associated with the error indication from being transmitted from the device.
13. The apparatus of claim 12, wherein the logic gate is configured to: Receive input signals with the following characteristics: The first input signal corresponds to the timeout signal (226-1); and The second input signal (226-2); and The trigger signal is output when each of the input signals is driven to correspond to the first logic value.
14. The device according to any one of claims 12 to 13, wherein the timer circuitry is configured to prevent the timeout signal from being provided to the logic gate in response to receiving the reset signal within the specified time period.
15. The device according to any one of claims 12 to 13, wherein: The device is an autonomous vehicle (550); and The various types of error indications correspond to errors in the data obtained from the corresponding sensors (544-1, 544-2, 544-3, 544-4, 544-5, 544-6, 544-7, 544-8, 544-X) of the autonomous vehicle, or to the malfunction of the corresponding sensors.
16. The device according to any one of claims 12 to 13, wherein the particular type of error indication is configured to be routed to the timer circuitry during the initialization phase of the device.
17. The device according to any one of claims 12 to 13, wherein the timer circuitry is configured to: Store information associated with the error indication; and In response to receiving a request for information associated with the error indication, the information is transmitted via the sideband channel (223).