Dynamic control of error management and messaging

By configuring error thresholds and determining the error set in the memory device, the error management and detection problems in the memory system are solved, dynamic error management and signal transmission are realized, and data integrity and system reliability are improved.

CN113227980BActive Publication Date: 2025-06-20MICRON TECHNOLOGY INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980085634.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-11
Filing Date
2019-12-12
Publication Date
2025-06-20
Estimated Expiration
2039-12-12

AI Technical Summary

Technical Problem

The prior art is difficult to effectively manage and detect errors in memory systems, resulting in the impact of data integrity and system reliability.

Method used

By configuring an error threshold, the error set in the data retrieved by the memory device is determined, and the indication of the error set is transmitted to the host device based on the threshold, realizing dynamic error management and signal transmission.

Benefits of technology

Improve error detection and correction capabilities in memory systems, ensure data integrity and system reliability, and adapt to error management needs of different applications and operating conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113227980B_ABST
    Figure CN113227980B_ABST
Patent Text Reader

Abstract

This application is directed to dynamic control of error management and messaging. Programmable thresholds can be configured for a memory device based on the type of data or the location of the stored data, among other aspects. For example, a host device can configure the number of error thresholds for data at the memory device. When retrieving the data, the memory device can track or count errors in the data and determine whether the threshold has been met. The memory device can emit an indication of whether the threshold has been met (e.g., emit to the host device), and the system can perform functions to correct the error and / or prevent additional errors. The memory device can also identify errors in received commands or can identify errors introduced into the data after receipt of the data (e.g., using an error detection code associated with the command or bus).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] This patent application claims priority to PCT Application No. PCT / US2019 / 065875, filed on December 12, 2019, by Richter et al. and titled "DYNAMIC CONTROL OF ERROR MANAGEMENT AND SIGNALING", which claims priority to U.S. Patent Application No. 16 / 711,354, filed on December 11, 2019, by Richter et al. and titled "DYNAMIC CONTROL OF ERROR MANAGEMENT AND SIGNALING", and U.S. Provisional Patent Application No. 62 / 779,024, filed on December 13, 2018, by Richter et al. and titled "DYNAMIC CONTROL OF ERROR MANAGEMENT AND SIGNALING", all of which are assigned to the assignee hereof and are hereby incorporated by reference in their entireties. Technical Field

[0003] This technical field relates to the dynamic control of error management and signaling. Background Art

[0004] The following generally relates to systems that include at least one memory device, and more particularly, to the dynamic control of error management and signaling.

[0005] Memory devices are widely used to store information in various electronic devices such as computers, wireless communication devices, cameras, digital displays, and the like. Information is stored by programming different states of the memory device. For example, binary devices most often store one of two states, often represented by a logic 1 or a logic 0. In other devices, more than two states can be stored. To access the stored information, components of the device can read or sense at least one of the stored states in the memory device. To access information, components of the device can write to or program the states in the memory device.

[0006] There are various types of memory devices, including magnetic hard disks, random access memory (RAM), read-only memory (ROM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), ferroelectric RAM (FeRAM), magnetic RAM (MRAM), resistive RAM (RRAM), flash memory, phase change memory (PCM), etc. Memory devices can be volatile or non-volatile. Non-volatile memory, such as FeRAM, can maintain the logical states it stores for a long time, even in the absence of an external power supply. Volatile memory devices (e.g., DRAM) can lose their stored states over time unless they are periodically refreshed by an external power supply.

[0007] The integrity of data received by and stored in a memory device can be affected by various conditions that can cause errors. For example, a memory device can receive data from another device using one or more channels, and transient electronic noise can introduce one or more errors into the received data. In other cases, the data can be affected after it is received at the memory device. Techniques for error detection and correction in data and memory systems may need to be improved. Summary of the Invention

[0008] Describes a method. In some instances, the method can include receiving, from a host device, a configuration of an error threshold for data at a memory device; determining a set of errors in data retrieved from the memory device; determining, at least in part based on the configuration, that the set of errors meets the threshold; and transmitting, at least in part based on determining that the set of errors meets the threshold, an indication that the threshold has been met to the host device.

[0009] Describes a method. In some instances, the method can include receiving, from a host device, an access command to access data at a memory device; determining, at least in part based on an error detection code associated with the access command or a bus of the memory device, whether the access command includes an error, the error detection code including one or more parity bits; and transmitting, at least in part based on the access command and the determination, an indication of whether an error has been detected in the access command.

[0010] Describes a method. In some instances, the method can include determining a configuration of an error threshold for data at a memory device; transmitting, at least in part based on an access operation associated with the data, the configuration of the threshold to the memory device; and receiving, at least in part based on retrieving the data at the memory device, an indication that the threshold has been met from the memory device.

[0011] Describe a method. In some instances, the method may include generating an error detection code associated with an access command for accessing data at a memory device; transmitting the access command including the error detection code to the memory device; and receiving an indication of whether an error is detected in the access command from the memory device, at least in part based on the error detection code. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 Illustrate examples of systems that support dynamic control of error management and signaling as disclosed herein.

[0013] Figure 2 Illustrate examples of memory dies that support dynamic control of error management and signaling as disclosed herein.

[0014] Figure 3 Illustrate a block diagram of a memory device that supports dynamic control of error management and signaling as disclosed herein.

[0015] Figure 4 Illustrate a block diagram of an error management component that supports dynamic control of error management and signaling as disclosed herein.

[0016] Figure 5 and 6 Illustrate examples of process flows in systems that support dynamic control of error management and signaling as disclosed herein.

[0017] Figures 7 to 10 Show flowcharts illustrating one or more methods that support dynamic control of error management and signaling as disclosed herein. DETAILED DESCRIPTION

[0018] Memory systems and memory devices (e.g., dynamic random access memory (DRAM)) can be essential devices in modern computers, which include personal computers (PCs), laptop computers, servers, smart phones, automobiles, and so on. However, such systems can be prone to errors, which can lead to failures within the system. Such failures can be attributed to unpredictable transient electronic noise within the system, defects in the operating system, or can be caused by other factors (e.g., electrostatic shock, electromechanical failure), manufacturing defects, and so on. Additionally, a memory device can communicate with another device (e.g., a host device, such as a controller, a graphics processing unit (GPU), a general-purpose GPU (GPGPU), a central processing unit (CPU), or other device) via one or more communication channels, and errors can be introduced (e.g., randomly) when transmitting data via the channels (e.g., from electronic noise in the system and its surroundings).

[0019] In any case, corruption or errors within a memory system can reduce the value of the information stored or transmitted. While some memory devices can reliably receive and store data with a low bit error rate, due to the physical limitations inherent in electronic devices, the bit error rate is unlikely to be zero. Errors in a memory device can correspond to a single memory cell (e.g., a bit flip), a portion of a memory array (e.g., a row and / or column defect), or the loss of data stored in an entire memory array, and such errors can be intermittent or permanent.

[0020] In addition, some types of data in a memory system can be more critical than other types of data, and failures associated with data of higher operational importance can thus have a greater impact on the system. As an example, failures in mission-critical applications (such as autonomous driving systems and air navigation, to name just a few) can result in more severe consequences compared to failures in a video streaming application, where a failure may result in, for example, a few color-changing pixels. Some systems (e.g., systems for self-driving cars) may accordingly include design parameters (e.g., when the risk of injury or death is extremely high) that tolerate memory failures to the extent that allows the system to remain operational (due to the importance of continuous operation). Thus, different applications and operating conditions of a memory device can require different error management techniques. Similarly, there is a need to further improve techniques for error management, detection, correction, and reporting in memory systems.

[0021] As described herein, various techniques can be used to adapt the fault tolerance to the design parameters and specific applications of a memory system. For example, the security circuitry in a memory device can be enhanced through programmable error counting, error flag output, on-die address bus protection, command / address error detection codes (e.g., parity bits, CRC), or temperature-controlled internal refresh rates, or a combination thereof.

[0022] In some instances, a memory device may use error detection and correction (EDC) codes to prevent memory failures from occurring, where techniques that monitor for errors and take action to correct failures (e.g., those that include software algorithms-based ones) may be used. An internal counter may be used to count the number of detected errors, and when an error threshold (e.g., based on the count) has been reached, the memory device may send a signal to the host device. In some aspects, the memory device may be configured with a threshold number of errors that triggers reporting to the host device. Different thresholds may be configured for corresponding types of data, or for data stored in different portions of the memory device (e.g., each memory bank and / or other portions of the device), or both, and other instances. In some instances, the configurable threshold may be a threshold number of errors, a threshold frequency of errors, a threshold type of errors, a threshold of errors detected at one or more locations of the memory device, or any combination thereof. Thus, the memory device may selectively report corrupted bits to the host device. In some cases, whenever a cell (or word, row, etc.) is refreshed or when an access command is issued, or a combination thereof, the memory device may check (e.g., autonomously check) for and correct errors.

[0023] In some cases, the host device may use the information received from the memory device to avoid portions of the memory array that have a high failure rate (compared to other portions of the array). In such cases, the host device may continue to use the other non-failed portions of the memory array and thus retain most of the available memory capacity while ensuring continued operation. In some aspects, the signal received from the memory device may trigger an interruption at the host device, and the host device may analyze the failure and respond accordingly (e.g., by reissuing a command). Thus, a system with a high safety margin that allows for a low error rate (if any) may set programmable error thresholds to trigger interruptions and prevent errors from propagating in the system.

[0024] Additional techniques described herein include identifying errors in information received at and transmitted within the memory device. For example, the memory device may include logic for checking the integrity of received commands, where command transmissions may be protected by parity or other check bits. The memory device may respond with an indication of whether the received information contains an error. In the case where the parity or checksum for the received command fails, the memory device may refrain from executing the received command (e.g., due to the detected error) and report the error back to the host device. By avoiding executing erroneous commands, additional errors in the system may be avoided. Additionally or alternatively, the memory device may enter a locked state to stop executing instructions received after an error has been identified. The host device may identify the error, send an instruction to the memory device to unlock the state and reissue the failed (and any later) commands.

[0025] In other instances, random errors (e.g., bit flips) may occur in the transmitted addresses (e.g., memory bank, row, and / or column addresses), which results in incorrect sections of the open, accessed, or closed array. Thus, the memory device may apply EDC protection (e.g., parity, CRC) to internal buses such as row and column address buses. The EDC protection applied to these buses can guard against errors in the information transmitted within the device and further increase the level of protection against errors occurring in the memory system.

[0026] In the context of Figure 1 the exemplary memory system hierarchy herein describes features of the present disclosure, and features of the present disclosure are further described with respect to Figure 2 the exemplary memory device in the context of Figure 3 and 4 Specific examples of memory devices and component block diagrams are described in the context of Figure 5 and 6 These and other features of the present disclosure are further illustrated by the process flows of Figures 7 to 10 and

[0027] Figure 1 FIG. 22 shows an example of a system 100 that utilizes one or more memory devices in accordance with aspects disclosed herein. The system 100 may include an external memory controller 105, a memory device 110, and a plurality of channels 115 that couple the external memory controller 105 to the memory device 110. The system 100 may include one or more memory devices, but for ease of description, the one or more memory devices may be described as a single memory device 110.

[0028] The system 100 may include aspects of an electronic device such as a computing device, a mobile computing device, a wireless device, or a graphics processing device. The system 100 may be an example of a portable electronic device. The system 100 may be an example of a computer, a laptop computer, a tablet computer, a smart phone, a cellular phone, a wearable device, an Internet-connected device, and the like. The memory device 110 may be a component of a system configured to store data for one or more other components of the system 100. In some instances, the system 100 is configured for two-way wireless communication with other systems or devices using a base station or an access point. In some instances, the system 100 is capable of machine type communication (MTC), machine-to-machine (M2M) communication, or device-to-device (D2D) communication.

[0029] At least a portion of system 100 can be an example of a host device. Such a host device can be an example of a device that uses memory to execute processes, such as a computing device, a mobile computing device, a wireless device, a graphics processing device, a computer, a laptop computer, a tablet computer, a smart phone, a cellular phone, a wearable device, an Internet-connected device, a vehicle, some other stationary or portable electronic device, etc. In some cases, the host device can refer to hardware, firmware, software, or a combination thereof that implements the functions of the external memory controller 105. In some cases, the external memory controller 105 can be referred to as the host or host device. In some instances, system 100 is a graphics card. In some cases, the host device can determine the degree of error tolerance in the system for data (e.g., based on the application, data type) and set the error threshold number based on the tolerance.

[0030] For example, a memory system for autonomous driving may have a relatively low tolerance for errors compared to other applications or implementations. The error threshold can thus be set to allow very few (or no) errors in the data accessed and stored by the system. Alternatively, the host device can set a higher error threshold for data associated with, for example, video streaming or 3D graphics applications. In either case, the error threshold numbers for different types of data can be determined by the host device and enabled dynamically within the system.

[0031] In some cases, the memory device 110 can be an independent device or component configured to communicate with other components of system 100 and provide a physical memory address / space that can be used or referenced by system 100. In some instances, the memory device 110 can be configured to cooperate with at least one or more different types of system 100. The signaling between the components of system 100 and the memory device 110 can be used to support modulation schemes for modulating signals, different pin designs for transmitting signals, different packages for system 100 and memory device 110, clock signaling and synchronization between system 100 and memory device 110, timing conventions, and / or other factors.

[0032] The memory device 110 can be configured to store data for the components of the system 100. In some cases, the memory device 110 can act as a slave device of the system 100 (e.g., respond to and execute commands provided by the system 100 through the external memory controller 105). Such commands can include access commands for access operations, such as write commands for write operations, read commands for read operations, refresh commands for refresh operations, or other commands. The memory device 110 can include two or more memory dies 160 (e.g., memory chips) that support the desired or specified capacity for data storage. A memory device 110 that includes two or more memory dies can be referred to as a multi-die memory or package (also referred to as a multi-chip memory or package).

[0033] The system 100 can additionally include a processor 120, a basic input / output system (BIOS) component 125, one or more peripheral components 130, and an input / output (I / O) controller 135. The components of the system 100 can communicate electronically with each other using a bus 140.

[0034] The processor 120 can be configured to control at least part of the system 100. The processor 120 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware component, or it can be a combination of these types of components. In such cases, the processor 120 can be an instance of a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose GPU (GPGPU), or a system-on-chip (SoC), among other instances.

[0035] The BIOS component 125 can be a software component that includes the BIOS operating as firmware, which can initialize and run various hardware components of the system 100. The BIOS component 125 can also manage the data flow between the processor 120 and various components of the system 100, such as peripheral components 130, I / O controller 135, etc. The BIOS component 125 can include programs or software stored in read-only memory (ROM), flash memory, or any other non-volatile memory.

[0036] The peripheral component 130 can be any input device or output device, or an interface to such devices, which can be integrated into or integrated with the system 100. Examples can include a disk controller, a sound controller, a graphics controller, an Ethernet controller, a modem, a Universal Serial Bus (USB) controller, a serial or parallel port, or a peripheral card slot, such as a Peripheral Component Interconnect (PCI) or Accelerated Graphics Port (AGP) slot. The peripheral component 130 can be other components understood by those skilled in the art as peripheral devices.

[0037] The I / O controller 135 can manage data communication between the processor 120 and the peripheral component 130, the input device 145, or the output device 150. The I / O controller 135 can manage peripheral devices not integrated into or not integrated with the system 100. In some cases, the I / O controller 135 can represent a physical connection or port to an external peripheral component.

[0038] The input 145 can represent a device or signal external to the system 100 that provides information, signals, or data to the system 100 or its components. This can include a user interface or an interface with or between other devices. In some cases, the input 145 can be a peripheral device interfaced with the system 100 via one or more peripheral components 130, or can be managed by the I / O controller 135.

[0039] The output 150 can represent a device or signal external to the system 100 that is configured to receive output from the system 100 or any of its components. Examples of the output 150 can include a display, an audio speaker, a printing device, or another processor on a printed circuit board, etc. In some cases, the output 150 can be a peripheral device interfaced with the system 100 via one or more peripheral components 130, or can be managed by the I / O controller 135.

[0040] The components of the system 100 can be composed of general-purpose or special-purpose circuits designed to perform their functions. This can include various circuit elements configured to perform the functions described herein, such as wires, transistors, capacitors, inductors, resistors, amplifiers, or other active or passive elements.

[0041] The memory device 110 may include a device memory controller 155 and one or more memory dies 160. Each memory die 160 may include a local memory controller 165 (e.g., local memory controller 165-a, local memory controller 165-b, and / or local memory controller 165-N) and a memory array 170 (e.g., memory array 170-a, memory array 170-b, and / or memory array 170-N). The memory array 170 may be a collection (e.g., a grid) of memory cells, where each memory cell is configured to store at least one bit of digital data. Refer to Figure 2 The characteristics of the memory array 170 and / or the memory cells are described in more detail.

[0042] In some cases, different types of data may be stored in different parts of the memory device 110. Additionally, some data may have a higher or lower priority within the system, and the corresponding data sets may be stored accordingly in, for example, different memory arrays 170 (e.g., based on the reliability of the location where the data is stored relative to the data priority). As an illustrative example, data associated with mission-critical operations (e.g., highly sensitive data, data necessary for system operation, data necessary for enterprise or organizational operation) may be stored in memory array 170-a, while data associated with less critical operations (e.g., data associated with media streaming, data for three-dimensional (3D) graphics rendering or gaming) may be stored in memory array 170-b. Similarly, a first type of data may be stored at memory die 160-a and a second type of data may be stored at memory die 160-b. Additionally or alternatively, different types of data may be stored at different dual in-line memory modules (DIMMs). Other instances of storing different types of data in corresponding parts of the memory device 110 that are not explicitly described herein are also contemplated and fall within the scope of the concepts disclosed herein.

[0043] Memory device 110 may be an example of a two-dimensional (2D) memory cell array or may be an example of a 3D memory cell array. For example, a 2D memory device may include a single memory die 160. A 3D memory device may include two (2) or more memory dies 160 (e.g., memory die 160-a, memory die 160-b, and / or any number of memory dies 160-N). In a 3D memory device, multiple memory dies 160-N may be stacked on top of each other. In some cases, the memory dies 160-N in a 3D memory device may be referred to as stacks, levels, tiers, or dies. A 3D memory device may include any number of stacked memory dies 160-N (e.g., two-high stacked memory dies, three-high stacked memory dies, four-high stacked memory dies, five-high stacked memory dies, six-high stacked memory dies, seven-high stacked memory dies, eight-high stacked memory dies). This may increase the number of memory cells that can be positioned on a substrate compared to a single 2D memory device, which in turn may reduce production costs or increase the performance of the memory array, or both. In some 3D memory devices, different stacks may share at least one common access line such that some stacks may share at least one of word lines, digit lines, and / or plate lines.

[0044] Device memory controller 155 may include circuitry or components configured to control the operation of memory device 110. Thus, device memory controller 155 may include hardware, firmware, and software that enable memory device 110 to execute commands and may be configured to receive, transmit, or execute commands, data, or control information regarding memory device 110. Device memory controller 155 may be configured to communicate with external memory controller 105, one or more memory dies 160, or processor 120. In some cases, memory device 110 may receive data and / or commands from external memory controller 105. For example, memory device 110 may receive a write command indicating that memory device 110 is to store certain data on behalf of a component of system 100 (e.g., processor 120), or a read command indicating that memory device 110 is to provide certain data stored in memory die 160 to a component of system 100 (e.g., processor 120). In some cases, device memory controller 155 may control the operation of memory device 110 described herein in conjunction with local memory controller 165 of memory die 160. Examples of components included in device memory controller 155 and / or local memory controller 165 may include a receiver for demodulating signals received from external memory controller 105, a decoder for modulating and transmitting signals to external memory controller 105, logic, decoders, amplifiers, filters, etc.

[0045] In some cases, the device memory controller 155 may perform the function of detecting and / or correcting errors in data at the memory device 110. As an example, the device memory controller 155 may use an error detection code to identify an error in a command received from the external memory controller 105. In other cases, the device memory controller 155 may use an error detection code associated with the bus set of the memory device 110 to identify an error in an address (e.g., a bank, row, and / or column address) used to store and retrieve data. In some cases, the device memory controller 155 may also use an error correction code to correct an error identified in data retrieved from the memory array 170. The device memory controller may count the number of errors and / or the number of corrections performed and may report the count to the external memory controller 105. The device memory controller 155 may also determine if a threshold number of errors in the data has been met and, if so, may place the memory device 110 in a locked state to prevent the execution of additional commands. In some instances, the device memory controller 155 may refrain from reporting correctable errors when reporting uncorrectable errors to the external memory controller 105.

[0046] The local memory controller 165 (e.g., local to the memory die 160) may be configured to control the operation of the memory die 160. Further, the local memory controller 165 may be configured to communicate with the device memory controller 155 (e.g., receive and transmit data and / or commands). The local memory controller 165 may support the device memory controller 155 in controlling the operation of the memory device 110 as described herein. In some cases, the memory device 110 does not include the device memory controller 155, and the local memory controller 165 or the external memory controller 105 may perform the various functions described herein. Accordingly, the local memory controller 165 may be configured to communicate with the device memory controller 155, communicate with other local memory controllers 165, or communicate directly with the external memory controller 105 or the processor 120. In some instances, the local memory controller 165 may detect and count errors in the data contained in the respective memory array 170. For example, each error identified in the data at the memory device 110 may increment a counter, and the memory device 110 may maintain an error count. In some instances, the error count may be transmitted to the device memory controller 155, and a summary or individual error count from each memory die 160 or each memory array 170 or a combination thereof may be reported to the external memory controller 105.

[0047] The external memory controller 105 may be configured to implement the transfer of information, data, and / or commands between components of the system 100 (e.g., the processor 120) and the memory device 110. The external memory controller 105 may act as a liaison between the components of the system 100 and the memory device 110 such that the components of the system 100 may not need to know the operational details of the memory device. The components of the system 100 may present requests (e.g., read commands or write commands) to the external memory controller 105 that the external memory controller 105 satisfies. The external memory controller 105 may translate or transcribe the communications exchanged between the components of the system 100 and the memory device 110. In some cases, the external memory controller 105 may include a system clock that generates a common (source) system clock signal. In some cases, the external memory controller 105 may include a common data clock that generates a common (source) data clock signal.

[0048] In some cases, the external memory controller 105 or other components of the system 100 or their functions described herein may be implemented by the processor 120. For example, the external memory controller 105 may be hardware, firmware, software, or some combination thereof implemented by the processor 120 or other components of the system 100. Although the external memory controller 105 is depicted as being external to the memory device 110, in some cases, the external memory controller 105 or its functions described herein may be implemented by the memory device 110. For example, the external memory controller 105 may be hardware, firmware, software, or some combination thereof implemented by the device memory controller 155 or one or more local memory controllers 165. In some cases, the external memory controller 105 may be distributed across the processor 120 and the memory device 110 such that portions of the external memory controller 105 are implemented by the processor 120 and other portions are implemented by the device memory controller 155 or the local memory controller 165. Similarly, in some cases, one or more functions attributed to the device memory controller 155 or the local memory controller 165 herein may, in some cases, be performed by the external memory controller 105 (separate from or included in the processor 120). In some instances, the external memory controller 105 may perform the function of managing and correcting errors within the memory device 110. For example, if the memory device 110 reports that a programmed error count threshold has been met, the external memory controller 105 may determine portions of the memory device 110 containing error data that may be avoided in later access operations. Thus, the external memory controller 105 may write data to different portions of the memory device 110 / read data from different portions of the memory device 110.

[0049] The components of system 100 may exchange information with the memory device 110 using multiple channels 115. In some instances, the channels 115 may enable communication between the external memory controller 105 and the memory device 110. Each channel 115 may include one or more signal paths or transmission media (e.g., conductors) between terminals associated with components of system 100. For example, a channel 115 may include a first terminal that includes one or more pins or pads at the external memory controller 105 and one or more pins or pads at the memory device 110. Pins may be examples of conductive input or output points of devices of system 100, and the pins may be configured to act as part of the channel.

[0050] In some cases, the pins or pads of the terminals may be part of the signal path of the channel 115. Additional signal paths may be coupled to the terminals of the channel for routing signals within the components of system 100. For example, the memory device 110 may include signal paths (e.g., signal paths internal to the memory device 110 or its components, such as internal to the memory die 160) that route signals from the terminals of the channel 115 to various components of the memory device 110 (e.g., the device memory controller 155, the memory die 160, the local memory controller 165, the memory array 170).

[0051] The channels 115 (and associated signal paths and terminals) may be dedicated to transmitting a particular type of information. In some cases, a channel 115 may be an aggregate channel and may thus include multiple individual channels. For example, a data channel 190 may be x4 (e.g., include four signal paths), x8 (e.g., include eight signal paths), x16 (include sixteen signal paths), and so on.

[0052] In some cases, the channel 115 may include one or more command and address (CA) channels 186. The CA channels 186 may be configured to transmit commands between the external memory controller 105 and the memory device 110, including control information (e.g., address information) associated with the commands. For example, a CA channel 186 may include a read command with the address of the desired data. In some cases, the CA channels 186 may be latched on the rising clock signal edge and / or the falling clock signal edge. In some cases, the CA channels 186 may include multiple signal paths.

[0053] In some cases, channel 115 may include one or more clock signal (CK) channels 188. The CK channels 188 may be configured to transmit one or more common clock signals between the external memory controller 105 and the memory device 110. Each clock signal may be configured to oscillate between a high state and a low state and to coordinate the operations of the external memory controller 105 and the memory device 110. In some cases, the clock signals may be differential outputs (e.g., CK_t signal and CK_c signal) and the signal paths of the CK channels 188 may be configured accordingly. In some cases, the clock signals may be single-ended. The CK channels 188 may include any number of signal paths. In some cases, the clock signals CK (e.g., CK_t signal and CK_c signal) may provide a timing reference for command and addressing operations of the memory device 110 or other system-wide operations of the memory device 110. The clock signal CK may thus be alternatively referred to as a control clock signal CK, a command clock signal CK, or a system clock signal CK. The system clock signal CK may be generated by a system clock, which may include one or more hardware components (e.g., oscillators, crystals, logic gates, transistors, etc.).

[0054] In some cases, channel 115 may include one or more data (DQ) channels 190. The data channels 190 may be configured to transmit data and / or control information between the external memory controller 105 and the memory device 110. For example, the data channels 190 may transmit information to be written to the memory device 110 (e.g., bidirectionally) or information read from the memory device 110. The data channels 190 may transmit signals that may be modulated using a variety of different modulation schemes (e.g., NRZ, PAM4). In some cases, channel 115 may include one or more other channels 192 that may be dedicated to other purposes. These other channels 192 may include any number of signal paths.

[0055] In some cases, other channels 192 may include one or more write clock signal (WCK) channels. Although the 'W' in WCK may nominally represent "write", the write clock signals WCK (e.g., WCK_t signal and WCK_c signal) may provide a timing reference generally used for access operations of the memory device 110 (e.g., a timing reference for both read and write operations). Thus, the write clock signal WCK may also be referred to as a data clock signal WCK. The WCK channels may be configured to transmit a common data clock signal between the external memory controller 105 and the memory device 110. The data clock signal may be configured to coordinate access operations (e.g., write operations or read operations) of the external memory controller 105 and the memory device 110. In some cases, the write clock signal may be a differential output (e.g., WCK_t signal and WCK_c signal), and the signal paths of the WCK channels may be configured accordingly. The WCK channels may include any number of signal paths. The data clock signal WCK may be generated from a data clock, which may include one or more hardware components (e.g., oscillators, crystals, logic gates, transistors, etc.). In certain cases, other channels 192 may include one or more EDC channels. The EDC channels may be configured to transmit error detection signals, such as checksums, to improve system reliability. The EDC channels may include any number of signal paths. In some cases, the EDC channels may be used to emit error flags based on an error threshold being met.

[0056] Channel 115 may couple the external memory controller 105 to the memory device 110 using a variety of different architectures. Examples of various architectures may include buses, point-to-point connections, crossbars, high-density interposers such as silicon interposers, or channels formed in organic substrates, or some combination thereof. For example, in some cases, the signal paths may at least partially include high-density interposers, such as silicon interposers or glass interposers.

[0057] Various different modulation schemes may be used to modulate the signals transmitted on channel 115. In some cases, a binary symbol (or binary level) modulation scheme may be used to modulate the signals communicated between the external memory controller 105 and the memory device 110. The binary symbol modulation scheme may be an example of an M-ary modulation scheme, where M equals two. Each symbol of the binary symbol modulation scheme may be configured to represent one bit of digital data (e.g., the symbol may represent a logic 1 or a logic 0). Examples of binary symbol modulation schemes include, but are not limited to, non-return-to-zero (NRZ), unipolar coding, bipolar coding, Manchester coding, pulse amplitude modulation (PAM) with two symbols (e.g., PAM2), and the like.

[0058] In some cases, a multi-symbol (or multi-level) modulation scheme can be used to modulate the signals transmitted between the external memory controller 105 and the memory device 110. The multi-symbol modulation scheme can be an instance of an M-ary modulation scheme, where M is greater than or equal to three. Each symbol of the multi-symbol modulation scheme can be configured to represent digital data of more than one bit (e.g., the symbol can represent logic 00, logic 01, logic 10, or logic 11). Examples of the multi-symbol modulation scheme include but are not limited to PAM4, PAM8, etc., quadrature amplitude modulation (QAM), quadrature phase shift keying (QPSK), etc. The multi-symbol signal or PAM4 signal can be a signal modulated using a modulation scheme that includes at least three levels for encoding information of more than one bit. The multi-symbol modulation scheme and symbols can alternatively be referred to as non-binary, multi-bit, or high-order modulation schemes and symbols.

[0059] The memory device 110 can be an essential component in modern computers or devices with computers, including PCs, laptops, servers, smartphones, vehicles, and so on. However, such devices can be prone to errors, resulting in data failures. Data failures can be attributed to unpredictable transient electronic noise, defects in the operating system, electrostatic shocks, electromechanical failures, manufacturing defects, etc. Additionally, the memory device 110 can communicate with a host device (e.g., the external memory controller 105) via one or more channels 115, and errors can be introduced when transmitting data via the channels 115 (e.g., from the electronic noise in and around the system 100). In any case, the errors can reduce the value of the information stored or transmitted at the memory device 110. Errors in the memory device 110 can correspond to the loss of data stored in a single memory cell (e.g., a bit flip), a portion of the memory array (e.g., row and / or column defects), or the entire memory array 170 or a combination thereof. Such errors can be intermittent or permanent. Additionally, although some memory devices 110 can receive and store data with a relatively low bit error rate, the physical limitations of the memory device 110 can cause a certain degree of transient errors during operation.

[0060] In some cases, if an error or data loss occurs in the memory, the host device may not be aware of the error, and the system may continue to operate with the corrupted data. Errors can also occur as the memory device 110 ages and its components degrade. In other instances, errors can be caused by the nature of the internal buses associated with the memory device 110 (e.g., the size of the buses, construction defects) and the charge that can accumulate on these internal buses. Additionally, errors in access commands (e.g., read commands, write commands) can lead to a cascade of failures in the system 100. For example, an error in an access command can cause a series of problems, ranging from bit flips to a write command being incorrectly executed as a read command. Further, if a later command is received at the memory device 110 after a command containing an error, some data may be overwritten, and the host device (or the memory device 110) may not be able to identify the error or its cause.

[0061] In some cases, based on the use cases (e.g., applications) of the memory system, certain failures can be more severe than others. For example, the system 100 can be used for mission-critical operations with a low tolerance for errors. Autonomous driving applications (e.g., self-driving cars, robotic cars, driverless vehicles) can be an example of such mission-critical operations, and data integrity in the system as well as the continued operation of the system can be of high importance (compared to a personal computer (PC) or a tablet computer used for web browsing). These systems can have complex operations that continuously manage data for various sensors, inputs, and components, where the data can be used for real-time calculations and decision-making regarding vehicle operations (e.g., acceleration, speed, relative orientation to other objects and people, navigation). Additionally, even within the same application, there can be different levels of impact at different times (e.g., driving on an interstate highway versus waiting at a traffic light). Therefore, different applications not only affect the way errors are handled in the system, but different operating conditions may require different error management techniques. It should be noted that autonomous driving is used herein as an illustrative example of a system configured for mission-critical operations, but other such systems and operations are also considered.

[0062] System 100 may support techniques for implementing dynamic management of data errors at memory device 110. For example, programmable thresholds may be configured for memory device 110 based on data type or the orientation of the stored data or a combination thereof. A host device may program an error threshold for data at memory device 110, and memory device 110 may count errors detected in the data (e.g., after reading data from a memory array) to determine if the threshold has been met. Memory device 110 may transmit an indication that the threshold has been met to the host device and may then perform functions to correct the error and / or prevent additional errors. For example, memory device 110 may enter a locked state immediately upon determining that the error threshold has been met. In this state, memory device 110 may not execute any commands received later from the host. However, the host device may use the received error indication to determine if the error is correctable (e.g., whether the corrupted data is stored in another redundant orientation). If so, then the host device may transmit an indication to memory device 110 to exit the locked state and retransmit the data. In other instances, the host device may receive an indication of the junction temperature from memory device 110, and the host device may adjust the operating parameters of memory device 110 to help mitigate errors, such as increasing the refresh rate of memory device 110 or decreasing the operating frequency of memory device 110. In other cases, the memory device may autonomously adjust the refresh rate based on the detected errors.

[0063] Memory device 110 may also identify errors in received commands. In such cases, error detection codes (e.g., parity bits, cyclic redundancy check (CRC) bits) may be included with the commands transmitted by the host device. Memory device 110 may use the error detection codes to determine if the command contains one or more errors. In cases where no errors are detected, memory device 110 may indicate this to the host device (where the host device may store the set of transmitted commands and may remove a particular command from storage once it receives an indication that the command is error-free). However, if the command contains an error, then memory device 110 may refrain from executing the received command, enter a locked state, and may transmit an indication of the error to the host device. Accordingly, the host device may retransmit the command based on the received indication.

[0064] Additionally or alternatively, memory device 110 may use EDC to identify errors associated with commands for the internal bus of memory device 110. That is, signals generated in response to commands (e.g., bank / row / column addresses) may also utilize a protection mechanism, such as an error detection code, to identify errors. Using EDC protection for the internal bus of memory device 110 may ensure that even if error-free commands, data, or both may have been received, later errors that occur when the associated signals propagate through memory device 110 may be identified and mitigated in a timely and efficient manner.

[0065] System 100 can thus support flexible error counting and reporting using memory device 110 based on the type of data accessed. In this way, errors in data associated with some operations can be processed with a higher priority compared to other operations, and system 100 can be adaptively configured with different thresholds for different operations and applications. In some aspects, programmable thresholds and corresponding techniques for correction to prevent future errors can enable continued functionality in system 100, where actions can be taken to ensure storing and retrieving data from portions of the memory that operate with few errors. For example, instead of stopping operation or resetting system 100 when a failure occurs, steps can be taken to slow down execution to recover and / or correct the failure to avoid a complete system failure. For complex operations, this can mean adjusting the memory bandwidth to allow extra time to perform functions for retuning the system to a stable (error-free) state while still processing data from memory device 110. In the example of autonomous driving, this can mean that the vehicle temporarily slows down while the memory system mitigates detected errors, which can prevent a sudden complete loss of one or more systems or components of the vehicle due to a system reset triggered by a failure. In other instances, signaling of errors identified in system 100 can also trigger user intervention.

[0066] Additionally, system 100 can be configured to selectively enable and disable the described data protection mechanisms. Thus, the same memory device 110 can be used for different applications and / or in different systems when achieving a customizable level of reliability and memory bandwidth (e.g., based on the application / system). For example, there can be a trade-off between techniques for enhancing data integrity and the speed at which memory device 110 stores and retrieves data, and this trade-off can be configured based on a given application. For example, some applications may allow an acceptable performance degradation in the absence of errors in the memory, while for other applications, performance may be more critical than the number of errors. However, for a given system, a certain combination of the described aspects can be dynamically enabled to achieve a programmable balance of error management, memory bandwidth, and other factors.

[0067] Figure 2 An example of a memory die 200 that supports dynamic control of error management and signaling according to aspects disclosed herein is described. Memory die 200 can be a reference Figure 1An example of the described memory die 160. In some cases, the memory die 200 may be referred to as a memory chip, a memory device, or an electronic memory device. The memory die 200 may include one or more memory cells 205 programmable to store different logic states. Each memory cell 205 may be programmable to store two or more states. For example, a memory cell 205 may be configured to store digital logic (e.g., logic 0 and logic 1) of one bit at a time. In some cases, a single memory cell 205 (e.g., a multi-level memory cell) may be configured to store digital logic (e.g., logic 00, logic 01, logic 10, or logic 11) of more than one bit at a time.

[0068] The memory cell 205 may store charge representing a programmable state in a capacitor. A DRAM architecture may include a capacitor that includes a dielectric material to store charge representing a programmable state. In other memory architectures, other storage devices and components are possible. For example, non-linear dielectric materials may be used.

[0069] Operations such as reading and writing may be performed on the memory cell 205 by activating or selecting access lines such as word lines 210 and / or digit lines 215. In some cases, the digit lines 215 may also be referred to as bit lines. References to access lines, word lines, and digit lines or the like may be interchangeable and do not affect understanding or operation. Activating or selecting the word line 210 or the digit line 215 may include applying a voltage to the corresponding line.

[0070] The memory die 200 may arrange access lines (e.g., word lines 210 and digit lines 215) in a grid pattern. The memory cells 205 may be located at the intersections of the word lines 210 and the digit lines 215. By biasing the word lines 210 and the digit lines 215 (e.g., applying a voltage to the word line 210 or the digit line 215), a single memory cell 205 may be accessed at their intersection.

[0071] Access to the memory cell 205 can be controlled by either the row decoder 220 or the column decoder 225. For example, the row decoder 220 can receive a row address from the local memory controller 260 and activate the word line 210 based on the received row address. The column decoder 225 can receive a column address from the local memory controller 260 and can activate the digit line 215 based on the received column address. For example, the memory die 200 can include a plurality of word lines 210 labeled WL_1 to WL_M and a plurality of digit lines 215 labeled DL_1 to DL_N, where M and N depend on the size of the memory array. Thus, by activating the word line 210 and the digit line 215, such as WL_1 and DL_3, the memory cell 205 at their intersection can be accessed. The intersection of the word line 210 and the digit line 215 in a two-dimensional or three-dimensional configuration can be referred to as the address of the memory cell 205. In some cases, information such as CRC or parity bits can be included with the address received by the row decoder 220 or the column decoder 225. In such cases, the address can be checked to identify any errors, thus ensuring that the correct word line 210 and digit line 215 are accessed in response to the received address. Such techniques can reduce or minimize errors in accessing the memory cell 205 with minimal overhead.

[0072] The memory cell 205 can include logic storage components, such as a capacitor 230 and a switching component 235. The capacitor 230 can be an example of a dielectric capacitor or a ferroelectric capacitor. The first node of the capacitor 230 can be coupled to the switching component 235, and the second node of the capacitor 230 can be coupled to a voltage source 240. In some cases, the voltage source 240 can be a cell plate reference voltage, such as Vpl, or can be grounded, such as Vss. In some situations, the voltage source 240 can be an example of a cell line coupled to a cell line driver. The switching component 235 can be an example of a transistor or any other type of switching device that selectively establishes or breaks an electronic communication between two components.

[0073] Selecting or deselecting the memory cell 205 can be achieved by activating or deactivating the switching component 235. The capacitor 230 can communicate electronically with the digit line 215 using the switching component 235. For example, when the switching component 235 is deactivated, the capacitor 230 can be isolated from the digit line 215, and when the switching component 235 is activated, the capacitor 230 can be coupled to the digit line 215. In some cases, the switching component 235 is a transistor, and its operation can be controlled by applying a voltage to the transistor gate, where the voltage difference between the transistor gate and the transistor source can be greater than or less than the threshold voltage of the transistor. In some situations, the switching component 235 can be a p-type transistor or an n-type transistor. The word line 210 can communicate electronically with the gate of the switching component 235, and the switching component 235 can be activated / deactivated based on the voltage applied to the word line 210.

[0074] The word line 210 can be a conductive wire that is in electronic communication with the memory cell 205 and is used to perform access operations on the memory cell 205. In some architectures, the word line 210 can be in electronic communication with the gate of the switching component 235 of the memory cell 205 and can be configured to control the switching component 235 of the memory cell. In some architectures, the word line 210 can be in electronic communication with the node of the capacitor of the memory cell 205, and the memory cell 205 may not include a switching component.

[0075] The digit line 215 can be a wire connecting the memory cell 205 and the sensing component 245. In some architectures, the memory cell 205 can be selectively coupled to the digit line 215 during part of an access operation. For example, the word line 210 and the switching component 235 of the memory cell 205 can be configured to couple and / or isolate the capacitor 230 of the memory cell 205 and the digit line 215. In some architectures, the memory cell 205 can be in electronic communication with the digit line 215 (e.g., constantly).

[0076] The sensing component 245 can be configured to detect the state (e.g., charge) stored on the capacitor 230 of the memory cell 205, and determine the logical state of the memory cell 205 based on the stored state. In some situations, the charge stored by the memory cell 205 may be extremely small. Therefore, the sensing component 245 may include one or more sense amplifiers to amplify the signal output by the memory cell 205. The sense amplifier can detect small changes in the charge of the digit line 215 during a read operation, and can generate a signal corresponding to the logical state 0 or the logical state 1 based on the detected charge. During a read operation, the capacitor 230 of the memory cell 205 can output a signal (e.g., release charge) to its corresponding digit line 215. The signal can cause the voltage of the digit line 215 to change. The sensing component 245 can be configured to compare the signal received from the memory cell 205 across the digit line 215 with a reference signal 250 (e.g., a reference voltage). The sensing component 245 can determine the stored state of the memory cell 205 based on the comparison. For example, in binary signaling, if the digit line 215 has a voltage higher than the reference signal 250, then the sensing component 245 can determine that the stored state of the memory cell 205 is the logical state 1, and if the digit line 215 has a voltage lower than the reference signal 250, then the sensing component 245 can determine that the stored state of the memory cell 205 is the logical state 0. The sensing component 245 can include various transistors or amplifiers to detect and amplify the difference in signals. The detected logical state of the memory cell 205 can be output as an output 255 via the column decoder 225. In certain cases, the sensing component 245 can be part of another component (e.g., the column decoder 225, the row decoder 220). In some situations, the sensing component 245 can communicate electronically with the row decoder 220 or the column decoder 225.

[0077] The local memory controller 260 can control the operation of the memory cell 205 via various components (e.g., the row decoder 220, the column decoder 225, and the sensing component 245). The local memory controller 260 can be an example of the local memory controller 165 described in Figure 1 In some cases, one or more of the row decoder 220, the column decoder 225, and the sensing component 245 can be in the same location as the local memory controller 260. The local memory controller 260 can be configured to receive commands or data from an external memory controller 105 (or Figure 1The described device memory controller 155) receives commands and / or data, translates the commands and / or data into information usable by the memory die 200, performs one or more operations on the memory die 200, and conveys data from the memory die 200 to the external memory controller 105 (or the device memory controller 155) in response to performing the one or more operations. The local memory controller 260 may generate row and column address signals to activate the target word line 210 and the target digit line 215. The local memory controller 260 may also generate and control various voltages or currents used during the operation of the memory die 200. Generally, the amplitude, shape, or duration of the applied voltage or current discussed herein may be adjusted or varied and may be different for the various operations discussed in operating the memory die 200.

[0078] In some cases, the local memory controller 260 may be configured to perform a write operation (e.g., a programming operation) on one or more memory cells 205 of the memory die 200. During the write operation, the memory cells 205 of the memory die 200 may be programmed to store a desired logic state. In some cases, multiple memory cells 205 may be programmed during a single write operation. The local memory controller 260 may identify the target memory cells 205 on which the write operation will be performed. The local memory controller 260 may identify the target word line 210 and the target digit line 215 that are in electronic communication with the target memory cells 205 (e.g., the addresses of the target memory cells 205). The local memory controller 260 may activate the target word line 210 and the target digit line 215 (e.g., apply a voltage to the word line 210 or the digit line 215) to access the target memory cells 205. The local memory controller 260 may apply a specific signal (e.g., a voltage) to the digit line 215 during the write operation to store a specific state (e.g., a charge) in the capacitor 230 of the memory cell 205, and the specific state (e.g., the charge) may indicate the desired logic state.

[0079] In some cases, the local memory controller 260 may be configured to perform a read operation (e.g., a sense operation) on one or more memory cells 205 of the memory die 200. During the read operation, the logical state stored in the memory cells 205 of the memory die 200 may be determined. In some cases, multiple memory cells 205 may be sensed during a single read operation. The local memory controller 260 may identify the target memory cells 205 on which the read operation will be performed. The local memory controller 260 may identify the target word line 210 and the target digit line 215 (e.g., the address of the target memory cell 205) that are in electronic communication with the target memory cell 205. The local memory controller 260 may activate the target word line 210 and the target digit line 215 (e.g., apply a voltage to the word line 210 or the digit line 215) to access the target memory cell 205. The target memory cell 205 may pass a signal to the sense component 245 in response to the biased access lines. The sense component 245 may amplify the signal. The local memory controller 260 may initiate the sense component 245 (e.g., latch the sense component) and thereby compare the signal received from the memory cell 205 with the reference signal 250. Based on the comparison, the sense component 245 may determine the logical state stored on the memory cell 205. As part of the read operation, the local memory controller 260 may transfer the logical state stored on the memory cell 205 to the external memory controller 105 (or the device memory controller 155). In some cases, the number of incorrect values in the data accessed during the read operation may be counted. Additionally, and as described herein, an error programmable threshold (e.g., a dynamic threshold, a non-static threshold) may be configured for the memory device. Thus, the local memory controller 260 may determine whether the counted number of errors has been met, and if so, may issue an indication that the threshold has been reached. In some cases, the programmable threshold may be selectively adjusted based on the type of data accessed during the read operation. Additionally or alternatively, multiple thresholds may be configured for the respective types of data stored in the memory cells 205.

[0080] In some memory architectures, accessing the memory cells 205 may degrade or destroy the logical state stored in the memory cells 205. For example, a read operation performed in a DRAM architecture may cause the capacitor of the target memory cell to discharge partially or completely. The local memory controller 260 may perform a rewrite operation or a refresh operation to restore the memory cell to its original logical state. The local memory controller 260 may rewrite the logical state to the target memory cell after the read operation. In some cases, the rewrite operation may be considered part of the read operation.

[0081] Figure 3Block diagram 300 illustrating a memory device 310 that supports dynamic control of error management and messaging, in accordance with aspects disclosed herein. Memory device 310 may be an example of memory device 110 described with reference to Figure 1 In some examples, memory device 310 may include one or more components that support error detection and correction enhancements (e.g., in data). Memory device 310 may include a command decoder 315, a memory array 320, a data input / output (I / O) 325, error correction code (ECC) logic 335, and an error counter 340. In some cases, the memory device may be configured to perform error detection and correction based on the operating conditions of a system that includes memory device 310, where different conditions may be associated with different degrees of error management.

[0082] Command decoder 315 may be configured to control processes performed by memory device 310. As an example, command decoder 315 may be configured to receive signals such as a clock signal and one or more commands, where the commands may include a write command or a read command transmitted by a controller. In some examples, command decoder 315 may be coupled to other components of memory device 310, such as a column decoder (e.g., column decoder 225 described with reference to Figure 2 ), or a row decoder (e.g., row decoder 220 described with reference to Figure 2 ), or a combination thereof. Command decoder 315 may decode the received commands and transmit signals to other components of memory device 310. For example, command decoder 315 may generate and transmit addresses (e.g., memory bank address, row address, column address) to store or retrieve data in memory array 320 according to a write or read command. Additionally, data may be transmitted from memory device 310 and / or received at memory device 310 on data I / O 325. For example, the command may include a write command, and data may be received at memory device 310 via data I / O 325 and stored at memory array 320. Alternatively, the command may include a read command, and data may be read from memory array 320 and later transmitted to another device (e.g., a controller) via data I / O 325.

[0083] In some cases, memory device 310 may be configured to detect and / or correct various errors (e.g., using an EDC code). For example, ECC logic 335 may perform various functions on data at memory device 310, and the ECC logic may manage error detection and / or correction using signals received from command decoder 315 and data from memory array 320. ECC logic 335 may include an ECC generation and check component 345 and an ECC correction component 350, which perform corresponding functions on data received and transmitted via data I / O 325.

[0084] For example, in a case where a write command can be received at the command decoder 315, data received via the data I / O 325 can be written to the memory array 320 without change. Additionally, the ECC generation and checking component 345 can generate a set of ECC bits for the data received via the data I / O 325, and the ECC bits can be written to and stored at the memory array 320. In some cases, the memory array 320 can include different portions for storing different types of information, or can include redundant memory cells, groups, and / or arrays to enhance data management and error protection (e.g., by providing locations where backup or support data can be stored). For example, the memory array 320 can include a first portion 330-a for storing the generated ECC bits, while a second portion 330-b can be used to store the data received via the data I / O 325.

[0085] In response to a read command, data can be read from the memory array 320 (e.g., array 320-b), and the ECC generation and checking component 345 and the ECC correction component 350 can check for errors in the read data by calculating a checksum using the ECC bits stored in the array 320-b. More specifically, the ECC bits can be read from the array 320-b, where the ECC bits can correspond to the data that can be read from the array 320-b. The error logic 335 can then check the read data against the ECC bits to identify any errors in the read data. In some cases, any identified errors can be corrected using the ECC bits before the read data is transmitted via the data I / O 325. Such techniques are sometimes referred to as on-die ECC.

[0086] In some instances, the memory device 310 can be part of a system for mission-critical applications (e.g., servers, autonomous vehicles), and an EDC code can be used to prevent memory failures from occurring. For example, the memory device 310 can use software algorithms to monitor for memory failures and take actions to correct the failures, up to and including recommending replacement of malfunctioning devices. In such cases, the memory device 310 can correct the failures (e.g., a single cell failure) and provide corrected data during an access operation (e.g., a read operation). Additionally, the memory device 310 can have an error counter 340 to count the number of detected and / or corrected errors. The memory device 310 can report this count upon request (e.g., using an indication or flag), for example, after receiving a command. In other instances, the memory device 310 can autonomously transmit an indication of the number of errors counted at the error counter 340, for example, when an error threshold has been met.

[0087] In some aspects, on-die ECC implemented by the memory device 310 can be enhanced by a programmable error threshold or error flag output or on-die address bus protection or command / address parity or CRC or temperature-controlled internal refresh rate or a combination thereof. In some cases, the memory device 310 can function in response to detected errors, such as error correction, signaling to the host device, command termination, etc., where the functions can be based on the operating conditions of the memory device 310. In some cases, the memory device 310 can selectively emit a damaged bit report, where instead of reporting a single failure, the number of failures exceeding the error threshold can trigger signaling to the host device. In some cases, the memory device 310 can implement a protocol that allows data error correction, where the error correction can be associated with additional signaling. For example, the memory device can identify an error and send a signal to the memory controller or host device to take corrective action.

[0088] In some cases, the programmable error threshold can be configured for the memory device 310. For example, the host device can determine the number of error thresholds allowable based on a particular application or operation. The memory device 310 can be configured with the number of error thresholds, and the memory device 310 can determine that once the error threshold number has been reached (e.g., based on the count generated by the error counter 340), an error flag can be signaled to the host device to indicate that the threshold has been met. Additionally or alternatively, the programmable error threshold can be an error threshold frequency, for example, where the frequency at which a set of errors counted by the error counter 340 occur (e.g., within a given time period) can be sufficient to meet the error threshold frequency. In some instances, the error threshold can be the maximum error count for a type of error (e.g., a single memory cell error, multiple memory cell errors). In other cases, the error threshold can be an error threshold identified at the location of the memory device 310 (e.g., an error threshold occurring at a particular group of the memory device 310). The error threshold can also be configured by the host device to include a combination of different thresholds. For example, the error threshold can be the threshold number for a first type of error, or can be an error threshold that exceeds the threshold frequency at a first group of the memory device 310. Additionally or alternatively, the error threshold can be the type of error threshold detected at a particular location of the memory device 310. In any case, the memory device 310 can include components other than or different from those illustrated to determine that a set of errors meets the error threshold.

[0089] Additionally, the error threshold can be programmable by the host device or the memory controller, where different types of data or data in different locations in the storage memory or both can be configured with different error thresholds. For example, a system with a strict safety margin may tolerate a low error rate (if any), and the programmable error count for data associated with the operation of the system can be set based on the safety margin. The memory device 310 can also be programmed to suppress reporting of some errors that meet the threshold when other errors are reported. For example, the memory device 310 can be configured to suppress reporting of correctable failures (e.g., single cell failures) while still reporting uncorrectable failures (e.g., double cell or multi - cell errors).

[0090] A signal transmitted to the host device when the error threshold has been reached can trigger one or more actions in the host device, such as an interrupt. The interrupt can allow the host device to analyze the failure and respond accordingly (e.g., by re - issuing a command). After the interrupt, the host can know the location of the previously identified error and can determine which parts of the memory to write data to. Alternatively, if writing to a previously error - prone part of the memory, the host device can determine to use redundant parts of the memory device 310 to ensure there is a backup / protection for the data being written to locations that are prone to failure.

[0091] In addition to counting the number of errors in the data, the memory device 310 can count the number of failures within a portion of the memory array 320 (e.g., per memory bank and / or other parts of the memory device 310). The memory device 310 can report the number of errors in a given part of the memory array 320, and the host can use this information, for example, to ignore or avoid regions of the memory array 320 that have a high failure rate (compared to other parts of the memory array 320). The host device can continue to use other non - failed parts (or parts with a relatively low bit error rate) of the memory array 320 and thus preserve most of the available capacity of the memory device 310. In some cases, the memory device 310 can check and correct errors whenever a unit (or word, row, etc.) is refreshed. Additionally or alternatively, error checking and correction can occur when a read command is received.

[0092] In some cases, the memory device 310 may use an error threshold to change (e.g., autonomously increase or autonomously adjust) an internal refresh rate, which may be based on the number of errors counted at the memory device 310 (e.g., satisfying a configured threshold). For example, the memory device 310 may refresh twice as many cells each time a refresh command is initiated, which may provide a degree of error management and prevention. In some aspects, the memory device 310 may transmit the temperature of one or more portions or components to the host device to assist in error management. As an example, the memory device 310 may transmit an indication of a junction temperature, wherein the temperature indication may be transmitted to the host device within an error indication / flag (e.g., in response to reaching a threshold) or in a separate signal. The host device may take steps accordingly to reduce the reported junction temperature. For example, a temperature sensor may be used to determine whether a portion of the memory device 310 has become too hot (e.g., above an optimal temperature) during operation of the memory device 310. In such cases, the refresh rate may be modified (e.g., doubled or otherwise increased) based on the junction temperature reported at the memory device 310. Additionally or alternatively, the operating frequency of the memory device 310 may be reduced to reduce the temperature of the memory device 310. Other techniques may also be used to reduce the reported temperature of the memory device 310.

[0093] Additional techniques described herein may be used to identify errors in information received from a host device. For example, the ECC logic 335 of the memory device 310 may be configured to check the integrity of the command received at the command decoder 315. In such cases, the command transmission may be protected by a parity or other check bit. That is, additional bits associated with an error detection code may be included in the transmitted command. If the received command is error-free, the memory device 310 may execute the command and may also signal to the host device an indication that the command does not contain an error.

[0094] Alternatively, in the event that the parity check or checksum fails, the memory device 310 may not execute the received command (e.g., due to the detected error) and report the error to the host device. In some cases, the memory device 310 may also enter a different state (e.g., a locked state) and stop executing instructions later received from the host device. Upon identifying the error, the host device may issue a command to the memory device 310 to release the state (e.g., locked state), and the host device may resend the failed (and any later) command. Such techniques may maintain the integrity of the data stored at the memory device 310.

[0095] In other cases, the memory device 310 may apply error detection codes (e.g., parity or CRC protection) to internal buses, such as row and column address buses. Random errors (e.g., bit flips) may occur in the transmitted set, row, and column addresses, which can result in incorrect sections of the open, accessed, or closed array. The error detection codes applied to these buses can protect the buses from such errors and add an additional level of error protection. In some instances, the trigger for applying the error detection code to the internal buses may be based on a signal from the host device or other factors, including the number of error thresholds detected in the data read from a particular application or a combination thereof in which the memory device 310 is in use.

[0096] The host device may analyze the failures and errors reported by the memory device and determine the corrective actions to take. As an example, the host device may determine to refrain from using a particular memory bank in the memory or may determine to use only a portion of the memory device 310. In cases where the corrective action of avoiding or deactivating the memory device 310 that causes or regularly experiences errors results in a system performance degradation, the system may gradually deactivate optional features and keep the mandatory (e.g., mission-critical) features present. In some cases, the host device may use the received indication as a trigger to perform additional operations, such as a service routine to identify the cause and / or solution of the indicated failure.

[0097] In some instances, checking and reporting errors can take time and reduce bandwidth (e.g., average memory bandwidth). Thus, the logic in the memory device 310 can be used to enable and disable the various security features described herein, which may be based on the error tolerance of the application or data. For example, critical program code may not tolerate any failures, and each of the security features described herein may be fully enabled. In contrast, pixel data may not be considered critical (pixel failures may be difficult to detect and pixel data can be overwritten or replaced in the memory device 310 at a relatively high rate (e.g., 60 frames per second)). In some cases, an application may utilize the maximum bandwidth supported by the memory device 310 and may thus disable the error detection, correction, and management schemes described herein.

[0098] Additionally or alternatively, a certain combination of the described techniques may be dynamically enabled or disabled at the memory device 310. More generally, some applications may use more bandwidth and thus may implement fewer error management techniques, while the need for reliability and data integrity may be more important for other applications, and thus the memory bandwidth may be less important. Thus, the same memory device 310 can be used for various applications or systems without the need to customize the memory device 310 for each application or each system. The memory device 310 can thus be used in a variety of scenarios (from PCs to motor vehicles), which can avoid the manufacturing and design costs for specialized memory systems.

[0099] Figure 4 A block diagram of an error management component 400 that supports dynamic control of error management and messaging according to aspects disclosed herein. The error management component 400 can be an instance of a component of a memory device, such as the memory device 110 described with reference to Figure 1 or the memory device 310 described with reference to Figure 3 In some instances, the error management component 400 can be an instance of the error counter 340 described with reference to Figure 3 The error management component 400 can support using corresponding error thresholds for different memory banks. Thus, the error management component 400 can include a plurality of error logic groups 405 (e.g., error logic groups 405-a to 405-n) corresponding to respective portions (e.g., memory banks, arrays) of the memory device.

[0100] Each error logic group 405 can include an error counter 410, a comparator 415, and an error count register 420 (which can store a value of one or more bits representing a maximum error count (i.e., a threshold)). As an example, the first error logic group 405-a can include a first error counter 410-a, a first comparator 415-a, and a first error count register 420-a, where the first error counter 410-a, the first comparator 415-a, and the first error count register 420-a perform an error counting function for data retrieved from the first memory bank. Similarly, the second error logic group 405-b can include a second error counter 410-b, a second comparator 415-b, and a second error count register 420-b for the second memory bank. The maximum error count (e.g., error threshold) for each error logic group 405 can be written to the corresponding error count register 420, for example, by a command decoder (such as the command decoder 315 described with reference to Figure 3 Alternatively, a central register can hold the corresponding threshold for each error logic group 405.

[0101] In some instances, the maximum error count can be individually configured for each error logic group 405 based on the sensitivity of the data stored in the corresponding portion of the memory. As an example, the host device can configure the first error logic group 405-a with a first error threshold and the second error logic group 405-b with a second error threshold. In some cases, the configured error thresholds can be based on the data type stored at the corresponding portion of the memory device. As an example, the first error logic group 405-a can be associated with a memory bank that contains mission-critical data (e.g., data associated with high-priority applications / operations), while the second error logic group 405-b can be associated with a memory bank that contains image, music, or video data. Thus, due to the more stringent tolerance for mission-critical data, the first threshold can be configured to be lower than the second threshold. In some cases, by using the described error thresholds for each group, each group can be dynamically adjusted based on the application of the memory device.

[0102] In some cases, each error logic group 405 can have a first (e.g., default) maximum error count, and the host device can modify the first (e.g., default) maximum error count of the corresponding error logic group 405 based on the data written to the corresponding memory bank. For example, the default maximum error count can be a certain pre-configured amount. In other cases, the default configuration can be that the memory device does not have an error threshold amount for each memory bank and the configuration can update each error count register 420 with an available error threshold amount. In some instances, a lookup table can be used to configure the maximum error count of each error count register 420 and other instances. As an example, the configuration can notify the memory device of a specific lookup table that indicates which thresholds can be set for each error logic group 405. In some cases, different lookup tables can be used for different applications.

[0103] When retrieving data from a portion of the memory, the error counter 410 can detect and count the number of errors in the data. The count can be incremented with each error detected. Additionally, the number of errors can be signaled by the error counter 410 to the comparator 415, and the comparator 415 can compare the error count with the maximum error count threshold received from the error count register 420 (e.g., by calculating the difference between the current error count and the maximum error count). The result of the comparison can then be sent to the error identification manager 425. Thus, each comparator 415-a, 415-b through 415-n can emit the result of its own comparison to the error identification manager 425.

[0104] The error identification manager 425 can determine whether one or more of the error logic groups 405 have an error count that meets the maximum error threshold of the error logic group 405 based on the signals received from each comparator 415. In some cases, the error identification manager 425 can implement Boolean logic, such as an OR operation, to determine whether one or more of the comparators 415 have detected an error count that meets the error threshold. For example, the outputs of comparators 415-a through 415-n are ORed, and if the maximum error count of any of the N error logic groups 405 is exceeded, then the error identification manager 425 can determine that the error set meets the configured threshold. It should be noted that the OR operation is an example of the logic that can be implemented by the error identification manager 425, and different logics or algorithms can be used to determine whether one or more error logic groups 405 have met the error threshold.

[0105] In cases where the maximum error count has been reached, the error identification manager 425 can transmit an indication that the threshold has been met to the host device and can provide other information, such as the error logic group (and memory bank) associated with the error. For example, the information sent to the host device can include various degrees of granularity, such as information for each failed address or a general indication that a failure has occurred at the memory device. The host device can use this information to analyze the failure and determine how to correct the failure, or to minimize additional failures, or both. In other cases, the memory device can correct errors in the data retrieved from the memory bank, but the host device can be aware of the corrected errors based on the messaging from the error identification manager 425.

[0106] In some instances, the memory device can be configured to determine which errors to report to the host device. For example, the memory device can identify an error set in one or more of the error logic groups 405, and based on the data type, data orientation, or a combination thereof, and other instances, the memory device can determine which errors to report to the host device. Additionally or alternatively, in cases where any of the error logic groups 405 (e.g., error logic groups 405-a through 405-n) have errors that meet the corresponding threshold, the host can be notified. The host device can also configure the memory device such that certain errors (e.g., double-bit errors, errors of a first type) are notified to the host and other errors (e.g., errors of a second type) are not. Thus, the configuration can dynamically modify which errors the memory device messages, which can reduce the messaging overhead within the system.

[0107] Such techniques can also be used to avoid locking a memory device when lower priority data fails, especially in cases where data of the type can be corrected. In some cases, if an error is detected in a particular group initially, then the threshold can be adjusted based on the error, for that group, one or more other groups, or both. For example, if a memory group appears to be failing based on error detection, then a lower threshold can be set to ensure that errors are captured earlier, such that the host device can determine to avoid using the memory group.

[0108] Figure 5 Illustrates a process flow 500 in a system supporting dynamic control of error management and messaging in accordance with aspects disclosed herein. In some instances, aspects of the process flow 500 may be implemented by a host device (e.g., a controller) and a memory device, which may be examples of the corresponding devices described in reference Figure 1 However, the operations and aspects described herein are not limited to using these components, and specifically contemplate other alternatives and fall within the scope of the concepts disclosed herein. The process flow may illustrate programmable error thresholds and features of a process for correcting errors and preventing additional errors in the system.

[0109] At 505, the host device may determine a value of an error count, for example, as one instance, a maximum error count, and configure the memory device with the value. For example, the host device may write the value of an error threshold to a register of the memory device (e.g., one or more mode registers or error count register 420 described in reference Figure 4 The maximum error count may correspond to an error threshold for a particular type of data (e.g., mission-critical data), or for data stored within a portion of the memory device (e.g., a particular memory group). In some cases, the configuration of the thresholds and behavior of the memory device may be adjustable during operation of the memory device.

[0110] At 510, the memory device may check for errors in data retrieved from the memory array. For example, as part of a read operation, the memory device may retrieve data from one or more memory arrays and may utilize EDC to identify one or more errors in the retrieved data. In some cases, the memory device may utilize ECC to identify and correct errors detected in the data (e.g., using the ECC logic described in reference Figure 3 When an error is identified, the memory device may increment an error counter. Specifically, whenever data is read from the memory, the memory device may check the data for errors. If an error exists, then the error count may be compared to a maximum allowable error count (e.g., an error threshold). In some cases, the memory device may report the error count in response to a command (e.g., from the host device).

[0111] At 515, the memory device may determine whether an error count meets a configured threshold received from the host device at 505. If the error count does not meet the threshold (e.g., the number of errors is less than the threshold), then the memory device may continue to monitor data retrieved from the memory array and detect errors within the data. The memory device may accordingly continue to increment the error count in the event that any other errors are detected.

[0112] If the error count does meet the configured threshold (e.g., the number of errors is greater than or equal to the threshold), then the memory device may transmit an indication that the threshold has been met to the host device. In some cases, the signal sent to the host device may include information associated with the errors detected at the memory device.

[0113] As an example, the information may include an indication of the data where the error was encountered, or the location of the retrieved data, the time of failure occurrence (e.g., based on a clock signal), etc. In some cases, at 520, the memory device may enter a locked state based on the maximum error count being met. The locked state may preserve the state of the memory such that errors can be identified and corrective actions can be taken to prevent additional errors. The locked state may include performing a self-refresh to ensure that data in the memory is not otherwise lost. In some instances, the memory device may report to the host device that the memory device has entered the locked state. The indication of the locked state may be transmitted together with the indication that the threshold has been met or may be transmitted separately to the host device. In some cases, the locked state may include a self-refresh state. In the locked state, the memory device may turn off all memory banks and stop executing commands.

[0114] At 525, the host device may perform an analysis of the reported errors and determine the cause of failure. In some instances, the host may analyze the status of the errors, for example, by reading the information received from the memory device. The failure analysis performed by the host device may assist in determining how to manage the errors encountered (e.g., adjusting the refresh rate or the operating frequency or both). Additionally, at 530, the host device may determine whether the errors are recoverable based on the information received from the memory device. Recoverable errors may be errors that can be corrected by retransmitting the failed data. In other instances, recoverable errors may be data corruption that can be corrected by ECC. Recoverable data may also include data that is stored at a redundant level and can be retrieved from another location (or another memory device) of the memory device.

[0115] If the host detects an irrecoverable failure, then the memory device (or memory device portion) may be deactivated. For example, at 535, the memory device may be deactivated based on an irrecoverable error identified in signaling from the memory device, and the host device may later use a different memory device for additional access operations. Such techniques may enable the system to remain in operation while avoiding portions of the memory associated with data errors.

[0116] Alternatively, the host device may determine that the error is recoverable, and the host device may reset the error count. For example, at 540, the host device may reset the error count in one or more error groups (such as the error groups described with reference to Figure 4 ). At 545, the host device may reset the lock based on the recoverable error. In such cases, the host device may transmit signaling or commands transitioning out of the locked state to the memory device.

[0117] At 550, the host device may reissue the command (e.g., from the time of the failure, for a duration before and after the failure). For example, an error may have been detected in response to a first access command, and the host device may have later issued additional access commands. However, because the memory device entered the locked state at 520, the additional access commands may have been ignored or not executed based on the previously detected error meeting a threshold. Thus, the host device may reissue the first command and / or additional commands at 550.

[0118] At 555, the system and host device may continue operation. For example, the memory device may continue to receive and execute commands from the host device. When a read command is received, the memory device may check for errors as it did at 510 and may determine, for example, at 515 whether a maximum error count has been met, etc.

[0119] Figure 6 A process flow 600 in a system supporting error management and dynamic control of signaling in accordance with aspects disclosed herein is shown. In some instances, aspects of the process flow 600 may be implemented by a controller 605 and a memory device 610, which may be examples of the corresponding devices described with reference to Figure 1 . The disclosure herein is not limited to examples including a controller or a memory device or both. The operations and aspects described herein are not limited to the use of these components, and other alternatives are contemplated. The process flow 600 may illustrate the use of EDC protection for commands received at the memory device.

[0120] At 615, the controller 605 may transmit an access command to the memory device 610. In some cases, the transmitted command may include one or more bits for checking the integrity of the command when the memory device 610 receives the command. For example, the access command may include an error detection code that includes one or more parity or other check bits.

[0121] At 620, the controller 605 may optionally store the command transmitted to the memory device 610. As an example, the controller 605 may store the last M commands sent to the memory device 610, where the commands may be retained until the controller 605 receives an indication that the received command is error free (e.g., based on the error detection code included with each command). In some cases, the commands may be buffered by the controller 605 when transmitted to the memory device 610.

[0122] At 625, the memory device 610 may determine whether the access command includes an error based on the error detection code associated with the access command. For example, the memory device may perform a parity check or a checksum on the received command. If the parity check or checksum passes, then it may be determined that the received command is error free. Alternatively, if the parity check fails, then it may be determined that the received command has an error.

[0123] In the case where the access command includes an error, the memory device may optionally enter a locked state at 630. Thus, the memory device may refrain from executing any further access commands. For example, the controller may transmit a second access command at 635, but the second access command may not be executed because the memory device is in the locked state. In some cases, the controller 605 may not know about the error until it receives an indication from the memory device 610.

[0124] At 640, the memory device 610 may transmit to the controller 605 an indication of whether an error has been detected in the access command based on the access command and the determination. For example, the memory device 610 may transmit an indication that the access command is error free at 625 based on the determination. In such cases, the memory device may execute the access command based on the determination that the access command does not include an error. In this way, the controller may receive an indication that the access command received at the memory device 610 is error free, in which case the stored version of the transmitted command may be removed from storage at the controller 605.

[0125] Alternatively, the memory device 610 may transmit an indication that the access command includes one or more errors at 625 based on the determination. In such a case, the controller 605 may identify the error in the access command at 645, which may be based on the indication received from the memory device 610. For example, the memory device 610 may include information about the error, how the error was detected in the received access command, when the access command was received, etc. in the indication. Thus, the controller 605 may determine that the error can be corrected by, for example, re-transmitting the same command. Therefore, at 650, the controller 605 may transmit a request to exit the locked state to the memory device 610.

[0126] After the locked state at 655, the memory device 610 may immediately be able to receive additional instructions from the controller 605. At 660, the controller 605 may re-transmit the access command (e.g., the command sent at 615) and any later commands (e.g., the command sent at 635) that may not have been executed due to the locked state of the memory device 610. In such a case, the memory device 610 may continue to operate, including checking the received access command for errors using the parity bit set, and the memory device may provide a signal indicating that an error-free access command has been received.

[0127] Figure 7 A flowchart illustrating a method 700 for supporting dynamic control of error management and signaling in accordance with aspects disclosed herein is shown. Operations of method 700 may be implemented by the memory device or its components described with reference to Figures 1 - 6 described. For example, the operations of method 700 may be performed by the memory device 310 described with reference to Figure 3 or the memory device 610 described with reference to Figure 6 described. In some instances, the memory device may execute an instruction or set of codes to control functional elements of the memory device to perform the functions described herein.

[0128] At 705, the memory device may receive a configuration of an error threshold for data at the memory device from a host device (e.g., a controller). For example, the configuration of the error threshold may indicate a maximum number of errors for a particular type of data, or for data stored at a particular location in the memory device, or a combination thereof. In this way, programmable thresholds may be set differently for data associated with different operations, thereby enabling dynamically controllable techniques for managing data failures in the memory device. The operations of 705 may be performed according to the method described with reference to Figures 1 - 6 described.

[0129] At 710, the memory device may determine a set of errors in data retrieved from the memory device. For example, when data is obtained from one or more memory arrays (e.g., in response to an access command), the memory device may identify one or more errors in the retrieved data (e.g., using an error detection code). The memory device may accordingly count the one or more identified errors. Additionally or alternatively, the memory device may identify and correct errors (e.g., using an error correction code) and count the number of corrected errors in the data. The operation of 710 may be performed according to the method described in reference Figures 1 - 6 The operation of 710 may be performed according to the method described in reference

[0130] At 715, the memory device may determine that the set of errors meets a threshold based on the configuration. In such cases, the memory device may compare the counted errors in the data with a configured threshold received from the host device. Thus, if the identified set of errors meets the configured threshold, the memory device may determine that the errors have exceeded an allowable number of failures in the data or the location storing the data. The operation of 715 may be performed according to the method described in reference Figures 1 - 6 The operation of 715 may be performed according to the method described in reference

[0131] At 720, the memory device may transmit an indication that the threshold has been met to the host device based on determining that the set of errors meets the threshold. In some instances, the memory device may transmit the indication in response to determining that the error threshold has been met. Additionally or alternatively, the memory device may transmit the indication in response to an associated command received from the host device. The operation of 720 may be performed according to the method described in reference Figures 1 - 6 The operation of 720 may be performed according to the method described in reference

[0132] In some instances, a device as described herein may perform one or more methods, such as method 700. The device may include features, means, or instructions for the following operations (e.g., instructions executable by a processor stored on a non-transitory computer-readable medium): receiving a configuration of an error threshold for data at the memory device from a host device; determining a set of errors in data retrieved from the memory device; determining that the set of errors meets the threshold based on the configuration; and transmitting an indication that the threshold has been met to the host device based on determining that the set of errors meets the threshold.

[0133] Some instances of the method 700, device, and non-transitory computer-readable medium described herein may further include operations, features, means, or instructions for the following operation: receiving a command from the host device to transmit the indication that the threshold may have been met, wherein the indication may be transmitted in response to the command.

[0134] Some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein may additionally include operations, features, apparatus, or instructions for: detecting the set of errors for a set (e.g., memory bank) in the memory device, wherein determining that the set of errors meets the threshold may be based on detecting that the set has the set of errors; and transmitting an indication of the set of errors to the host device.

[0135] Some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein may additionally include operations, features, apparatus, or instructions for: incrementing a counter for each error in the set of errors in the data based on determining the set of errors; and operating the memory device in a first mode that inhibits execution of access commands based on incrementing the counter for each error in the set of errors.

[0136] Some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein may additionally include operations, features, apparatus, or instructions for: receiving a command to reset the counter from the host device based on the set of errors being recoverable; and operating the memory device in a second mode in which access commands are executed based on the set of errors being recoverable.

[0137] Some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein may additionally include operations, features, apparatus, or instructions for: deactivating the memory device (e.g., by entering a locked state) based on the set of errors being non-recoverable.

[0138] Some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein may additionally include operations, features, apparatus, or instructions for: determining the set of errors based on an error correction code associated with the data retrieved from the memory device; correcting the set of errors in the data using the error correction code, wherein correcting the set of errors may be based on receiving an access command or a command to refresh a portion of the memory device; and transmitting the data including the corrected set of errors to the host device.

[0139] Some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein may additionally include operations, features, apparatus, or instructions for: determining that a first subset of the set of errors is a correctable error based on an error correction code; and inhibiting transmission of an indication of the first subset of correctable errors based on the determination that the first subset of the set of errors may be a correctable error.

[0140] Some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein may additionally include operations, features, apparatus, or instructions for: determining, based on the error correction code, that a second subset of the error set may be uncorrectable errors; and transmitting an indication of the uncorrectable errors to the host device based on the determination that the second subset of the error set may be uncorrectable errors.

[0141] Some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein may additionally include operations, features, apparatus, or instructions for: adjusting a refresh rate of the memory device based on detecting the error set. Some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein may additionally include operations, features, apparatus, or instructions for: identifying a temperature of the memory device (e.g., a junction temperature); and transmitting an indication of the temperature of the memory device to the host device based on detecting the error set. In some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein, the error threshold includes an error threshold quantity, or an error threshold type, or an error threshold at an orientation of the memory device, or an error threshold frequency, or a combination thereof.

[0142] Some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein may additionally include operations, features, apparatus, or instructions for: receiving a signal indicating whether to perform the determination that the error set meets the threshold, or whether to perform the transmission of the indication, or a combination thereof, where the signal may be based on the data stored at the memory device; and determining whether to perform the determination that the error set meets the threshold, or whether to perform the transmission of the indication, or a combination thereof, based on the signal. In some examples of the method 700, apparatus, and non-transitory computer-readable medium described herein, the configuration of the threshold may be based on a type of the data, or an orientation of the data at the memory device, or a combination thereof.

[0143] Figure 8 A flowchart illustrating a method 800 for supporting dynamic control of error management and messaging in accordance with aspects disclosed herein. Operations of method 800 may be implemented by the memory device or components thereof described with reference to Figures 1 - 6 described. For example, operations of method 800 may be performed by the memory device 310 described with reference to Figure 3 or the memory device 610 described with reference to Figure 6 described. In some examples, the memory device may execute an instruction or set of codes to control functional elements of the memory device to perform the functions described herein.

[0144] At 805, a memory device may receive an access command from a host device to access data at the memory device. The access command may include instructions to read or write data to a memory array, for example. The operations at 805 may be performed according to the method described in reference Figures 1 - 6 The operations at 810 may be performed according to the method described in reference

[0145] At 810, the memory device may determine whether the access command includes an error based on an error detection code associated with the access command or a bus of the memory device. In some instances, the error detection code includes one or more check bits. For example, the error detection code may be a parity bit appended to the access command, and the memory device may determine whether a parity check passes or fails based on the parity bit included. In some cases, the error detection code may be associated with a particular bus of the memory device and may be used to determine whether there is an error in an address (e.g., a row address, a column address) used in response to the received access command (where an access command without an error may have been received). That is, the propagation of the command and address to the memory array may have an error that can be detected by using the error detection code. The operations at 810 may be performed according to the method described in reference Figures 1 - 6 The operations at 815 may be performed according to the method described in reference

[0146] At 815, the memory device may transmit an indication of whether an error is detected in the access command or the bus based on the access command and the determination. In such cases, if the access command or the bus includes an error identified based on the error detection code, then the memory device may notify the host device that the command includes an error. Additionally or alternatively, the indication may signal that an error has been introduced into an address, which may have caused a conflict in storing or obtaining data in response to the access command. The operations at 815 may be performed according to the method described in reference Figures 1 - 6 The operations at 815 may be performed according to the method described in reference

[0147] In some instances, a device as described herein may perform one or more methods, such as method 800. The device may include features, means, or instructions (e.g., instructions executable by a processor stored on a non-transitory computer-readable medium) for: receiving an access command from a host device to access data at the memory device; determining whether the access command includes an error based on an error detection code associated with the access command or a bus of the memory device, the error detection code including one or more check bits; and transmitting an indication of whether an error is detected in the access command based on the access command and the determination.

[0148] Some examples of the method 800, apparatus, and non-transitory computer-readable medium described herein may further include operations, features, apparatus, or instructions for: determining that the access command does not include an error based on the error detection code, and performing the access command based on the determination. Some examples of the method 800, apparatus, and non-transitory computer-readable medium described herein may further include operations, features, apparatus, or instructions for: transmitting an indication that the access command is error-free to the host device, at least in part based on the determination.

[0149] Some examples of the method 800, apparatus, and non-transitory computer-readable medium described herein may further include operations, features, apparatus, or instructions for: determining that the access command includes an error based on the error detection code, and refraining from performing the access command based on the determination. Some examples of the method 800, apparatus, and non-transitory computer-readable medium described herein may further include operations, features, apparatus, or instructions for: transmitting an indication that the access command includes the error to the host device, at least in part based on the determination.

[0150] Some examples of the method 800, apparatus, and non-transitory computer-readable medium described herein may further include operations, features, apparatus, or instructions for: transitioning, by the memory device, to a state of refraining from performing the access command and one or more additional commands received after the access command; receiving, based on the error being correctable, a command to exit the state from the host device; and exiting the state based on the command and receiving a retransmission of the access command with the corrected error.

[0151] Figure 9 FIG. 900 is a flow diagram illustrating a method 900 that supports dynamic control for error management and messaging in accordance with aspects disclosed herein. The operations of method 900 may be implemented by the controller or components thereof described with reference to Figures 1 - 6 For example, the operations of method 900 may be performed by the controller 605 described with reference to Figure 6 In some examples, the controller may execute a set of code to control the functional elements of a device (e.g., a memory device, which may include the memory device 110 described with reference to Figure 1 to perform the functions described herein.

[0152] At 905, the controller may determine a configuration of an error threshold for data at a memory device. In such cases, the controller may determine an application, or a type of data associated with an access command. For example, the data may be associated with a mission-critical operation, and the controller may configure a relatively low threshold based on the tolerance of the mission-critical operation. Alternatively, the data may be associated with image rendering or video streaming, and the threshold may be relatively high. The operation of 905 may be performed according to the method described in reference Figures 1 - 6 The operation of 905 may be performed according to the method described in reference

[0153] At 910, the controller may transmit the configuration of the threshold to the memory device based on an access operation associated with the data. The configuration may be transmitted to the memory device via a channel coupled to the memory device. The operation of 910 may be performed according to the method described in reference Figures 1 - 6 The operation of 910 may be performed according to the method described in reference

[0154] At 915, the controller may receive an indication that a threshold has been met from the memory device based on data retrieved at the memory device. In such cases, the memory device may determine that the threshold has been met when retrieving data from a memory array (or a portion of the memory array) in response to a read command. Additionally, based on the met threshold, the memory device may enter a locked state in which future commands are not executed. The controller may identify an error, signal the memory device to exit the locked state and retransmit a command to obtain corrected data from the memory device. The operation of 915 may be performed according to the method described in reference Figures 1 - 6 The operation of 915 may be performed according to the method described in reference

[0155] In some instances, a device as described herein may perform one or more methods, such as method 900. The device may include features, means, or instructions for the following operations (e.g., instructions executable by a processor stored on a non-transitory computer-readable medium): determining a configuration of an error threshold for data at a memory device; transmitting the configuration of the threshold to the memory device based on an access operation associated with the data; and receiving an indication that a threshold has been met from the memory device based on data retrieved at the memory device.

[0156] Some instances of the method 900, device, and non-transitory computer-readable medium described herein may additionally include operations, features, means, or instructions for the following operation: determining the configuration of the threshold based on the type of the data, or the location of the data at the memory device, or a combination thereof.

[0157] Some examples of the method 900, apparatus, and non-transitory computer-readable medium described herein may further include operations, features, means, or instructions for: receiving an indication of one or more errors associated with a first orientation of the memory device from the memory device; and based on the indication of the one or more errors, writing additional data to a second orientation of the memory device, the second orientation being different from the first orientation.

[0158] Some examples of the method 900, apparatus, and non-transitory computer-readable medium described herein may further include operations, features, means, or instructions for: analyzing, based on the received indication, one or more errors associated with the threshold in the data; and based on analyzing the one or more errors, transmitting a command including corrected data to the memory device.

[0159] Figure 10 A flowchart illustrating a method 1000 for supporting dynamic control of error management and messaging in accordance with aspects disclosed herein. Operations of method 1000 may be implemented by a controller or components thereof as described with reference to Figures 1 - 6 described. For example, operations of method 1000 may be performed by a controller 605 as described with reference to Figure 6 described. In some instances, the controller may execute a set of code to control functional elements of a device (e.g., a memory device, which may include a memory device 110 as described with reference to Figure 1 described) to perform the functions described herein.

[0160] At 1005, the controller may generate an error detection code associated with an access command for accessing data at a memory device. For example, the error detection code may include one or more check bits that enable a receiving device (e.g., the memory device) to determine whether the access command was affected by an error during or after transmission. The operation of 1005 may be performed in accordance with the method described with reference to Figures 1 - 6 described.

[0161] At 1010, the controller may transmit the access command including the error detection code to the memory device. In such a case, the controller may use one or more channels coupled to the memory device to transmit the access command. The operation of 1010 may be performed in accordance with the method described with reference to Figures 1 - 6 described.

[0162] At 1015, the controller may receive an indication of whether an error has been detected in an access command from the memory device based on error detection. For example, the controller may receive an indication that one or more previously issued commands contained an error when received. The controller may reissue any command that contains an error. In a case where repeated commands always have an error, the controller may take steps to deactivate a particular memory device and use another memory device for access operations. The operations of 1015 may be performed according to the method described with reference to Figures 1 - 6 The method described in

[0163] In some instances, a device as described herein may perform one or more methods, such as method 1000. The device may include features, means, or instructions for the following operations (e.g., instructions executable by a processor stored on a non-transitory computer-readable medium): generating an error detection code associated with an access command for accessing data at a memory device; transmitting the access command including the error detection code to the memory device; and receiving, based on the error detection code, an indication of whether an error has been detected in the access command from the memory device.

[0164] Some instances of method 1000, devices, and non-transitory computer-readable media described herein may additionally include operations, features, means, or instructions for the following operations: receiving, based on the error detection code, an indication that an error has been detected in the access command; identifying the error in the access command based on the indication; and retransmitting the access command with the identified corrected error.

[0165] Some instances of method 1000, devices, and non-transitory computer-readable media described herein may additionally include operations, features, means, or instructions for the following operations: transmitting an indication to exit a state that inhibits execution of commands to the memory device, where the indication to exit the state may be based on an error set in the access command being correctable. In some instances of the methods, devices, and non-transitory computer-readable media described herein, the error detection code includes one or more parity bits associated with the access command.

[0166] Note that the method descriptions herein describe possible implementations, and the operations and steps may be rearranged or otherwise modified, and other implementations are possible. Additionally, aspects from two or more of the methods may be combined.

[0167] Any of a variety of different technologies and techniques may be used to represent the information and signals described herein. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof. Some of the figures may illustrate a signaling as a single signal; however, those of ordinary skill in the art will understand that a signal may represent a bus of signals, where the bus may have various bit widths.

[0168] As used herein, the term "virtual ground" refers to a node of a circuit that is held at a voltage approximately zero volts (0V) but is not directly coupled to ground. Thus, the voltage of the virtual ground may fluctuate over time and return to approximately 0V in a steady state. A virtual ground may be implemented using various electronic circuit components such as, for example, a voltage divider consisting of an operational amplifier and resistors. Other implementations are possible. "Virtual ground" or "virtual earth ground" refers to being connected to approximately 0V.

[0169] The terms "electronically communicate", "electrically contact", "connect", and "couple" may refer to a relationship between components that supports the flow of electrons between the components. Components are considered to be in electronic communication (or in electrical contact, or connected, or coupled) with each other if there is any conductive path between the components that can support the flow of signals between the components at any time. At any given time, based on the operation of the device that includes the connected components, the conductive path between components that are in electronic communication (or in electrical contact or connected or coupled) with each other may be an open circuit or a closed circuit. The conductive path between the connected components may be a direct conductive path between the components, or the conductive path between the connected components may be an indirect conductive path that may include intermediate components such as switches, transistors, or other components. In some cases, the flow of signals between the connected components may be interrupted for a period of time using, for example, one or more intermediate components such as switches or transistors.

[0170] The term "couple" refers to the condition of moving from an open circuit relationship between components, where a signal currently cannot be transmitted between the components through a conductive path, to a closed circuit relationship between the components, where a signal can be transmitted between the components through a conductive path. When a component such as a controller couples other components together, the component initiates a change that allows a signal to flow between the other components via a conductive path that was previously not allowed to carry the signal.

[0171] The term "isolated" refers to the relationship between components where a signal cannot currently flow between the components. If there is an open circuit between components, the components are isolated from each other. For example, components separated by a switch positioned between two components are isolated from each other when the switch is open. When a controller isolates two components, the controller effects the following change: preventing a signal from flowing between the components using a conductive path that previously permitted signal flow.

[0172] As used herein, the term "layer" refers to a layer or sheet of a geometric structure. Each layer can have three dimensions (e.g., height, width, and depth) and can cover at least a portion of a surface. For example, a layer can be a three-dimensional structure where two dimensions are greater than the third dimension, such as a thin film. A layer can contain different elements, components, and / or materials. In some cases, a layer can be composed of two or more sub-layers. In some of the figures, two dimensions of a three-dimensional layer are depicted for illustrative purposes. However, those skilled in the art will recognize that a layer is three-dimensional in nature.

[0173] As used herein, the term "electrode" can refer to an electrical conductor and, in some cases, can serve as an electrical contact to other components of a memory cell or memory array. An electrode can include traces, wires, conductive lines, conductive layers, etc., which provide a conductive path between elements or components of a memory array.

[0174] As used herein, the term "lithography" can refer to a process of patterning using a photoresist material and exposing such material using electromagnetic radiation. For example, a photoresist material can be formed on a substrate material by, for example, spin-coating the photoresist on the substrate material. A pattern can be created in the photoresist by exposing the photoresist to radiation. For example, the pattern can be defined by a photomask that spatially depicts where the radiation exposes the photoresist. For example, the exposed photoresist regions can then be removed by chemical treatment, leaving the desired pattern. In some cases, the exposed regions can be retained and the unexposed regions can be removed.

[0175] As used herein, the term "shorted" refers to the relationship between components where a conductive path is established between the components by activating a single intermediate component between the two components under discussion. For example, a first component shorted to a second component can exchange signals with the second component when a switch between the two components is closed. Thus, shorting can be a dynamic operation that enables charge flow between components (or lines) in an electronic communication.

[0176] The devices discussed in this document include a memory array and can be formed on a semiconductor substrate such as silicon, germanium, silicon-germanium alloy, gallium arsenide, gallium nitride, etc. In some cases, the substrate is a semiconductor wafer. In other situations, the substrate can be a silicon-on-insulator (SOI) substrate, such as silicon-on-glass (SOG) or silicon-on-sapphire (SOP), or an epitaxial layer of semiconductor material on another substrate. The conductivity of the substrate or a sub-region of the substrate can be controlled by doping with various chemicals including but not limited to phosphorus, boron, or arsenic. Doping can be performed by ion implantation or by any other doping method during the initial formation or growth of the substrate.

[0177] The switching components or transistors discussed in this document can represent field-effect transistors (FETs) and include three-terminal devices comprising a source, a drain, and a gate. The terminals can be connected to other electronic components by conductive materials such as metals. The source and the drain can be conductive and can include heavily doped, e.g., degenerate, semiconductor regions. The source and the drain can be separated by a lightly doped semiconductor region or channel. If the channel is n-type (e.g., the majority carriers are electrons), then the FET can be called an n-type FET. If the channel is p-type (i.e., the majority carriers are holes), then the FET can be called a p-type FET. The channel can be capped by an insulating gate oxide. The channel conductivity can be controlled by applying a voltage to the gate. For example, applying a positive voltage or a negative voltage to an n-type FET or a p-type FET respectively can cause the channel to become conductive. When a voltage greater than or equal to the threshold voltage of the transistor is applied to the transistor gate, the transistor can be "turned on" or "activated". When a voltage less than the threshold voltage of the transistor is applied to the transistor gate, the transistor can be "turned off" or "deactivated".

[0178] The description presented in this document in conjunction with the figures describes example configurations and does not represent all examples that can be implemented or that are within the scope of the claims. The term "exemplary" as used in this document means "serving as an example, instance, or illustration" and is not "preferred over" or "superior to" other examples. The detailed description includes specific details to provide an understanding of the described technology. However, these technologies can be practiced without these specific details. In some cases, well-known structures and devices are shown in block diagram form to avoid obscuring the concepts of the described examples.

[0179] In the figures, similar components or features can have the same reference numeral. Additionally, various components of the same type can be distinguished by following the reference numeral with a dash and a second numeral that differentiates among the similar components. If only the first reference numeral is used in the specification, the description applies to any one of the similar components having the same first reference numeral, regardless of the second reference numeral.

[0180] Any of a variety of different technologies and techniques can be used to represent the information and signals described herein. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.

[0181] The various illustrative blocks and modules described in connection with the disclosure herein can be implemented or executed using a general purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration).

[0182] The techniques described herein can be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions can be stored on or transmitted through a computer-readable medium as one or more instructions or code. Other examples and implementations are within the scope of the present disclosure and the appended claims. For example, due to the nature of software, the functions described herein can be implemented using software, hardware, firmware, hardwiring, or any combination of these. The features implementing the functions can also be physically located at various positions, including being distributed such that portions of the functions are implemented at different physical locations. And, as used herein, including in the claims, the "or" in a list of items (e.g., a list of items that begins with a phrase such as "at least one of" or "one or more of") indicates an inclusive list, such that a list of at least one of A, B, or C means A or B or C or AB or AC or BC or ABC (i.e., A and B and C). Additionally, as used herein, the phrase "based on" should not be construed as referring to a closed set of conditions. For example, without departing from the scope of the present disclosure, an exemplary step described as "based on condition A" can be based on both condition A and condition B. In other words, as used herein, the phrase "based on" should be interpreted in the same way as the phrase "at least partially based on".

[0183] A computer-readable medium includes both a non-transitory computer storage medium and a communication medium including any medium that facilitates transfer of a computer program from one place to another. The non-transitory storage medium can be any available medium that can be accessed by a general purpose or special purpose computer. By way of example, and not limitation, non-transitory computer-readable media can include RAM, ROM, electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD) ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to carry or store the desired program code in the form of instructions or data structures and that can be accessed by a general purpose or special purpose computer or a general purpose or special purpose processor. Also, any connection is properly termed a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include CD, laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above are also included within the scope of computer-readable media.

[0184] The description provided herein enables one skilled in the art to make or use the present disclosure. Various modifications to the present disclosure will be apparent to those skilled in the art, and the general principles defined herein can be applied to other variations without departing from the scope of the present disclosure. Thus, the present disclosure is not limited to the examples and designs described herein, but is accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method, comprising: Receive a configuration of an error threshold for data at a memory device; Use error correction code logic to generate an error correction code for the data at least in part based on a received command; Use the error correction code logic to detect one or more errors in the data at least in part based on generating the error correction code; Increment an error counter at least in part based on detecting the one or more errors; Determine that the error counter meets the threshold at least in part based on the configuration; and Transmit an indication that the error counter meets the threshold at least in part based on determining that the error counter meets the threshold.

2. The method according to claim 1, further comprising: Receive a second command to transmit the indication that the error counter meets the threshold, wherein the indication is transmitted in response to the second command.

3. The method according to claim 1, further comprising: Transmit an indication of the one or more errors.

4. The method according to claim 1, further comprising: Cause the memory device to operate in a first mode that inhibits execution of access commands at least in part based on incrementing the error counter.

5. The method according to claim 4, further comprising: Receive a second command to reset the error counter at least in part based on the one or more errors being recoverable; and Cause the memory device to operate in a second mode in which access commands are executed at least in part based on the one or more errors being recoverable.

6. The method according to claim 4, further comprising: Disable the memory device at least in part based on the one or more errors being non-recoverable.

7. The method according to claim 1, further comprising: Correct the one or more errors in the data using the error correction code, wherein correcting the one or more errors is at least in part based on receiving an access command or a second command to refresh a portion of the memory device; and Transmit the data including the corrected one or more errors.

8. The method according to claim 1, further comprising: Determine that a first error among the one or more errors is a correctable error at least in part based on the error correction code; and Suppress transmission of an indication of the first error at least in part based on determining that the first error is a correctable error.

9. The method according to claim 8, further comprising: Determine that a second error among the one or more errors is an uncorrectable error at least in part based on the error correction code; and Transmit an indication of the second error at least in part based on determining that the second error is an uncorrectable error.

10. The method according to claim 1, further comprising: Adjust a refresh rate of the memory device at least in part based on detecting the one or more errors.

11. The method according to claim 1, further comprising: Identify a temperature of the memory device; and Transmit an indication of the temperature of the memory device at least in part based on detecting the one or more errors.

12. The method according to claim 1, further comprising: Receive a signal indicating whether to perform the determination that the error counter meets the threshold, or whether to perform the transmission of the indication, or a combination thereof, wherein the signal is at least in part based on the data stored at the memory device; and Determine whether to perform the determination that the error counter meets the threshold, or whether to perform the transmission of the indication, or a combination thereof, at least in part based on the signal.

13. The method according to claim 1, wherein the threshold includes an error threshold quantity, or an error threshold type, or an error threshold at the orientation of the memory device, or an error threshold frequency, or a combination thereof.

14. The method according to claim 1, wherein the configuration of the threshold is at least partially based on the type of the data, or the orientation of the data at the memory device, or a combination thereof.

15. A method, comprising: Receive an access command to access data at a memory device; Determine that the access command contains an error at least in part based on an error detection code associated with the access command or a bus of the memory device, the error detection code including one or more parity bits; Transmit an indication that the error is detected in the access command, at least partially based on the access command and the determination; and Perform one or more self - refresh operations, at least partially based on transmitting the indication that the error is detected in the access command.

16. The method according to claim 15, further comprising: Receive a second access command for accessing second data at the memory device; Determine that the second access command does not contain an error, at least partially based on an error - detection code associated with the second access command or the bus of the memory device; Execute the second access command, at least partially based on determining that the second access command does not contain an error, at least partially based on the error - detection code, wherein transmitting the indication includes: Transmit an indication that the second access command is error - free, at least partially based on the determination.

17. The method according to claim 15, further comprising: Suppress execution of the access command, at least partially based on determining that the access command contains the error, at least partially based on the error - detection code, wherein the indication is transmitted at least partially based on the determination.

18. The method according to claim 17, further comprising: Transition, by the memory device, to a state of suppressing execution of the access command and one or more additional commands received after the access command; Receive a command to exit the state, at least partially based on the error being correctable; and Receive a re - transmission of the access command for which the error has been corrected, at least partially based on exiting the state based on the command.

19. A method, comprising: Determine a first threshold number of errors for first data at a memory device, wherein the first threshold number of errors is based on a first application associated with the first data, and wherein the first threshold number of errors is different from a second threshold number of errors based on a second application associated with second data at the memory device; Generate an error - correction code for the first data; Transmit an indication of the first threshold number of errors to the memory device, at least partially based on an access operation associated with the first data; and Receive an indication that the first threshold number of errors has been satisfied from the memory device, at least partially based on retrieving the first data and the error - correction code at the memory device.

20. The method according to claim 19, further comprising: Determine the first threshold number of errors, at least partially based on the type of the first data, or the orientation of the first data at the memory device, or a combination thereof.

21. The method according to claim 19, further comprising: Receive an indication of one or more errors associated with a first orientation of the memory device from the memory device; and Write additional data to a second orientation of the memory device, the second orientation being different from the first orientation, at least partially based on the indication of the one or more errors.

22. The method according to claim 19, further comprising: Analyze one or more errors in the first data associated with the first threshold number of errors, at least partially based on the received indication; and Transmit a command including the corrected first data to the memory device, at least partially based on analyzing the one or more errors.

23. A method, comprising: Generate an error - detection code associated with an access command for accessing data at a memory device; Transmit the access command including the error - detection code to the memory device; and Receive an indication that an error has been detected in the access command from the memory device, at least in part based on the error detection code, wherein the memory device is configured to perform one or more self-refresh operations at least in part based on detecting the error in the access command.

24. The method according to claim 23, further comprising: Identify the error in the access command, at least in part based on the indication; and Re-transmit the access command for which the identified error has been corrected.

25. The method according to claim 23, further comprising: Transmit an indication to the memory device to exit a state that inhibits execution of commands, wherein the indication to exit the state is at least in part based on the error set in the access command being correctable.

Citation Information

Patent Citations

  • Apparatus, method and system to determine memory access command timing based on error detection

    US20140211579A1

  • Hybrid Memory System With Configurable Error Thresholds And Failure Analysis Capability

    US20140281661A1

  • Semiconductor memory devices including error correction circuits and methods of operating the semiconductor memory devices

    US20170192721A1