Techniques for handling errors in persistent storage
By introducing controllers in NVDIMMs, scanning and processing uncorrected errors and converting them into accessible SPAs, the data loss and system inaccessibility caused by persistent memory errors in NVDIMMs are solved, and higher data reliability and system stability are achieved.
Patent Information
- Application Number
- CN202010757667.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2014-06-30
- Filing Date
- 2015-05-28
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2036-03-12
AI Technical Summary
Uncorrected errors in persistent memory in NVDIMMs can lead to data loss and system inaccessibility, and prior art is difficult to effectively handle these errors.
By introducing a controller in NVDIMM, receiving error scan requests, scanning the DPA range of nonvolatile memory, identifying uncorrected errors, and converting them into accessible SPAs, stored in the data structure of the computing platform so that the operating system or device driver can avoid mapping to the error address.
It effectively reduces the risk of data loss, ensures that the system can correctly handle uncorrected errors after power outage or reset, and avoids the problem that the system cannot access the persistent memory address.
Smart Images

Figure CN112131031B_ABST
Abstract
Description
Technical Field
[0001] Examples described herein generally relate to handling errors of persistent memory in a non-volatile dual in-line memory module (NVDIMM). Background Art
[0002] Memory modules coupled to a computing platform or system (e.g., those configured as servers) may include dual in-line memory modules (DIMMs). DIMMs may include volatile memory types of such dynamic random access memory (DRAM) or other types of memory, such as non-volatile memory. As DRAM and other types of memory technologies have evolved to include memory cells with increasingly higher densities, the memory capacity of DIMMs has also increased greatly. Because DRAM is a volatile memory, if not all data is maintained in the DRAM during power failure or reset, a power failure or reset may result in the loss of most of the data. Some non-volatile memory technologies may also utilize encryption schemes, which may result in the loss of encrypted information after power failure or reset and because the encrypted data may not be accessible due to the loss of the encrypted information, the non-volatile memory for these technologies may thus act as some kind of volatile memory. In addition, the large memory capacity of these types of memory technologies poses challenges to the operating system (OS) or application (e.g., device driver) in sensing power failure and attempting to prevent or reduce data loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Figure 1 An example first system is shown.
[0004] Figure 2 An example second system is shown.
[0005] Figure 3 The third system is shown as an example.
[0006] Figure 4 Diagram showing an example process.
[0007] Figure 5 An example first state machine is shown.
[0008] Figure 6 An example second state machine is shown.
[0009] Figure 7 An example block diagram for a first device is illustrated.
[0010] Figure 8 An example of a first logic flow is illustrated.
[0011] Fig. 9 An example of the first storage medium is illustrated.
[0012] Fig.10 An example block diagram of a second device is illustrated.
[0013] Fig.11 An example of the second logic flow is illustrated.
[0014] Fig.12 An example of the second storage medium is illustrated.
[0015] Fig.13 Illustration of an example computing platform.
[0016] Fig.14 An example non-volatile dual in-line memory module controller is shown. DETAILED DESCRIPTION
[0017] In some examples, in order to mitigate or reduce data loss in the event of a power outage or reset, a memory module of this type including both volatile memory and non-volatile memory is developed. This type of memory module is collectively referred to as a non-volatile DIMM (NVDIMM). Typically, an NVDIMM can be a combination of DRAM and a certain type of non-volatile memory (e.g., NAND flash memory) or can be a combination of DRAM and a certain type of non-volatile memory technology (e.g., 3D crosspoint memory), which can act as both volatile and non-volatile memory. NVDIMMs can provide persistent memory by using non-volatile memory to capture an image of the volatile memory contents (e.g., a checksum indication) after the NVDIMM is powered off due to an intentional or unintentional system or component / device reset. A supercapacitor package can be coupled to the NVDIMM to maintain power to the NVDIMM for a period of time long enough to capture an image of the volatile memory contents to the non-volatile memory.
[0018] According to some examples, an NVDIMM that includes non-volatile memory to provide persistent memory for a memory type at the NVDIMM, such as DRAM or some types of non-volatile memory (e.g., 3D crosspoint memory), may have substantially unrecoverable uncorrected errors in the volatile memory. However, because the error is due to a data error or the basic input / output system (BIOS) of the computing platform coupled to the NVDIMM may cause the memory address that is considered to be an error to be easily mapped outside the system memory, these uncorrected errors may disappear when the NVDIMM is restarted or reset. However, uncorrected errors for non-volatile memory used as persistent memory may not disappear upon restart or reset. Because data loss may occur, the BIOS may not be able to clear these uncorrected errors. Therefore, uncorrected errors in non-volatile memory used as persistent memory may have the possibility of causing a "Groundhog Day" scenario, which may cause the system at the computing system to be unable to access non-volatile memory address locations with uncorrected errors. Software such as an operating system (OS) or device drivers requires a priori knowledge of uncorrected errors in non-volatile memory used as persistent storage so that decisions can be made not to map or access persistent storage address locations having uncorrected errors.
[0019] In some examples, an OS or device driver for a computing platform coupled to an NVDIMM may not be able to access device physical addresses (DPAs) because these addresses are maintained within the NVDIMM and are not exposed to the OS or device driver. System physical addresses (SPAs) may also include multiple NVDIMMs and / or cross-NVDIMMs, while DPAs are device-specific. Therefore, identified uncorrected errors with DPAs need to be converted to SPAs so that the OS or device driver can decide not to map or access persistent memory address locations with uncorrected errors. These and other challenges require the examples described herein.
[0020] Techniques for handling errors in persistent memory included in an NVDIMM may be implemented via one or more example methods. A first example method may include a controller resident on the NVDIMM receiving an error scan request. In response to the error scan request, a DPA range of non-volatile memory at the NVDIMM that can provide persistent memory to the NVDIMM may be scanned by the controller. Uncorrected errors may be identified in the DPA range and the DPA address with the identified uncorrected errors may be indicated to a BIOS of a computing platform coupled to the NVDIMM.
[0021] A second example may include a BIOS configured to be implemented by circuitry at a host computing platform. The BIOS may send an error scan request to a controller residing on an NVDIMM coupled to the host computing platform. The NVDIMM may have a non-volatile memory capable of providing persistent memory to the NVDIMM. For this second example method, the BIOS may be able to determine that the controller completes a scan of a DPA range of the non-volatile memory in response to the error scan request and accesses a first data structure residing at the NVDIMM to read a DPA with an identified uncorrected error identified by the controller during the scan of the DPA range. For this second example method, the BIOS may also be able to convert a DPA with an identified uncorrected error into a SPA with an identified uncorrectable error.
[0022] Figure 1 The first system is shown in the figure. Figure 1 As shown in FIG. 1 , an example first system includes system 100. In some examples, such as in Figure 1 As shown in FIG. 1 , system 100 includes a host computing platform 110 coupled to NVDIMMs 120-1 through 120-n, where “n” is any positive integer having a value greater than 3. For these examples, NVDIMMs 120-1 through 120-n may be coupled to host computing platform 110 via communication channels 115-1 through 115-n, as shown in FIG. Figure 1 The host computing platform 110 may include, but is not limited to, a server, a server array or server farm, a web server, a network server, an Internet server, a workstation, a microcomputer, a mainframe computer, a supercomputer, a network appliance, a web appliance, a distributed computing system, a multiprocessor system, a processor-based system, or a combination thereof.
[0023] In some examples, such as in Figure 1 As shown in FIG. 1 , NVDIMMs 120-1 to 120-n may include respective controllers 122-1 to 122-n, volatile memories 124-1 to 124-n, and non-volatile memories 126-1 to 126-n. Communication channels 115-1 to 115-n may include memory channels, communication and / or control links coupled between elements of the host computing platform 110 and respective NVDIMMs 120-1 to 120-n to enable communication with and / or access to elements of these respective NVDIMMs. As described more below, communication between elements of a host computing platform (e.g., host computing platform 110) may allow for handling of uncorrected errors in a non-volatile memory (e.g., non-volatile memory 126-1) that is capable of providing persistent memory of a type of volatile memory (e.g., volatile memory 124-1) resident on the NVDIMM.
[0024] In some examples, although not in Figure 1 As shown in FIG. 1 , at least some of the memory channels included in the communication channels 115 - 1 to 115 - n may be interleaved to provide increased bandwidth for elements of the host computing platform to access the volatile memories 124 - 1 to 124 - n .
[0025] Figure 2 The second system is shown in the figure. Figure 2 As shown in FIG. 1 , the example second system includes system 200. In some examples, such as in Figure 2 As shown in FIG. 2 , system 200 includes a host computing platform 210 coupled to NVDIMM 205 via a communication channel 215. Also in FIG. Figure 2 As shown in FIG. 1 , capacitor bank 270 can be coupled to NVDIMM 205 via power link 277. In some examples, such as in Figure 2 As shown in FIG. 2 , the NVDIMM 205 may further include a host interface 220 , a controller 230 , a control switch 240 , a volatile memory 250 , or a non-volatile memory 260 .
[0026] In some examples, the host computing platform 210 may include circuitry 212 that is capable of executing various functional elements of the host computing platform 110, which may include, but are not limited to, a basic input / output system (BIOS) 214, an operating system (OS) 215, a device driver 216, or an application (App) 218. The host computing platform 210 may also include a data structure 213. As described more below, the data structure 213 may be accessible to an element of the host computing platform 210 (e.g., the BIOS 214 or the OS 215) and may be configured to at least temporarily store the SPA associated with the uncorrected site information for persistent storage provided by the non-volatile memory 260 to the volatile memory 250.
[0027] According to some examples, such as Figure 2, the host interface 220 at the NVDIMM 205 may include a controller interface 222 and a memory interface 224. In some examples, the controller interface 222 may be an SMBus interface designed or operated in compatibility with the SMBus Specification, Version 2.0, August 2000 ("SMBus Specification"). For these examples, elements of the host computing platform 210 may communicate with the controller 230 through the controller interface 222. Elements of the host computing platform 210 may also access the volatile memory 250 through the memory interface 224 through the control channel 227, through the control switch 240, and then through the control channel 247. In some examples, in response to a power loss or reset of the NVDIMM 205, access to the volatile memory 250 may be switched by the control switch 240 to the controller 230 (via the control channel 237) to use the memory channel 255 coupled between the volatile memory 250 and the non-volatile memory 260 to restore or save the contents of the volatile memory 250 from or to the non-volatile memory 260. By saving or restoring the contents of the volatile memory 250 prior to the power loss or reset, the non-volatile memory 260 may be able to provide persistent memory to the NVDIMM 205.
[0028] According to some examples, such as Figure 2 2, the controller 230 may include a data structure 232 and a circuit 234. The circuit 234 may be capable of executing a component or feature to receive an error scan request from an element of the host computing platform 210 (e.g., the BIOS 214). As described more below, the error scan request may be about scanning a DPA range of the non-volatile memory 260. The feature or component may also be capable of identifying uncorrected errors in the DPA range and then indicating to the BIOS 214 the DPA with the identified uncorrected errors.
[0029] In some examples, data structure 232 may include registers that may be selectively asserted to indicate DPAs with identified uncorrected errors to BIOS 214. For these examples, BIOS 214 may access these registers via controller interface 222. The registers may also be selectively asserted to indicate error scan status to BIOS 214 and / or to indicate whether the capacity of data structure 232 is sufficient to indicate all of the DPAs with identified uncorrected errors.
[0030] In some examples, BIOS 214 may include logic and / or features implemented by circuitry 212 to send an error scan request to controller 230. For these examples, the logic and / or features of BIOS 214 may be able to access data structure 232 to determine whether controller 230 has completed scanning the DPA range of non-volatile memory 260 and also access the data structure to read the DPA with identified uncorrected errors identified by the controller during scanning the DPA range.
[0031] According to some examples, BIOS 214 may access the DPA from data structure 232 and / or have the ability to read the DPA from data structure 232. Other elements of host computing platform 210 (e.g., OS 215, device driver 216, or App 218) may not have the ability to access or read the DPA from data structure 232. For these examples, logic and / or features of BIOS 214 may be able to convert the DPA with identified uncorrected errors into SPAs with identified uncorrected errors. BIOS 214 may then cause the SPAs with identified uncorrected errors to be stored in data structure 213 at host computing platform 210. Data structure 213 may be accessible to other elements of host computing platform 210 and the SPAs may also be readable by these other elements. For example, OS 215, device driver 216, or App 218 may be able to use the SPAs with identified uncorrected errors to avoid mapping system memory of host computing platform 210 to those SPAs with identified uncorrected errors.
[0032] In some examples, BIOS 214 may also read the DPAs from data structure 232 and use identified uncorrected errors for these DPAs to avoid mapping any BIOS-related activities to these DPAs for a current or subsequent boot or reset of system 200 and / or NVDIMM 205 .
[0033] In some examples, data structure 213 may include registers that may be selectively asserted by BIOS 214 to indicate a SPA with identified uncorrected errors to OS 215, device driver 216, or App 218. Registers may also be selectively asserted to indicate an error scan status of controller 230, the status of any conversion of a DPA with identified uncorrected errors to a SPA with identified uncorrected errors by BIOS 214, and / or to indicate whether the converted SPA does not include at least some uncorrected errors (e.g., because data structure 232 has insufficient capacity to indicate all of the DPAs with identified uncorrected errors).
[0034] In some examples, the volatile memory 250 may include volatile memory that is designed or operated compatibly with one or more standards or specifications (including descendants or variants) associated with various types of volatile memory (e.g., DRAM). For example, a DRAM type (e.g., synchronous double data rate DRAM (DDR DRAM)) may be included in the volatile memory 250 and the standards or specifications associated with DDR DRAM may include those published by the JEDEC Solid State Technology Association ("JEDEC") for various DDR generations (e.g., DDR2, DDR3, DDR4, or future DDR generations). Some example standards or specifications may include, but are not limited to, JESD79-3F-"DDR3 SDRAM Standard" published in July 2012, or JESD79-4-"DDR4 SDRAM Standard" published in September 2012.
[0035] According to some examples, non-volatile memory 260 may include one or more types of non-volatile memory to include, but are not limited to, NAND flash memory, NOR flash memory, 3D crosspoint memory, ferroelectric memory, silicon-oxide-nitride-oxide-silicon (SONOS) memory, polymer memory (e.g., ferroelectric polymer memory), ferroelectric transistor random access memory (FeTRAM or FeRAM), Osram memory, or nanowires. In some examples, non-volatile memory 260 may also include sufficient memory capacity to receive the entire contents of volatile memory 250 or possibly multiple copies of the contents of volatile memory 250.
[0036] In some examples, capacitor bank 270 may include one or more capacitors to provide at least temporary power to NVDIMM 205 via power link 277. The one or more capacitors may be capable of storing sufficient energy to power NVDIMM 205 for a sufficient time to enable controller 230 to cause data maintained in volatile memory 250 to be saved to non-volatile memory 260 if a sudden power outage or system reset causes the main power supply to NVDIMM 205 to be cut off or shut down. Saving of data content to non-volatile memory 260 due to a sudden power outage or system reset may be referred to as a "catastrophic save."
[0037] Figure 3 The third system is shown in the figure. Figure 3 As shown in FIG. 1 , the example third system includes system 300. In some examples, such as in Figure 3 As shown in FIG. 3 , system 300 includes a host computing platform 310 coupled to NVDIMM 305 via a communication channel 315 . Figure 3 As shown in FIG. 1 , capacitor bank 370 can be coupled to NVDIMM 305 via power link 377. In some examples, such as in Figure 3 As shown in FIG. 3 , the NVDIMM 305 may also include a host interface 320 , a controller 330 , a control switch 340 , or a memory pool 350 (which has a volatile memory 352 portion and a persistent memory 354 portion).
[0038] In some examples, the host computing platform 310 may include circuitry 312 that is capable of executing various functional elements of the host computing platform 110, which may include, but are not limited to, BIOS 314, OS 315, device drivers 316, or App 318. The host computing platform 310 may also include a data structure 313. As described more below, the data structure 313 may be accessible to an element of the host computing platform 310 (e.g., BIOS 314 or OS 315) and may be configured to at least temporarily store a SPA associated with uncorrected error location information of the persistent memory 354.
[0039] According to some examples, such as Figure 3 As shown in FIG. 3 , the host interface 320 at the NVDIMM 305 may include a controller interface 322 and a memory interface 324. In some examples, the controller interface 322 may be an SMBus interface. For these examples, elements of the host computing platform 310 may communicate with the controller 330 through the controller interface 322. Elements of the host computing platform 310 may also access the volatile memory 352 through the memory interface 324 through the control channel 327, through the control switch 340, and then through the control channel 347. In some examples, in response to the NVDIMM 305 being powered off or reset, access to the volatile memory 352 may be switched by the control switch 340 to the controller 330 (through the control channel 337) to restore or save the contents of the volatile memory 352 from the persistent memory 354 to the persistent memory 354. By saving or restoring the contents of the volatile memory 352 before the power failure or reset, the persistent memory 354 may be able to provide persistent memory to the NVDIMM 305.
[0040] According to some examples, such as Figure 3 3, controller 330 may include data structure 332 and circuit 334. Circuit 334 may be capable of executing a component or feature to receive an error scan request from an element of host computing platform 310 (e.g., BIOS 314). As described more below, the error scan request may be about scanning a DPA range of persistent memory 354. The component or feature may also be capable of identifying uncorrected errors in the DPA range and then indicating to BIOS 314 the DPA with the identified uncorrected errors.
[0041] In some examples, data structure 332 may include registers that may be selectively asserted to indicate DPAs with identified uncorrected errors to BIOS 314. For these examples, BIOS 314 may access these registers through controller interface 322. The registers may also be selectively asserted to indicate error scan status to BIOS 314 and / or to indicate whether the capacity of data structure 332 is sufficient to indicate all of the DPAs with identified uncorrected errors.
[0042] In some examples, BIOS 314 may include logic and / or features implemented by circuitry 312 to send an error scan request to controller 330. For these examples, the logic and / or features of BIOS 314 may be able to access data structure 332 to determine whether controller 330 has completed scanning the DPA range of persistent memory 354 and also access data structure 332 to read the DPA with identified uncorrected errors identified by controller 330 during scanning the DPA range.
[0043] According to some examples, BIOS 314 may access the DPA from data structure 332 and / or have the ability to read the DPA from data structure 332. Other elements of host computing platform 310 (e.g., OS 315, device driver 316, or App 318) may not have the ability to access or read the DPA from data structure 332. For these examples, logic and / or features of BIOS 314 may be able to convert the DPA with identified uncorrected errors into SPAs with identified uncorrected errors. BIOS 314 may then cause the SPAs with identified uncorrected errors to be stored in data structure 313 at host computing platform 310. Data structure 313 may be accessible to other elements of host computing platform 310 and the SPAs may also be readable by these other elements. For example, OS 315, device driver 316, or App 318 may be able to use the SPAs with identified uncorrected errors to avoid mapping system memory of host computing platform 310 to those SPAs with identified uncorrected errors.
[0044] In some examples, BIOS 314 may also read DPAs from data structure 332 and use identified uncorrected errors for these DPAs to avoid mapping any BIOS-related activities to these DPAs for a current or subsequent boot or reset of system 300 and / or NVDIMM 305 .
[0045] In some examples, data structure 313 may include registers that may be selectively asserted by BIOS 314 to indicate a SPA with identified uncorrected errors to OS 315, device driver 316, or App 318. Registers may also be selectively asserted to indicate the error scan status of controller 330, the status of any conversion of a DPA with identified uncorrected errors by BIOS 214 to a SPA with identified uncorrected errors, and / or to indicate whether the converted SPA does not include at least some uncorrected errors (e.g., because data structure 232 has insufficient capacity to indicate all of the DPAs with identified uncorrected errors).
[0046] In some examples, memory pool 350 may include some type of non-volatile memory, such as, but not limited to, 3D cross point memory. In some examples, persistent memory 354 may also include sufficient memory capacity to receive the entire contents of volatile memory 352 or possibly multiple copies of the contents of volatile memory 352.
[0047] In some examples, capacitor bank 370 may include one or more capacitors to provide at least temporary power to NVDIMM 305 via power link 377. The one or more capacitors may be capable of storing sufficient energy to power NVDIMM 305 for a sufficient time to enable controller 330 to cause data maintained in volatile memory 352 to be saved to persistent memory 354 if a sudden power outage or system reset causes the main power supply to NVDIMM 305 to be cut off or shut down. Saving of data content to persistent memory 354 due to a sudden power outage or system reset may be referred to as a "catastrophic save."
[0048] Figure 4 An example process 400 is shown. In some examples, a system (e.g. Figure 1-3 Elements of the system 100, 200, or 300 shown in FIG. 4 may implement the process 400 to utilize a data structure (e.g., data structure 232 / 332 or 213 / 313) to identify and communicate error locations at a non-volatile or persistent memory (e.g., non-volatile memory 260 or persistent memory 354) that can provide persistent memory to an NVDIMM (e.g., NVDIMM 120-1 to 120-n or NVDIMM 205 / 305). Examples are not limited to systems (e.g., Figure 1-3 The elements of the system 100, 200 or 300 shown in Figure 1-3 The host computing platform or other elements of the NVDIMM are not shown.
[0049] According to some examples, in response to an error scan request from a BIOS (e.g., BIOS 214 / 314), logic and / or features of a controller (e.g., controller 230 / 330) may be able to scan a DPA range of a non-volatile or persistent memory (e.g., non-volatile memory 260 / persistent memory 254) and identify uncorrected errors in the DPA range. The logic and / or features of the controller may then be able to detect uncorrected errors via Figure 4 420-1, a second DPA for a second identified uncorrected error site may be indicated in field 420-2, and an mth DPA for an mth identified uncorrected error site may be indicated in field 420-m, where "m" is equal to any positive integer greater than 2. For these examples, data structure 323 / 332 may have a capacity limited to a number of "m" fields to indicate DPAs with identified uncorrected errors. In response to a DPA having identified uncorrected errors exceeding a number of "m" fields, logic and / or features of the controller may utilize overflow flag field 430 to indicate that the capacity of the first data structure is insufficient to indicate at least some of the DPAs with identified uncorrected errors.
[0050] In some examples, such as in Figure 4 As shown in FIG. 4 , the data structure 232 / 332 may also include an error scan status 410. For these examples, the logic and / or features of the controller may be capable of indicating the error scan status to the BIOS via the error scan status field 410. The error scan status may include, but is not limited to, an idle, in progress, or completed error scan status. The scan status included in the status field 410 may be periodically polled by the BIOS and / or the controller may send a system management interrupt (SMI) to the BIOS to indicate the scan status.
[0051] According to some examples, the logic and / or features of the BIOS may be able to read various fields of the data structure 232 / 332 to determine whether the controller has completed an error scan request, a DPA with an identified uncorrected error, or whether the controller indicates an overflow flag. For these examples, upon determining that the controller has completed an error scan, the logic and / or features of the BIOS may then convert the DPAs in fields 420-1 to 420-m to SPAs and cause these SPAs to be stored in the data structure 213 / 313. The SPAs may be stored in fields 450-1 to 450-m of the data structure 213 / 313 to indicate the SPAs with the identified uncorrected errors. In some examples, an OS, device driver, or application at a host computing platform may access these fields and use the access to avoid mapping system memory to those SPAs with the identified uncorrected errors.
[0052] In some examples, such as in Figure 4 As shown in FIG. 4 , data structure 213 / 313 may also include an error scan status field 440. For these examples, logic and / or features of the BIOS may be able to indicate an error scan status to the OS, device driver, or application via the error scan status field 440. The error scan status may include, but is not limited to, an idle state, an indication that the controller is currently performing an error scan, an indication that the controller has completed an error scan and a DPA to SPA conversion is in progress, or an indication that the controller has completed an error scan and the results are ready for reading.
[0053] According to some examples, logic and / or features of the BIOS may be able to indicate in the overflow flag field 460 whether the controller indicates that at least some uncorrected errors are not included for SPAs with identified uncorrected errors stored in the data structure 213 / 313. As mentioned above, at least some of the DPAs with identified uncorrected errors may not be included in the fields 420-1 to 420-m due to the DPAs having identified uncorrected errors exceeding a count value greater than the "m" value. For these examples, if the overflow flag field 460 indicates that some uncorrected errors are not included for SPAs in the fields 450-1 to 450-m, the OS, device driver, or application may read or copy the SPAs in these fields. The OS, device driver, or application may then request that the BIOS cause the controller to perform one or more additional error scans to indicate those DPAs with identified uncorrected errors that were not previously indicated due to overflow problems of the data structure 232 / 332 or to provide those DPAs if they are stored by the controller but not indicated due to overflow problems. The BIOS may convert these additional DPAs into SPAs, as mentioned above.
[0054] Figure 5 The first state machine of the illustrated example. In some examples, such as in Figure 5 As shown in FIG. 5 , the first state machine includes state machine 500. In some examples, an element of an NVDIMM (eg, NVDIMM 205 / 305) may have a controller, such as in Figure 2-3 Controller 230 / 330 shown in. For example, controller 230 / 330 may include circuit 234 / 334 to implement logic and / or features to indicate the status of error scanning operation according to state machine 500. Error scanning can identify and communicate error locations at non-volatile memory (e.g., non-volatile / persistent memory 260 / 354 that can provide persistent memory to NVDIMM 205 / 305).
[0055] According to some examples, the state machine 500 starts at state 510 (idle), and the logic and / or features implemented by the circuit to perform error scanning of the non-volatile memory can have an idle state. For these examples, as in Figure 5 As shown in , the idle state may be after an NVDIMM reset (eg, due to a power cycle) of the NVDIMM. The idle state may be indicated in a data structure (eg, data structure 232 / 332) at the NVDIMM.
[0056] Moving from state 510 to state 520 (error scan in progress), the logic and / or features implemented by the circuit may have an error scan in progress state, which includes scanning the DPA range for the non-volatile memory in response to a BIOS (e.g., BIOS 214 / 314) requesting an error scan. In some embodiments, the error scan in progress state may be indicated in a data structure.
[0057] From state 520, moving to state 530 (error scan complete), the logic and / or features implemented by the circuit may have an error scan complete state. In some examples, upon reaching this state, the logic and / or features may trigger a system management interrupt (SMI) to the BIOS. For these systems, the completion state at state 530 may be reached after indicating DPAs with uncorrected errors in a data structure so that the BIOS can access those DPAs in response to the triggered SMI. The logic and / or features implemented by the circuit may then move back to the idle state of state 510 in response to another NVDMM reset.
[0058] Figure 6 The second state machine of the diagram example. In some examples, such as in Figure 6 As shown in FIG. 1 , the second state machine includes state machine 600. In some examples, the BIOS of the host computing platform (e.g., Figure 2-3The BIOS 214 / 314 of the host computing platform 210 / 310 shown in FIG. 2 may have logic and / or features to indicate the status of the error scanning work of the controller, such as the controller 230 / 330 resident on an NVDIMM (e.g., NVDIMM 205 / 305) having non-volatile memory (e.g., non-volatile or persistent memory 260 / 354) capable of providing persistent memory to the NVDIMM. The indicated status may also include the status of the DPA to SAP conversion after the error scanning work is completed and when the results are received by other elements of the host computing platform (e.g., OS 215 / 315, device driver 216 / 316, or App 218 / 318, as shown in FIG. 2 ). Figure 2-3 ) read instructions.
[0059] According to some examples, the state machine 600 begins in state 610 (idle), and logic and / or features of the BIOS may indicate that error scanning of non-volatile or persistent memory at the NVDIMM is in an idle state. For these examples, as in Figure 6 As shown in , the idle state may be after an NVDIMM reset of the NVDIMM. The idle state may be indicated by these logics and / or features in a first data structure (eg, data structure 213 / 313) at the computing platform.
[0060] From state 610 to state 620 (NVDIMM error scan in progress), logic and / or features of the BIOS may indicate that an NVDIMM error scan is in progress by the controller. In some examples, such as in Figure 6 As shown in , this state can be entered after the BIOS requests an error scan. For these examples, the logic and / or features can indicate that the NVDIMM error scan is in progress in the first data structure.
[0061] Moving from state 620 to state 630 (NVDIMM error scan complete, DPA to SPA transition in progress), logic and / or features of the BIOS may indicate that the controller has completed the NVDIMM error scan and the DPA to SPA transition is in progress. In some examples, this state may be entered when the BIOS polls a second data structure (e.g., data structure 232 / 332) at the controller to determine that the error scan is complete or when the BIOS receives an SMI from the controller indicating completion.
[0062] Moving from state 630 to state 640 (NVDIMM Complete, Results Ready for Reading), logic and / or features of the BIOS may indicate that the NVDIMM error scan has been completed and the results are ready for reading by an element of the host computing platform (e.g., OS 215 / 315, device driver 216 / 316, or App 218 / 318). According to some examples, this state may be entered after completing a conversion of a DPA with identified uncorrected errors to a SPA with identified uncorrected errors. For these examples, the SPAs with identified uncorrected errors may be stored to a first data structure that may be accessible to the OS, device driver, or App of these elements of the host computing platform to read and then possibly use to avoid mapping system memory to those SPAs with identified uncorrected errors. The logic and / or features of the BIOS may then indicate an idle state of the error scan of the NVDIMM in response to another NVDIMM reset.
[0063] Figure 7 An example block diagram of the first device 700 is shown. Figure 7 As shown in FIG. 1 , the first device includes device 700. Although Figure 7 While apparatus 700 is shown in FIG. 7 as having a limited number of elements in a certain topology, it will be appreciated that apparatus 700 may include more or fewer elements in alternative topologies as desired for a given implementation.
[0064] The device 700 may be supported by a circuit 720 maintained at a controller for an NVDIMM that may be coupled to a host computing platform. The circuit 720 may be configured to execute one or more software or firmware implemented components 722-a. It is noteworthy that "a" and "b" and "c" and similar indicators as used herein are defined as variables representing any positive integer. Thus, for example, if the implementation sets a value of a=5, the complete set of software or firmware for component 722-a may include components 722-1, 722-2, 722-3, 722-4, or 722-5. The examples presented are not limited in this context and different variables used throughout may represent the same or different integer values.
[0065] According to some examples, the circuit 720 may include a processor or processor circuit. The processor or processor circuit may be any of a variety of commercially available processors, including without limitation AMD® Athlon®, Duron®, and Opteron® processors; ARM® application, embedded, and security processors; IBM® and Motorola® DragonBall® and PowerPC® processors; IBM and Sony® Cell processors; Intel® Atom®, Celeron®, Core (2) Duo®, Core i3, Core i5, Core i7, Itanium®, Pentium®, Xeon®, Xeon Phi®, and XScale® processors; and similar processors. According to some examples, the circuit 720 may also be an application specific integrated circuit (ASIC) and at least some of the components 722-a may be implemented as hardware elements of the ASIC.
[0066] According to some examples, the apparatus 700 may include a receiving component 722-1. The receiving component 722-1 may be executed by the circuit 720 to receive an error scan request. For these examples, the error scan request may be included in the scan request 710 and may be received from the BIOS of the host computing platform. The BIOS may be communicatively coupled to the controller via a controller interface (which may include an SMBus interface). The error scan request may be directed to scanning non-volatile memory at the NVDIMM, which may provide persistent memory to the NVDIMM.
[0067] In some examples, the apparatus 700 may also include an error component 722-2. The error component 722-2 may be executed by the circuit 720 to scan a DPA range of the non-volatile memory in response to an error scan request and identify uncorrected errors in the DPA range. For these examples, the scan may include scanning through the DPA range to generate data values that may be used to determine whether uncorrected errors occur at one or more DPAs included in the range. Uncorrected errors and their associated DPAs may then be identified based on the inability of the error correction circuit to correct errors that may be encoded using one or more types of error correction codes (e.g., Reed-Solomon codes). The DPA with the identified uncorrected errors may be maintained at least temporarily with the error location information 723-a (e.g., maintained in a lookup table (LUT)).
[0068] According to some examples, the apparatus 700 may also include an indication component 722-3. The indication component 722-3 may be executed by the circuit 720 to indicate to the BIOS of the host computing platform the DPA with the identified uncorrected error. For these examples, the indication component 722-3 may access the error location information 723-a and may indicate the DPA with the uncorrected error in a data structure that is also accessible to the BIOS at the controller. The location 730 may include those DPAs with the identified uncorrected errors. The indication component 722-3 may also indicate the status of the error scan of the error component 722-2 in the data structure at the controller. The status 740 may include a status indication.
[0069] In some examples, the apparatus 700 may also include an interrupt component 722 - 4 . The interrupt component 722 - 4 may be executed by the circuit 720 to send an SMI to the BIOS in response to completing the scan of the DPA range. For these examples, the SMI may be included in the SMI 750 .
[0070] In some examples, the apparatus 700 may also include a flag component 722-5. The flag component 722-5 may be executed by the circuit 720 to set a flag to indicate that the capacity of the data structure is insufficient to indicate at least some of the DPAs with the identified uncorrected errors. For these examples, the flag may indicate to the BIOS that some of the DPAs with the identified uncorrected errors are not indicated in the data structure. The flag may be included in the flag 760.
[0071] This article includes a logical flow set of example methods for performing novel aspects of the disclosed architecture. Although one or more methods shown herein are shown and described as a series of actions for the purpose of simplicity of explanation, those skilled in the art will understand and appreciate that these methods are not limited by the order of actions. Some actions may accordingly adopt a different order from other actions shown and described herein and / or occur simultaneously with other actions shown and described herein. For example, those skilled in the art will understand and appreciate that the method can alternatively be represented as a series of interrelated states or events, such as in a state diagram. In addition, all actions illustrated in the method may not be required for novel implementation.
[0072] The logic flow may be implemented in software, firmware, and / or hardware. In software and firmware embodiments, the logic flow may be implemented by computer executable instructions stored on at least one non-transitory computer readable medium or machine readable medium (e.g., optical, magnetic, or semiconductor storage). The embodiments are not limited in this context.
[0073] Figure 8 An example of the first logic flow is shown. Figure 8As shown in FIG. 8 , the first logic flow includes a logic flow 800. The logic flow 800 may represent some or all of the operations performed by one or more logics, features, or devices described herein (e.g., the apparatus 700). More specifically, the logic flow 800 may be implemented by a receiving component 722-1, an error component 722-2, an indication component 722-3, an interrupt component 722-4, or a flag component 722-5.
[0074] According to some examples, the logic flow 800 may receive an error scan request at a controller resident on the NVDIMM at block 802. For these examples, the receiving component 722-1 may receive the error scan request.
[0075] In some examples, the logic flow 800 may scan a DPA range of a non-volatile memory at the NVDIMM that can provide persistent memory to the NVDIMM in response to the error scan request at block 804. For these examples, the error component 722-2 may scan the DPA range.
[0076] According to some examples, the logic flow 800 may identify uncorrected errors in the DPA range at block 806. For these examples, the error component 722-2 may identify the DPA for the uncorrected errors in the DPA range.
[0077] In some examples, the logic flow 800 may indicate the DPA with the identified uncorrected error to a BIOS of a host computing platform coupled to the NVDIMM at block 808. For these examples, the indicating component 722-3 may indicate the DPA in a data structure accessible to the BIOS.
[0078] Fig. 9 An example of a first storage medium is shown. Fig. 9 As shown in , the first storage medium includes storage medium 900. The storage medium 900 may include an article of manufacture. In some examples, the storage medium 900 may include any non-temporary computer-readable medium or machine-readable medium, such as optical, magnetic or semiconductor storage. The storage medium 900 may store various types of computer-executable instructions, such as instructions that implement the logic flow 800. Examples of computer-readable or machine-readable storage media may include any tangible medium capable of storing electronic data, including volatile memory or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. Examples of computer-executable instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, object-oriented code, visual code and the like. Examples are not limited in this context.
[0079] Fig.10An example block diagram of the second device is shown. Fig.10 As shown in FIG. 1 , the second device includes device 1000. Although Fig.10 While apparatus 1000 is shown in FIG. 1 with a limited number of elements in a certain topology or configuration, it will be appreciated that apparatus 1000 may include more or fewer elements in alternative configurations as desired for a given implementation.
[0080] The device 1000 may be supported by a circuit 1020 maintained at a host computing platform to implement logic and / or features of a BIOS of the host computing platform. The circuit 1020 may be configured to execute one or more software or firmware implemented components 1022-a. It is noteworthy that "a" and "b" and "c" and similar indicators as used herein are defined as variables representing any positive integer. Thus, for example, if the implementation sets a value of a=7, a complete set of software or firmware for component 1022-a may include components 1022-1, 1022-2, 1022-3, 1022-4, 1022-5, 1022-6, or 1022-7. The examples presented are not limited in this context and different variables used throughout may represent the same or different integer values.
[0081] In some examples, such as in Fig.10 As shown in , device 1000 includes circuit 1020. Circuit 1020 can generally be configured to execute one or more software and / or firmware components 1022-a. Circuit 1020 can be part of the circuit of a host computing platform, which includes a processing core (e.g., used as a central processing unit (CPU)). Alternatively, circuit 1020 can be part of the circuit in a chipset of the host computing platform. In either scenario, circuit 1020 can be part of any of a variety of commercially available processors to include, but are not limited to, those mentioned previously for circuit 720 of device 700. Circuit 1020 can also be part of a dual microprocessor, a multi-core processor, and other multi-processor architectures. According to some examples, circuit 1020 can also be an ASIC and component 1022-a can be implemented as a hardware element of the ASIC.
[0082] According to some examples, the device 1000 may include a request component 1022-1. The request component 1022-1 may be executed by the circuit 1020 to send an error scan request to a controller resident on an NVDIMM coupled to a host computing platform (which includes the device 1000). The NVDIMM may have a non-volatile memory that can provide persistent memory to the NVDIMM. The error scan request may be for scanning the non-volatile memory for uncorrected errors. For these examples, the request component 1022-1 may access a data structure (e.g., a register) at the NVDIMM via a controller interface at the NVDIMM and may send a scan request 1010 in response to a NVDIMM reset 1005, which includes an error scan request.
[0083] In some examples, the apparatus 1000 may also include a status component 1022-2. The status component 1022-2 may be executed by the circuit 1020 to determine whether the controller has completed the error scan based on polling a field or register of a data structure at the NVDIMM. For these examples, the NVDIMM status 1015 may include the results of such polling of the data structure.
[0084] According to some examples, apparatus 1000 may also include SMI component 1022-3. SMI component 1022-3 may be executed by circuit 1020 to receive an SMI sent by the controller when error scanning is completed and a DPA with uncorrected errors is identified. For these examples, the SMI may be included in SMI 1040.
[0085] In some examples, the apparatus 1000 may also include a read component 1022-4. The read component 1022-4 may be executed by the circuit 1020 to access a data structure at the NVDIMM to read the DPA with the identified uncorrected errors identified by the controller during the scan of the DPA range. For these examples, the read DPA may be included in the DPA site 1035.
[0086] According to some examples, the apparatus 1000 may also include a conversion component 1022-5. The conversion component 1022-5 may be executed by the circuit 1020 to convert the DPA with the identified uncorrected errors into the SPA with the identified uncorrected errors. For these examples, the conversion component 1022-5 may at least maintain these SPAs together with the error location information 1023-a (e.g., in the LUT).
[0087] According to some examples, the apparatus 1000 may further include a storage component 1022-6. The storage component 1022-6 may be executed by the circuit 1020 to store the SPA with the identified uncorrected error to a data structure at the host computing platform. The data structure at the host computing platform may be accessible to the OS, device driver, or application of the host computing platform. For these examples, the storage component 1022-6 may first obtain the SPA from the error location information 1023-a and include the SPA with the identified uncorrected error in the SPA location 1045.
[0088] According to some examples, the apparatus 1000 may also include a flag component 1022-7. The flag component 1022-7 may be executed by the circuit 1020 to determine that the controller at the NVDIMM sets a flag in a data structure at the NVDIMM via a flag 1050 indicating that the capacity of the data structure at the NVDIMM is insufficient to indicate at least some of the DPAs with uncorrected errors. The flag component 1022-7 may then set another flag at the host computing platform via a flag 1055 indicating the same information. For these examples, the OS, device driver, or application may use the flag 1055 to determine that not all uncorrected errors are included for the SPA site 1045 and these elements of the host computing platform may react accordingly.
[0089] The various components of the device 1000 and the host computing platform (which includes the device 1000) can be communicatively coupled to each other through various types of communication media to coordinate operations. Coordination can involve unidirectional or bidirectional information exchange. For example, the components can convey information in the form of signals conveyed through the communication medium. The information can be implemented as signals assigned to various signal lines. In such an assignment, each message is a signal. However, other embodiments may alternatively use data messages. Such data messages can be sent across various components. Example connections include parallel interfaces, serial interfaces, and bus interfaces.
[0090] This article includes a logical flow set of example methods for performing novel aspects of the disclosed architecture. Although one or more methods shown herein are shown and described as a series of actions for the purpose of simplicity of explanation, those skilled in the art will understand and appreciate that these methods are not limited by the order of actions. Some actions may accordingly adopt a different order from other actions shown and described herein and / or occur simultaneously with other actions shown and described herein. For example, those skilled in the art will understand and appreciate that the method can alternatively be represented as a series of interrelated states or events, such as in a state diagram. In addition, all actions illustrated in the method may not be required for novel implementation.
[0091] The logic flow may be implemented in software, firmware, and / or hardware. In software and firmware embodiments, the logic flow may be implemented by computer executable instructions stored on at least one non-transitory computer readable medium or machine readable medium (e.g., optical, magnetic, or semiconductor storage). The embodiments are not limited in this context.
[0092] Fig.11 An example of the second logic flow is shown. Fig.11 , the second logic flow includes logic flow 1100. Logic flow 1100 may represent some or all of the operations performed by one or more logics, features, or devices described herein (e.g., apparatus 1000). More specifically, logic flow 1100 may be implemented by request component 1022-1, status component 1022-2, SMI component 1022-3, read component 1022-4, conversion component 1022-5, storage component 1022-6, or flag component 1022-7.
[0093] exist Fig.11 In the illustrated example shown in FIG, logic flow 1100 may send an error scan request from a circuit configured to implement a BIOS of a host computing platform to a controller resident on an NVDIMM coupled to the host computing platform at block 1102, the NVDIMM having non-volatile memory capable of providing persistent memory to the NVDIMM. For these examples, request component 1022-1 may cause the status request to be sent.
[0094] According to some examples, logic flow 1100 can determine that the controller has completed scanning the DPA range of the nonvolatile memory in response to the error scan request at block 1104. For these examples, status component 1022-2 can make this determination.
[0095] According to some examples, the logic flow 1100 can access a first data structure resident at the NVDIMM to read the DPA with the identified uncorrected errors identified by the controller during scanning the DPA range at block 1106. For these examples, the read component 1022-3 can access the first data structure to read the DPA.
[0096] In some examples, the logic flow 1100 may convert the DPA with the identified uncorrected errors to the SPA with the identified uncorrectable errors at block 1108. For these examples, the conversion component 1022-5 may convert the DPA to the SPA.
[0097] According to some examples, logic flow 1100 may store the SPA with the identified uncorrectable errors to a second data structure accessible to an operating system or device driver of the host computing device at block 1110. For these examples, storage component 1022-6 may cause the SPA to be stored to the second data structure.
[0098] Fig.12 An example of a second storage medium is shown. Fig.12 As shown in , the second storage medium includes storage medium 1200. Storage medium 1200 may include an article of manufacture. In some examples, storage medium 1200 may include any non-temporary computer-readable medium or machine-readable medium, such as optical, magnetic or semiconductor storage. Storage medium 1200 may store various types of computer-executable instructions, such as instructions that implement logic flow 1100. Examples of computer-readable or machine-readable storage media may include any tangible medium capable of storing electronic data, including volatile memory or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. Examples of computer-executable instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, object-oriented code, visual code and the like. Examples are not limited in this context.
[0099] Fig.13 An example computing platform 1300 is shown. In some examples, such as in Fig.13 As shown in FIG. 1 , computing platform 1300 may include processing component 1340, other platform components, or communication interface 1360. According to some examples, computing platform 1300 may be part of a host computing platform, as mentioned above.
[0100] According to some examples, the processing component 1340 can perform processing operations or logic for the device 1000 and / or the storage medium 1200. The processing component 1340 may include various hardware elements, software elements, or a combination of the two. Examples of hardware elements may include devices, logic devices, components, processors, microprocessors, circuits, processor circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), memory cells, logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. Examples of software elements may include software components, programs, applications, computer programs, applications, device drivers, system programs, software development programs, machine programs, operating system software, middleware, firmware, software components, routines, subroutines, functions, methods, procedures, software interfaces, application program interfaces (APIs), instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination thereof. Determining whether an example is implemented using hardware elements and / or software elements can vary depending on many factors, such as desired computing rates, power levels, thermal tolerances, processing cycle budgets, input data rates, output data rates, memory resources, data bus speeds, and other design or performance constraints, as desired for a given example.
[0101] In some examples, other platform components 1350 may include common computing elements, such as one or more processors, multi-core processors, co-processors, memory units, chipsets, controllers, peripherals, interfaces, oscillators, timing devices, video cards, audio cards, multimedia input / output (I / O) components (e.g., digital displays), power supplies, etc. Examples of memory cells may include, without limitation, various types of computer-readable and machine-readable storage media in the form of one or more higher-speed memory cells, such as read-only memory (ROM), random access memory (RAM), dynamic RAM (DRAM), double data rate DRAM (DDRAM), synchronous DRAM (SDRAM), static RAM (SRAM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, polymer memory (e.g., ferroelectric polymer memory), Austenite memory, phase change or ferroelectric memory, silicon-oxide-nitride-oxide-silicon (SONOS) memory, magnetic or optical cards, device arrays such as redundant array of independent disks (RAID) drives, solid-state memory devices (e.g., USB memory), solid-state drives (SSDs), and any other type of storage medium suitable for storing information.
[0102] In some examples, the communication interface 1360 may include logic and / or features to support the communication interface. For these examples, the communication interface 1360 may include one or more communication interfaces that operate according to various communication protocols or standards to communicate through direct or network communication links. Direct communication may occur via the use of communication protocols or standards described in one or more industrial standards (including progeny and variants), such as those associated with the SMBus specification, the PCI Express specification. Network communication may occur via the use of communication protocols or standards, such as those described in one or more Ethernet standards promulgated by the Institute of Electrical and Electronics Engineers (IEEE). For example, one such Ethernet standard may include IEEE 802.3-2008, issued in December 2008: Carrier Sense Multiple Access Collision Detection (CSMA / CD) Access Method and Physical Layer Specifications (hereinafter "IEEE 802.3").
[0103] The computing platform 1300 may be part of a computing device, which may be, for example, a server, a server array or server farm, a web server, a network server, an Internet server, a workstation, a microcomputer, a mainframe computer, a supercomputer, a network appliance, a web appliance, a distributed computing system, a multiprocessor system, a processor-based system, or a combination thereof. Thus, the functionality and / or specific configuration of the computing platform 1300 described herein may be included in or omitted from various embodiments of the computing platform 1300, as appropriately desired.
[0104] The components and features of computing platform 1300 may be implemented using any combination of discrete circuits, application specific integrated circuits (ASICs), logic gates, and / or single chip architectures. In addition, where appropriate, features of computing platform 1300 may be implemented using microcontrollers, programmable logic arrays, and / or microprocessors, or any combination of the foregoing. Note that hardware, firmware, and / or software elements may be collectively or individually referred to herein as "logic" or "circuitry."
[0105] It should be realized that Fig.13 The example computing platform 1300 shown in the block diagram of the embodiment can represent one functional description example of many potential implementations. Therefore, the division, omission or inclusion of block functions depicted in the drawings does not infer that the hardware components, circuits, software and / or elements used to implement these functions will necessarily be divided, omitted or included in the embodiment.
[0106] Fig.14 An example NVDIMM controller 1400 is shown. In some examples, such as in Fig.14As shown in FIG. 1 , NVDIMM controller 1400 may include processing component 1440, other platform components 1450, or communication interface 1460. According to some examples, NVDIMM controller 1400 may be implemented in an NVDIMM controller resident on or with an NVDIMM coupled to a host computing platform.
[0107] According to some examples, the processing component 1440 can perform processing operations or logic for the device 700 and / or the storage medium 900. The processing component 1440 may include various hardware elements, software elements, or a combination of the two. Examples of hardware elements may include devices, logic devices, components, processors, microprocessors, circuits, processor circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), memory cells, logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. Examples of software elements may include software components, programs, applications, computer programs, applications, device drivers, system programs, software development programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application program interfaces (APIs), instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination thereof. Determining whether an example is implemented using hardware elements and / or software elements can vary depending on many factors, such as desired computing rates, power levels, thermal tolerances, processing cycle budgets, input data rates, output data rates, memory resources, data bus speeds, and other design or performance constraints, as desired for a given example.
[0108] In some examples, other controller components 1450 may include common computing elements, such as one or more processors, multi-core processors, co-processors, memory units, interfaces, oscillators, timing devices, etc. Examples of memory units may include, without limitation, various types of computer-readable and machine-readable storage media in the form of one or more higher-speed memory units, such as ROM, RAM, DRAM, DDRAM, SDRAM, SRAM, PROM, EPROM, EEPROM, flash memory, or any other type of storage medium suitable for storing information.
[0109] In some examples, the communication interface 1460 may include logic and / or features to support the communication interface. For these examples, the communication interface 1460 may include one or more communication interfaces that operate according to various communication protocols or standards to communicate through a communication link or channel. Communication may occur via the use of communication protocols or standards described in one or more industry standards (including progeny and variants), such as those associated with the PCI Express specification or the SMBus specification.
[0110] The components and features of the NVDIMM controller 1400 may be implemented using any combination of discrete circuits, application specific integrated circuits (ASICs), logic gates, and / or single chip architectures. Additionally, where appropriate, the features of the NVDIMM controller 1400 may be implemented using a microcontroller, a programmable logic array, and / or a microprocessor, or any combination of the foregoing. Note that hardware, firmware, and / or software elements may be collectively or individually referred to herein as "logic" or "circuitry."
[0111] It should be realized that Fig.14 The example NVDIMM controller 1400 shown in the block diagram may represent one functional description example of many potential implementations. Therefore, the division, omission, or inclusion of block functions depicted in the drawings does not infer that the hardware components, circuits, software, and / or elements used to implement these functions will necessarily be divided, omitted, or included in the embodiments.
[0112] One or more aspects of at least one example may be implemented by representative instructions stored on at least one machine-readable medium, which represents various logic within the processor, which when read by a machine, computing device, or system causes the machine, computing device, or system to make logic to perform the techniques described herein. Such representations may be stored on tangible machine-readable media and supplied to various customers or manufacturing facilities to load into manufacturing machines that actually make the logic or processor.
[0113] Various examples can be implemented using hardware elements, software elements or a combination of the two. In some examples, hardware elements may include devices, components, processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, ASICs, PLDs, DSPs, FPGAs, memory cells, logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. In some examples, software elements may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, APIs, instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols or any combination thereof. Determining whether an example is implemented using hardware elements and / or software elements may vary according to many factors, such as desired computing rate, power level, heat resistance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed and other design or performance constraints, as desired for a given implementation.
[0114] Some examples may include an article of manufacture or at least one computer-readable medium. The computer-readable medium may include a non-transitory storage medium to store logic. In some examples, the non-transitory storage medium may include one or more types of computer-readable storage media capable of storing electronic data, including volatile memory or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. In some examples, the logic may include various software elements, such as software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, APIs, instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination thereof.
[0115] According to some examples, a computer-readable medium may include a non-transitory storage medium to store or maintain instructions that, when executed by a machine, a computing device, or a system, cause the machine, the computing device, or the system to perform methods and / or operations according to the described examples. Instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, and the like. Instructions may be implemented according to a predefined computer language, manner, or syntax to instruct a machine, a computing device, or a system to perform a function. Instructions may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.
[0116] Some examples may be described using the expression "in one example" or "example" along with their derivatives. These terms mean that a particular feature, structure, or characteristic described in connection with the example is included in at least one example. The appearances of the phrase "in one example" in various places in the specification are not necessarily all referring to the same example.
[0117] Some examples may be described using the expressions "coupled" and "connected" along with their derivatives. These terms are not necessarily intended to be synonymous with each other. For example, descriptions using the terms "connected" and / or "coupled" may indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but still operate or interact with each other.
[0118] The following examples pertain to additional examples of the techniques disclosed herein.
[0119] Example 1. An example apparatus may include circuitry at a controller resident on an NVDIMM. The apparatus may also include a receiving component for execution by the circuit to receive an error scan request. The apparatus may also include an error component for execution by the circuit. In response to the error scan request, the error component may scan a device physical address range of a non-volatile memory at the NVDIMM and identify uncorrected errors in the device physical address range. For this example, the non-volatile memory may be capable of providing persistent memory to the NVDIMM. The apparatus may also include an indication component for execution by the circuit to indicate to a BIOS of a computing platform coupled to the NVDIMM a device physical address with an identified uncorrected error.
[0120] Example 2. The apparatus of Example 1 may further include an interrupt component for the circuit to execute to send a system management interrupt to the BIOS in response to completing the scan of the device physical address range.
[0121] Example 3. The apparatus of Example 1, wherein the indicating component indicates a device physical address having an identified uncorrected error may include indicating via a first data structure resident at the NVDIMM. For this example, the BIOS may be able to access the first data structure to read the indicated device physical address and convert the device physical address to a system physical address, which is then maintained in a second data structure accessible to an operating system or device driver of the computing platform.
[0122] Example 4. The apparatus of Example 3, the first data structure may include a first register residing on the NVDIMM, which is accessible to the controller and the BIOS. For this example, the second data structure may include a second register residing on the computing platform, which is accessible to the BIOS and the operating system or device driver.
[0123] Example 5. The apparatus of Example 3, wherein the indicating component may indicate an error scan state to the BIOS via the first data structure, the error scan state comprising one of an idle state, an ongoing state, or a completed state.
[0124] Example 6. The apparatus of Example 3 may further include a flag component for execution by the circuitry to set a flag to indicate that the capacity of the data structure is insufficient to indicate at least some of the device physical addresses having the identified uncorrected errors.
[0125] Example 7. In the apparatus of Example 1, the receiving component may receive the error scan request after a power cycle of the computing platform or a reset of the NVDIMM.
[0126] Example 8. The apparatus of Example 1, the nonvolatile memory may include at least one of a 3-dimensional cross point memory, a flash memory, a ferroelectric memory, a SONOS, a polymer memory, a nanowire, a FeTRAM, or a FeRAM.
[0127] Example 9. A method may include receiving an error scan request at a controller resident on the NVDIMM. The method may also include scanning a device physical address range of a non-volatile memory at the NVDIMM in response to the error scan request. For this example, the non-volatile memory may be capable of providing persistent memory to the NVDIMM. The method may also include identifying an uncorrected error in the device physical address range and indicating a device physical address having the identified uncorrected error to a BIOS of a computing platform coupled to the NVDIMM.
[0128] Example 10. The method of Example 9 may further include sending a system management interrupt to the BIOS in response to completing the scan of the device physical address range.
[0129] Example 11. The method of Example 9, indicating the device physical address having the identified uncorrected error may include indicating via a first data structure resident at the NVDIMM. For this example, the BIOS may be able to access the first data structure to read the indicated device physical address and convert the device physical address to a system physical address, which is then maintained in a second data structure accessible to an operating system or device driver of the computing platform.
[0130] Example 12. The method of Example 11, the first data structure may include a first register resident on the NVDIMM, which is accessible to the controller and the BIOS. For this example, the second data structure may include a second register resident on the computing platform, which is accessible to the BIOS and the operating system or device driver.
[0131] Example 13. The method of Example 11 may further include indicating, via the first data structure, to the BIOS an error scan status comprising one of an idle state, an in-progress state, or a completed state.
[0132] Example 14. The method of Example 11 may further include setting a flag to indicate that the capacity of the data structure is insufficient to indicate at least some of the device physical addresses having identified uncorrected errors.
[0133] Example 15. Example method, receiving an error scan request after a power cycle of a computing platform or a reset of an NVDIMM.
[0134] Example 16. The method of Example 9, the non-volatile memory may include at least one of a 3-dimensional cross point memory, a flash memory, a ferroelectric memory, a SONOS, a polymer memory, a nanowire, a FeTRAM, or a FeRAM.
[0135] Example 17. At least one machine-readable medium may include a plurality of instructions which, in response to being executed by a controller residing on an NVDIMM, may cause the controller to perform the method according to any one of Examples 9 to 16.
[0136] Example 18. An apparatus may include means for performing the method of any of Examples 9 to 16.
[0137] Example 19. At least one machine-readable medium comprising a plurality of instructions that, in response to being executed by a controller residing on a non-volatile dual in-line memory module (NVDIMM), may cause the controller to receive an error scan request. The instructions may also cause the controller to scan a device physical address range of a non-volatile memory at the NVDIMM in response to the error scan request. For this example, the non-volatile memory may be capable of providing persistent memory to the NVDIMM. The instructions may also cause the controller to identify uncorrected errors in the device physical address range and indicate to a BIOS of a computing platform coupled to the NVDIMM the device physical address with the identified uncorrected error.
[0138] Example 20. The at least one machine-readable medium of Example 19, the instructions further causing the controller to send a system management interrupt to the BIOS in response to completing the scan of the device physical address range.
[0139] Example 21. The at least one machine-readable medium of Example 19, indicating the device physical address having the identified uncorrected error may include indicating via a first data structure resident at the NVDIMM. For this example, the BIOS may be able to access the first data structure to read the indicated device physical address and convert the device physical address to a system physical address, which is then maintained in a second data structure accessible to an operating system or device driver of the computing platform.
[0140] Example 22. The at least one machine-readable medium of Example 21, wherein the first data structure may include a first register resident on the NVDIMM, which is accessible to the controller and the BIOS. The second data structure may include a second register resident on the computing platform, which is accessible to the BIOS and the operating system or device driver.
[0141] Example 23. The at least one machine-readable medium of Example 21, the instructions further causing the controller to indicate, via the first data structure, to the BIOS an error scan status comprising one of an idle state, an in-progress state, or a completed state.
[0142] Example 24. The at least one machine-readable medium of Example 21, the instructions further causing the controller to set a flag to indicate that a capacity of the data structure is insufficient to indicate at least some of the device physical addresses having identified uncorrected errors.
[0143] Example 25. The at least one machine-readable medium of Example 19, receiving an error scan request after a power cycle of the computing platform or a reset of the NVDIMM.
[0144] Example 26. The at least one machine-readable medium of Example 19, the non-volatile memory comprising at least one of 3-D cross point memory, flash memory, ferroelectric memory, SONOS, polymer memory, nanowire, FeTRAM, or FeRAM.
[0145] Example 27. A method may include sending an error scan request from a circuit configured to implement a BIOS of a computing platform to a controller resident on an NVDIMM coupled to the computing platform. For this example, the NVDIMM may have a non-volatile memory that can provide persistent memory to the NVDIMM. The method may also include determining that the controller has completed a scan of a device physical address range of the non-volatile memory in response to the error scan request. The method may also include accessing a first data structure resident at the NVDIMM to read a device physical address having an identified uncorrected error identified by the controller during the scan of the device physical address range and converting the device physical address having the identified uncorrected error to a system physical address having the identified uncorrected error.
[0146] Example 28. The method of Example 27 may further include storing the system physical address having the identified uncorrected error to a second data structure accessible to an operating system or a device driver of the computing platform.
[0147] Example 29. The method of Example 28, the operating system or device driver may be able to use system physical addresses having identified uncorrected errors to avoid mapping system memory for the computing platform to those system physical addresses having identified uncorrected errors.
[0148] Example 30. The method of Example 28, the first data structure may include a first register resident on the NVDIMM, which is accessible to the controller and the BIOS. For this example, the second data structure may include a second register resident on the computing platform, which is accessible to the BIOS and the operating system or device driver.
[0149] Example 31. The method of Example 28, determining that the controller has completed the scan can be based on polling the first data structure. For this example, the first data structure can be capable of indicating an error scan status to the BIOS when polled. For this example, the error scan status can also include one of an ongoing state or a completed state.
[0150] Example 32. The method of Example 28 may also include determining that the controller sets a flag in the first data structure indicating that the capacity of the first data structure is insufficient to indicate at least some of the device physical addresses having uncorrected errors. The method may also include indicating to an operating system or a device driver that the system physical addresses having identified uncorrected errors stored in the second data structure do not include at least some of the uncorrected errors.
[0151] Example 33. The method of Example 27, sending the error scan request may be after a power cycle of the computing platform or a reset of the NVDIMM.
[0152] Example 34. The method of Example 27, the non-volatile memory may include at least one of a 3-dimensional cross point memory, a flash memory, a ferroelectric memory, a SONOS, a polymer memory, a nanowire, a FeTRAM, or a FeRAM.
[0153] Example 35. At least one machine-readable medium may include a plurality of instructions which, in response to being executed by a system at a computing platform, may cause the system to perform a method according to any one of Examples 27 to 34.
[0154] Example 36. An apparatus may include means for performing the method of any one of Examples 27 to 34.
[0155] Example 37. At least one machine-readable medium may include a plurality of instructions that, in response to being executed by a system, may cause the system to send an error scan request to a controller residing on an NVDIMM coupled to a computing platform, the system having circuitry configured to implement a BIOS of the computing platform. For this example, the NVDIMM may have a non-volatile memory that is capable of providing persistent memory to the NVDIMM. The instructions may also cause the system to determine that the controller has completed a scan of a device physical address range of the non-volatile memory in response to the error scan request. The instructions may also cause the system to access a first data structure residing at the NVDIMM to read a device physical address having an identified uncorrected error identified by the controller during a scan of the device physical address range and convert the device physical address having the identified uncorrected error into a system physical address having the identified uncorrected error.
[0156] Example 38. The at least one machine-readable medium of Example 37, the instructions further causing the system to store the system physical address having the identified uncorrected error to a second data structure accessible to an operating system or a device driver of the computing platform.
[0157] Example 39. The at least one machine-readable medium of Example 38, an operating system or a device driver may be capable of using system physical addresses having identified uncorrected errors to avoid mapping system memory for the computing platform to those system physical addresses having identified uncorrected errors.
[0158] Example 40. The at least one machine-readable medium of Example 38, wherein the first data structure may include a first register resident on the NVDIMM, which is accessible to the controller and the BIOS. For this example, the second data structure may include a second register resident on the computing platform, which is accessible to the BIOS and the operating system or device driver.
[0159] Example 41. The at least one machine-readable instruction of Example 38, the instruction may further cause the system to determine that the controller has completed the scan based on polling the first data structure. For this example, the first data structure may be capable of indicating an error scan state to the BIOS when polling. The error scan state may also include one of an ongoing state or a completed state.
[0160] Example 42. The at least one machine-readable medium of Example 38, the instructions further causing the system to determine that the controller sets a flag in the first data structure indicating that the capacity of the first data structure is insufficient to indicate at least some of the device physical addresses having uncorrected errors. The instructions further causing the system to indicate to an operating system or a device driver that the system physical addresses having identified uncorrected errors stored in the second data structure do not include at least some of the uncorrected errors.
[0161] Example 43. The at least one machine-readable medium of Example 37, may send the error scan request after a power cycle of the computing platform or a reset of the NVDIMM.
[0162] Example 44. The at least one machine-readable medium of Example 37, the non-volatile memory may include at least one of a 3-dimensional cross point memory, a flash memory, a ferroelectric memory, a SONOS, a polymer memory, a nanowire, a FeTRAM, or a FeRAM.
[0163] It is emphasized that the abstract of the present disclosure is provided to comply with 37 CFR Section 1.72 (b), thereby requiring an abstract that will allow the reader to quickly ascertain the essence of the present technical disclosure. The abstract is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the preceding detailed description, it can be seen that various features are grouped together in a single embodiment for the purpose of simplifying the disclosure. The disclosed method is not interpreted as reflecting the intention that the claimed examples require more features than the features explicitly detailed in each claim. On the contrary, as reflected in the following claims, the inventive subject matter is less than all the features of a single disclosed example. Thus, the following claims are hereby incorporated into the detailed description, wherein each claim stands on its own as an independent example. In the attached claims, the terms "including" and "in..." are used as the plain language equivalents of the corresponding terms "including" and "wherein", respectively. In addition, the terms "first", "second", "third", etc. are used only as labels and are not intended to impose numerical requirements on their objects.
[0164] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A memory device, comprising: A first data structure; as well as A controller comprising modules to: initiating, in response to a requestor, an error scan of a device physical address range of a non-volatile memory at the memory device; identifying uncorrectable errors in data stored in the non-volatile memory at a physical address range of the device; as well as Indicating, via the first data structure, a device physical address within the device physical address range having an uncorrectable error, the requestor being able to access the first data structure to read the indicated device physical address having the uncorrectable error to facilitate mapping of a system physical address to a device physical address of a non-volatile memory that does not have an uncorrectable error, wherein the requestor is to convert the device physical address having the uncorrectable error into a system physical address, which is then maintained in a second data structure accessible to an application hosted by a computing platform, the computing platform also hosting the requestor, the application being used to avoid mapping to the system physical address converted from the device physical address having the uncorrectable error.
2. The memory device of claim 1, comprising the first data structure, the first data structure comprising one or more fields accessible to the requestor, the one or more fields indicating the device physical address having the uncorrectable error.
3. The memory device of claim 1 , further comprising the module to: An error scan status is indicated to the requestor via the first data structure, the error scan status comprising a completion status.
4. The memory device of claim 1 , further comprising the module to: A flag is set to indicate that the capacity of the first data structure is insufficient to indicate at least some of the device physical addresses having uncorrectable errors.
5. The memory device of claim 1 , further comprising: Volatile memory; as well as The non-volatile memory is used to provide a persistent memory for the volatile memory.
6. The memory device of claim 5, comprising the memory device being a non-volatile dual in-line memory module (NVDIMM).
7. The memory device of claim 1, wherein the non-volatile memory comprises at least one of a 3-dimensional cross point memory, a flash memory, a ferroelectric memory, a silicon-oxide-nitride-oxide-silicon SONOS memory, a polymer memory, a nanowire memory, or a ferroelectric transistor random access memory.
8. A method comprising: Initiating, at a controller of a memory device, an error scan of a device physical address range of a non-volatile memory at the memory device, the error scan being initiated in response to a requester; identifying uncorrectable errors in data stored in the non-volatile memory at a physical address range of the device; as well as Indicating a device physical address with an uncorrectable error within the device physical address range via a first data structure, the requestor being able to access the first data structure to read the indicated device physical address with the uncorrectable error to facilitate mapping of a system physical address to a device physical address range of a non-volatile memory that does not have an uncorrectable error, wherein the requestor is to convert the device physical address with the uncorrectable error into a system physical address, which is then maintained in a second data structure accessible to an application hosted by a computing platform, the computing platform also hosting the requestor, the application being used to avoid mapping to the system physical address converted from the device physical address with the uncorrectable error.
9. The method of claim 8, comprising the first data structure comprising one or more fields accessible to the requestor, the one or more fields indicating the physical address of the device having the uncorrectable error.
10. The method of claim 8, further comprising: An error scan status is indicated to the requestor via the first data structure, the error scan status comprising a completion status.
11. The method of claim 8, further comprising: A flag is set to indicate that the capacity of the first data structure is insufficient to indicate at least some of the device physical addresses having identified uncorrected errors.
12. The method of claim 8, wherein the memory device further comprises a volatile memory, the non-volatile memory being configured to provide persistent storage for the volatile memory.
13. The method of claim 12, the memory device comprising a non-volatile dual in-line memory module (NVDIMM).
14. A device comprising: means for initiating, at a controller of a memory device, an error scan of a device physical address range of a non-volatile memory at the memory device, the error scan being initiated in response to a requestor; means for identifying uncorrectable errors in data stored in said non-volatile memory at a physical address range of said device; as well as Means for indicating, via a first data structure, a device physical address within the device physical address range having an uncorrectable error, the requestor being able to access the first data structure to read the indicated device physical address having the uncorrectable error, to facilitate mapping of a system physical address to a device physical address range of a non-volatile memory that does not have the uncorrectable error, wherein the requestor is to convert the device physical address having the uncorrectable error to a system physical address, the system physical address then being maintained in a second data structure accessible to an application hosted by a computing platform, the computing platform also hosting the requestor, the application being configured to avoid mapping to the system physical address converted from the device physical address having the uncorrectable error.
15. The device of claim 14, comprising the first data structure, the first data structure comprising one or more fields accessible to the requestor, the one or more fields indicating the device physical address having the uncorrectable error.
16. The apparatus of claim 14, further comprising: Means for indicating an error scan status to the requestor via the first data structure, the error scan status comprising a completion status.
17. The apparatus of claim 14, further comprising: Means for setting a flag to indicate that capacity of the first data structure is insufficient to indicate at least some of the device physical addresses having identified uncorrected errors.
18. The device of claim 14, the memory device further comprising a volatile memory, the non-volatile memory being configured to provide persistent storage for the volatile memory.
19. The device of claim 18, the memory device comprising a non-volatile dual in-line memory module (NVDIMM).
20. A computer readable medium having stored thereon instructions which, when executed, cause a computing device to perform the method of claims 8-13.
Citation Information
Patent Citations
System and Method for Increased System Availability In Virtualized Environments
US20090248949A1