A method for handling memory errors and a computing device

By identifying and modifying the error indication information in the SRAO error object, the problem of system reset on the CPU platform that does not support SRAO error handling is solved, which improves the operating system's operating efficiency and serviceability, and reduces hardware costs.

CN116382958BActive Publication Date: 2025-06-13XFUSION DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310332958.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2025-06-13
Estimated Expiration
2043-03-30

AI Technical Summary

Technical Problem

On CPU platforms that do not support SRAO error handling, when an SRAO error occurs in the memory page, the operating system can only perform system reset, affecting operation efficiency.

Method used

By identifying whether the SRAO errors occurring in the target memory page include processor context errors (PCCs), and generating the corresponding error object. The first error indication information is modified to the second error indication information, indicating that the SRAO error does not include the PCC, thereby avoiding system reset.

Benefits of technology

It improves the operating efficiency of the operating system, extends the online operation time, reduces business interruption time, improves the serviceability of the system, and reduces hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116382958B_ABST
    Figure CN116382958B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a method for handling memory errors and a computing device. Obtain an error object corresponding to a target memory page; wherein, the error object is used to indicate the error type of the target memory page; in the case that the error object includes first error indication information, modify the first error indication information to second error indication information; wherein, the first error indication information is used to indicate that a selected processing SRAO error occurs in the target memory page, and the SRAO error includes a processor context error PCC; the second error indication information is used to indicate that an SRAO error occurs in the target memory page, and the SRAO error does not include PCC; isolate the target memory page according to the second error indication information. In the present application, when an SRAO error occurs in a memory page, the second error indication information is used to indicate that the SRAO error does not include PCC, so that the operating system will not perceive that a PCC event has occurred, and thus will not perform a system reset, improving the operating efficiency of the operating system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of servers, and in particular, to a method for handling memory errors and a computing device. Background Art

[0002] With the progress of computer technology, the memory capacity used by the central processing unit (CPU) has been continuously increasing, and memory failures have become a frequent problem in system hardware failures.

[0003] Memory devices are designed to support error detection and correction mechanisms at the hardware level. When a correctable error (CE) occurs in the memory, usually the memory controller can detect the error and correct it. However, when an uncorrectable error (UCE) occurs in the memory, the parity algorithm cannot restore the correct value. Since the hardware cannot correct this error, it will trigger a process interruption to notify the hardware error handling module in the operating system to continue processing at the software level. Among them, recoverable errors in UCE can be classified into uncorrected no action (UCNA) errors, software recoverable action optional (SRAO) errors, and software recoverable action required (SRAR) errors.

[0004] On some CPU platforms that do not support SRAO error handling, when an SRAO error is detected, the hardware will trigger an interruption, and the operating system (OS) will perform a system reset according to the indication of the hardware, thus affecting the operating efficiency of the operating system. Summary of the Invention

[0005] Embodiments of this application provide a method for handling memory errors and a computing device, which are used to improve the operating efficiency of the system.

[0006] In a first aspect, embodiments of this application provide a method for handling memory errors. The method for handling memory errors in this application is applicable to CPU platforms that do not support SRAO error handling, and can avoid the problem that when an SRAO error occurs in a memory page, the operating system can only perform a system reset. Among them, for such CPU platforms that do not support SRAO error handling, when the hardware detects that an SRAO error occurs in a target memory page, it will indicate that the SRAO error includes a processor context corrupted (PCC). After the operating system perceives the PCC, it will perform a system reset.

[0007] In the embodiment of the present application, an SRAO error occurs in the target memory page, and moreover, the SRAO error includes a processor context corrupted (PCC). First, identify the SRAO error including PCC that occurs in the target memory page (for example, identify the SRAO error with the Bit value of PCC being 1), and then generate a corresponding error object for the SRAO error of the target memory page, so as to feedback the SRAO error to the operating system at the software layer. Among them, the error object includes first error indication information, and the first error indication information indicates that an SRAO error occurs in the target memory page, and the SRAO error includes PCC.

[0008] Since the first error indication information indicates that the SRAO error that occurs in the target memory page includes PCC. At this time, if the operating system perceives a PCC event, it will perform a system reset. Therefore, in the present application, the first error indication information is updated to second error indication information, that is, the error object corresponding to the target memory page includes the second error indication information, and the second error indication information indicates that an SRAO error occurs in the target memory page, and the SRAO error does not include PCC. Therefore, the operating system will not perceive a PCC event and will not perform a system reset.

[0009] In the embodiment of the present application, after an SRAO error occurs in the memory page, the second error indication information is used to indicate that the SRAO error does not include PCC, so that the operating system will not perceive a PCC event and will not perform a system reset, improving the operating efficiency of the operating system. On the other hand, the operating system will not perform a system reset, thereby increasing the online running duration of the operating system, reducing the service interruption time, and improving the serviceability of the operating system. And, for a CPU that supports handling SRAO errors, its procurement cost is higher. The method for handling memory errors in the present application expands the application scenarios of CPUs that do not support SRAO errors, reduces the use of CPUs that support handling SRAO errors, and reduces the hardware cost overhead of users.

[0010] Based on the first aspect, in an optional implementation manner, whether a PCC occurs can be represented by a first indicator and a second indicator, where the first indicator indicates that a PCC occurs, and the second indicator indicates that a PCC does not occur. Therefore, in the embodiment of the present application, the first error indication information includes the first indicator describing that the SRAO error occurring in the target memory page includes PCC, and the second error indication information includes the second indicator describing that the SRAO error occurring in the target memory page does not include PCC. Among them, the first indicator and the second indicator are different.

[0011] Based on the first aspect, in an alternative implementation, the first indicator is 1 and the second indicator is 0.

[0012] Based on the first aspect, in an alternative implementation, the first indicator in the first error indication information can be modified to the second indicator to generate the second error indication information.

[0013] Based on the first aspect, in an alternative implementation, obtain the error information stored in the error register; generate an error object corresponding to the target memory page based on the error information stored in the error register.

[0014] Based on the first aspect, in an alternative implementation, when the memory controller is independently set from the CPU, the memory controller performs detecting the status information of the target memory page to obtain the error detection information of the target memory page.

[0015] Based on the first aspect, in an alternative implementation, when the memory controller is integrated inside the CPU, the CPU performs detecting the status information of the target memory page to obtain the error detection information of the target memory page.

[0016] In a second aspect, an embodiment of the present application provides a memory error handling device, including:

[0017] An obtaining unit, configured to obtain an error object corresponding to a target memory page; wherein, the error object is used to indicate the error type of the target memory page;

[0018] A processing unit, configured to modify the first error indication information to the second error indication information when the error object includes the first error indication information; wherein, the first error indication information is used to indicate that a selective processing SRAO error occurs in the target memory page, and the SRAO error includes a processor context error PCC; the second error indication information is used to indicate that an SRAO error occurs in the target memory page, and the SRAO error does not include PCC;

[0019] An isolation unit, isolating the target memory page according to the second error indication information.

[0020] Based on the second aspect, in an alternative implementation, the first error indication information includes a first indicator describing that the SRAO error includes PCC; the second error indication information includes a second indicator describing that the SRAO error does not include PCC; the first indicator is different from the second indicator.

[0021] Based on the second aspect, in an alternative implementation, the first indicator is 1; the second indicator is 0.

[0022] Based on the second aspect, in an optional implementation, a processing unit is configured to, when an error object includes first error indication information, modify the first error indication information to second error indication information, including:

[0023] The processing unit is configured to modify a first indication word in the first error indication information to a second indication word to generate second error indication information.

[0024] Based on the second aspect, in an optional implementation, the processing unit is further configured to match the error object with first error information to determine whether the error object includes the first error indication information.

[0025] Based on the second aspect, in an optional implementation, an obtaining unit is specifically configured to: obtain error information stored in an error register;

[0026] Generate an error object corresponding to a target memory page based on the error information stored in the error register.

[0027] Based on the second aspect, in an optional implementation, the processing unit is further configured to trigger a memory controller to detect error information of a target memory page,

[0028] And write the error information into the error register.

[0029] Based on the second aspect, in an optional implementation, the processing unit is further configured to detect error information of a target memory page and write the error information into the error register.

[0030] Based on the second aspect, in an optional implementation, the error register includes a status register and a global status register;

[0031] The processing unit is specifically configured to: update bits of the status register and bits of the global status register based on the obtained error information.

[0032] The content such as the information interaction and execution process of the embodiments shown in this aspect is based on the same concept as the embodiments shown in the first aspect. Therefore, for the description of the beneficial effects shown in this aspect, please refer to the above-mentioned first aspect, and specific details are not elaborated here.

[0033] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a communication interface, and a processor coupled to the memory and the communication interface; the memory is used for storing instructions, the processor is used for executing the instructions, and the communication interface is used for communicating with other devices under the control of the processor; wherein, the processor executes the instructions to enable the computing device to execute the method in the first aspect and its related implementations.

[0034] Fourthly, an embodiment of the present application provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program runs on a processor, it enables a computing device to implement the methods in the above-mentioned first aspect and its related implementation manners. Description of the Drawings

[0035] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.

[0036] Figure 1 It is a schematic structural diagram of a computing device;

[0037] Figure 2 It is a schematic diagram of the types of memory errors;

[0038] Figure 3A It is a schematic flowchart of a method for handling memory errors provided by an embodiment of the present application;

[0039] Figure 3B It is a schematic flowchart of obtaining an error object of a target memory page provided by an embodiment of the present application;

[0040] Figure 4 It is the combined intention of each Bit in the register describing different error types in an embodiment of the present application;

[0041] Figure 5 It is a schematic diagram of the interpretation of different Bit positions in the status register;

[0042] Figure 6 It is a schematic diagram of the interpretation of different Bit positions in the global status register;

[0043] Figure 7 It is a schematic diagram of the description corresponding to the SRAO error in the status register;

[0044] Figure 8 It is a schematic diagram of the description corresponding to the SRAO error in the global status register;

[0045] Figure 9 It is a schematic structural diagram of a memory error handling device provided by an embodiment of the present application. Detailed Embodiments

[0046] An embodiment of the present application provides a method for handling memory errors and related devices, which are used to improve the operating efficiency of an operating system.

[0047] The embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, rather than intended to limit the present application. As is known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0048] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. "At least one (item)" or similar expressions thereof refer to any combination of these items, including any combination of single items or plural items. For example, at least one (item) of a, b, or c can mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0049] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0050] Please refer to Figure 1 , which Figure 1 is a schematic structural diagram of a computing device according to an embodiment of the present application.

[0051] The computing device 100 includes, but is not limited to, electronic devices with computing functions such as servers, switches, and minicomputers. Among them, when the computing device is a server, the server can be any type of server that supports memory page isolation technology, such as a server with an X86 architecture, specifically, various types of servers such as blade servers, high-density servers, rack servers, or high-performance servers.

[0052] In the following, the server will be taken as an example to describe the various solutions in the embodiments of the present application.

[0053] The server 100 may include a processor 101, a memory controller 102, and a memory 103. In practical applications, the server also includes a bus (not shown in the figure), and the bus can implement a path for transmitting information between various components of the server (for example, the processor 101, the memory controller 102, and the memory 103). The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0054] Among them, the processor 101 and the memory controller 102 can be integrated together or independently set. The processor 101 can be a central processing unit (CPU), a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary method flows described in connection with the disclosed content of the embodiments of the present application. The processor 101 can also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on.

[0055] The memory 103 is a cache space for storing the operation data in the processor 101 and the data exchanged with an external memory such as a hard disk. It is a bridge for communication between the external storage or peripheral storage and the processor 101. The memory 103 generally uses semiconductor storage units, including but not limited to random access memory (RAM), read-only memory (ROM), and cache. The memory 103 includes the target memory page in the embodiments of the present application.

[0056] The memory controller 102 is used to manage the access to data / programs in the memory 103. In the embodiments of the present application, the memory controller can detect an error that occurs in the target memory page and feedback the error to the error register in the processor 101.

[0057] It should be noted that Figure 1 The shown server is only a schematic structural diagram of a server applicable to the embodiments of the present application, and it does not limit the servers applicable to the embodiments of the present application. For example, the server can also include a persistent storage medium, a communication interface, a communication line, etc.

[0058] The memory capacity used by the central processing unit of the server is constantly increasing, and memory failures have become a frequent problem in the hardware failures of the memory failure system.

[0059] Memory devices are designed in hardware to support error detection and correction mechanisms. When a correctable error occurs in the memory, usually the memory controller can detect the error and correct it. However, when an uncorrectable error occurs in the memory, the parity algorithm cannot restore the correct value. Since the hardware cannot correct the error, it will trigger a process interruption, notifying the hardware error handling module in the operating system to continue processing at the software level.

[0060] Please refer to Figure 2 , Figure 2 for a schematic diagram of the types of memory errors. As Figure 2 shown, UCE includes fatal errors and recoverable errors. Among them, recoverable errors mean that the error may be fixable at the software level, but not necessarily. If not, the final result is to terminate the process corresponding to the faulty memory or reset the system. Specifically, recoverable errors can be classified into uncorrected no action (UCNA) errors, software recoverable action optional (SRAO) errors, and software recoverable action required (SRAR) errors. Among them, SRAO errors mean that when an error occurs in the memory, the error data generated has not been loaded into the cache in the CPU and is not in the execution path of the CPU.

[0061] In the related art, for such SRAO errors, according to different types of CPUs, it can be roughly divided into the following two processing situations. They will be described separately below.

[0062] On a CPU platform that supports SRAO error handling, when an SRAO error is detected, the hardware will trigger an interruption, and the operating system will try to repair the memory error according to the instructions of the hardware. For example, try to isolate the faulty memory page at the software level. The system does not directly reset.

[0063] On some CPU platforms that do not support SRAO error handling, when an SRAO error is detected, the hardware will trigger an interruption, and the operating system will perform a system reset according to the instructions of the hardware, thereby affecting the operating efficiency of the operating system.

[0064] In view of this, the embodiments of the present application disclose a method for handling memory errors to improve the operating efficiency of the operating system.

[0065] Please refer to Figure 3A , Figure 3AThe flowchart of the method for handling memory errors in the embodiments of the present application. The method for handling memory errors in the embodiments of the present application includes:

[0066] 201. Obtain an error object of a target memory page, where the error object is used to indicate the error type of the target memory page.

[0067] The method for handling memory errors in the embodiments of the present application is applicable to a CPU platform that does not support SRAO error handling (for example, it can be Figure 1 the CPU installed in the corresponding server in the embodiment), which can avoid the problem that when an SRAO error occurs in a memory page, the operating system can only perform a system reset. Among them, such a CPU platform that does not support SRAO error handling can be deployed in Figure 1 the servers, network devices, or terminals shown, and specifically, it is not limited here. In the embodiments of the present application, only taking the deployment of such a CPU platform that does not support SRAO error handling in a server as an example for introduction.

[0068] Please refer to Figure 3B , Figure 3B which is the flowchart of obtaining the error object of the target memory page in the embodiments of the present application.

[0069] As Figure 3B shown, obtaining the error object of the target memory page may specifically include the following steps:

[0070] 2011. Obtain the error information of the target memory page.

[0071] In one implementation, when the memory controller is independently set from the CPU, the memory controller performs detecting the status information of the target memory page to obtain the error detection information of the target memory page.

[0072] In one implementation, when the memory controller is integrated inside the CPU, the CPU performs detecting the status information of the target memory page to obtain the error detection information of the target memory page.

[0073] 2012. Write the error information of the target memory page into the error register corresponding to the target memory page.

[0074] The CPU / memory controller writes the detected error information of the target memory page into the error register, and the error register is used to store the error detection information of the target memory page, where each error information corresponds to one error type.

[0075] Specifically, in practical applications, error information for describing various error types is stored in an error register. Exemplarily, the error information for describing various error types stored in the error register can be a combination of one or more bits. In this example, after the memory controller or CPU detects an error in a memory page, the memory controller / CPU can write to the corresponding bit in the error register. When the combined situation of the bits in the error register after writing matches the condition of the bits corresponding to a certain error type, the CPU can determine that the memory page has an error of that error type.

[0076] 2013. Determine whether a target memory page has an SRAO error based on error information.

[0077] Exemplarily, please refer to Figure 4 , Figure 4 as a possible combined situation of each bit in the error register for describing different error types. As Figure 4 shown, SRAO errors, UCNA errors, and CE errors all need to be obtained through the combined situation of multiple bits. Among them, for a CPU that does not support SRAO error handling, when an SRAO error occurs, the value of the bit describing PCC will be set to 1, indicating that the SRAO error includes PCC, thereby triggering the operating system to perform a system reset.

[0078] Generally speaking, there are clear definitions for the key bits corresponding to each type of error in the error register, and from the perspective of compatibility, chip manufacturers will not change the original semantics. In practical applications, the error register of the CPU generally includes a status register (IA32_MCi_STATUS) and a global status register (IA32_MCG_STATUS).

[0079] Exemplarily, Figure 5 are the definitions of different bits in the status register (IA32_MCi_STATUS); Figure 6 are the definitions of different bits in the global status register (IA32_MCG_STATUS). Among them, Figure 4 the bits used to indicate SRAO errors in Figure 5 include the VAL field in the 63rd bit, the UC field in the 61st bit, the PCC field in the 57th bit, the Service field in the 56th bit, the AR field in the 55th bit, the ADDRV field in the 58th bit, the MISCV field in the 59th bit, and, as well as Figure 6 the RIPV field in the 0th bit and the EIPV field in the 1st bit in the global status register of

[0080] From Figure 4It can be known that the SRAO error needs to be jointly indicated by the Bit in the status register (IA32_MCi_STATUS) and the Bit in the global status register (IA32_MCG_STATUS). Specifically, Figure 7 is the description corresponding to the SRAO error in the status register (IA32_MCi_STATUS); Figure 8 is the description corresponding to the SRAO error in the global status register (IA32_MCG_STATUS).

[0081] Exemplarily, in the Linux architecture, the function __mc_scan_banks can be called. This function is the main function for obtaining hardware error information. After reading the hardware error, the corresponding error object (MCE Error Object) is generated through the mce_severity function, and the type of the error object is obtained (such as determining that the error type occurring in the target memory page is the SRAO error including PCC).

[0082] In 2014. When an SRAO error occurs in the target memory page, an error object corresponding to the SRAO error is generated based on the error information.

[0083] In the embodiment of the present application, the error register of the CPU determines that an SRAO error has occurred in the target memory page, and moreover, the SRAO error includes a processor context corrupted (PCC). The CPU generates a corresponding error object for the SRAO error of the target memory page for feedback to the operating system at the software layer. Among them, the error object includes first error indication information, and the first error indication information indicates that an SRAO error has occurred in the target memory page, and the SRAO error includes PCC.

[0084] In practical applications, the error object corresponding to the generated SRAO error can be stored in the form of a sequential list, linked list, stack, queue, tree structure, or graph storage structure, etc. Exemplarily, the error object is a structure. The present application does not make a limitation on this.

[0085] In 202. When the error object includes the first error indication information, the first error indication information is modified to the second error indication information. The second error indication information indicates that an SRAO error has occurred in the target memory page, and the SRAO error does not include a processor context corrupted (PCC).

[0086] Since the first error indication information indicates that the SRAO error occurring in the target memory page includes PCC. At this time, if the operating system perceives that a PCC event has occurred, it will perform a system reset. Therefore, the first error indication information is modified to the second error indication information, and the error object corresponding to the target memory page includes the second error indication information. The second error indication information indicates that the target memory page has an SRAO error, and this SRAO error does not include PCC. Therefore, the operating system will not perceive that a PCC event has occurred, and thus will not perform a system reset.

[0087] In this embodiment, after an SRAO error occurs in the memory page, the second error indication information is used to indicate that this SRAO error does not include PCC, so that the operating system will not perceive that a PCC event has occurred, and thus will not perform a system reset, improving the operating efficiency of the operating system. On the other hand, since the operating system will not perform a system reset, the online running duration of the operating system is increased, the service interruption time is reduced, and the serviceability of the operating system is improved. Moreover, the application scenario of CPUs that do not support SRAO errors is extended, the use of CPUs that support processing SRAO errors is reduced, and the hardware cost overhead of users is reduced.

[0088] In a possible implementation manner, whether PCC has occurred can be represented by a first indicator and a second indicator. Among them, the first indicator indicates that PCC has occurred, and the second indicator indicates that PCC has not occurred. Therefore, in the embodiment of the present application, the first error indication information includes the first indicator describing that the SRAO error occurring in the target memory page includes PCC, and the second error indication information includes the second indicator describing that the SRAO error occurring in the target memory page does not include PCC. Among them, the first indicator and the second indicator are different. Therefore, the first indicator in the first error indication information can be modified to the second indicator to generate the second error indication information.

[0089] In a possible implementation manner, the first indicator is 1 and the second indicator is 0. Exemplarily, the first error indication information can be the Bit bits shown above Figure 4 Assume that the first error information is 11110111010. Among them, the first indicator in the description field of PCC is 1, indicating that the value of the Bit bit representing PCC has been set to 1, indicating that this SRAO error includes PCC.

[0090] It should be understood that in practical applications, in addition to using the first indicator and the second indicator to represent whether the SRAO error includes PCC, other methods can also be used to represent whether the SRAO error includes PCC. For example, it can also be that the first indicator is M indicating that the SRAO error includes PCC, while the second indicator is N indicating that the SRAO error does not include PCC. The present application does not make any limitations in this regard.

[0091] Isolate the target memory page according to the second error indication information.

[0092] In the error object corresponding to the target memory page, if the second error indication information indicates that the target memory page has a SRAO error and this SRAO error does not include PCC, the operating system will not perceive that a PCC event has occurred, and thus will not perform a system reset. At this time, the target memory page can be isolated to prevent the error data generated by the target memory page from affecting other processes, improving the system execution efficiency.

[0093] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a memory error handling device 300 provided by an embodiment of the present application.

[0094] As Figure 9 shown, the memory error handling device 300 includes:

[0095] An acquisition unit 301, configured to acquire an error object corresponding to a target memory page; wherein, the error object is used to indicate the error type of the target memory page;

[0096] A processing unit 302, configured to modify the first error indication information to the second error indication information when the error object includes the first error indication information; wherein, the first error indication information is used to indicate that the target memory page has a selected processing SRAO error, and the SRAO error includes a processor context error PCC; the second error indication information is used to indicate that the target memory page has a SRAO error, and the SRAO error does not include PCC;

[0097] An isolation unit, which isolates the target memory page according to the second error indication information.

[0098] In a possible design, the first error indication information includes a first indicator word describing that the SRAO error includes PCC; the second error indication information includes a second indicator word describing that the SRAO error does not include PCC; the first indicator word is different from the second indicator word.

[0099] In a possible design, the first indicator word is 1; the second indicator word is 0.

[0100] In a possible design, modifying the first error indication information to the second error indication information includes:

[0101] Modifying the first indicator word in the first error indication information to the second indicator word to generate the second error indication information.

[0102] In a possible design, the processing unit 302 is further configured to match the error object with the first error information to determine whether the error object includes the first error indication information.

[0103] In a possible design, the obtaining unit 301 is specifically configured to: obtain the error information stored in the error register;

[0104] Generate an error object corresponding to the target memory page based on the error information stored in the error register.

[0105] In a possible design, the processing unit 302 is further configured to trigger the memory controller to detect the error information of the target memory page,

[0106] And write the error information into the error register.

[0107] In a possible design, the processing unit 302 is further configured to detect the error information of the target memory page,

[0108] And write the error information into the error register.

[0109] In a possible design, the error register includes a status register and a global status register;

[0110] The processing unit 302 is specifically configured to: update the Bit bits of the status register and the Bit bits of the global status register based on the obtained error information.

[0111] It should be noted that the information interaction, execution process, etc. between the modules / units in the memory error handling device 300 are based on the same concept as the corresponding method embodiments in this application. For specific content, reference can be made to the descriptions in the method embodiments shown above in this application, and details are not elaborated here. Figure 3A The embodiments of the present application further provide a computer program product including instructions. The computer program product can be software or a program product including instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computer device, at least one computer device is caused to execute the method described in the foregoing

[0112] shown embodiments. Figure 3A shown embodiments.

[0113] Embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the method described in the foregoing Figure 3A as shown in the embodiments described.

[0114] The remote access device provided by the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit, etc. The processing unit can execute the computer-executable instructions stored in the storage unit to enable the chip to execute the method described in the foregoing Figure 3A as shown in the embodiments described. Optionally, the storage unit is a storage unit within the chip, such as a register, a cache, etc. The storage unit may also be a storage unit outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0115] It should be further noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that they have a communication connection, which can specifically be implemented as one or more communication buses or signal lines.

[0116] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, software program implementation is a better embodiment in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0117] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0118] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

Claims

1. A method for handling memory errors, characterized in that, applied to a central processing unit (CPU) platform that does not support selective recovery and abort option (SRAO) error handling, including: obtaining an error object corresponding to a target memory page; wherein, the error object is used to indicate the error type of the target memory page; when the error object includes first error indication information, modifying the first error indication information to second error indication information; wherein, the first error indication information is used to indicate that an SRAO error occurs in the target memory page, and the SRAO error includes a processor context error (PCC); the second error indication information is used to indicate that an SRAO error occurs in the target memory page, and the SRAO error does not include PCC; isolating the target memory page according to the second error indication information.

2. The method according to claim 1, characterized in that, the first error indication information includes a first indicator describing that the SRAO error includes PCC; the second error indication information includes a second indicator describing that the SRAO error does not include PCC; the first indicator is different from the second indicator.

3. The method according to claim 2, characterized in that, the first indicator is 1; the second indicator is 0.

4. The method according to claim 2 or 3, characterized in that, the modifying the first error indication information to second error indication information includes: modifying the first indicator in the first error indication information to the second indicator to generate the second error indication information.

5. The method according to any one of claims 1-3, characterized in that, after obtaining the error object corresponding to the target memory page, the method further includes: matching the error object with the first error indication information to determine whether the error object includes the first error indication information.

6. The method according to any one of claims 1-3, characterized in that, the obtaining the error object corresponding to the target memory page includes: obtaining error information stored in an error register; generating an error object corresponding to the target memory page based on the error information stored in the error register.

7. The method according to claim 6, characterized in that, before obtaining the error information stored in the error register, the method further includes: triggering a memory controller to detect the error information of the target memory page, writing the error information into the error register.

8. The method according to claim 6, characterized in that, before obtaining the error information stored in the error register, the method further includes: detecting the error information of the target memory page, writing the error information into the error register.

9. The method according to claim 7 or 8, characterized in that, the error register includes a status register and a global status register; the writing the error information into the error register includes: updating the Bit bits of the status register and the global status register based on the obtained error information.

10. A computer device, characterized in that, It includes a processor and a memory, wherein the processor is coupled to the memory; The memory is used for storing program instructions; The processor is used for executing the program instructions, so that the computer device executes the method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Memory error processing method and device and server

    CN111625387A

  • Machine inspection error processing method and device

    CN115858211A