Memory error processing method and system, electronic equipment and storage medium
By collecting and persisting memory error information when the target operating system restarts, the problem of memory error isolation being limited to a single startup cycle is solved, memory error management across system restart cycles is realized, system stability and hardware error reliability are improved, and downtime frequency is reduced.
Patent Information
- Application Number
- CN202511023988.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-08-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, memory error isolation is limited to a single startup cycle. The system after restart cannot sense the previous memory error status, resulting in the memory error address being allocated again, which cannot effectively reduce repeated downtime or service interruption, affecting the stability of the data center and cloud service system.
When the target operating system restarts, memory error information is collected and persists storage is performed, error memory is isolated based on storage results, application access is denied, memory error management is realized across system restart cycles.
By persisting memory error information, we ensure that the system can timely isolate the error memory after restarting, reduce the frequency of system downtime, improve system stability and reliability of hardware errors, and avoid data inconsistency and service interruptions.
Smart Images

Figure CN120523640A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a memory error handling method, system, electronic device, and storage medium. Background Art
[0002] In hyperscale cloud computing data centers, the reliability of dynamic random access memory (DRAM) is crucial to maintaining stable server operations. With increasing server resource density, a single physical machine often hosts hundreds of virtual machines and thousands of container instances, dramatically increasing the potential impact and cost of hardware failures. Related technologies can immediately isolate the erroneous memory address when a memory error occurs while the system remains operational, preventing subsequent program access and thus minimizing the direct impact of uncorrectable memory errors on programs. However, memory error isolation in related technologies is limited to a single boot cycle. After a reboot, the system is unaware of the previous memory error state, resulting in known erroneous memory addresses being reassigned to programs. This fails to effectively reduce repeated downtime or service interruptions caused by memory errors, further impacting the stability of data centers and cloud service systems.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The embodiments of the present application provide a memory error handling method, system, electronic device, and storage medium to at least solve the technical problems of high system downtime frequency and poor stability in memory error isolation methods in related technologies.
[0005] According to one aspect of an embodiment of the present application, a memory error handling method is provided, including: in response to restarting a target operating system, collecting memory error information, wherein the memory error information is used to record historical memory errors generated before the target operating system crashes; persistently storing the memory error information to obtain a storage result; and isolating the erroneous memory based on the storage result to deny applications under the target operating system from applying for or accessing the erroneous memory.
[0006] According to another aspect of an embodiment of the present application, a memory error handling method is also provided, including: responding to an input instruction acting on an operation interface, running an application under a target operating system on the operation interface; responding to a processing instruction acting on the operation interface, displaying a rejection prompt on the operation interface; wherein the rejection prompt is used to reject the application's application for or access to erroneous memory, and the erroneous memory is isolated based on a storage result, and the storage result is obtained by persistently storing memory error information, and the memory error information is used to record historical memory errors generated before the target operating system crashes.
[0007] According to another aspect of an embodiment of the present application, a memory error handling system is also provided, including: a memory error collection component, used to collect memory error information in response to restarting the target operating system, wherein the memory error information is used to record historical memory errors generated before the target operating system crashes; a memory error storage component, used to persistently store the memory error information to obtain a storage result; and a memory error isolation component, used to isolate the erroneous memory based on the storage result to deny the application under the target operating system from applying for or accessing the erroneous memory.
[0008] According to another aspect of an embodiment of the present application, a memory error handling system is also provided, including: a client, configured to collect and send memory error information in response to restarting a target operating system, wherein the memory error information is used to record historical memory errors generated before the target operating system crashes; a server, connected to the client, configured to persistently store the memory error information to obtain a storage result, and to isolate the erroneous memory based on the storage result; the client is also configured to output a rejection prompt, wherein the rejection prompt is used to reject an application under the target operating system from applying for or accessing the erroneous memory.
[0009] According to another aspect of the embodiments of the present application, an electronic device is provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in the various embodiments of the present application when running.
[0010] According to another aspect of an embodiment of the present application, a computer-readable storage medium is also provided, which includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the methods in various embodiments of the present application.
[0011] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a computer program, which implements the methods in various embodiments of the present application when executed by a processor.
[0012] According to another aspect of an embodiment of the present application, a computer program product is further provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method in each embodiment of the present application is implemented.
[0013] According to another aspect of the embodiments of the present application, a computer program is further provided, which implements the methods in various embodiments of the present application when executed by a processor.
[0014] In an embodiment of the present application, by responding to restarting the target operating system, collecting memory error information, and then persistently storing the memory error information to obtain a storage result, and finally isolating the erroneous memory based on the storage result to deny the application under the target operating system from applying for or accessing the erroneous memory, thereby achieving continuous management of memory errors across system restart cycles. By actively collecting historical memory error information before the crash when the target operating system is restarted, it is possible to break through the limitation that traditional memory error isolation is limited to the current operating cycle, ensuring that the target operating system can take isolation measures in time after the restart, thereby significantly enhancing system stability and reliability protection against hardware errors, and reducing the frequency of system crashes caused by memory errors. The persistently stored memory error information is used to guide the isolation decision of the erroneous memory after the target operating system is restarted, effectively avoiding the reuse of faulty memory resources, reducing the risk of resource conflicts and error propagation, and ensuring that the application does not access memory known to have errors, thereby effectively avoiding the risk of data inconsistency and service interruption, thereby achieving the technical effect of reducing the frequency of system crashes and improving system stability, and thus solving the technical problem of high system crash frequency and poor stability in the memory error isolation method in the related art.
[0015] It is easy to notice that the above general description and the following detailed description are merely for the purpose of exemplifying and explaining the present application, and do not constitute a limitation of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0017] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a memory error handling method according to an embodiment of the present application;
[0018] Figure 2 is a structural block diagram of a computing environment according to an embodiment of the present application;
[0019] Figure 3 This is a structural block diagram of a service grid according to an embodiment of the present application;
[0020] Figure 4 is a flowchart of a memory error handling method according to an embodiment of the present application;
[0021] Figure 5 is a schematic diagram of a memory error type according to an embodiment of the present application;
[0022] Figure 62 is a schematic diagram of a memory isolation determination process according to an embodiment of the present application;
[0023] Figure 7 is a structural diagram of a memory error handling system according to an embodiment of the present application;
[0024] Figure 8 It is a process diagram of a memory error handling method according to related art;
[0025] Figure 9 This is a process diagram of a memory error handling method according to an embodiment of the present application;
[0026] Figure 10 is a flowchart of another memory error handling method according to an embodiment of the present application;
[0027] Figure 11 is a structural block diagram of a memory error handling device according to an embodiment of the present application;
[0028] Figure 12 is a structural block diagram of another memory error handling device according to an embodiment of the present application;
[0029] Figure 13 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:
[0033] Reliability, Availability and Serviceability (RAS): Usually used to measure the stability and maintainability of a computer system during operation.
[0034] Service Level Agreement (SLA): A formal agreement between a service provider and a customer that defines the quality of service and the scope of responsibilities.
[0035] Correctable Error (CE): refers to a minor error in the system that can be automatically corrected by hardware or software and usually does not cause system crashes or interruptions.
[0036] Uncorrected Correctable Error (UCE): An error that could have been corrected but could not be successfully corrected due to some reasons (such as insufficient resources) and may escalate to a more serious error.
[0037] Machine-Check Architecture (MCA): is a mechanism in the x86 architecture for detecting and reporting hardware errors, particularly in memory and processors.
[0038] Corrected Machine Check Interrupt (CMCI): An interrupt triggered when the system detects a correctable error, used to notify the operating system to record or handle it.
[0039] Machine-Check Exception (MCE): A hardware-generated exception used to report a non-negligible hardware error that may cause the system to crash or reboot.
[0040] Error Correction Code (ECC): is a technology used to detect and correct errors in data transmission or storage, commonly used in memory and storage devices.
[0041] Deferred Error (DE): refers to errors that do not immediately affect system operation but need to be handled at some point in the future.
[0042] ACPI Error Record Serialization Table (ACPI ERST): Part of the Advanced Configuration and Power Interface specification, it defines how to persist and retrieve hardware error information.
[0043] According to an embodiment of the present application, a memory error handling method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0044] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The following is a hardware block diagram of a computer terminal (or mobile device) for implementing a memory error handling method. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (illustrated as 102a, 102b, ..., 102n in the figure) (the processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of the BUS), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0045] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." This data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be fully or partially integrated into any of the other components of the computer terminal 10 (or mobile device). As discussed in the embodiments of this application, this data processing circuitry functions as a processor control (e.g., selecting a variable resistor terminal path connected to an interface).
[0046] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the methods in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the methods in the above embodiments. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0047] Transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of computer terminal 10. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module configured to communicate with the Internet wirelessly.
[0048] The display may be, for example, a touch screen liquid crystal display (LCD), which enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0049] Figure 1 The hardware structure block diagram shown can be used not only as an exemplary block diagram of the computer terminal 10 (or mobile device), but also as an exemplary block diagram of the server. In an optional embodiment, Figure 2 The block diagram shows the use of the above Figure 1 The computer terminal 10 (or mobile device) shown is an embodiment of a computing node in the computing environment 201 . Figure 2 A structural diagram of a computing environment is shown, such as Figure 2As shown, computing environment 201 includes multiple computing nodes (e.g., servers) (illustrated as 210-1, 210-2, ...) running on a distributed network. Each computing node contains local processing and memory resources, allowing end users 202 to remotely run applications or store data in computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 in computing environment 201, representing services "A," "D," "E," and "H," respectively.
[0050] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or request of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 can include a corresponding agent to handle the provisioning and / or request for services (one or more services provided in computing environment 201).
[0051] Services are provided or deployed based on various virtualization technologies supported by computing environment 201. In some embodiments, services can be provided based on virtual machine (VM) virtualization, container-based virtualization, and / or similar approaches. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While VMs virtualize machines, container-based virtualization can launch containers to virtualize entire operating systems (OSs), allowing multiple workloads to run on a single OS instance.
[0052] In an embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). Figure 2 As shown, service 220-2 can be deployed with one or more Pods 240-1, 240-2, ..., 240-N (collectively, "Pods"). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively, "Containers"). One or more containers in a Pod handle requests related to one or more corresponding functions of the service. Proxy 245 typically controls network functions related to the service, such as routing and load balancing. Similar Pods can also be deployed with other services.
[0053] During operation, executing a user request from the end user 202 may require calling one or more services in the computing environment 201, and executing one or more functions of a service may require calling one or more functions of another service. Figure 2As shown, service “A” 220 - 1 receives a user request from end user 202 from ingress gateway 230 , service “A” 220 - 1 may call service “D” 220 - 2 , and service “D” 220 - 2 may request service “E” 220 - 3 to perform one or more functions.
[0054] This computing environment can be a cloud computing environment, where resource allocation is managed by the cloud service provider, allowing for feature development without having to worry about implementing, adjusting, or scaling servers. This computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be partitioned to perform a set of functions that can scale independently and automatically, rather than scaling a single hardware device to handle the potential load.
[0055] In another optional embodiment, Figure 3 The block diagram shows the use of the above Figure 1 The computer terminal 10 (or mobile device) is shown as an embodiment of a service grid. Figure 3 A structural diagram of a service grid is shown, Figure 3 As shown, the service grid 300 is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to decomposing an application into multiple smaller services or instances and distributing them to run on different clusters / machines.
[0056] like Figure 3 As shown, microservices may include application service instance A and application service instance B, which form the functional application layer of service mesh 300. In one embodiment, application service instance A runs as a container / process 308 on a machine / workload container group 314 (Pod), and application service instance B runs as a container / process 310 on a machine / workload container group 316 (Pod).
[0057] In one implementation, application service instance A may be a product query service, and application service instance B may be a product ordering service.
[0058] like Figure 3As shown, application service instance A and grid proxy (sidecar) 303 coexist in machine / workload container group 314, while application service instance B and grid proxy 305 coexist in machine / workload container group 316. Grid proxy 303 and grid proxy 305 form the data plane layer (dataplane) of service grid 300. Grid proxy 303 and grid proxy 305 run as container / process 304 and container / process 306, respectively, and can receive requests 312 for product query services. Bidirectional communication is possible between grid proxy 303 and application service instance A, and between grid proxy 305 and application service instance B. Furthermore, bidirectional communication is possible between grid proxy 303 and grid proxy 305.
[0059] In one embodiment, all traffic from application service instance A is routed to the appropriate destination via grid proxy 303, and all network traffic from application service instance B is routed to the appropriate destination via grid proxy 305. It should be noted that network traffic referred to herein includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), the high-performance, general-purpose open source framework (Google Remote Procedure Call, gRPC), and the open source in-memory data structure storage system (Redis).
[0060] In one embodiment, the data plane layer's functionality can be extended by writing custom filters for the proxy (Envoy) in service mesh 300. The service mesh proxy configuration can be designed to enable the service mesh to correctly proxy service traffic, enabling service interoperability and service governance. Mesh proxy 303 and mesh proxy 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.
[0061] like Figure 3 As shown, the service grid 300 also includes a control plane layer. The control plane layer can be a group of services running in a dedicated namespace, and these services are hosted by a hosting control plane component 301 in a machine / workload container group (machine / Pod) 302. Figure 3 As shown, managed control plane component 301 communicates bidirectionally with mesh proxy 303 and mesh proxy 305. Managed control plane component 301 is configured to perform certain control and management functions. For example, managed control plane component 301 receives telemetry data transmitted by mesh proxy 303 and mesh proxy 305 and can further aggregate this telemetry data. Managed control plane component 301 also provides user-oriented application programming interfaces (APIs) to facilitate manipulation of network behavior and to provide configuration data to mesh proxy 303 and mesh proxy 305.
[0062] Under the above operating environment, this application provides Figure 4 Memory error handling method shown. Figure 4 is a flow chart of a memory error handling method according to an embodiment of the present application, such as Figure 4 As shown, the method includes the following steps:
[0063] Step S41, in response to restarting the target operating system, collecting memory error information, wherein the memory error information is used to record historical memory errors generated before the target operating system crashes;
[0064] Step S42, persistently storing the memory error information to obtain a storage result;
[0065] Step S43: Isolate the error memory based on the stored result to deny the application program under the target operating system from applying for or accessing the error memory.
[0066] The target operating system described above can be an operating system running in environments such as data centers and server clusters, such as open-source operating systems used in servers, workstations, mobile devices, and embedded systems. It possesses a high degree of flexibility and customizability, capable of optimization and adjustment based on diverse hardware architectures and application scenarios. The target operating system is capable of effectively managing large-scale computing resources, handling highly concurrent network requests, and providing reliable data storage and transmission services. In data centers and server clusters, the target operating system typically serves as the foundational platform supporting various application services, such as database services, network services, and cloud computing platforms.
[0067] This memory error information is detailed data about memory hard errors generated and recorded by the hardware error detection mechanism during the target operating system's operation, especially before a system crash. This memory error information includes, but is not limited to, the error type, the timestamp of the error, the physical memory address, the error severity, and details about the hardware component that triggered the error. Specifically, the error type distinguishes between correctable errors (CE) and uncorrected errors (UE). In computer systems, particularly in the context of memory management and error detection, correctable errors refer to minor errors in the target operating system that can be automatically corrected by hardware or software and typically do not cause system crashes or outages. Uncorrectable errors are data errors that the target operating system can detect but cannot automatically correct. The error timestamp records the exact moment the error occurred. The physical memory address indicates the specific memory location where the error occurred. The error severity is used to assess the potential impact of the error on system operations. Details about the hardware component that triggered the error include the location and module identifier of the Dual In-line Memory Module (DIMM).
[0068] Dynamic random access memory (DRAM), the main memory of modern computing systems, is a core component enabling fast data storage and retrieval. As DRAM ages, irreversible failures gradually occur. When accessing the faulty location in the DRAM chip, the read content may deviate from the stored content.
[0069] To prevent data inconsistencies caused by DRAM failures, error correction code (ECC) technology was developed to detect and correct errors in memory accesses at the cache line granularity. For example, chip-level error correction technology can correct any erroneous data bits within a single DRAM chip during a single access. Errors that can be corrected by ECC are called correctable errors. When erroneous data bits in a single access span two or more chips, the data inconsistency caused by the DRAM failure will exceed the error correction capability of chip-level ECC, resulting in uncorrectable errors, which often cause system crashes.
[0070] ECC memory errors are primarily categorized as soft errors and hard errors. Soft errors are typically temporary and can be caused by factors such as electromagnetic interference and temperature fluctuations, while hard errors are caused by physical damage to the memory module itself. Soft errors are typically random, meaning the memory address where the error occurs is not fixed. Hard errors, on the other hand, are repetitive and persistent, meaning uncorrectable memory errors are repeatedly reported at the same memory address.
[0071] Figure 5 is a schematic diagram of a memory error type according to an embodiment of the present application, such as Figure 5 As shown in the figure, memory errors can be categorized into the following categories based on their impact on computer systems: correctable errors, deferred errors, uncorrected correctable errors, uncorrected thread errors, uncorrected system errors, and node errors. When an uncorrectable memory error is consumed by an operating system process, it causes the process to terminate. When consumed by the operating system kernel, it causes a kernel panic. Correctable memory errors may trigger a System Management Interrupt (SMI). An SMI is a global / broadcast event that forcibly suspends normal operation of all processors in the system. When an SMI is triggered, all CPU threads enter System Management Mode (SMM) immediately after completing the current instruction. Once in SMM, CPU threads are no longer available for operating system scheduling, and thread execution is suspended until the SMI handler returns control to the original execution context. This can cause unpredictable system performance jitter. Therefore, both correctable and uncorrectable errors can have a noticeable impact on computer systems.
[0072] When collecting memory error information, for non-system crash errors, since the target operating system is still running, the error can be reported through a framework such as MCE, and then collected and stored in a file / database by a user-mode program. For system crash errors, since the open source operating system kernel enters an emergency stop state upon receiving a system crash error, persistent storage is required after the memory error information is collected to prevent it from being lost after a system restart.
[0073] For example, ERST technology can be used to collect system crash errors. ERST technology is defined in the ACPI specification and is used to persistently store hardware error information across system restarts. When the target operating system detects an uncorrectable memory error and is about to crash, the error record is serialized and stored in non-volatile memory (such as flash memory or a solid-state drive). After restarting, the target operating system can read these error records stored in ERST to obtain memory error information before the crash. System crash errors can also be collected using kernel crash dump (Kdump) technology. Kdump technology can capture kernel memory snapshots when the system crashes. Through pre-configured crash dump kernel parameters, the target operating system reserves a portion of memory for kernel dumps at the time of a crash, thereby providing detailed error information, including the memory error that caused the system crash.
[0074] Furthermore, the memory error information is persistently stored to obtain a storage result. Persistent storage is the process of storing the memory error information in a non-volatile storage medium so that the memory error information can be retained and restored even after the system is shut down or restarted. Non-volatile storage media include but are not limited to hard disk drives, solid-state drives, flash drives, tapes, etc. Non-volatile storage media can still retain data after power is lost. Persistent storage can be implemented in a variety of ways. The specific method selected depends on the requirements of the target operating system, performance requirements and hardware limitations. For example, persistent storage can be implemented through file system storage, database storage, error log table storage, etc.
[0075] During the restart or subsequent operation of the target operating system, the stored memory error information can be used to identify and mark which memory areas are unstable or damaged, thereby preventing applications from allocating or accessing such areas. This can significantly improve system stability and data integrity, and avoid program crashes or system failures caused by memory errors.
[0076] When the target operating system restarts, the previously recorded memory error information can be read from the non-volatile storage. By analyzing the memory error information, it is possible to determine which memory addresses are repeatedly erroneous. Based on the analysis results, the target operating system can mark the memory addresses with fault records as "unavailable" or "isolated", that is, the faulty memory will not be allocated to any application. In open source operating systems, memory isolation can be achieved by modifying the memory management module. For example, in the early stages of kernel startup, the use of problematic memory can be avoided by adjusting the properties of the memory area or using a specific isolation list. When an application requests memory allocation, the memory manager checks whether the requested memory address is in the isolation list. If the requested memory has been marked as erroneous memory, the memory allocation request will be rejected and the application will be allocated to other healthy memory areas.
[0077] Based on the above steps S41 to S43, by responding to restarting the target operating system, collecting memory error information, and then persistently storing the memory error information to obtain a storage result, and finally isolating the erroneous memory based on the storage result to deny the application under the target operating system from applying for or accessing the erroneous memory, thereby achieving continuous management of memory errors across system restart cycles. By actively collecting historical memory error information before the crash when the target operating system is restarted, it is possible to break through the limitation that traditional memory error isolation is limited to the current operating cycle, ensuring that the target operating system can take isolation measures in time after the restart, thereby significantly enhancing system stability and reliability protection against hardware errors, and reducing the frequency of system crashes caused by memory errors. The persistently stored memory error information is used to guide the isolation decision of the erroneous memory after the target operating system is restarted, effectively avoiding the reuse of faulty memory resources, reducing the risk of resource conflicts and error propagation, and ensuring that the application does not access memory known to have errors, thereby effectively avoiding the risk of data inconsistency and service interruption, thereby achieving the technical effect of reducing the frequency of system crashes and improving system stability, and thus solving the technical problem of high system crash frequency and poor stability in the memory error isolation method in the related art.
[0078] The following further introduces the memory error handling method in the embodiment of the present application.
[0079] In an optional embodiment, the memory error information includes system-level memory error information. In step S41, collecting the memory error information includes:
[0080] System-level memory error information is read from a preset non-volatile storage area, wherein the preset non-volatile storage area is used to persistently record system-level memory errors generated before the target operating system crashes.
[0081] The above-mentioned preset non-volatile storage area can be a storage area pre-defined in the computer system that can still save data after power failure. The preset non-volatile storage area usually includes flash memory, non-volatile random access memory or a specific hard disk partition, which can continue to retain previously written data after the machine is restarted or powered off. In an embodiment of the present application, the preset non-volatile storage area is set as a place to store system-level memory error information to ensure that even after the target operating system crashes, the relevant error information will not be lost, providing a basis for subsequent error analysis and isolation. The preset non-volatile storage area can be ERST, which is a technology for persisting error information across system restarts. The data platform needs to implement some form of non-volatile storage to save error records. The size of the required storage space depends on the processor architecture of the data platform.
[0082] The above-mentioned system-level memory error information can be generated by memory hardware and pose a threat to the stability of the entire operating system. System-level memory error information typically includes uncorrectable memory errors, such as machine check exceptions, and other serious memory errors that may trigger system crashes. System-level memory error information includes the specific location of the error (memory address), error type, possible severity, and relevant system information at the time of the error, such as the Central Processing Unit (CPU) status and timestamp.
[0083] Exemplary methods for reading system-level memory error information include, but are not limited to, reading from firmware logs, reading from dedicated hardware interfaces, and reading from bootloaders. Specifically, system-level memory error information can be recorded in the firmware's log area. The firmware log can be read using specific management tools or system call interfaces to retrieve previous system-level memory error information. Some server hardware may provide dedicated hardware interfaces or command-line tools that allow system-level memory error information stored in onboard non-volatile memory to be read after the target operating system is restarted. For example, the health status and error logs of server hardware can be read through the Intelligent Platform Management Interface (IPMI). The bootloader can access certain non-volatile storage areas of the target operating system, such as flash memory or specific partitions on a hard drive, during early boot, thereby obtaining system-level memory error information before the target operating system starts. It should be noted that each reading method has different applicable scenarios and limitations. The appropriate system-level memory error information reading method can be selected based on the characteristics of the system architecture and available resources, and this embodiment of the present application is not limited thereto.
[0084] Based on the above optional embodiment, by reading system-level memory error information from a preset non-volatile storage area, the preset non-volatile storage area is used to persistently record system-level memory errors generated before the target operating system crashes, thereby ensuring that even after a system crash, the system-level memory error information can still be fully retained and read, providing a solid foundation for real-time isolation and prevention of memory errors.
[0085] In an optional embodiment, reading system-level memory error information from a preset non-volatile storage area includes:
[0086] System-level memory error information is read from a preset non-volatile storage area through a preset virtual file system.
[0087] The preset virtual file system can be a virtual file system used within the target operating system kernel to store and restore critical system information. This preset virtual file system is primarily used to save and restore various system status data, including hardware error information, kernel logs, and error codes, after a system crash or reboot. By interacting with a preset non-volatile storage area, the preset virtual file system ensures that memory error information remains accessible after the target operating system reboots, thereby assisting system administrators and developers in locating and troubleshooting issues.
[0088] Based on the above optional embodiment, by presetting a virtual file system and reading system-level memory error information from a preset non-volatile storage area, after the target operating system is restarted, the memory address where a hard error occurred before the restart can be re-isolated based on the memory error information read by the preset virtual file system to prevent the program from accessing the damaged memory again, thereby reducing the risk of system instability and service interruption caused by memory errors.
[0089] In an optional embodiment, the preset non-volatile storage area includes: serialized error records, and the serialized error records are encoded according to a universal platform error record format.
[0090] Serialized error records can be encoded in the Common Platform Error Record (CPER) format, defined in an appendix to the Unified Extensible Firmware Interface specification. This format provides a standardized method for describing and storing hardware error information, ensuring consistency across platforms and across reboots. Serialization converts memory error information into a storable and transportable format for persistent storage and ease of reading and processing by error collection components. The design of the error record serialization interface allows hardware vendors flexibility in implementing their error record serialization hardware. Data platforms can provide the detailed information necessary to communicate with their serialization hardware by populating an ERST with a set of serialization instruction entries. One or more serialization instruction entries constitute a serialization action, which the target operating system executes through a series of serialization actions to complete the serialization operation.
[0091] Based on the above optional embodiment, serialized error records are obtained by encoding according to the general platform error record format, and then stored in a preset non-volatile storage area. When the target operating system is restarted, the serialized error records can be read to obtain the serialized error records, thereby providing a basis for subsequent error memory isolation.
[0092] In an optional embodiment, the serialized error record encapsulates a first memory error stored in a structure format, and / or a second memory error stored in a text format, wherein the first memory error is a memory error detected by a processor in a computer system adopting a preset instruction set architecture, and the second memory error is a memory error reported by different hardware modules in multiple computer systems, and the multiple computer systems adopt different types of instruction set architectures.
[0093] Specifically, the first memory error may be a machine check exception error, which may be a memory error detected by the CPU in a computer system with a preset instruction set architecture, which may be an x86 microprocessor architecture system. The second memory error may be a Generic Hardware Error Source (GHES) memory error, which may be a memory error reported by different hardware modules in various computer systems with different instruction set architectures, including but not limited to x86 microprocessor architecture systems, Advanced RISC Machine (ARM) architecture systems, and fifth-generation RISC-V architecture systems. The first memory error may be stored in the ERST using the MCE structure format, while the second memory error may be stored in the ERST using the text format.
[0094] Based on the above optional embodiment, the serialized error record encapsulates a first memory error stored in a structure format and / or a second memory error stored in a text format, thereby enhancing the applicability of the isolation solution of the present application in computer systems across different architectures, making the memory hard error isolation solution more versatile and applicable to a wider range of hardware platforms. At the same time, by distinguishing and adapting to different error storage formats, the complexity of error handling is simplified, and the overall error management and response efficiency of the system is improved.
[0095] In an optional embodiment, in step S42, the memory error information is persistently stored, and the storage result obtained includes:
[0096] The memory error information is persistently stored according to a preset storage format to obtain a storage result, wherein the preset storage format includes: multiple storage fields, and the multiple storage fields include at least:
[0097] The first storage field is used to store the timestamp of the memory error;
[0098] The second storage field is used to store the memory address where the memory error occurs;
[0099] The third storage field is used to store the error type of the memory error;
[0100] The fourth storage field is used to store the number of memory error occurrences.
[0101] This pre-defined storage format is designed to efficiently and clearly record key memory error details for subsequent analysis and isolation. The resulting data is typically stored in one or more files, preserving the memory error information in a structured format. This allows the memory error history isolation component in the memory error handling system to quickly read and parse the memory error information.
[0102] The first storage field mentioned above is a timestamp field, which is used to record the exact time when the memory error occurred, usually expressed in milliseconds or microseconds, which helps in the sorting and time series analysis of memory errors. The second storage field mentioned above is a memory address field, which contains the physical memory address where the error occurred, and is key information for identifying and isolating faulty memory. The third storage field mentioned above is an error type field, which is used to identify the type of memory error, such as whether it is a correctable error, an uncorrected correctable error, or an uncorrectable error. For MCE errors, the timestamp is indicated in the MCE format and can be directly converted into a standard time format. For GHES errors, there is no direct timestamp, which only contains the timestamp after power-on. The last boot timestamp can be obtained from the log to calculate the overall timestamp.
[0103] In addition to the above basic fields, the preset storage format can also include other fields, such as the error count field, which is used to record the total number of errors that occur at a specific memory address; the hardware information field: contains the hardware context when the error occurs, such as the CPU number, memory module information, etc.; the software context field: contains information about the software or process running when the error occurs.
[0104] After collecting memory error information, the memory error information can be organized according to a preset storage format, including filling in the above-mentioned multiple storage fields, and then persistently stored in a non-volatile storage medium, such as a hard disk, SSD, or a dedicated error log storage area. The storage result is usually one or more files containing all serialized error records. After the target operating system is restarted, the memory history error isolation component can read the persistently stored files, parse the error information therein, and decide which memory addresses need to be isolated based on a preset algorithm (such as error type, number of errors, most recent occurrence time, etc.).
[0105] Based on the above optional embodiments, by persistently storing memory error information in a preset storage format, the integrity and availability of the memory error information are ensured. Even after the system crashes or restarts, the memory error information can be quickly restored, thereby achieving a smarter and more effective memory error management strategy.
[0106] In an optional embodiment, the memory error information further includes runtime memory error information. In step S41, collecting the memory error information includes:
[0107] Runtime memory error information is collected through a preset structure recording interface and / or a preset event notification interface, wherein the runtime memory error information includes: recoverable memory error information and thread-level memory error information.
[0108] Recoverable memory errors refer to errors that occur during system operation but can be automatically corrected or recovered through other means. While recoverable memory errors won't cause a system crash, frequent occurrences can gradually deplete system resources and impact performance. By collecting recoverable memory error information through a pre-defined structure recording interface, we can count error frequency, assess memory health, and take appropriate action, such as isolating memory areas with frequent errors.
[0109] These thread-level memory errors are associated with specific threads or processes and can directly impact the integrity of the running program or task. By collecting thread-level memory error information through a pre-defined event notification interface, we can pinpoint the scope of the error and assess whether thread-level error recovery or isolation is necessary to prevent the error from propagating and further impacting system stability.
[0110] Based on the above optional embodiment, runtime memory error information is collected through a preset structure recording interface and / or a preset event notification interface, which significantly enhances the memory error management capability and reduces the risk of unforeseen system failures.
[0111] In an optional embodiment, in step S43, isolating the error memory based on the storage result includes:
[0112] Step S431, performing a memory isolation determination on the storage result based on a preset memory isolation condition to obtain a determination result;
[0113] Step S432: Persistently isolate the erroneous memory from the memory resources to be allocated based on the determination result.
[0114] The above-mentioned preset memory isolation conditions include but are not limited to: historical occurrence number conditions, error type conditions, etc.
[0115] The preset memory isolation conditions described above can be formulated based on the characteristics of memory errors to determine which memory errors require attention and isolation. Specifically, these conditions include, but are not limited to, error type, error frequency, and time span. Based on the preset memory isolation conditions, a memory isolation determination is performed on the stored results to determine which erroneous memory addresses should be marked as requiring isolation, thereby obtaining a determination result.
[0116] After the judgment result is obtained, the faulty memory addresses that need to be isolated can be identified from the memory resources to be allocated based on the judgment result, and the persistent isolation operation is performed. In other words, the faulty memory will no longer be considered as an available memory pool resource by the operating system and will not be allocated to any running or upcoming programs. This prevents the program from accessing these faulty memory areas again during runtime, preventing potential recurrence of the fault and service interruption.
[0117] Persistent isolation typically involves modifying the operating system's memory management data structures, such as memory mapping tables or bitmaps, to ensure that isolated memory addresses remain marked as unallocatable after the target operating system restarts. This ensures that the isolation state of faulty memory persists across system restarts, effectively preventing system failures and performance degradation caused by hard memory errors and improving the overall stability of data centers or cloud services.
[0118] Based on the above optional embodiment, a memory isolation judgment is performed on the storage result based on the preset memory isolation condition to obtain a judgment result, and then the erroneous memory is persistently isolated from the memory resources to be allocated based on the judgment result, thereby effectively preventing the redistribution of the faulty memory, reducing the system downtime rate caused by memory failure, and optimizing the availability and reliability of the computer system.
[0119] In an optional embodiment, in step S432, persistently isolating the erroneous memory from the memory resources to be allocated according to the determination result includes:
[0120] In response to the determination result indicating that the error type of the erroneous memory is an uncorrectable error, and the interval between two adjacent occurrences of the erroneous memory is less than a preset time length, the erroneous memory is persistently isolated from the memory resources to be allocated.
[0121] Specifically, uncorrectable errors usually mean failures at the hardware level. These failures cannot be automatically repaired by software means, and if left untreated, they may cause data corruption, calculation errors, or system crashes. When the judgment result shows that the error type of the erroneous memory is an uncorrectable error, the timestamp information of the memory error is further analyzed to evaluate whether the time interval between two adjacent erroneous memories is less than a preset duration. The preset duration can be a threshold set by the system administrator based on historical data and service requirements, such as a few seconds, minutes, or hours, to determine whether the uncorrectable error is a frequent or potential persistent failure. When the judgment result shows that the error type of the erroneous memory is an uncorrectable error, and the interval between two adjacent occurrences of the erroneous memory is less than the preset duration, the erroneous memory is persistently isolated from the memory resources to be allocated.
[0122] Based on the above optional embodiments, in response to the judgment result indicating that the error type of the erroneous memory is an uncorrectable error, and the interval between two adjacent occurrences of the erroneous memory is less than a preset time length, the erroneous memory is persistently isolated from the memory resources to be allocated, thereby significantly reducing the probability of repeated system downtime and service interruption caused by memory failure, and improving the reliability, availability, and serviceability of the data center or cloud computing environment.
[0123] In an optional embodiment, in step S432, persistently isolating the erroneous memory from the memory resources to be allocated according to the determination result includes:
[0124] In response to the determination result indicating that the error type of the erroneous memory is a correctable error, the number of errors of the erroneous memory exceeds a preset number, and the interval between two adjacent occurrences of the erroneous memory is less than a preset time length, the erroneous memory is persistently isolated from the memory resources to be allocated.
[0125] Specifically, when the error type of the error memory is a correctable error, check whether the number of errors in the error memory exceeds the preset number threshold. The preset number is a value set by a system administrator based on historical data and service reliability requirements, and is used to determine whether a memory unit has an abnormally high frequency of errors. For example, the preset number can be set to 50 times. Further analyze the timestamp of the error to determine whether the time interval between two adjacent correctable errors is less than the preset duration. The preset duration is also an important judgment factor, used to evaluate whether correctable errors show a pattern of frequent occurrence. The duration can be several minutes, hours or days, depending on the sensitivity and fault tolerance requirements of the service. When the error type of the error memory is a correctable error, the number of errors exceeds the preset number, and the interval between two adjacent occurrences is less than the preset duration, a persistent isolation operation is performed during the system initialization phase or before memory allocation.
[0126] Based on the above optional embodiments, in response to the judgment result indicating that the error type of the erroneous memory is a correctable error, the number of errors of the erroneous memory exceeds a preset number, and the interval between two adjacent occurrences of the erroneous memory is less than a preset time length, the erroneous memory is persistently isolated from the memory resources to be allocated, thereby being able to intelligently identify and isolate memory units that exhibit potential dangerous trends, thereby reducing the risk of downtime and other system failures at the source and further improving system stability.
[0127] Figure 6 Schematic diagram of a memory isolation determination process according to an embodiment of the present application. Figure 6As shown, the memory error record is first read. In response to the determination result indicating that the error type of the erroneous memory is an uncorrectable error and the interval between two consecutive occurrences of the erroneous memory is less than a preset time length, the erroneous memory is persistently isolated from the memory resources to be allocated. In response to the determination result indicating that the error type of the erroneous memory is a correctable error, the number of errors in the erroneous memory exceeds a preset number, and the interval between two consecutive occurrences of the erroneous memory is less than a preset time length, the erroneous memory is persistently isolated from the memory resources to be allocated.
[0128] Figure 7 is a structural diagram of a memory error handling system according to an embodiment of the present application, such as Figure 7 As shown, in the memory error handling system, the memory error collection component is used to collect memory error information in response to restarting the target operating system. The memory error information is used to record historical memory errors generated before the target operating system crashes. The memory error storage component is used to persistently store the memory error information to obtain a storage result. The erroneous memory is isolated based on the storage result to deny the application under the target operating system from applying for or accessing the erroneous memory. In the memory error handling system, runtime memory error information can be collected through a preset structure recording interface and / or a preset event notification interface. The runtime memory error information includes: recoverable memory error information and thread-level memory error information. Serialized error records are stored in the preset non-volatile storage area. The serialized error records include: memory errors of the machine check exception type, and / or memory errors of the general hardware error source type stored in text format.
[0129] Figure 8 This is a process diagram of a memory error handling method according to related technology, such as Figure 8 As shown, the method includes the following steps:
[0130] Step S801: The application accesses an erroneous memory during operation;
[0131] Step S802: a memory error occurs and memory error information is recorded in a register;
[0132] Step S803, notifying the CPU of a memory error via an interrupt;
[0133] Step S804, the central processing unit routes the memory error to the kernel of the target operating system for processing;
[0134] Step S805: reading memory error information from the register and reporting it;
[0135] Step S806: Restart the system if it encounters a system-level error.
[0136] Step S807, after the target operating system is started, the application program accesses the error memory again;
[0137] Step S808: a memory error occurs again;
[0138] Step S809: The target operating system crashes again.
[0139] Depend on Figure 8 As shown, the memory error isolation in the related technology is limited to a single boot cycle. The target operating system after restart cannot perceive the previous memory error state, which causes the known erroneous memory address to be assigned to the application again, thereby failing to effectively reduce the repeated downtime or service interruption caused by memory errors, resulting in a high frequency of system downtime and low service stability.
[0140] Figure 9 FIG. 1 is a process diagram of a memory error handling method according to an embodiment of the present application. Figure 9 As shown, the method includes the following steps:
[0141] Step S901: The application accesses an erroneous memory during operation;
[0142] Step S902: a memory error occurs and memory error information is recorded in a register;
[0143] Step S903, notifying the CPU of a memory error via an interrupt;
[0144] Step S904, the central processing unit routes the memory error to the kernel of the target operating system for processing;
[0145] Step S905: reading memory error information from the register and reporting it;
[0146] Step S906: Restart the system if it encounters a system-level error.
[0147] Step S907, collecting memory error information after the target operating system is restarted;
[0148] Step S908: persistently store the memory error information to obtain a storage result, and isolate the error memory based on the storage result;
[0149] Step S909: denying the application under the target operating system from applying for or accessing the wrong memory.
[0150] Depend on Figure 9As shown, the embodiment of the present application can achieve continuous management of memory errors across system restart cycles when processing memory errors. By actively collecting historical memory error information before the crash when the target operating system is restarted, it can break through the limitation of traditional memory error isolation being limited to the current operating cycle, ensuring that the target operating system can take isolation measures in time after the restart, thereby significantly enhancing system stability and reliability protection against hardware errors, and reducing the frequency of system crashes caused by memory errors. The persistently stored memory error information is used to guide the isolation decision of the erroneous memory after the target operating system is restarted, effectively avoiding the reuse of faulty memory resources, reducing the risk of resource conflicts and error propagation, and ensuring that applications do not access memory known to have errors, thereby effectively avoiding the risk of data inconsistency and service interruption, thereby achieving the technical effect of reducing the frequency of system crashes and improving system stability, and thus solving the technical problem of high system crash frequency and poor stability in the memory error isolation method in the related art.
[0151] Figure 10 is a flowchart of another memory error handling method according to an embodiment of the present application. Figure 10 As shown, the method includes the following steps:
[0152] Step S101, responding to an input instruction on the operation interface, running an application under a target operating system on the operation interface;
[0153] Step S102, in response to the processing instruction acting on the operation interface, a rejection prompt is displayed on the operation interface; wherein, the rejection prompt is used to reject the application's application for or access to the error memory, the error memory is isolated based on the storage result, and the storage result is obtained by persistently storing the memory error information, and the memory error information is used to record historical memory errors generated before the target operating system crashes.
[0154] Based on the above steps S101 to S102, by responding to the input instructions acting on the operation interface, the application under the target operating system is run on the operation interface, and then responding to the processing instructions acting on the operation interface, a rejection prompt is displayed on the operation interface; wherein, the rejection prompt is used to reject the application application to apply for or access the error memory, the error memory is isolated based on the storage result, and the storage result is obtained by persistently storing the memory error information. The memory error information is used to record the historical memory errors generated before the target operating system crashes, thereby realizing continuous management of memory errors across system restart cycles. By actively collecting the historical memory error information before the crash when the target operating system is restarted, it is possible to break through the limitation of traditional memory error isolation being limited to the current operation cycle, ensuring that the target operating system can take isolation measures in time after restarting, thereby significantly enhancing the system stability and reliability protection against hardware errors, and reducing the frequency of system crashes caused by memory errors. The persistently stored memory error information is used to guide the isolation decision of the erroneous memory after the target operating system is restarted, effectively avoiding the reuse of faulty memory resources, reducing the risk of resource conflicts and error propagation, and ensuring that applications do not access memory that is known to have errors. This effectively avoids the risk of data inconsistency and service interruption, thereby achieving the technical effect of reducing the frequency of system crashes and improving system stability, and thus solving the technical problems of high system crash frequency and poor stability in the memory error isolation method in related technologies.
[0155] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0156] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0157] Through the description of the above implementation methods, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus the necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of this application.
[0158] According to an embodiment of the present application, a memory error handling device for implementing the above-mentioned memory error handling method is also provided. Figure 11 is a structural block diagram of a memory error handling device according to an embodiment of the present application, such as Figure 11 As shown, the device includes:
[0159] A collecting module 1101 is configured to collect memory error information in response to restarting the target operating system, wherein the memory error information is used to record historical memory errors generated before the target operating system crashes;
[0160] The storage module 1102 is used to persistently store the memory error information and obtain a storage result;
[0161] The isolation module 1103 is used to isolate the error memory based on the storage result, so as to deny the application program under the target operating system from applying for or accessing the error memory.
[0162] Optionally, the collecting module 1101 is further configured to read system-level memory error information from a preset non-volatile storage area, wherein the preset non-volatile storage area is used to persistently record system-level memory errors generated before the target operating system crashes.
[0163] Optionally, the collecting module 1101 is further configured to read the system-level memory error information from a preset non-volatile storage area through a preset virtual file system.
[0164] Optionally, the preset non-volatile storage area includes: serialized error records, and the serialized error records are encoded according to a universal platform error record format.
[0165] Optionally, the serialized error record encapsulates a first memory error stored in a structure format, and / or a second memory error stored in a text format, wherein the first memory error is a memory error detected by a processor in a computer system adopting a preset instruction set architecture, and the second memory error is a memory error reported by different hardware modules in multiple computer systems, and the multiple computer systems adopt different types of instruction set architectures.
[0166] Optionally, the storage module 1102 is also used to: persistently store the memory error information according to a preset storage format to obtain a storage result, wherein the preset storage format includes: multiple storage fields, and the multiple storage fields include at least: a first storage field, used to store the timestamp of the memory error; a second storage field, used to store the memory address of the memory error; a third storage field, used to store the error type of the memory error; and a fourth storage field, used to store the number of times the memory error occurs.
[0167] Optionally, the collection module 1101 is further used to collect runtime memory error information through a preset structure recording interface and / or a preset event notification interface, wherein the runtime memory error information includes: recoverable memory error information and thread-level memory error information.
[0168] Optionally, the isolation module 1103 is further used to: perform a memory isolation determination on the storage result based on a preset memory isolation condition to obtain a determination result; and perform persistent isolation on the erroneous memory from the memory resources to be allocated based on the determination result.
[0169] Optionally, the isolation module 1103 is also used to: in response to the judgment result indicating that the error type of the erroneous memory is an uncorrectable error, and the interval between two adjacent occurrences of the erroneous memory is less than a preset time length, persistently isolate the erroneous memory from the memory resources to be allocated.
[0170] Optionally, the isolation module 1103 is also used to: in response to the judgment result indicating that the error type of the erroneous memory is a correctable error, the number of errors of the erroneous memory exceeds a preset number, and the interval between two adjacent occurrences of the erroneous memory is less than a preset time length, persistently isolate the erroneous memory from the memory resources to be allocated.
[0171] It should be noted that the collection module 1101, storage module 1102, and isolation module 1103 correspond to steps S41 to S43 in the above-mentioned embodiment. The examples and application scenarios implemented by these three modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned embodiment. It should be noted that the above-mentioned modules or units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above-mentioned modules may also be part of an apparatus and run in the computer terminal provided in the above-mentioned embodiment.
[0172] Figure 12 is a structural block diagram of another memory error handling device according to an embodiment of the present application, such as Figure 12 As shown, the device includes:
[0173] The running module 1201 is used to respond to the input instruction on the operation interface and run the application program under the target operating system on the operation interface;
[0174] Display module 1202, for responding to the processing instruction acting on the operation interface and displaying a rejection prompt on the operation interface;
[0175] Among them, the rejection prompt is used to reject the application's application for or access to the error memory. The error memory is isolated based on the storage result. The storage result is obtained by persistently storing the memory error information. The memory error information is used to record the historical memory errors generated before the target operating system crashes.
[0176] It should be noted that the aforementioned execution module 1201 and display module 1202 correspond to steps S101 and S102 in the aforementioned embodiment. The examples and application scenarios implemented by these two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the aforementioned embodiment. It should be noted that the aforementioned modules or units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The aforementioned modules may also be executed as part of an apparatus in the computer terminal provided in the aforementioned embodiment.
[0177] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in the above embodiments, as well as the application scenario and implementation process, but is not limited to the scheme provided in the above embodiments.
[0178] An embodiment of the present application can also provide a memory error handling system, including: a memory error collection component, used to collect memory error information in response to restarting the target operating system, wherein the memory error information is used to record historical memory errors generated before the target operating system crashes; a memory error storage component, used to persistently store the memory error information and obtain a storage result; a memory error isolation component, used to isolate the erroneous memory based on the storage result, so as to deny the application under the target operating system from applying for or accessing the erroneous memory.
[0179] An embodiment of the present application may also provide a memory error handling system, including: a client, for collecting and sending memory error information in response to restarting the target operating system, wherein the memory error information is used to record historical memory errors generated before the target operating system crashes; a server, connected to the client, for persistently storing the memory error information to obtain a storage result, and isolating the erroneous memory based on the storage result; the client is also used to output a rejection prompt, wherein the rejection prompt is used to reject an application under the target operating system from applying for or accessing the erroneous memory.
[0180] The embodiment of the present application may provide an electronic device, which may be any electronic device in a group of electronic devices. Optionally, in this embodiment, the electronic device may also be replaced by a terminal device such as a mobile terminal.
[0181] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0182] In this embodiment, the computer terminal can execute the program code in the method.
[0183] Optionally, Figure 13 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 13 As shown, the electronic device may include: one or more (only one is shown in the figure) processors 132, a memory 134, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module, and a display.
[0184] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0185] The processor may call the executable program stored in the memory through the transmission device to execute the method described in any one of the above embodiments.
[0186] Those skilled in the art will appreciate that the structure shown in the figure is merely illustrative, and the electronic device may also be a smartphone (such as an Android phone, iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. This figure does not limit the structure of these electronic devices. For example, the electronic device may include more or fewer components (such as a network interface, a display device, etc.) than shown in this figure, or have a configuration different from that shown in this figure.
[0187] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0188] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the method provided in the above embodiment.
[0189] Optionally, in this embodiment, the storage medium may be located in any electronic device in a group of electronic devices in a computer network, or in any mobile terminal in a group of mobile terminals.
[0190] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, and when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the method described in any one of the above embodiments.
[0191] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the computer program product may include a computer program, and when the computer program is executed by a processor, the method provided in the embodiment is implemented.
[0192] The embodiments of the present application further provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which may be used to store a computer program that, when executed by a processor, implements the method provided in the embodiments above.
[0193] The embodiment of the present application further provides a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, the method provided in the above embodiment is implemented.
[0194] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0195] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0196] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0197] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0198] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program code.
[0199] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A memory error handling method, characterized in that: include: In response to restarting the target operating system, collecting memory error information, wherein the memory error information is used to record historical memory errors generated before the target operating system crashes; Persistently storing the memory error information to obtain a storage result; The error memory is isolated based on the storage result to deny the application program under the target operating system from applying for or accessing the error memory.
2. The memory error handling method according to claim 1, wherein: The memory error information includes: system-level memory error information, and collecting the memory error information includes: The system-level memory error information is read from a preset non-volatile storage area, wherein the preset non-volatile storage area is used to persistently record the system-level memory error generated before the target operating system crashes.
3. The memory error handling method according to claim 2, wherein: Reading the system-level memory error information from the preset non-volatile storage area includes: The system-level memory error information is read from the preset non-volatile storage area through a preset virtual file system.
4. The memory error handling method according to claim 2, wherein: The preset non-volatile storage area includes: serialized error records, and the serialized error records are encoded according to a universal platform error record format.
5. The memory error handling method according to claim 4, characterized in that: The serialized error record encapsulates a first memory error stored in a structure format and / or a second memory error stored in a text format, wherein the first memory error is a memory error detected by a processor in a computer system adopting a preset instruction set architecture, and the second memory error is a memory error reported by different hardware modules in multiple computer systems, and the multiple computer systems adopt different types of instruction set architectures.
6. The memory error handling method according to claim 1, wherein: The memory error information is persistently stored, and the storage result obtained includes: The memory error information is persistently stored according to a preset storage format to obtain the storage result, wherein the preset storage format includes: multiple storage fields, and the multiple storage fields include at least: The first storage field is used to store the timestamp of the memory error; The second storage field is used to store the memory address where the memory error occurs; The third storage field is used to store the error type of the memory error; The fourth storage field is used to store the number of memory error occurrences.
7. The memory error handling method according to claim 1, wherein: The memory error information further includes runtime memory error information, and collecting the memory error information includes: The runtime memory error information is collected through a preset structure recording interface and / or a preset event notification interface, wherein the runtime memory error information includes: recoverable memory error information and thread-level memory error information.
8. The memory error handling method according to any one of claims 1 to 7, characterized in that: Isolating the error memory based on the storage result includes: Based on a preset memory isolation condition, performing a memory isolation determination on the storage result to obtain a determination result; The erroneous memory is persistently isolated from the memory resources to be allocated according to the determination result.
9. The memory error handling method according to claim 8, characterized in that: Persistently isolating the erroneous memory from the memory resources to be allocated according to the determination result includes: In response to the determination result indicating that the error type of the erroneous memory is an uncorrectable error, and an interval between two adjacent occurrences of the erroneous memory is less than a preset time length, the erroneous memory is persistently isolated from the memory resources to be allocated.
10. The memory error handling method according to claim 8, wherein: Persistently isolating the erroneous memory from the memory resources to be allocated according to the determination result includes: In response to the determination result indicating that the error type of the erroneous memory is a correctable error, the number of errors of the erroneous memory exceeds a preset number, and the interval between two adjacent occurrences of the erroneous memory is less than a preset time length, the erroneous memory is persistently isolated from the memory resources to be allocated.
11. A memory error handling method, characterized in that: include: In response to an input instruction applied to an operation interface, running an application program under a target operating system on the operation interface; In response to a processing instruction acting on the operation interface, a rejection prompt is displayed on the operation interface; Among them, the rejection prompt is used to reject the application's application for or access to the error memory, the error memory is isolated based on the storage result, and the storage result is obtained by persistently storing the memory error information. The memory error information is used to record historical memory errors generated before the target operating system crashes.
12. A memory error handling system, characterized in that: include: a memory error collection component, configured to collect memory error information in response to restarting the target operating system, wherein the memory error information is used to record historical memory errors generated before the target operating system crashes; A memory error storage component is used to persistently store the memory error information and obtain a storage result; A memory error isolation component is used to isolate the error memory based on the storage result to deny the application program under the target operating system from applying for or accessing the error memory.
13. A memory error handling system, characterized in that: include: The client is configured to collect and send memory error information in response to restarting the target operating system, wherein the memory error information is used to record historical memory errors generated before the target operating system crashes; A server, connected to the client, configured to persistently store the memory error information to obtain a storage result, and isolate the error memory based on the storage result; The client is further configured to output a rejection prompt, wherein the rejection prompt is used to reject an application program under the target operating system from applying for or accessing the erroneous memory.
14. An electronic device, characterized in that: include: a memory storing an executable program; A processor is used to run the program, wherein the program executes the memory error handling method according to any one of claims 1 to 11 when running.
15. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the memory error handling method according to any one of claims 1 to 11.
16. A computer program product, characterized in that The invention comprises a computer program, which implements the memory error handling method according to any one of claims 1 to 11 when executed by a processor.
Citation Information
Patent Citations
Method and device for isolating internal memories
CN103279406A
Fault memory isolation method and device, equipment and storage medium
CN117472622A
Cited By
Memory management unit fault diagnosis method and device, computer device and medium
CN122363968A