Memory error processing method and system, electronic equipment, computer readable storage medium and computer program product

Through the operating system kernel, the memory error information is directly read from the target register and parsed, which solves the problem of low completeness of information collection in the firmware-first method, and realizes more efficient memory error information processing.

CN120523635AInactive Publication Date: 2025-08-22ALIBABA CLOUD COMPUTING CO LTD

Patent Information

Application Number
CN202511014226.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-08-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the memory error handling method that is preferred by firmware leads to low completeness of memory error information collection, high communication overhead, and serious information loss.

Method used

By directly reading memory error information from the target register in the operating system kernel, the communication link length is reduced, and the operating system kernel is used for parsing to ensure the completeness of information.

Benefits of technology

It improves the completeness of the collection of memory error information, reduces the overall communication overhead of the system, and improves the reliability and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523635A_ABST
    Figure CN120523635A_ABST
Patent Text Reader

Abstract

The invention discloses a memory error processing method and system, electronic equipment, a computer readable storage medium and a computer program product, and relates to the technical field of computers. The method is applied to an operating system kernel of a target computer architecture and comprises the steps that in response to a reading event corresponding to a memory error reading mechanism, memory error information is read from a target register, and the memory error reading mechanism is obtained according to instruction set type configuration corresponding to the target computer architecture; the memory error information is obtained by detecting a hardware module corresponding to the target register in a process of accessing a memory space; the memory error information is analyzed, an analysis result is obtained, and the analysis result is used for representing the memory position corresponding to the hardware module with the memory error in the target computer architecture. According to the method and the device, the technical problem of low collection integrity of memory error information in a firmware priority scheme in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a memory error handling method, system, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] In practical computer architecture applications, memory errors can occur, potentially causing system downtime and reducing system security and reliability. Therefore, accurately and effectively detecting and locating memory errors has become a pressing issue in the field.

[0003] In related technologies, a firmware-first approach is often used to handle memory errors in computer architectures. Specifically, when a hardware module detects a memory error, it counts the errors. When the number of errors reaches a threshold, the system firmware reads the memory error information, processes it, and reports it to the operating system kernel, which then locates and analyzes the memory error. This approach has several drawbacks: First, the communication overhead between the firmware and the kernel is high, resulting in a high overall cost for memory error handling. Second, the error count-based reporting method can lead to significant loss of memory error information, resulting in a low collection rate for memory error information.

[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0005] Embodiments of the present application provide a memory error handling method, system, electronic device, computer-readable storage medium, and computer program product to at least address the technical problem of low integrity of memory error information collection in firmware-first solutions in related technologies.

[0006] According to one aspect of an embodiment of the present application, a memory error handling method is provided, which is applied to an operating system kernel of a target computer architecture, comprising: reading memory error information from a target register in response to a read event corresponding to a memory error reading mechanism, wherein the memory error reading mechanism is configured according to an instruction set type corresponding to the target computer architecture, and the memory error information is detected by a hardware module corresponding to the target register during access to memory space; and parsing the memory error information to obtain a parsing result, wherein the parsing result is used to characterize a memory location corresponding to the hardware module where the memory error occurs in the target computer architecture.

[0007] According to another aspect of an embodiment of the present application, a memory error handling system is also provided, including: a hardware module for detecting and obtaining memory error information during access to memory space; a target register for storing memory error information; an operating system kernel for responding to a read event corresponding to a memory error reading mechanism, reading the memory error information from the target register, and parsing the memory error information to obtain a parsing result, wherein the memory error reading mechanism is configured according to an instruction set type corresponding to a target computer architecture, and the parsing result is used to characterize the memory location corresponding to the hardware module where a memory error occurs in the target computer architecture.

[0008] According to another aspect of an embodiment of the present application, an electronic device is further provided, including: a memory storing an executable program; and a processor for running the program, wherein any one of the above-mentioned memory error handling methods is executed when the program is running.

[0009] According to another aspect of an embodiment of the present application, a computer-readable storage medium is also provided, which includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned memory error handling methods.

[0010] According to another aspect of an embodiment of the present application, a computer program product is further provided, including a computer program, which implements any of the above-mentioned memory error handling methods when executed by a processor.

[0011] In an embodiment of the present application, first, the hardware module corresponding to the target register detects and obtains memory error information during the process of accessing the memory space. When the operating system kernel detects that the read event corresponding to the memory error reading mechanism is triggered, the operating system kernel can directly read the memory error information from the target register without relying on the firmware to report the memory error information to the operating system kernel, thereby shortening the link length for collecting memory error information and reducing the overall communication overhead of the system. Therefore, while consuming the same amount of communication overhead, more memory error information can be read and the lost memory error information can be reduced, thereby ensuring the completeness of the read memory error information. In addition, the above-mentioned memory error reading mechanism is configured according to the instruction set type corresponding to the target computer architecture, and can be targeted to configure the system with a memory error reading mechanism corresponding to the target computer architecture, ensuring that the memory error information can be read more completely; further, the operating system kernel is used to locate and analyze the memory error information to obtain an analysis result, so that the user can use the analysis result to determine the memory location of the hardware module where the memory error occurs in the target computer architecture.

[0012] From the above, the embodiments of the present application achieve the purpose of improving the collection completeness of memory error information by directly collecting memory error information through the operating system kernel, thereby achieving the technical effect of improving the collection completeness of memory error information in the target computer architecture, and further solving the technical problem of low memory error information collection completeness of the firmware-first solution in the related art.

[0013] It is easy to notice that the above general description and the following detailed description are merely for the purpose of exemplifying and explaining the present application, and do not constitute a limitation of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0015] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) according to a memory error handling method of an embodiment of the present application;

[0016] Figure 2 is a structural block diagram of a computing environment according to an embodiment of the present application;

[0017] Figure 3 This is a structural block diagram of a service grid according to an embodiment of the present application;

[0018] Figure 4 is a flowchart of a memory error handling method according to an embodiment of the present application;

[0019] Figure 5 is a schematic diagram of an optional memory error handling method according to an embodiment of the present application;

[0020] Figure 6 is a schematic diagram of the architecture of an optional memory error handling method according to an embodiment of the present application;

[0021] Figure 7 is a schematic diagram of another optional memory error handling method according to an embodiment of the present application;

[0022] Figure 8 is a schematic diagram of another optional memory error handling method according to an embodiment of the present application;

[0023] Figure 9 is a schematic diagram of the architecture of another optional memory error handling method according to an embodiment of the present application;

[0024] Figure 10 is a schematic diagram of an optional target computer architecture according to an embodiment of the present application;

[0025] Figure 11 is a schematic diagram of another optional memory error handling method according to an embodiment of the present application;

[0026] Figure 12 is a schematic diagram of an optional update polling cycle duration according to an embodiment of the present application;

[0027] Figure 13 is a structural block diagram of a memory error handling device according to an embodiment of the present application;

[0028] Figure 14 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0032] First, some nouns or terms that appear in the process of describing the embodiments of the present application are subject to the following explanations.

[0033] Reduced Instruction Set Computer (RISC): This simplifies the instruction set and addressing methods, making instructions easier to implement and enhancing the performance of parallel instruction execution. Compared to complex instruction sets, RISC can use fewer, simpler instructions to complete the same function, thereby improving overall computer performance.

[0034] Advanced RISC Machines (ARM): The ARM architecture is a RISC-based computer architecture suitable for devices requiring high performance and low power consumption, such as smartphones, tablets, and laptops. ARM-based computers include registers that store interrupt signals used to report memory errors. Therefore, in this application, the ARM architecture can use these interrupt signals to trigger read events.

[0035] Complex Instruction Set Computer (CISC): This provides a set of complex and powerful instructions that implement advanced operations directly in hardware. CISC reduces the complexity of code writing for programmers and compilers.

[0036] Advanced Micro Devices Architecture (AMD): The AMD architecture is a CISC-based computer architecture. AMD processors are compatible with the Intel x86 architecture. AMD processors typically utilize a multi-core design, allowing multiple threads to run simultaneously, improving concurrency and multitasking efficiency. Compared to the ARM architecture, the AMD architecture does not store interrupt signals in its registers. Therefore, in this application, the AMD architecture can set an information read cycle to actively trigger a read event.

[0037] Exception Levels (EL): Exception levels can be used to describe the severity of exceptions that occur during system operation. In this application, exception levels can be divided into four levels, namely EL0, EL1, EL2, and EL3. Among them, exception level EL3 represents the highest level of exception privileges, which is used to handle low-level hardware exceptions, such as security exceptions and hardware resets. When the exception level is EL3, it has full access rights to the target computer architecture.

[0038] According to an embodiment of the present application, an embodiment of a memory error handling method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0039] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The following is a hardware block diagram of a computer terminal (or mobile device) for implementing a memory error handling method. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (illustrated as 102a, 102b, ..., 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microcontroller unit (MCU) or a field programmable gate array (FPGA)), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, the computer terminal 10 may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a computer bus), a network interface, a cursor control device (such as a mouse, touchpad, etc.), a keyboard, a power supply, and / or a camera.

[0040] It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0041] It should be noted that the one or more processors 102 and / or other data processing circuits mentioned above may generally be referred to herein as "other data processing circuits". The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be fully or partially integrated into any of the other components in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0042] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the memory error handling method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned memory error handling method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0043] Transmission device 106 is configured to connect to a network via a network interface to receive or transmit data. Specific examples of the aforementioned network may include a wired and / or wireless network provided by the communications provider of computer terminal 10. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module configured to communicate with the Internet wirelessly.

[0044] like Figure 1 The display shown may be, for example, a touch screen liquid crystal display (LCD), which enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0045] It should be noted that, in some optional embodiments, the above Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the aforementioned computer device (or mobile device).

[0046] Figure 1 The hardware structure block diagram shown can be used not only as an exemplary block diagram of the computer terminal 10 (or mobile device), but also as an exemplary block diagram of the server. In an optional embodiment, Figure 2 The block diagram shows the use of the above Figure 1The computer terminal 10 (or mobile device) shown is an embodiment of a computing node in the computing environment 201 . Figure 2 A structural diagram of a computing environment is shown, such as Figure 2 As shown, computing environment 201 includes multiple computing nodes (e.g., servers) (illustrated as 210-1, 210-2, ...) running on a distributed network. Each computing node contains local processing and memory resources, allowing end users 202 to remotely run applications or store data in computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 in computing environment 201, representing services "A," "D," "E," and "H," respectively.

[0047] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or request of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 can include a corresponding agent to handle the provisioning and / or request for services (one or more services provided in computing environment 201).

[0048] Services are provided or deployed based on various virtualization technologies supported by computing environment 201. In some embodiments, services can be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar approaches. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While VMs virtualize machines, container-based virtualization can launch containers to virtualize the entire operating system (OS), allowing multiple workloads to run on a single OS instance.

[0049] In an embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). Figure 2 As shown, service 220-2 can be deployed with one or more Pods 240-1, 240-2, ..., 240-N (collectively, "Pods"). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively, "Containers"). One or more containers in a Pod handle requests related to one or more corresponding functions of the service. Proxy 245 typically controls network functions related to the service, such as routing and load balancing. Similar Pods can also be deployed with other services.

[0050] During operation, executing a user request from the end user 202 may require calling one or more services in the computing environment 201, and executing one or more functions of a service may require calling one or more functions of another service. Figure 2 As shown, service “A” 220 - 1 receives a user request from end user 202 from ingress gateway 230 , service “A” 220 - 1 may call service “D” 220 - 2 , and service “D” 220 - 2 may request service “E” 220 - 3 to perform one or more functions.

[0051] This computing environment can be a cloud computing environment, where resource allocation is managed by the cloud service provider, allowing for feature development without having to worry about implementing, adjusting, or scaling servers. This computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be partitioned to perform a set of functions that can scale independently and automatically, rather than scaling a single hardware device to handle the potential load.

[0052] In another optional embodiment, Figure 3 The block diagram shows the use of the above Figure 1 The computer terminal 10 (or mobile device) is shown as an embodiment of a service grid. Figure 3 A structural diagram of a service grid is shown, Figure 3 As shown in the figure, the service grid is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to decomposing an application into multiple smaller services or instances and distributing them across different clusters / machines.

[0053] like Figure 3 As shown, the microservices may include application service instance A and application service instance B, which form the functional application layer of the service mesh. In one embodiment, application service instance A runs as a container / process 308 on a machine / workload container group 314 (pod), and application service instance B runs as a container / process 310 on a machine / workload container group 316 (pod).

[0054] In one implementation, application service instance A may be a product query service, and application service instance B may be a product ordering service.

[0055] like Figure 3As shown, application service instance A and grid proxy (sidecar) 303 coexist in machine / workload container group 314, while application service instance B and grid proxy 305 coexist in machine / workload container group 316. Grid proxy 303 and grid proxy 305 form the data plane layer of the service mesh. Grid proxy 303 and grid proxy 305 run as container / process 304 and container / process 306, respectively, and can receive requests 312 for product query services. Bidirectional communication is possible between grid proxy 303 and application service instance A, and between grid proxy 305 and application service instance B. Furthermore, bidirectional communication is possible between grid proxy 303 and grid proxy 305.

[0056] In one embodiment, all traffic from application service instance A is routed to the appropriate destination via grid proxy 303, and all network traffic from application service instance B is routed to the appropriate destination via grid proxy 305. It should be noted that network traffic referred to herein includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), the high-performance, general-purpose open source framework (Google Remote Procedure Call, gRPC), and the open source in-memory data structure storage system (Redis).

[0057] In one embodiment, the data plane layer's functionality can be extended by writing custom filters for the proxy (Envoy) in the service mesh. The service mesh proxy configuration can be designed to enable the service mesh to correctly proxy service traffic, enabling service interoperability and service governance. Mesh proxy 303 and mesh proxy 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.

[0058] like Figure 3 As shown, the service grid also includes a control plane layer. The control plane layer can be a group of services running in a dedicated namespace, and these services are hosted by a hosting control plane component 301 in a machine / workload container group (machine / Pod) 302. Figure 3As shown, managed control plane component 301 communicates bidirectionally with mesh proxy 303 and mesh proxy 305. Managed control plane component 301 is configured to perform certain control and management functions. For example, managed control plane component 301 receives telemetry data transmitted by mesh proxy 303 and mesh proxy 305 and can further aggregate this telemetry data. Managed control plane component 301 also provides user-oriented application programming interfaces (APIs) to facilitate manipulation of network behavior and provide configuration data to mesh proxy 303 and mesh proxy 305.

[0059] Under the above operating environment, this application provides Figure 4 Memory error handling method shown. Figure 4 is a flow chart of a memory error handling method according to an embodiment of the present application, such as Figure 4 As shown, the memory error handling method is applied to the operating system kernel of the target computer architecture and includes the following steps.

[0060] Step S41, in response to a read event corresponding to a memory error reading mechanism, memory error information is read from a target register, wherein the memory error reading mechanism is configured according to an instruction set type corresponding to a target computer architecture, and the memory error information is detected by a hardware module corresponding to the target register during access to memory space.

[0061] The target computer architecture may refer to a specific type of computer system, including hardware components, instruction sets, data paths, and processing mechanisms. For example, the target computer architecture may include, but is not limited to, a first architecture (ARM architecture), a second architecture (AMD architecture), and a third architecture (Very Long Instruction Word architecture, or VLIW architecture).

[0062] The operating system kernel may be responsible for managing the operating system, for example, system memory management, system network management, processor scheduling, etc. The operating system kernel may include multiple components, for example, a scheduler, a memory manager, a device driver, and a system call interface.

[0063] The information corresponding to the target computer architecture (e.g., instruction set type) can be determined by querying the hardware configuration, identifying the processor identifier, or configuring a program based on information loaded about the target computer architecture. In particular, the instruction set type corresponding to the target computer architecture can also be determined during the construction of the target computer architecture.

[0064] The aforementioned hardware module may refer to a specific component for monitoring and reporting memory errors. The aforementioned target register may be used to store memory error information. When the hardware module detects a memory error while accessing memory space, it may collect memory error information corresponding to the memory error and store the memory error information in the target register. During the memory access process, the hardware module may detect whether the data stored in the memory space contains erroneous data bits. If it is determined that the data stored in the memory space contains erroneous data bits, a memory error is determined to have occurred.

[0065] The target register can also be used to interact with the operating system kernel. When the operating system kernel detects a read event corresponding to the memory error read mechanism (i.e., in response to a read event corresponding to the memory error read mechanism), the operating system kernel can directly read the memory error information from the target register. This memory error information may include, but is not limited to, the error type (e.g., correctable memory error, uncorrectable memory error), the device physical address where the error occurred, the error bit location, and the time when the error occurred (e.g., a timestamp).

[0066] In an exemplary application scenario, a memory error reading program is designed based on the operating characteristics of the target computer architecture, and the memory error reading program is stored in the operating system kernel. When the operating system kernel detects that a read event corresponding to the above-mentioned memory error reading mechanism is triggered, the above-mentioned memory error reading program is driven to support the operating system kernel to directly read the memory error information from the target register without relying on the system firmware to read the memory error information.

[0067] In another exemplary application scenario, the Input / Output Memory Management Unit (IOMMU) of the operating system kernel can be configured so that the hardware module triggers an interrupt when accessing a specific memory area (e.g., a target register), thereby allowing the operating system kernel to read memory error information directly from the target register.

[0068] It should be noted that when different memory error reading mechanisms are configured in the operating system kernel of the target computer architecture, the triggering mechanism of the above-mentioned reading event is also different. The triggering mechanism of the reading event will affect the collection rate of memory error information. For example, for a computer architecture that supports immediate interrupts, setting a real-time interrupt triggering mechanism can trigger the above-mentioned reading event every time the hardware module detects memory error information, thereby collecting memory error information in real time and improving the collection rate of memory error information.

[0069] It should also be noted that the above-mentioned target registers are associated with the above-mentioned target computer architecture. That is, for different types of target computer architectures, the above-mentioned target registers can also be set to different (different numbers, different types) of registers. In this application, there is no specific limitation on the type and number of target registers.

[0070] It is easy to understand that different trigger mechanisms are set for different memory error reading mechanisms. When supported by the target computer architecture, more read events are triggered to collect memory error information more comprehensively. In response to the read events corresponding to the memory error reading mechanism, the operating system kernel can directly read the memory error information from the target register without relying on the firmware to read the memory error information, and then the firmware reports the memory error information to the operating system kernel, which shortens the link length for collecting memory error information and can reduce the overall communication overhead of the system. Therefore, while consuming the same amount of communication overhead, more memory error information can be read and the lost memory error information can be reduced, thereby ensuring the completeness of the read memory error information.

[0071] Step S42 , parsing the memory error information to obtain a parsing result, wherein the parsing result is used to characterize a memory location corresponding to a hardware module where a memory error occurs in a target computer architecture.

[0072] The above-mentioned parsing results may include memory physical layout addresses (e.g., memory row and column addresses, system physical addresses). The parsing results can be used to detect and locate memory errors. The parsing results may be in a data format recognizable by the operating system kernel. It is understood that the memory error information read above is in a data format recognizable by the hardware device. This memory error information needs to be parsed and converted into a data format recognizable by the operating system kernel (i.e., to obtain the parsing results) to support the operating system kernel in further processing of the above-mentioned memory error information.

[0073] The above-mentioned analysis results can be obtained by utilizing data analysis methods. The above-mentioned data analysis methods may include but are not limited to: analysis methods based on machine learning models (for example, inputting memory error information into a pre-trained machine learning model, using the machine learning model to infer the memory error information, thereby converting the memory error information into a memory physical layout address to obtain an analysis result), page table conversion methods (for example, a page table contains multiple entries, each entry stores the correspondence between memory error information and memory location information, and by searching the page table, the address information in the memory information can be converted into the corresponding memory physical layout address to obtain an analysis result), hardware-assisted address translation methods, and memory mapping information table methods (for example, describing the mapping relationship between memory error information and memory location information in the memory mapping information to achieve the conversion of address information in the memory information into the corresponding memory physical layout address). The above-mentioned memory location may refer to the physical location of the hardware module where the memory error occurs in the target computer architecture.

[0074] It should be noted that for different target computer architectures, an appropriate data parsing method can be selected to convert memory error information into a data format recognizable by the operating system kernel of the target computer architecture to obtain more accurate parsing results.

[0075] It is easy to understand that the memory error information is parsed and converted into a data form that can be recognized by the operating system kernel to obtain a parsing result. Then, the operating system kernel can use the parsing result to determine the memory location of the hardware module where the memory error occurs in the target computer architecture, thereby realizing memory error prediction, reducing the system downtime rate, and enhancing the reliability and stability of the system. In the related art, since the memory error information is collected by the firmware, the firmware needs to parse the memory error information before uploading it to the operating system kernel, and then the operating system kernel performs a second parsing of the memory error information before the operating system kernel can recognize the memory error information. Compared with the related art, the present application does not need to rely on the firmware to collect the memory error information. Therefore, there is no need to parse the memory error information multiple times. The operating system kernel can directly parse the memory error information, which can save the overall resource cost of the system.

[0076] In an embodiment of the present application, first, the hardware module corresponding to the target register detects and obtains memory error information during the process of accessing the memory space. When the operating system kernel detects that the read event corresponding to the memory error reading mechanism is triggered, the operating system kernel can directly read the memory error information from the target register without relying on the firmware to report the memory error information to the operating system kernel, thereby shortening the link length for collecting memory error information and reducing the overall communication overhead of the system. Therefore, while consuming the same amount of communication overhead, more memory error information can be read and the lost memory error information can be reduced, thereby ensuring the completeness of the read memory error information. In addition, the above-mentioned memory error reading mechanism is configured according to the instruction set type corresponding to the target computer architecture, and can be targeted to configure the system with a memory error reading mechanism corresponding to the target computer architecture, ensuring that the memory error information can be read more completely; further, the operating system kernel is used to locate and analyze the memory error information to obtain an analysis result, so that the user can use the analysis result to determine the memory location of the hardware module where the memory error occurs in the target computer architecture.

[0077] From the above, the embodiments of the present application achieve the purpose of improving the collection completeness of memory error information by directly collecting memory error information through the operating system kernel, thereby achieving the technical effect of improving the collection completeness of memory error information in the target computer architecture, and further solving the technical problem of low memory error information collection completeness of the firmware-first solution in the related art.

[0078] In an optional embodiment, in the above-mentioned memory error handling method, memory error information is collected by the hardware module and reported to the target register when a target type of memory error is detected in the memory space.

[0079] During the application of this application, memory errors occurring in the memory space may be classified into multiple types according to classification criteria. The classification criteria may include, but are not limited to, classification criteria based on whether the memory error is correctable, classification criteria based on security level, classification criteria based on time, and classification criteria based on the number of bits.

[0080] After the memory errors are divided into multiple types, the target type of memory error can be selected from the multiple types of memory errors according to the actual application requirements, so that when the hardware module detects that a memory error of the target type occurs in the memory space during the process of accessing the memory space, the memory error information corresponding to the target type of memory error is collected and reported to the target register. In this way, the hardware module can collect memory error information in a targeted manner and enhance the accuracy of subsequent analysis of the memory error information. In addition, based on the above-mentioned reporting mechanism, not only can unnecessary data transmission be reduced and the waste of system resources be reduced, in particular, frequent write operations on the target register are avoided, the loss of the target register is reduced, and the situation that all types of memory errors are stored in the target register, resulting in a large amount of storage space of the target register being occupied and the memory error of concern cannot be stored, can also be avoided.

[0081] In particular, based on the classification standard of whether the memory errors are correctable, memory errors occurring in the memory space can be divided into correctable errors (Correctable Error, referred to as CE) and uncorrected correctable errors (Uncorrected Correctable Error, referred to as UE).

[0082] The aforementioned correctable memory error may refer to a data bit misalignment detected in a single dynamic random-access memory (DRAM) chip during a single memory access by a hardware module. The aforementioned uncorrectable error may refer to a data bit misalignment detected in multiple DRAM chips during a single memory access by a hardware module.

[0083] In an exemplary application scenario, Figure 5 As shown, the target type of memory error may include a correctable error. During a single access of the memory space by the hardware module, when the hardware module detects a data bit misalignment in a single DRAM chip, it can be considered that the hardware module has detected a target type of memory error (i.e., a correctable error) in the memory space. At this time, the hardware module will collect the memory error information corresponding to the correctable error and report the collected memory error information to the target register. Furthermore, when the operating system kernel detects a read event trigger, it can read the memory error information corresponding to the correctable error from the target register to ensure that the correctable error can be subsequently parsed in a targeted manner, thereby achieving the effect of accurately predicting uncorrectable errors based on correctable errors in application scenarios, making it easier for users to monitor the health of the memory space and take timely measures to improve the overall stability of the target computer architecture.

[0084] It is easy to understand that the above optional embodiments of the present application can achieve the following technical effects: when the hardware module detects a memory error of the target type, it collects and reports it to the target register, that is, the hardware module can specifically collect the memory error of the target type, thereby enhancing the accuracy of subsequent analysis of the memory error. In addition, based on the above-mentioned reporting mechanism, the embodiment of the present application can reduce unnecessary data transmission and reduce the waste of system resources. In particular, when the memory error of the target type is a correctable error, by reporting the correctable error to the target register for reading by the operating system kernel, it can further support the subsequent process of performing targeted parsing and positioning of the correctable error, thereby achieving accurate prediction of uncorrectable errors based on correctable errors in application scenarios, making it easy to take measures for uncorrectable errors to improve the overall stability of the target computer architecture.

[0085] In an optional embodiment, the above memory error handling method further includes the following method steps:

[0086] Step S43: configuring a memory error reading mechanism based on the instruction set type corresponding to the target computer architecture.

[0087] The above-mentioned instruction set may refer to a set of instructions that can be recognized and executed by the processor of the target computer architecture. It is understandable that each computer architecture corresponds to a unique instruction format, addressing mode and operand type, that is, each computer architecture corresponds to a unique instruction set, and thus the above-mentioned instruction set type can be used to determine the operating characteristics of the target computer architecture. The above-mentioned instruction set types may include a reduced instruction set type, a complex instruction set type, a very long instruction word type (for example, a very long instruction word type instruction set may be used in a VLIW architecture. The instructions in this instruction set contain multiple operations. The compiler determines the parallelism of the instructions, and the hardware is responsible for execution to achieve instruction-level parallelism and improve processing efficiency), a graphics processor instruction set type, and a hybrid instruction set type (for example, a hybrid instruction set constructed by combining the characteristics of a reduced instruction set and a complex instruction set. The use of this hybrid instruction set helps to deploy applications in the memory of an operating system supported by multiple architectures).

[0088] The memory error reading mechanism described above can be used to read memory error information. The memory error reading mechanism may include, but is not limited to, an interrupt response reading mechanism, a polling reading mechanism, a manual reading mechanism, and a threshold-based reading mechanism.

[0089] By utilizing the instruction set type corresponding to the target computer architecture, the operating characteristics of the target computer architecture are determined, and then, based on the operating characteristics of the target computer architecture, an adaptive memory error reading mechanism is configured for the operating system kernel of the target computer architecture, ensuring that the subsequent operating system kernel can utilize the memory error reading mechanism to directly read memory error information from the target register.

[0090] In particular, the registers of the ARM architecture store interrupt signals used for reporting memory errors. Therefore, when the target computer architecture is the ARM architecture, an interrupt response reading mechanism can be configured for the operating system kernel of the target computer architecture. Compared with the ARM architecture, the registers of the AMD architecture do not store interrupt signals used for reporting memory errors. Therefore, when the target computer architecture is the AMD architecture, a polling reading mechanism can be configured for the operating system kernel of the target computer architecture. Furthermore, when the target computer architecture is another architecture, a corresponding memory error reading mechanism can also be configured to ensure that the configured memory error reading mechanism can read memory error information from the target registers.

[0091] It is easy to understand that the above-mentioned optional embodiments of the present application can achieve the following technical effects: by configuring the memory error reading mechanism based on the instruction set type corresponding to the target computer architecture, the memory error reading mechanism corresponding to the target computer architecture can be configured for the operating system kernel of the target computer architecture in a more targeted manner, thereby ensuring that the operating system kernel can effectively read memory error information from the target register.

[0092] In an optional embodiment, the memory error reading mechanism is an interrupt response reading mechanism. In step S43, the memory error reading mechanism is configured based on the instruction set type, including the following method steps:

[0093] Step S431, in response to the instruction set type being the first type, obtaining interrupt mapping information and base address information corresponding to the target register from an error source table, wherein the error source table is reported by the target firmware to the operating system kernel when the target computer architecture is started;

[0094] Step S432 : registering an interrupt response program corresponding to the interrupt response reading mechanism based on the interrupt mapping information and the base address information, wherein the interrupt response program is used to respond to a read event triggered by an interrupt signal to read memory error information.

[0095] The first type can be set to a reduced instruction set computing (RISC) type. When the instruction set type is a RISC type, the target computer architecture can be the first architecture (ARM architecture). The target firmware can be used to initialize and configure the operating system kernel. The target firmware has a higher execution privilege level than the operating system kernel. The error source table can be obtained by converting configuration information of the target computer architecture.

[0096] The target firmware has access to the configuration information of the target computer architecture, so the target firmware can collect the configuration information of the target computer architecture. Figure 6 As shown, the target firmware can collect the interrupt numbers and base addresses of registers used by the register agent and then organize this information into an error source table according to a specific format. For example, the collected information can be organized into an error source table according to the Error Source Table (AEST) format in the Advanced Configuration and Power Interface (ACPI) standard. The error source table is represented in the AEST table format. Furthermore, the target firmware loads the AEST table into the ACPI table. When the ARM architecture boots, the error source table is reported to the operating system kernel via the ACPI table.

[0097] The error source table may include, but is not limited to, the interrupt number used by the register, the register's base address information, and the error source identifier. When the operating system kernel detects that the instruction set type is RISC, it parses the error source table reported by the target firmware and extracts the interrupt mapping information and base address information corresponding to the target register from the error source table. Methods for parsing the error source table may include, but are not limited to, parsing methods based on ACPI parsing tools, parsing methods based on dynamic link libraries, and parsing methods based on pre-set parsing files.

[0098] The interrupt mapping information may refer to configuration information linking the interrupt number of the target register with the interrupt response program. This interrupt mapping information may include, but is not limited to, the interrupt number used by the target register, the interrupt priority, the interrupt vector, and the interrupt mask configuration. The interrupt number may refer to a unique identifier for a hardware interrupt request. This interrupt number may indicate the specific manner in which memory error information is reported to the operating system kernel, such as when and what notification operation is performed. When the hardware module detects a memory error (particularly, a memory error of the target type) while accessing memory space, the hardware module triggers this interrupt number, thereby reporting the memory error information corresponding to the memory error to the operating system kernel via the target register. Based on this, by parsing the interrupt number in the error source table, the configuration information linking the interrupt number of the target register with the interrupt response program (i.e., the interrupt mapping information) can be obtained. The interrupt response program in the operating system kernel can then be registered based on this interrupt mapping information, ensuring that the operating system kernel can immediately respond to the interrupt signal corresponding to the interrupt number to trigger a read event.

[0099] The above-mentioned base address information can be used to identify the storage location of the memory error information. The base address information can represent the starting position of the target register. When the hardware module detects a memory error (especially, detects a target type of memory error) during the process of accessing the memory space, the hardware module will report the memory error information corresponding to the memory error to the target register. Therefore, registering the interrupt response program in the operating system kernel according to the base address information can guide the operating system kernel to access the storage location of the memory error information, thereby ensuring that the operating system kernel can use the base address information to access the starting position of the target register to directly read the memory error information from the target register, and collect the memory error information without relying on the firmware to upload the memory error information.

[0100] Still like Figure 6 As shown, after the operating system kernel parses the interrupt mapping information and base address information corresponding to the target register from the error source table, the interrupt response program corresponding to the interrupt response reading mechanism is registered in the operating system kernel based on the interrupt mapping information and base address information. Specifically, a pre-designed interrupt handler is stored in the operating system kernel, and the interrupt handler can be used to read the information stored in the register. After the operating system kernel parses the interrupt mapping information and base address information corresponding to the target register, the pre-designed interrupt handler is called and configured using the interrupt mapping information and base address information corresponding to the target register through the built-in interface in the operating system kernel to implement the registration of the interrupt response program. Furthermore, when the hardware module detects that a memory error has occurred, the interrupt response program can be used to enable the operating system kernel to respond to the interrupt signal to trigger a read event, thereby directly reading the memory error information from the target register using the base address information, thereby implementing the above-mentioned interrupt response reading mechanism.

[0101] It is easy to understand that the above-mentioned optional embodiments of the present application can achieve the following technical effects: for a computer architecture with an instruction set type of the first type, based on the error source table provided by the target firmware, the operating system kernel can automatically parse and obtain the interrupt mapping information and base address information corresponding to the target register, and then register the interrupt response program in the kernel based on the interrupt mapping information and base address information, to ensure that the operating system kernel can respond to the interrupt signal to trigger a read event and use the base address information to access the starting position of the target register, so as to support the operating system kernel to directly read the memory error information from the target register without relying on the firmware to upload the memory error information.

[0102] In an optional embodiment, the above memory error handling method further includes the following method steps:

[0103] Step S44 , in response to receiving an interrupt signal routed by a processor in the target computer architecture, determining a read event trigger corresponding to an interrupt response read mechanism, wherein the interrupt signal is generated by a target register according to memory error information and sent to the processor.

[0104] The interrupt signal can be used to indicate an event requiring an immediate processor response, requiring the processor to interrupt the current task to handle the event. Specifically, the interrupt signal can be used to indicate a target type of memory error in memory space. The interrupt signal can be a digital signal. The interrupt signal can include the interrupt type (e.g., secure interrupt, non-secure interrupt), an interrupt source identifier, and the base address of the target register.

[0105] The above-mentioned interrupt signal can be generated in the following ways: generating an interrupt signal based on software (for example, by setting a specific instruction, when memory error information is detected to be stored in the target register, using the instruction and the memory error information to generate an interrupt signal), generating an interrupt signal based on an error status register (for example, when the state of the error status register in the target computer architecture changes, generating an interrupt signal according to the memory error information).

[0106] Specifically, a predefined interrupt number can be set at a specific data bit within the target register. When the target register receives memory error information uploaded by the hardware module, the interrupt number is triggered, causing the target register to generate an interrupt signal based on the memory error information and send the interrupt signal to the processor, allowing the operating system kernel to accurately trigger a read event for the target register based on the interrupt source identification code in the interrupt signal. Based on this, the interrupt signal generated based on the memory error information is routed to the operating system kernel via the processor, and the operating system kernel directly responds to the interrupt signal and accurately reads the memory error information from the target register, reducing the consumption of system resources.

[0107] It should be noted that the ARM architecture supports the operating system kernel to directly process non-secure interrupt signals at the EL2 exception level (in particular, non-secure interrupt signals can be interrupt signals triggered by correctable errors). Therefore, in this solution, only non-secure interrupt signals at the EL2 exception level can be routed to the operating system kernel to support the operating system kernel to respond to interrupt signals immediately, while secure interrupt signals at the EL3 exception level are still routed to the target firmware for processing. In the related art, a firmware-first solution is adopted, which requires first triggering an EL3 exception-level interrupt signal so that the firmware responds to the EL3 exception-level interrupt signal. After the firmware collects and obtains memory error information, it triggers an EL2 exception-level interrupt signal for subsequent processing by the operating system kernel. Compared with the firmware-first solution, the present application can support the operating system kernel to directly process EL2 exception-level interrupt signals, reducing the length of the hardware correctable error collection link and reducing the overall communication overhead of the system, thereby consuming the same amount of communication overhead while being able to read more memory error information. Therefore, the present application supports triggering a read event corresponding to the interrupt response reading mechanism when each correctable error type memory error occurs, thereby achieving complete collection of correctable error type memory errors (i.e., achieving 100% collection of correctable error type memory errors).

[0108] In an exemplary application scenario, the target register includes multiple sub-registers, wherein the sub-registers can store function configuration information and other configuration information. Figure 5 As shown, each time the hardware module stores memory error information in a target register, the target register can generate an interrupt signal based on the memory error information and configuration data (including functional configuration information and other configuration information) and send the interrupt signal to the central processing unit (CPU). After the target register notifies the CPU, the CPU routes the interrupt signal to the operating system kernel. Then, when the operating system kernel receives the interrupt signal routed by the CPU, a read event is triggered. That is, each time the hardware detects a memory error of a correctable error type, a read event can be triggered immediately. The operating system kernel can collect more memory error information, reduce the loss of memory error information, and ensure the integrity of the memory error information, so as to support accurate parsing and positioning of the memory error information in subsequent processes.

[0109] It is easy to understand that the above optional embodiments of the present application can achieve the following technical effects: when the hardware module stores the memory error information to the target register, the target register can generate an interrupt signal based on the memory error information, and the interrupt signal is routed to the operating system kernel via the processor. The operating system kernel directly responds to the interrupt signal without relying on the firmware to process the interrupt signal, which can enhance the real-time performance of the memory error information method. In addition, the operating system kernel can accurately trigger a read event for the target register based on the interrupt signal to accurately read the memory error information from the target register. In particular, when the hardware module can generate an interrupt signal to immediately trigger a read event each time it stores the memory error information to the target register, the real-time performance of the memory error handling method can be enhanced, and the operating system kernel can collect more memory error information, reduce the lost memory error information, and ensure the integrity of the memory error information to support the accurate parsing and positioning of the memory error information in the subsequent process.

[0110] In an optional embodiment, the hardware module is a memory component configured in the target computer architecture, and the target register corresponding to the memory component includes multiple sub-registers associated with the storage agent. In step S41, memory error information is read from the target register, including the following method steps:

[0111] In step S411, the interrupt response program is driven to determine the register agent according to the base address information, and multiple data items in the memory error information are read from multiple sub-registers. The multiple data items include: the error type of the memory error, the device physical address corresponding to the memory error, and auxiliary information of the memory error.

[0112] The above-mentioned storage agent can be used to manage a specific set of sub-registers. The storage agent can be used to record memory error information. The storage agent can be regarded as a container, containing multiple sub-registers, and the multiple sub-registers contained in the storage agent can be regarded as target registers. The above-mentioned storage agent can include multiple types of sub-registers. In particular, each sub-register can be used to store a corresponding data item in the memory error information. By storing multiple data items of memory errors in multiple sub-registers, the organization and integrity of data storage can be enhanced, making the error information clearer and easier to read, and facilitating the subsequent reading of the memory error information by the operating system kernel.

[0113] When the operating system kernel detects a read event triggered by an interrupt signal, it drives the interrupt response program to read the base address information to determine the starting address of the hosting agent based on the base address information, thereby identifying the locations of multiple sub-registers associated with the hosting agent. This ensures that the operating system kernel can accurately access and read multiple data items of memory error information, thereby enabling the operating system kernel to directly read the memory error information from the target register without requiring firmware to communicate with the kernel, reducing data transfer steps and delays in the memory error handling method. Furthermore, the operating system kernel drives the interrupt response program to read multiple data items of memory error information from the multiple sub-registers associated with the hosting agent. The multiple data items may include: the error type of the memory error, the device physical address corresponding to the memory error, and auxiliary information of the memory error. This enables the operating system kernel to more comprehensively read the memory error information, providing more comprehensive data support for subsequent parsing and locating the memory error information, enhancing the accuracy of the memory error handling method, and facilitating timely implementation of measures to improve the overall stability of the target computer architecture.

[0114] The error type of the memory error is used to characterize the nature of the memory error that occurred in the memory space. For example, the error type may include correctable error or uncorrectable error. The device physical address can be used to identify the location where the memory error occurred. The device physical address may refer to an address at the memory hardware level. The auxiliary information can be used to provide additional characteristics of the memory error. This auxiliary information may include, but is not limited to, error flags, error severity, error context, error source identifier, error count, timestamp, error attributes, device status, etc.

[0115] In an exemplary application scenario, Figure 6 As shown, the target computer architecture is configured with multiple hardware modules ( Figure 6 ), in particular, the above-mentioned hardware module 1, hardware module 2, and hardware module 3 are all memory components, wherein there is a correspondence between the hardware module 1, the error source table 1, and the interrupt response program 1. Exemplarily, the target firmware can collect the interrupt number used by the register in the storage agent 1 and the base address information of the register, organize the collected information into the error source table 1 according to the format of the AEST table, and the operating system kernel parses the error source table 1 to obtain the interrupt mapping information and the base address information, and registers the interrupt response program 1 in the operating system kernel based on the interrupt mapping information and the base address information. Similarly, the correspondence between the hardware module 2, the error source table 2, and the interrupt response program 2, as well as the correspondence between the hardware module 3, the error source table 3, and the interrupt response program 3 can be determined.

[0116] Still in the above application scenario, the target register corresponding to each memory component includes a group of sub-registers (that is, multiple sub-registers associated with the hosting agent). For any group of sub-registers, it can include the first sub-register (ERRFHI), the second sub-register (ERRERI), the third sub-register (ERRCRI), the fourth sub-register (ERRCTRL), the fifth sub-register (ERRSTATUS), the sixth sub-register (ERRADDR), the seventh sub-register (ERRMISC), and the eighth sub-register (ERRFR). The first, second, and third subregisters described above can be used to record the interrupt number used to report memory errors. Specifically, the first subregister (ERRFHI) can store configuration information for the error detection types supported by the interrupt number (for example, a portion of the bit fields contained in the first subregister (ERRFHI) can be used to determine whether the interrupt number supports secure interrupts and / or non-secure interrupts). The second subregister (ERRERI) can store behavior configuration information for the interrupt number (for example, a portion of the bit fields contained in the second subregister can be used to determine whether the interrupt number triggers a read event). The third subregister (ERRCRI) can store the interrupt type used for memory errors (particularly, a portion of the bit fields contained in the third subregister (ERRCRI) can be used to determine whether the interrupt type used for correctable memory errors is a non-secure interrupt). The bit fields in the first, second, and third subregisters (ERRFHI), ERRERI, and ERRCRI can be set during the target register initialization process. The fourth subregister can be used to record configuration information for the hosting agent. The fifth subregister can be used to record the error type of the memory error. The sixth subregister can be used to record the device physical address corresponding to the memory error. The seventh subregister can be used to record auxiliary information about the memory error. The eighth subregister can be used to record descriptive information corresponding to the functions of the registrar. Specifically, three seventh subregisters (ERRMISC) can be included, designated ERRMISC0, ERRMISC1, and ERRMISC2, respectively.

[0117] Still in the above application scenarios, such as Figure 7As shown, when the operating system kernel detects a triggering read event, an interrupt response program is executed. Specifically, the operating system kernel driver interrupt response program reads the base address information corresponding to the target register to determine the starting address of the storage agent based on the base address information, thereby identifying the locations of multiple sub-registers associated with the storage agent; further, the operating system kernel driver interrupt response program reads the error type information of the memory error from the fifth sub-register (ERRSTATUS) to determine whether a memory error of the target type has occurred. When a memory error of the target type has not occurred, the interrupt response program ends. When a memory error of the target type has occurred, it is determined whether there is a value in the sixth sub-register (ERRADDR); further, when there is a value in the sixth sub-register (ERRADDR), the device physical address corresponding to the memory error is read from the sixth sub-register (ERRADDR) and a determination is made as to whether there is a value in the seventh sub-register (ERRMISC); when there is no value in the sixth sub-register (ERRADDR), it is directly determined whether there is a value in the seventh sub-register (ERRMISC); further, when there is a value in the seventh sub-register (ERRMISC), auxiliary information of the memory error is read from the seventh sub-register (ERRMISC), and the above-read data is passed to the error handling driver to parse the memory error information.

[0118] In some exemplary application scenarios, the hosting agent may be a reliability, availability, and serviceability (RAS) agent, which is used to measure the stability and maintainability of a computer system during operation.

[0119] It is easy to understand that the above optional embodiments of the present application can achieve the following technical effects: by driving the interrupt response program in the kernel, the storage agent can be determined based on the base address information, ensuring that the operating system kernel can accurately access and read multiple data items of the memory error information, thereby achieving the effect of the operating system kernel directly reading the memory error information from the target register, without the need for firmware to communicate with the kernel, which can reduce the steps of data transmission and reduce the delay of the memory error handling method; further, the operating system kernel drives the interrupt response program to read multiple data items of the memory error information from multiple sub-registers, which can read the memory error information more comprehensively, providing more comprehensive data support for the subsequent parsing and positioning of the memory error information. In addition, the multiple data items in the above memory error information are respectively stored in the corresponding sub-registers, which can enhance the clarity of data storage and help the operating system kernel to read the memory error information more easily.

[0120] In an optional embodiment, the parsing result includes memory row and column addresses. In step S42, the memory error information is parsed to obtain the parsing result, which includes the following method steps:

[0121] Step S421 : using the error handling driver, error type and auxiliary information, convert the device physical address into a memory row and column address.

[0122] The error handling driver can be used to parse the device physical address. The device physical address represents the physical location of the hardware device where the memory error occurred. The memory row and column addresses represent the location of the memory error within the internal structure of the memory system.

[0123] It is understandable that the device physical address represents the location of the storage unit on the memory component. The device physical address can be based on memory pages to facilitate the management of the device physical address by the operating system. However, at the internal structural layout level of the memory system, memory errors are not stored in the form of memory pages, but rather memory errors are stored according to memory row and column addresses. Therefore, two consecutive device physical addresses are not necessarily stored consecutively in the internal structural layout of the memory system. Converting the device physical address to memory row and column addresses is more helpful for the operating system kernel to perform subsequent memory error prediction.

[0124] In the above process of converting the device physical address into the memory row and column address, the error type can be used to guide the driver built into the error handling driver on how to parse and process the device physical address. For example, different error types can set different parsing logic or algorithms to ensure the accuracy of address conversion; auxiliary information can provide the error handling driver with detailed information on the occurrence of the memory error (such as error severity, error context, error count, timestamp of sending the error, etc.). The use of auxiliary information helps the error handling driver to more accurately locate the specific memory row and column address of the device physical address in the memory architecture, thereby enhancing the accuracy of the memory row and column address; the error handling driver performs address translation operations based on the error type and auxiliary information, combined with the system memory layout information, to convert the device physical address into a memory row and column address to ensure that the operating system kernel can process the memory row and column address, facilitate the operating system kernel to perform error analysis and prediction, and more accurately determine the specific location of the memory error.

[0125] In an optional application scenario, Figure 5 As shown, the address translation is performed using the error handling driver. Specifically, it is still as Figure 7As shown, the operating system kernel passes multiple data items of memory error information read from the target register to the error handling driver. Specifically, after the operating system kernel passes the error type, auxiliary information, and device physical address read from the target register to the error handling driver, the error handling driver can convert the device physical address into a memory row and column address based on the error type and auxiliary information, combined with system memory layout information, so that the operating system kernel can more accurately locate the memory error. After converting the device physical address to a memory row and column address, a write-back operation can also be performed on the fifth sub-register (ERRSTATUS). Specifically, the fifth sub-register (ERRSTATUS) can be restored to an initialized state to avoid affecting the next memory error information reading process.

[0126] In some application scenarios, the error handling driver described above can be the Error Detection and Correction (EDAC) driver within the operating system, which can be used to provide address translation functionality. Furthermore, the parsing results described above can include error bits, which can indicate the number of erroneous data bits detected in the memory space. These error bits can assist the operating system kernel in subsequent error prediction.

[0127] It is easy to understand that the above-mentioned optional embodiments of the present application can achieve the following technical effects: by utilizing the error handling driver in combination with the error type and auxiliary information, the kernel can directly convert the device address into the memory row and column address, so that the kernel can more accurately determine the specific location of the memory error, and by converting the device address into the memory row and column address, it is convenient for the operating system kernel to perform more accurate error analysis and prediction.

[0128] In an optional embodiment, the memory error reading mechanism is a polling reading mechanism. In step S43, the memory error reading mechanism is configured based on the instruction set type, including the following method steps:

[0129] Step S433: in response to the instruction set type being the second type, initializing enable bits of multiple machine check functions configured in the hardware module, wherein the instruction operation complexity corresponding to the second type of instruction set is higher than the instruction operation complexity corresponding to the first type of instruction set, and the multiple machine check functions are respectively used to detect multiple types of hardware errors corresponding to the hardware module;

[0130] Step S434 , in response to the completion of the enable bit initialization process, registering a polling reading program corresponding to the polling reading mechanism, wherein the polling reading program is used to respond to a read event triggered according to the polling cycle duration to read memory error information.

[0131] The second type mentioned above can be set to a complex instruction set (CISC) type. Compared to the first type of instruction set (RISC), the CISC instruction set has higher instruction operation complexity. When the instruction set type mentioned above is a complex instruction set type, the target computer architecture mentioned above can be the second architecture (AMD architecture). The enable bit mentioned above can refer to a binary bit used to turn on or off a specific function in the hardware module (specifically, a machine check function). In particular, when the enable bit is set to 1, the function corresponding to the enable bit is activated.

[0132] In an exemplary application scenario, Figure 8 As shown, when it is determined that the above-mentioned instruction set type is the second type, it is necessary to initialize the enable bits of the various machine check functions configured in the hardware module so that the machine check function is in an activated state to ensure that the hardware module can detect hardware errors as required and report the memory error information corresponding to the hardware error to the target register, so as to avoid the problem that the hardware module cannot detect hardware errors due to the inactivation of the machine check function, thereby ensuring that when the operating system kernel detects a read event triggered according to the polling cycle length, it drives the polling reading program registered in the operating system kernel to read the memory error information from the target register (that is, polling to read the memory error information).

[0133] Continuing with the aforementioned application scenario, the specific process for initializing the enable bits for the various machine check functions configured in the hardware module can be as follows. The CPU is configured with a machine check architecture register (MCG_CAP), which can be used to store machine check register information, the number of error reporting register sets, and other configuration information. The CPU's built-in instruction (Central Processing Unit Identification, CPUID) is used to check whether the CPU has the Machine Check Exception (MCE) function and whether the CPU is configured with a Machine Check Architecture (MCA). The CPUID instruction is used to check whether the CTLP bit in MCG_CAP is set to 1. If the CTLP bit in MCG_CAP is set to 1, it indicates that the CPU is configured with a machine check register (MCG_CTL). Therefore, the CPUID instruction is used to set all corresponding enable bits in the MCG_CTL register to 1, enabling the error reporting register set. Furthermore, the CPUID instruction is used to read the number information of the error reporting register groups stored in MCG_CAP (eg, the COUNT field stored in MCG_CAP) to determine the maximum number of error reporting register groups supported by the central processing unit that are enabled.

[0134] Still in the above application scenario, during the process of initializing the enable bits of the various machine check functions configured in the hardware module, for any error reporting register set, the CPUID instruction can be used to set which enable bits in the control register unit (denoted as MCi_CTL, where i = 0, 1, ..., n, i represents the i+1th error reporting register set of n+1 error reporting register sets, and n is an integer counting upwards from 0, i.e., at least one error reporting register set must exist) in the error reporting register set to be initialized, based on the various types of hardware errors corresponding to the hardware module to be detected. The CPUID instruction is then used to initialize the corresponding enable bits in the control register unit MCi_CTL. Specifically, all enable bits corresponding to the control register unit MCi_CTL can be set to 1 to activate all types of machine check functions, thereby improving the comprehensiveness of hardware error detection and thereby enhancing the reliability of the memory error handling method. Specifically, if the central processing unit is not configured with a machine check register, the enable bit initialization process for the various machine check functions configured in the hardware module is canceled.

[0135] Still in the above application scenario, before the enable bits of various machine check functions configured in the hardware module are initialized, the VAL data bits in the status register unit in the error reporting register group (denoted as MCi_STATUS, where i=0, 1,...,n, i represents the i+1th error reporting register group in the n+1 error reporting register groups, and n is an integer and counts upwards starting from 0) may have recorded error status information (for example, error status information corresponding to errors occurring before hot reset, error status information recorded during system power-on, error status information recorded during system startup, etc.). Therefore, when initializing the enable bits for the various machine check functions configured in the hardware module, it is also necessary to check the VAL data bit in the status register (MCi_STATUS) in the error reporting register group. Specifically, it is necessary to determine whether the status register MCi_STATUS contains a data value in MCi_STATUS[VAL]. If MCi_STATUS[VAL] contains a data value, the CPUID instruction is used to write all data bits of the status register MCi_STATUS to 0 to clear the status field in the status register MCi_STATUS. Furthermore, the CPUID instruction is used to set the MCE bit in the master control register (CR4), CR4[MCE], to 1 to enable the CPU's machine check exception function. Specifically, if a hardware error occurs while CR4[MCE] is set to 0, the CPU's processing core enters a shutdown state.

[0136] Still in the above application scenarios, in particular, Figure 9 As shown, the error reporting register group corresponding to the hardware module can be enabled and initialized to enable the various machine check functions configured in the hardware module. Figure 9 As shown, the target computer architecture may include hardware module 1, hardware module 2, and hardware module 3. For example, there is a correspondence between hardware module 1, error reporting register set 1, and polling read program 1. An enable bit initialization process is performed on error reporting register set 1 corresponding to hardware module 1. In response to completion of the enable bit initialization process for error reporting register set 1, polling read program 1 is registered. Similarly, a correspondence between hardware module 2, error reporting register set 2, and polling read program 2 can be determined, as can a correspondence between hardware module 3, error reporting register set 3, and polling read program 3.

[0137] The various machine check functions mentioned above may include, but are not limited to, correctable error checking, uncorrectable error checking, error interrupt checking, and register configuration checking. The various types of hardware errors mentioned above may include, but are not limited to, correctable memory errors, uncorrectable memory errors, register configuration errors, interrupt errors, cache errors, bus errors, clock errors, system power-on errors, and power-on errors.

[0138] Each of the above machine check functions can be used to detect a type of hardware error corresponding to a hardware module. It will be appreciated that by utilizing multiple machine check functions to detect multiple types of hardware errors corresponding to hardware modules, a more comprehensive error detection capability can be provided, ensuring that multiple types of hardware errors can be more comprehensively captured.

[0139] Since there is no interrupt mechanism in the AMD architecture, in order to ensure that the operating system kernel of the AMD architecture can directly read the memory error information from the target register, when the operating system kernel detects that the above-mentioned enable bit initialization processing is completed, the polling reading program corresponding to the polling reading mechanism is registered. The polling reading program can be used to respond to the read event triggered according to the polling cycle length to read the memory error information stored in the target register.

[0140] The polling cycle duration can be preset according to application requirements, for example, the polling cycle duration can be set to 5 minutes. In particular, the polling cycle duration can also be updated during the polling reading process to avoid excessively frequent polling operations that waste system resources.

[0141] By setting a timer in the operating system kernel and setting the timer countdown to the above-mentioned polling cycle length, when the timer countdown ends, the read event corresponding to the polling reading mechanism is automatically triggered, so that the polling reading program responds to the read event to read the memory error information, thereby providing the operating system kernel with a mechanism for regularly and actively reading memory error information. Even in the absence of an interrupt mechanism, the operating system kernel can directly read the memory error information from the target register without relying on the firmware to upload the memory error information.

[0142] In one application scenario, when the processor is reset, the enable bits corresponding to various machine check functions are set to a disabled state. Therefore, after resetting the processor, the enable bits of various machine check functions configured in the hardware module may also be initialized.

[0143] It is easy to understand that the above-mentioned optional embodiment of the present application can achieve the following technical effects: by registering a polling reading program, a mechanism is provided for the operating system kernel to periodically and proactively read memory error information. Even in the absence of an interrupt mechanism, the operating system kernel can directly read memory error information from the target register according to the polling cycle duration, without relying on firmware to upload memory error information. In addition, by initializing the enable bits of various machine check functions in the hardware module, various machine check functions in the hardware module can be activated, thereby ensuring that the hardware module can more comprehensively detect various types of hardware errors, thereby supporting the precise parsing and positioning of memory error information in subsequent processes.

[0144] In an optional embodiment, the above memory error handling method further includes the following method steps:

[0145] Step S45 : In response to the accumulated time reaching the polling cycle time, determining that a read event corresponding to the polling read mechanism is triggered, wherein the accumulated time is the accumulated time since the last triggering of the read event.

[0146] The cumulative duration can represent the time interval between the last read event and the current read event. The read event triggered in the previous polling cycle relative to the current polling cycle is the previous read event. The cumulative duration can be obtained using the timer function provided by the operating system kernel.

[0147] In an exemplary application scenario, Figure 8 As shown, a timer in the operating system kernel is used to count from the moment when the last read event was triggered to the current moment, and the difference between the moment when the last read event was triggered and the current moment is calculated, and the difference is determined as the cumulative duration. When the operating system kernel detects that the cumulative duration has reached the polling cycle duration, it automatically triggers the read event corresponding to the polling read mechanism. As a result, the operating system kernel can periodically read memory error information from the target register. This mechanism of periodically reading memory error information avoids the operating system kernel from frequently performing read operations on the target register, which can save system resources. In particular, the above-mentioned polling cycle duration can also be adjusted according to actual needs to enhance the flexibility of the memory error handling method.

[0148] It is easy to understand that the above-mentioned optional embodiments of the present application can achieve the following technical effects: when the operating system kernel detects that the cumulative time reaches the polling cycle length, it automatically triggers the read event corresponding to the polling read mechanism, which can support the operating system kernel to periodically read memory error information from the target register, avoid the operating system kernel frequently performing read actions on the target register, and reduce the waste of system resources.

[0149] In an optional embodiment, the target register includes an error reporting register group corresponding to the hardware module, the error reporting register group includes a plurality of machine check register units corresponding to a plurality of machine check functions, and the memory error information includes a device physical address and auxiliary information. In the above step S41, the memory error information is read from the target register, including the following method steps:

[0150] Step S412: driving the polling reading program to execute, so as to read the memory error status corresponding to the current polling cycle from the status register unit in the plurality of machine check register units;

[0151] Step S413, in response to determining that a new memory error has occurred in the current polling cycle based on the memory error status, read the device physical address corresponding to the memory error from the address register unit in the multiple machine check register units, and read auxiliary information corresponding to the memory error from the miscellaneous register unit in the multiple machine check register units.

[0152] The hardware module may refer to a processor core. The plurality of machine check register units may include a control register unit, a status register unit, an address register unit, and a miscellaneous register unit.

[0153] In an exemplary application scenario, Figure 10 As shown, the AMD architecture may include multiple central processing units, each central processing unit corresponds to a processor core. In particular, n+1 processors may be included, and these n+1 processors are sequentially denoted as CPU_0, CPU_1, ..., CPU_n, where CPU_0 represents the first processor, CPU_1 represents the second processor, and CPU_n represents the n+1th processor. Accordingly, the AMD architecture includes n+1 processor cores, and these n+1 processor cores are sequentially denoted as Core_0, Core_1, ..., Core_n, where Core_0 represents the processor core corresponding to the first processor, Core_1 represents the processor core corresponding to the second processor, and Core_n represents the processor core corresponding to the n+1th processor.

[0154] Continuing with the aforementioned application scenario, in the AMD architecture, each error reporting register set is typically associated with a specific execution unit. By configuring the error reporting register set, the hardware module can pass detected memory error information to the operating system kernel. Assuming that each processor core has n+1 sets of error reporting registers that support error reporting, these n+1 sets of error reporting registers are denoted as MC0, MC1, ..., MCn, in sequence. Specifically, for a certain processor model, there may be seven error reporting register sets, so n is 6, and these six sets of error reporting registers are denoted as MC0, MC1, ..., MC6, in sequence. MC0 can refer to a load-store unit, which can be used for data caching; MC1 can refer to an instruction fetch unit, which can be used for instruction caching; MC2 can refer to a combination unit; MC3 can refer to a retention unit; MC4 can refer to a northbridge unit; MC5 can refer to an execution unit, which can have mapping / scheduling / retirement / execution / fixed-issue reorder buffer functions; and MC6 can refer to a floating-point unit.

[0155] Still in the above application scenario, for any set of error reporting registers, the error reporting registers include multiple machine check registers corresponding to multiple machine check functions. The multiple machine check registers may include a control register (MCi_CTL), a status register (MCi_STATUS), an address register (MCi_ADDR), and a miscellaneous error information register (MCi_MISC), where i = 0, 1, ..., n, i represents the i+1th error reporting register in the n+1 error reporting registers. For example, for the first error reporting register MC0, the value of i is 0, the control register is denoted as MC0_CTL, the status register is denoted as MC0_STATUS, the address register is denoted as MC0_ADDR, and the miscellaneous error information register is denoted as MC0_M0SC.

[0156] The newly added memory error may refer to a memory error detected by the hardware module in the memory space during the current polling cycle. The memory error status may be used to indicate whether a newly added memory error has occurred during the current polling cycle.

[0157] In an exemplary application scenario, Figure 9 As shown in , when the operating system kernel detects a read event corresponding to the triggering polling read mechanism, the operating system kernel drives the polling read program to execute and read the memory error information from the target register. Specifically, Figure 11As shown, for the i+1th group of error reporting registers, when the operating system kernel detects a read event corresponding to the triggering polling read mechanism, the operating system kernel drives the polling read program to execute to read the memory error status corresponding to the current polling cycle from the status register unit MCi_STATUS, thereby determining whether a new memory error has occurred in the current polling cycle. In particular, if a new memory error has occurred in the current polling cycle, the VAL data bit in the status register unit MCi_STATUS is set to 1. In other words, the determination of whether a new memory error has occurred in the current polling cycle can be achieved by determining whether the value of MCi_STATUS[VAL] is 1. When the operating system kernel drives the polling read program to execute and the value of MCi_STATUS[VAL] read from the status register unit MCi_STATUS is 1, it is determined that a new memory error has occurred in the current polling cycle. When the operating system kernel drives the polling reading program to execute and the value of MCi_STATUS[VAL] read from the status register unit MCi_STATUS is not 1, it is determined that no new memory error occurs in the current polling cycle, and the polling reading program ends.

[0158] Still in the above application scenario, when the address register unit MCi_ADDR stores the device physical address corresponding to the memory error, the ADDRV data bit in the status register unit MCi_STATUS is set to 1. That is to say, it is possible to judge whether the address register unit MCi_ADDR stores the device physical address corresponding to the memory error by judging whether the value of MCi_STATUS[ADDRV] is 1. When the value of MCi_STATUS[ADDRV] is 1, it is considered that the address register unit MCi_ADDR stores the device physical address corresponding to the memory error.

[0159] Still in the above application scenario, when auxiliary information corresponding to a memory error is stored in the miscellaneous register unit MCi_MISC, the MISCV data bit in the status register unit MCi_STATUS is set to 1. That is, whether auxiliary information corresponding to a memory error is stored in the miscellaneous register unit MCi_MISC can be judged by judging whether the value of MCi_STATUS[MISCV] is 1. When the value of MCi_STATUS[MISCV] is 1, it is considered that auxiliary information corresponding to a memory error is stored in the miscellaneous register unit MCi_MISC.

[0160] Still in the above application scenario, when the operating system kernel determines that a new memory error has occurred in the current polling cycle based on the memory error status, the operating system kernel drives the polling reading program to read MCi_STATUS[ADDRV] from the status register MCi_STATUS. When the value of MCi_STATUS[ADDRV] is 1, the kernel reads the device physical address corresponding to the memory error from the address register MCi_ADDR in the multiple machine check registers and then determines whether the value of MCi_STATUS[MISCV] is 1. When the value of MCi_STATUS[ADDRV] is not 1, the kernel directly determines whether the value of MCi_STATUS[MISCV] is 1. When the value of MCi_STATUS[MISCV] is 1, the kernel reads auxiliary information corresponding to the memory error from the miscellaneous register MCi_MISC in the multiple machine check registers and passes the read memory error information to the address translation driver. When the value of MCi_STATUS[MISCV] is not 1, the kernel passes the read memory error information to the address translation driver.

[0161] In some application scenarios, the memory error information may also include an error occurrence frequency, which may be recorded using a timestamp counter in the operating system kernel.

[0162] It is easy to understand that the above-mentioned optional embodiments of the present application can achieve the following technical effects: the operating system kernel drives the polling reading program to preferentially read the memory error status in the status register unit, and can judge whether a new memory error occurs in the current polling cycle. By querying the error report register group, the kernel can determine the memory error status in each polling cycle. When the operating system kernel determines that a new memory error occurs in the current polling cycle based on the memory error status, the operating system kernel reads the device physical address and auxiliary information corresponding to the memory error from multiple machine check register units. Therefore, based on the above-mentioned reading mechanism, when no new memory error occurs in the current polling cycle, it can reduce the execution of redundant read operations on the machine check register unit and avoid unnecessary communication overhead.

[0163] In an optional embodiment, the above memory error handling method further includes the following method steps:

[0164] Step S46, in response to determining that a new memory error has occurred in the current polling cycle according to the memory error status, updating the polling cycle duration to the initial duration;

[0165] Step S47 , in response to determining according to the memory error status that no new memory error occurs in the current polling cycle and the polling cycle duration is less than a preset duration upper limit, the polling cycle duration is updated using a target increment algorithm.

[0166] The aforementioned initial duration may refer to the default polling cycle duration set when the polling read program is first executed. This initial duration may be determined by the configuration file corresponding to the polling read program. The aforementioned preset upper limit duration may refer to the maximum polling cycle time set in the memory error polling read mechanism. This preset upper limit duration may be set based on actual application requirements. By setting this preset upper limit duration, information loss due to excessively long polling cycles can be avoided.

[0167] The target incrementing algorithm may refer to a strategy for adjusting the polling cycle duration. The target incrementing algorithm may include, but is not limited to, an exponential incrementing algorithm, a linear incrementing algorithm, and an adaptive incrementing algorithm. If no new memory errors occur during the current polling cycle and the polling cycle duration is less than a preset upper limit, the target incrementing algorithm may be used to dynamically adjust the polling cycle duration by using the memory error status as feedback information on system performance, gradually extending the polling cycle duration to avoid wasting system resources due to overly frequent polling operations.

[0168] In an exemplary application scenario, Figure 8 As shown in , when the operating system kernel determines that a new memory error has occurred in the current polling cycle, the operating system kernel can pass the memory error information read from the target register to the address translation driver for address translation and update the polling cycle length. Specifically, the specific process of updating the polling cycle length is as follows Figure 12As shown, the above-mentioned initial duration is set to 1s, and the above-mentioned preset duration upper limit is set to 300s. The polling cycle duration is recorded as N. Assuming that the current polling cycle duration N is 300s, the value of N is 300. When the operating system kernel detects that the cumulative duration reaches the polling cycle duration, the operating system kernel drives the polling reading program to execute to read the memory error status corresponding to the current polling cycle from the status register unit. Furthermore, when the operating system kernel determines that a new memory error has occurred in the current polling cycle based on the memory error status, the polling cycle duration N is updated to 1s, and the value of N is updated to 1; when the operating system kernel determines that no new memory error has occurred in the current polling cycle based on the memory error status and the polling cycle duration is less than the preset duration upper limit, the target increment algorithm is used to update the polling cycle duration. In particular, the target increment algorithm adopts a linear increment algorithm. When the operating system kernel determines, based on the memory error status, that no new memory errors have occurred in the current polling cycle and the polling cycle duration is less than the preset upper limit, the value of the polling cycle duration multiplied by 2 is compared with the preset upper limit. When the value of the polling cycle duration multiplied by 2 is less than the preset upper limit, the polling cycle duration N is updated to 2N, that is, the value of N is updated to 2N. When the value of the polling cycle duration multiplied by 2 exceeds the preset upper limit, the polling cycle duration N is updated to the preset upper limit of 300s, that is, the value of N is updated to 300. Since memory errors are usually triggered in a centralized manner, by setting the above-mentioned dynamic polling cycle update mechanism, not only can the excessively frequent triggering of read operations that wastes system resources be avoided, but also the polling cycle duration can be updated to a smaller polling cycle duration when a new memory error is detected, thereby ensuring that more memory error information is collected to balance the completeness of memory error information collection and the system communication resource overhead.

[0169] In some application scenarios, when the operating system kernel determines based on the memory error status that no new memory errors have occurred in the current polling cycle and the polling cycle duration is equal to the preset upper limit, that is, when no new memory errors have occurred in the current polling cycle and N has been set to 300, there is no need to update the polling cycle duration (that is, keep N at 300) until the operating system kernel determines based on the memory error status that a new memory error has occurred in the current polling cycle, and then re-uses the above-mentioned dynamic update mechanism of the polling cycle duration to update the polling cycle duration.

[0170] It should be noted that for the AMD architecture, the characteristic that correctable memory errors are stored in the error reporting register group in the AMD architecture can be utilized to design a periodically triggered polling reading program. By updating the polling cycle length, the polling cycle length is made more consistent with the characteristics of centralized triggering of correctable errors, thereby improving the completeness of collecting correctable memory errors. Compared to the firmware-first solution in the prior art, the above-mentioned optional embodiment of the present application can increase the collection rate of correctable memory errors from 0.024% to over 50%.

[0171] It is easy to understand that the above optional embodiment of the present application can achieve the following technical effects: based on the above-mentioned mechanism of using the memory error status as feedback information to dynamically update the polling cycle length, it can not only avoid triggering the read operation too frequently and wasting system resources, but also update the polling cycle length to a shorter polling cycle length when a new memory error is detected, ensuring that more memory error information is collected to balance the completeness of memory error information collection and system communication resource overhead. In addition, by utilizing the target increment algorithm, when no new memory errors occur in the current polling cycle and the polling cycle length is less than the preset upper limit, the polling cycle length can be quickly increased to save system resources.

[0172] In an optional embodiment, the above memory error handling method further includes the following method steps:

[0173] Step S48 , in response to the completion of reading the memory error information from the target register, the corresponding register bit of the status register unit is initialized.

[0174] The register bits corresponding to the aforementioned status register units may be binary bits, which can be used to indicate the occurrence of a specific event or the satisfaction of a specific condition. These register bits may include, but are not limited to, error detection status register bits, address validity register bits, and auxiliary information validity register bits. When the operating system kernel reads memory error information from the target register, by reading the register bits corresponding to the status register units, it can quickly determine whether a specific event has occurred or whether a specific condition has been satisfied.

[0175] In an exemplary application scenario, Figure 11As shown, when the operating system kernel completes reading memory error information from the target register, the register bits corresponding to the status register unit are initialized and the values ​​stored in the register bits corresponding to the status register unit are cleared. This ensures that after each reading of memory error information from the target register by the operating system kernel, the register bits of the status register unit can be initialized, thereby avoiding the situation where an incorrect memory error status is read from the status register unit in the next training cycle due to residual information in the register bits corresponding to the status register unit. In particular, the register bits corresponding to the above-mentioned status register unit MCi_STATUS may include the error detection status register bit MCi_STATUS[VAL], the address valid register bit MCi_STATUS[ADDRV] corresponding to the physical address of the storage device in the address register unit, and the auxiliary information valid register bit MCi_STATUS[MISCV] corresponding to the auxiliary information stored in the miscellaneous register unit.

[0176] It is easy to understand that the above-mentioned optional embodiments of the present application can achieve the following technical effects: when the operating system kernel completes reading the memory error information from the target register, the register bits corresponding to the status register unit are initialized, ensuring that the register bits of the status register unit can be reset after each memory error information reading is completed, avoiding the situation where residual information in the register bits of the status register unit causes memory error status errors, ensuring that the operating system kernel can accurately read the memory error status in the next training cycle, thereby ensuring that the operating system kernel accurately reads the memory error information to accurately locate the error.

[0177] In an optional embodiment, the parsing result includes the system physical address of the hardware module where the memory error occurs in the target computer system. In the above step S42, the memory error information is parsed to obtain the parsing result, which includes the following method steps:

[0178] Step S422 : translating the device physical address into a system physical address using the auxiliary information and an address translation driver, wherein the address translation driver is configured based on a hardware space address mapping table corresponding to the target computer system.

[0179] The aforementioned address translation driver can be used to resolve device physical addresses. Specifically, the address translation driver can be built into the operating system kernel to provide address translation functionality. The aforementioned system physical address can be used to identify the physical memory address space managed globally by the system. The system physical address can be used to represent the exact storage location of data in actual physical memory.

[0180] The hardware space address mapping table can be used to describe the conversion relationship between device physical addresses and system physical addresses. The hardware space address mapping table can represent the actual location of memory error information within the system memory layout. The hardware space address mapping table can be generated during the construction of the target computer architecture and stored in the target computer system.

[0181] In particular, the AMD architecture has a data structure management unit (e.g., Data Fabric), which can be used to manage physical memory layout (i.e., system physical address). However, the hardware modules (e.g., memory modules, input / output interfaces, etc.) connected to the data structure management unit cannot directly use the hardware space address mapping table to convert the device physical address corresponding to the memory error information collected by the hardware module into the system physical address. Therefore, the address translation driver built into the operating system kernel can be used to convert the device physical address read from the target register using the polling reading program into the system physical address to ensure that the operating system kernel can perform subsequent processing.

[0182] During the aforementioned process of converting the device physical address to the system physical address, the auxiliary information can provide the address translation driver with detailed information about the memory error (e.g., error severity, error context, error count, error transmission timestamp, etc.), helping the address translation driver to more accurately translate the device physical address and improve the accuracy of the system physical address. The address translation driver, based on the hardware space address mapping table and in combination with the auxiliary information, converts the device physical address to a system physical address recognizable by the operating system kernel. This ensures that the operating system kernel can subsequently perform error analysis and prediction, and locate the specific location of the memory error. The specific process by which the address translation driver converts the device physical address to a system physical address recognizable by the operating system kernel, based on the hardware space address mapping table and in combination with the auxiliary information, can include: after the operating system kernel reads the device physical address and auxiliary information from the target register, the operating system kernel passes the device physical address and auxiliary information to the address translation driver. Furthermore, the address translation driver can utilize a built-in error parsing program (e.g., a program corresponding to a binary tree search algorithm) to locate the conversion relationship between the device physical address and the system physical address from the hardware space address mapping table. The address translation driver uses this conversion relationship to convert the device physical address to the system physical address and, based on the auxiliary information, adjusts the system physical address to improve the accuracy of the system physical address.

[0183] In an exemplary application scenario, Figure 8 As shown, the address translation driver is used to perform address translation. Specifically, it is still as follows Figure 9As shown, the operating system kernel passes the memory error information read from the target register (including the status register unit MCi_STATUS, the address register unit MCi_ADDR, the miscellaneous register unit MCi_MISC, and in particular, the control register unit MCi_CTL) to the address translation driver. In particular, after the operating system kernel passes the auxiliary information read from the target register and the device physical address to the address translation driver, the address translation driver can convert the device physical address into a system physical address that can be recognized by the operating system kernel based on the hardware space address mapping table and the auxiliary information, so that the operating system kernel can more accurately perform more precise error analysis and prediction of memory errors.

[0184] In some application scenarios, the above analysis results may also include error bits, which may refer to the number of erroneous data bits detected in the memory space. The error bits can be used to assist the operating system kernel in subsequent error prediction.

[0185] It is easy to understand that the above-mentioned optional embodiments of the present application can achieve the following technical effects: an address translation driver is obtained based on the hardware space address mapping table configuration corresponding to the target computer system; then, by utilizing auxiliary information and the address translation driver, the device physical address is accurately converted into a system physical address that can be recognized by the operating system kernel, so that the operating system kernel can more accurately locate the specific location of the memory error, which helps the operating system kernel to more accurately perform more precise error analysis and prediction of memory error information.

[0186] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0187] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0188] According to an embodiment of the present application, a device embodiment for implementing the above-mentioned memory error handling method is also provided. Figure 13 As shown, the memory error handling device includes:

[0189] The reading module 1301 is configured to: in response to a read event corresponding to the memory error reading mechanism, read memory error information from a target register, wherein the memory error information is detected by a hardware module corresponding to the target register during access to the memory space;

[0190] The parsing module 1302 is used to parse the memory error information to obtain a parsing result, wherein the parsing result is used to characterize the memory location corresponding to the hardware module where the memory error occurs in the target computer architecture.

[0191] Optionally, the memory error information is collected by the hardware module and reported to the target register when a memory error of the target type is detected in the memory space.

[0192] Optionally, the memory error handling apparatus further comprises a configuration module (not shown in the figure) for: configuring a memory error reading mechanism based on an instruction set type corresponding to a target computer architecture;

[0193] Optionally, the memory error reading mechanism is an interrupt response reading mechanism, and the above-mentioned configuration module is also used to: in response to the instruction set type being the first type, parse the interrupt mapping information and base address information corresponding to the target register from the error source table, wherein the error source table is reported to the operating system kernel by the target firmware when the target computer architecture is started; based on the interrupt mapping information and base address information, register the interrupt response program corresponding to the interrupt response reading mechanism, wherein the interrupt response program is used to respond to the read event triggered by the interrupt signal to read the memory error information.

[0194] Optionally, the above-mentioned memory error handling device also includes a first determination module (not shown in the figure) for: in response to receiving an interrupt signal routed by a processor in the target computer architecture, determining that a read event corresponding to an interrupt response reading mechanism is triggered, wherein the interrupt signal is generated by the target register according to the memory error information and sent to the processor.

[0195] Optionally, the hardware module is a memory component configured in the target computer architecture, and the target register corresponding to the memory component includes multiple sub-registers associated with the storage agent. The above-mentioned reading module 1301 is also used to: drive the execution of the interrupt response program to determine the storage agent based on the base address information, and read multiple data items in the memory error information from the multiple sub-registers, and the multiple data items include: the error type of the memory error, the device physical address corresponding to the memory error, and auxiliary information of the memory error.

[0196] Optionally, the parsing result includes memory row and column addresses, and the parsing module 1302 is further configured to convert the device physical address into a memory row and column address by using an error handling driver, an error type, and auxiliary information.

[0197] Optionally, the memory error reading mechanism is a polling reading mechanism, and the above-mentioned configuration module is also used to: in response to the instruction set type being the second type, perform enable bit initialization processing on multiple machine check functions configured in the hardware module, wherein the instruction operation complexity corresponding to the second type of instruction set is higher than the instruction operation complexity corresponding to the first type of instruction set, and the multiple machine check functions are respectively used to detect multiple types of hardware errors corresponding to the hardware module; in response to the execution of the enable bit initialization processing, register the polling reading program corresponding to the polling reading mechanism, wherein the polling reading program is used to respond to the read event triggered according to the polling cycle length to read memory error information.

[0198] Optionally, the above-mentioned memory error handling device also includes a second determination module (not shown in the figure) for: in response to the accumulated time reaching the polling cycle time, determining that a read event corresponding to the polling read mechanism is triggered, wherein the accumulated time is the accumulated time since the last triggering of the read event.

[0199] Optionally, the target register includes an error reporting register group corresponding to the hardware module, the error reporting register group includes multiple machine check register units corresponding to multiple machine check functions, the memory error information includes a device physical address and auxiliary information, and the above-mentioned reading module 1301 is also used to: drive the execution of the polling reading program to read the memory error status corresponding to the current polling cycle from the status register unit in the multiple machine check register units; in response to determining that a new memory error has occurred in the current polling cycle based on the memory error status, read the device physical address corresponding to the memory error from the address register unit in the multiple machine check register units, and read the auxiliary information corresponding to the memory error from the miscellaneous register unit in the multiple machine check register units.

[0200] Optionally, the above-mentioned memory error handling device also includes an update module (not shown in the figure) for: in response to determining that a new memory error has occurred in the current polling cycle according to the memory error status, updating the polling cycle duration to the initial duration; in response to determining that no new memory error has occurred in the current polling cycle according to the memory error status and the polling cycle duration is less than a preset duration upper limit, updating the polling cycle duration using a target increment algorithm.

[0201] Optionally, the memory error handling device further includes an initialization module (not shown in the figure) configured to initialize a register bit corresponding to the status register unit in response to completion of reading the memory error information from the target register.

[0202] Optionally, the analysis result includes the system physical address of the hardware module where the memory error occurs in the target computer system. The above-mentioned analysis module 1302 is also used to: use auxiliary information and an address translation driver to convert the device physical address into a system physical address, wherein the address translation driver is configured based on the hardware space address mapping table corresponding to the target computer system.

[0203] It should be noted here that the above-mentioned reading module 1301 and parsing module 1302 correspond to steps S41 to S42 in the embodiment, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned embodiment.

[0204] It should be noted that the above modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above modules may also be run in a computer terminal as part of a device.

[0205] According to an embodiment of the present application, an embodiment of a memory error handling system is also provided, which includes: a hardware module for detecting and obtaining memory error information during access to memory space; a target register for storing memory error information; an operating system kernel for responding to a read event corresponding to a memory error reading mechanism, reading the memory error information from the target register, and parsing the memory error information to obtain a parsing result, wherein the memory error reading mechanism is configured according to the instruction set type corresponding to the target computer architecture, and the parsing result is used to characterize the memory location corresponding to the hardware module where the memory error occurs in the target computer architecture.

[0206] According to an embodiment of the present application, an electronic device is also provided, including: a memory storing an executable program; and a processor for running the program, wherein any one of the aforementioned memory error handling methods is executed when the program is running.

[0207] Figure 14 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 14 As shown, the electronic device 140 may include: one or more (only one is shown in the figure) processors 142, a memory 144, a storage controller, and a peripheral interface.

[0208] The above-mentioned electronic device can be understood as an integrated intelligent terminal, including but not limited to a server, a desktop computer, a personal computer (PC), a model all-in-one machine, etc., and the electronic device can be pre-installed with the above-mentioned model in the above-mentioned embodiment of this application.

[0209] Specifically, the electronic device can pre-install multiple types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multimodal task processing, etc., so as to provide a variety of model choices. In different product forms, the electronic device can support one or more model usage methods, including but not limited to model training, model calling, model fine-tuning, model deployment, model reasoning and application, etc. In some product forms, the electronic device also supports model management, including but not limited to multi-type model management (supporting the management of multiple types of models such as discriminants and genesis), model version control (supporting the control of different model versions), model evaluation (based on model evaluation tools to evaluate the performance and effect of the model), etc. In other product forms, the electronic device can also create applications based on the model, provide application programming interface (API) calling capabilities, and can call the model into the created application through the API interface. At the same time, it provides application management tools to achieve management and monitoring of the application.

[0210] Furthermore, the electronic device can also include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn and master artificial intelligence (AI) technology), and basic management and control capabilities (providing enterprise-level basic management and control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive, integrated AI development, training, deployment and application device is provided.

[0211] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the memory error handling method and related devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the memory error handling method in the above embodiment. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories can be connected to the terminal A via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0212] The processor may call the executable program stored in the memory through the transmission device to execute any one of the above-mentioned memory error handling methods in the above-mentioned embodiments.

[0213] Those skilled in the art will understand that Figure 14 The structure shown is for illustration only, and the electronic device may also be a terminal device such as a smart phone (eg, an Android phone, an iOS phone, etc.), a tablet computer, a PDA, and a mobile Internet device (MID). Figure 14 The structure of the electronic device is not limited. For example, the electronic device 140 may also include Figure 14 More or fewer components (e.g., network interface, display device, etc.) shown in, or with Figure 14 Different configurations shown.

[0214] A person skilled in the art will understand that all or part of the steps in the various memory error handling methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, ROM, RAM, a magnetic disk or an optical disk, etc.

[0215] An embodiment of the present application further provides a computer-readable storage medium, which includes a stored executable program, wherein when the executable program runs, the device where the computer-readable storage medium is located is controlled to execute any one of the aforementioned memory error handling methods.

[0216] Optionally, in this embodiment, the above storage medium may be located in an electronic device.

[0217] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, and when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned memory error handling methods in the above-mentioned embodiments.

[0218] An embodiment of the present application further provides a computer program product, comprising a computer program, which implements any of the aforementioned memory error handling methods when executed by a processor.

[0219] Embodiments of the present application also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium that can be used to store a computer program that, when executed by a processor, implements the memory error handling method provided in the above embodiment.

[0220] The embodiment of the present application further provides a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, the memory error handling method provided in the above embodiment is implemented.

[0221] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0222] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0223] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0224] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0225] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0226] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: various media that can store program code, such as USB flash drives, ROM, RAM, mobile hard drives, magnetic disks, or optical disks.

[0227] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A memory error handling method, characterized in that: Applied to an operating system kernel of a target computer architecture, the memory error handling method includes: In response to a read event corresponding to a memory error reading mechanism, reading memory error information from a target register, wherein the memory error reading mechanism is configured according to an instruction set type corresponding to the target computer architecture, and the memory error information is detected by a hardware module corresponding to the target register during a process of accessing memory space; The memory error information is parsed to obtain a parsing result, wherein the parsing result is used to characterize a memory location corresponding to a hardware module where a memory error occurs in the target computer architecture.

2. The memory error handling method according to claim 1, wherein: The memory error information is collected by the hardware module and reported to the target register when a target type of memory error is detected in the memory space.

3. The memory error handling method according to claim 1 or 2, characterized in that: The memory error handling method further includes: The memory error reading mechanism is configured based on the instruction set type corresponding to the target computer architecture.

4. The memory error handling method according to claim 3, wherein: The memory error reading mechanism is an interrupt response reading mechanism. Based on the instruction set type, configuring the memory error reading mechanism includes: In response to the instruction set type being the first type, obtaining interrupt mapping information and base address information corresponding to the target register by parsing from an error source table, wherein the error source table is reported by target firmware to the operating system kernel when the target computer architecture is started; Based on the interrupt mapping information and the base address information, an interrupt response program corresponding to the interrupt response reading mechanism is registered, wherein the interrupt response program is used to respond to the read event triggered by the interrupt signal to read the memory error information.

5. The memory error handling method according to claim 4, characterized in that: The memory error handling method further includes: In response to receiving the interrupt signal routed by the processor in the target computer architecture, determining that the read event corresponding to the interrupt response read mechanism is triggered, wherein the interrupt signal is generated by the target register based on the memory error information and sent to the processor.

6. The memory error handling method according to claim 4, wherein: The hardware module is a memory component configured in the target computer architecture, the target register corresponding to the memory component includes a plurality of sub-registers associated with a register agent, and reading the memory error information from the target register includes: Drive the interrupt response program to execute to determine the storage agent according to the base address information, and read multiple data items in the memory error information from the multiple sub-registers, the multiple data items including: the error type of the memory error, the device physical address corresponding to the memory error, and auxiliary information of the memory error.

7. The memory error handling method according to claim 6, wherein: The parsing result includes memory row and column addresses, and the memory error information is parsed to obtain the parsing result including: The device physical address is converted into the memory row and column address using an error handling driver, the error type and the auxiliary information.

8. The memory error handling method according to claim 3, wherein: The memory error reading mechanism is a polling reading mechanism. Based on the instruction set type, configuring the memory error reading mechanism includes: In response to the instruction set type being the second type, performing enable bit initialization processing on multiple machine check functions configured in the hardware module, wherein the instruction operation complexity corresponding to the second type of instruction set is higher than the instruction operation complexity corresponding to the first type of instruction set, and the multiple machine check functions are respectively used to detect multiple types of hardware errors corresponding to the hardware module; In response to the completion of the enable bit initialization process, a polling reading program corresponding to the polling reading mechanism is registered, wherein the polling reading program is used to respond to the read event triggered according to the polling cycle length to read the memory error information.

9. The memory error handling method according to claim 8, characterized in that: The memory error handling method further includes: In response to the accumulated duration reaching the polling cycle duration, it is determined that the read event corresponding to the polling read mechanism is triggered, wherein the accumulated duration is the duration accumulated since the last triggering of the read event.

10. The memory error handling method according to claim 8, wherein: The target register includes an error reporting register group corresponding to the hardware module, the error reporting register group includes a plurality of machine check register units corresponding to the plurality of machine check functions, the memory error information includes a device physical address and auxiliary information, and reading the memory error information from the target register includes: driving the polling reading program to execute, so as to read the memory error status corresponding to the current polling cycle from the status register unit in the plurality of machine check register units; In response to determining, based on the memory error status, that a new memory error has occurred in the current polling cycle, the device physical address corresponding to the memory error is read from an address register unit among the multiple machine check register units, and the auxiliary information corresponding to the memory error is read from a miscellaneous register unit among the multiple machine check register units.

11. The memory error handling method according to claim 10, characterized in that: The memory error handling method further includes: In response to determining, according to the memory error status, that a new memory error occurs within the current polling cycle, updating the polling cycle duration to an initial duration; In response to determining, according to the memory error status, that no new memory error occurs in the current polling cycle and the polling cycle duration is less than a preset duration upper limit, the polling cycle duration is updated using a target increment algorithm.

12. The memory error handling method according to claim 10, characterized in that: The memory error handling method further includes: In response to the completion of reading the memory error information from the target register, a register bit corresponding to the status register unit is initialized.

13. The memory error handling method according to claim 10, wherein: The parsing result includes a system physical address of the hardware module where the memory error occurs in the target computer system. The memory error information is parsed to obtain the parsing result including: The device physical address is converted into the system physical address by using the auxiliary information and an address translation driver, wherein the address translation driver is configured based on a hardware space address mapping table corresponding to the target computer system.

14. A memory error handling system, characterized in that: include: A hardware module for detecting and obtaining memory error information during access to memory space; a target register, configured to store the memory error information; An operating system kernel is configured to read the memory error information from the target register in response to a read event corresponding to a memory error reading mechanism, and to parse the memory error information to obtain a parsing result, wherein the memory error reading mechanism is configured according to an instruction set type corresponding to a target computer architecture, and the parsing result is used to characterize a memory location corresponding to a hardware module in the target computer architecture where a memory error occurs.

15. An electronic device, characterized in that: include: a memory storing an executable program; A processor, configured to run the program, wherein the program executes the memory error handling method according to any one of claims 1 to 13 when running.

16. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the memory error handling method according to any one of claims 1 to 13.

17. A computer program product, characterized in that The invention comprises a computer program, which implements the memory error handling method according to any one of claims 1 to 13 when executed by a processor.

Citation Information

Patent Citations

  • Memory isolation method and device, electronic equipment and readable storage medium

    CN114780276A

  • Operating system memory fault processing method and system

    CN118689690A

  • Data processing method and device, storage medium and electronic equipment

    CN119396619A

  • Interrupt register processing method and device, product, equipment and medium

    CN119514478A

  • Memory isolation method and device, server, storage medium and program product

    CN120045365A

Cited By

  • Memory management unit fault diagnosis method and device, computer device and medium

    CN122363968A