Memory fault processing method and device, electronic equipment and readable storage medium

By isolating the target physical address in the event of a memory row failure, the system downtime problem caused by memory row failure cannot be quickly handled in the prior art, quickly respond and reduce the risk of downtime, and improve the stability and performance of the server.

CN120256185APending Publication Date: 2025-07-04INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510708177.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, memory line failures cannot be handled quickly, resulting in system downtime, and relying on manual intervention to respond slowly, which is susceptible to system load and isolation granularity.

Method used

When there is a memory row failure in the target system's memory, the target physical address is isolated by triggering an isolation operation, including obtaining the running status information of the memory, judging the fault row and isolating it, using BMC and BIOS to process it in a coordinated manner, combining machine learning algorithms and PECI protocol for real-time monitoring and early warning.

Benefits of technology

It effectively avoids performance degradation and data error accumulation caused by continued participation in data reading and writing, reduces the risk of system downtime, and improves server availability and overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256185A_ABST
    Figure CN120256185A_ABST
Patent Text Reader

Abstract

The invention discloses a memory fault processing method and device, electronic equipment and a readable storage medium, and relates to the technical field of computers.The method comprises the steps that when a memory line fault exists in a memory of a target system, an isolation operation is triggered for a target physical address, the target physical address at least comprises the physical address of the fault line with the memory line fault in the memory, the technical problem of system downtime caused by the fact that the memory line fault cannot be rapidly processed in the related technology is solved, the isolation process is started for the target physical address in time when the memory line fault is detected, and the system downtime is reduced. Performance reduction and data error accumulation caused by the fact that the fault line continues to participate in data reading and writing are avoided, and then the technical effect of reducing the risk of system downtime is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a method and apparatus for processing memory faults, an electronic device, and a readable storage medium. Background Art

[0002] In the field of computer technologies, especially in the operating environments of servers and data centers, as one of the core components, the stability of memory is directly related to the reliability of the entire system and business continuity. Memory row faults are likely to cause uncorrectable errors, which directly threaten the stable operation of the server. Therefore, achieving rapid identification and effective processing of memory row faults has become an urgent need for maintaining the high availability of servers. However, there are obvious defects in the current technical status in the industry. Although some servers are equipped with a fault warning mechanism, when dealing with memory row faults, they still rely on manual intervention, which not only has a slow response speed but is also easily affected by system load and isolation granularity, and cannot quickly respond when a fault first appears, easily causing uncorrectable errors and directly threatening the stable operation of the server. Summary of the Invention

[0003] This application provides a method and apparatus for processing memory faults, an electronic device, and a readable storage medium, so as to at least solve the problem in the related technologies that the system crashes because memory row faults cannot be quickly processed.

[0004] This application provides a method for processing memory faults, including: when there is a memory row fault in the memory of a target system, triggering an isolation operation on a target physical address, where the target physical address at least includes the physical address of the faulty row with the memory row fault in the memory.

[0005] This application further provides an apparatus for processing memory faults, including: a first triggering unit, configured to trigger an isolation operation on a target physical address when there is a memory row fault in the memory of a target system, where the target physical address at least includes the physical address of the faulty row with the memory row fault in the memory.

[0006] This application further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above-mentioned methods for processing memory faults when executing the computer program.

[0007] This application further provides a computer-readable storage medium, in which a computer program is stored, where the computer program, when executed by a processor, implements the steps of any of the above-mentioned methods for processing memory faults.

[0008] This application further provides a computer program product, including a computer program, where the computer program, when executed by a processor, implements the steps of any of the above-mentioned methods for processing memory faults.

[0009] Through the present application, when there is a memory line fault in the memory of the target system, the isolation process is started for the target physical address in a timely manner, avoiding performance degradation and data error accumulation caused by the faulty line continuing to participate in data reading and writing. Therefore, the technical problem in the related art that the memory line fault cannot be quickly processed, resulting in system downtime, can be solved, and the technical effect of reducing the risk of system downtime can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0011] Figure 1 is a hardware structure block diagram of memory fault handling provided according to an embodiment of the present application;

[0012] Figure 2 is the flow of a memory fault handling method provided according to an embodiment of the present application Figure 1 ;

[0013] Figure 3 is a schematic diagram of data acquisition provided according to an embodiment of the present application;

[0014] Figure 4 is a schematic diagram of memory fault handling provided according to an embodiment of the present application Figure 1 ;

[0015] Figure 5 is a schematic diagram of memory fault handling provided according to an embodiment of the present application Figure 2 ;

[0016] Figure 6 is the flow of a memory fault handling method provided according to an embodiment of the present application Figure 2 ;

[0017] Figure 7 is a schematic diagram of state transition provided according to an embodiment of the present application;

[0018] Figure 8 is the flow of a memory fault handling method provided according to an embodiment of the present application Figure 3 ;

[0019] Figure 9 is a schematic diagram of a memory fault handling device provided according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the protection scope of the present application.

[0021] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0022] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0023] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the memory fault handling method depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0024] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking the operation on a server device as an example, Figure 1 is the hardware structure block diagram of the execution of the system functions of the embodiments of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that, Figure 1 the structure shown is only for illustration and does not limit the structure of the above-mentioned server device. For example, the server device may further include more or fewer components than those shown in Figure 1 , or have a different configuration from that shown in Figure 1 .

[0025] The memory 104 can be used to store computer programs, for example, software programs of application software and modules, such as the computer program corresponding to the execution method of the system function in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0026] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0027] Embodiments of the present application provide a method for handling memory faults. The method is described in detail in combination with the execution process of the method for handling memory faults.

[0028] The following explains the professional terms appearing in the present application:

[0029] BIOS: Basic Input / Output System, the basic input / output system, is the first program to run when the computer starts up, responsible for initializing the hardware, detecting hardware devices, and loading the operating system. The BIOS is stored in the ROM (Read-Only Memory) chip on the motherboard, providing a basic interaction layer between the hardware and the operating system, ensuring that the computer system can start up and run smoothly.

[0030] BMC: Baseboard Management Controller, a dedicated microcontroller used to monitor and manage the hardware health and status in servers, workstations, and other computing platforms. The BMC can remotely monitor and control the server through the IPMI (Intelligent Platform Management Interface) protocol, including monitoring temperature, voltage, fan speed, power status, etc., and performing remote control operations such as power-on, power-off, restart, etc.

[0031] IPMI: Intelligent Platform Management Interface, an industry standard specification for remotely monitoring and managing system-level events, especially applicable to servers and other computing devices.

[0032] Redfish: "Redfish" protocol or "Redfish standard" or "Redfish technology", a standardized management interface for unified management of servers, storage, and network devices in modern data centers. It provides a RESTful (Representational State Transfer) API (Application Programming Interface) interface, making the management of hardware devices more flexible and efficient.

[0033] Memory row failure: Refers to a read / write error that occurs in the storage units of a specific row in a computer's memory.

[0034] In this embodiment, a method for handling memory failures is provided. Figure 2 It is the flow of the method for handling memory failures provided according to the embodiments of the present application. Figure 1 As Figure 2 shown, the method for handling memory failures includes:

[0035] Step S201, when there is a memory row failure in the memory of the target system, trigger an isolation operation on the target physical address, where the target physical address at least includes the physical address of the faulty row with a memory row failure in the memory.

[0036] Optionally, during the operation of the server system, to reduce the downtime probability and ensure business continuity, automatic diagnosis and isolation of equipment failures are required. As one of the key components of the server, memory row failures are likely to cause uncorrectable errors (UCE), leading to server downtime. Therefore, the diagnosis and isolation of memory row failures are crucial. In this application, if there is a memory row failure in the memory of the target system, the faulty row with the memory row failure in the memory will be located, the physical address of the faulty row will be determined, and the target physical address will be determined based on the physical address of the faulty row. Finally, an isolation operation will be triggered for the target physical address to prevent any application or system process from attempting to access the possibly corrupted data in that row, thereby preventing further data errors or system failures.

[0037] In an optional embodiment, the entire memory can be comprehensively tested through professional memory testing software to determine whether there is a memory row failure.

[0038] In an optional embodiment, a fault analysis algorithm based on machine learning can also be constructed to predict possible row-level errors in the memory by learning historical data.

[0039] In an optional embodiment, the physical address of the faulty row can be directly used as the above-mentioned target physical address, or the physical address of the memory page corresponding to the faulty row can be determined as the target physical address, or the physical address of the memory module corresponding to the faulty row can be determined as the target physical address.

[0040] By accurately determining the target physical address and triggering the isolation operation, the system can effectively prevent the spread of memory row-level errors, reduce the risk of system downtime caused by row-level errors, and improve the availability and overall performance of the server.

[0041] Optionally, in the memory fault handling method provided in the embodiments of this application, the method further includes: obtaining the operating status information of the memory; and determining whether there is a memory row failure based on the operating status information.

[0042] Optionally, the operating status information of the memory can be obtained through the BMC (Baseboard Management Controller). It should be noted that the operating status information can include the temperature and voltage data of the memory, and can also include detailed records of memory operations stored in the register, including information such as read and write times and access patterns.

[0043] In an optional embodiment, the BMC can actively poll the sensor data in the memory based on the PECI (Platform Environment Control Interface) protocol to obtain the above-mentioned operating status information.

[0044] After obtaining the above operating status information, the BMC can accurately determine whether there is a memory row fault based on this operating status information. For example, it can detect whether the memory temperature is within the normal range (for example, the normal temperature range is 0 - 85 degrees), detect whether the memory voltage is within the normal range (for example, the normal voltage range is 1.2V to 1.35V), can also detect the number of errors that occur in the memory, and can also identify the "hot spots" in the memory, that is, the areas that are frequently accessed, through access pattern analysis. If the rows in the hot spots have frequent error counts, it may indicate the existence of row-level faults, and further determine whether there is a memory row fault based on these detection results.

[0045] For example, it is possible to determine whether there is a memory row fault by combining the parameter information of the memory SPD (Serial Presence Detect) chip information.

[0046] For example, it is also possible to train a machine learning algorithm through historical operating status, and then use the trained machine learning algorithm to determine whether there is a memory row fault based on logs and real-time operating status data.

[0047] For example, the BMC can intelligently determine whether there is a row-level fault based on real-time status such as memory temperature, voltage, and memory access, in combination with historical data and the output of the machine learning model. For example, when the memory temperature continuously stays near the upper limit of the normal range (close to 85 degrees), even if it has not exceeded the threshold, it may indicate that the risk of row-level errors is accumulating; similarly, even if the voltage fluctuation is between 1.2V and 1.35V, but if the fluctuation is frequent, it may also have a negative impact on the row-level stability. Combining this real-time data, the BMC can give an early warning of potential row-level faults and take preventive measures instead of waiting until the fault occurs and then dealing with it.

[0048] Determining whether there is a memory row fault through the operating status information provides a more comprehensive and intelligent memory protection mechanism for the server, improving the stability of the system.

[0049] Optionally, in the memory fault handling method provided in the embodiments of the present application, obtaining the operating status information of the memory includes: obtaining the temperature information during the operation of the memory through a temperature sensor in the memory; obtaining the voltage information during the operation of the memory through a voltage sensor in the memory; obtaining the memory access information through a register in the memory; and obtaining the operating status information based on the temperature information, voltage information, and memory access information.

[0050] In an optional embodiment, during the operation cycle of the server (i.e., the aforementioned target system), the temperature sensor and voltage sensor built into the memory hardware continuously monitor the temperature and voltage status of the memory module. The temperature sensor ensures that the memory is within the recommended temperature range, avoiding performance degradation or row-level failures caused by overheating; the voltage sensor monitors the stability of the memory power supply in real time, preventing data errors or component damage caused by voltage fluctuations. The real-time acquisition of these physical status data provides a key basis for the preliminary judgment of the memory health status. The BMC can actively poll the sensor data in the memory based on the PECI protocol to obtain the above temperature information and voltage information.

[0051] In addition to the temperature and voltage information, the registers in the memory also record detailed memory access information, including but not limited to access frequency, access mode, data read / write speed, etc. Therefore, the BMC can also obtain the memory access information recorded in the registers of the memory. Finally, the temperature information, voltage information, and memory access information are determined as the above operation status information.

[0052] In an optional embodiment, at the beginning of system startup, the BIOS (Basic Input / Output System) reads the memory SPD chip information through the SMBus (System Management Bus) / I2C (Inter-Integrated Circuit Bus) bus for capacity and timing parameter initialization. During the operation phase, the BMC shares information with the BIOS to obtain configuration parameters and other information, and actively queries the temperature and voltage sensors through the PECI protocol. Combining the configuration parameters passed by the BIOS, it realizes the multi-source data fusion acquisition of the memory health status and obtains the above operation status information of the memory.

[0053] In an optional embodiment, it can also be through such as Figure 3The data acquisition schematic diagram shown realizes multi-source data acquisition, specifically including: during the system startup phase, the BIOS reads the information of the memory SPD (Serial Presence Detect) chip in the server through the SMBus / I2C bus to complete the initialization of basic parameters such as capacity and timing; after entering the running phase, the BIOS synchronizes real-time parameters with the BMC. SMI (System Management Interrupt) is a special interrupt mechanism mainly used for system management operations in a computer system. In the communication between the BIOS (Basic Input / Output System) and the BMC (Baseboard Management Controller), SMI plays a key role. When a system management event occurs, such as temperature overrun, power state change, or other hardware monitoring events, the hardware will trigger SMI, causing the processor to enter the System Management Mode (SMM). At the same time, the BMC actively polls sensor data based on the PECI protocol, and combines the real-time configuration parameters obtained from the BIOS to realize the data acquisition of the memory health status through multi-source data fusion.

[0054] Through the acquisition and analysis of diversified status information, the system can more precisely warn of potential row-level errors. Different from the simple threshold alarm of a single indicator, it can predict the probability of row-level faults based on the comprehensive performance of the memory's temperature, voltage, and access information, so as to take more targeted preventive measures.

[0055] Optionally, in the memory fault handling method provided in the embodiment of the present application, before triggering the isolation operation on the target physical address, the method further includes: obtaining a first log file in the target system, and extracting first logical address information with a memory row fault from the first log file; determining first physical address information corresponding to the first logical address information; in the case where the isolation policy is a page isolation policy, determining the physical address of the target memory page corresponding to the first physical address information as the target physical address.

[0056] In an optional embodiment, before triggering the isolation operation on the target physical address, the determination of the target physical address can be implemented by the following steps: during the operation of the server, any memory error will be recorded in the log. The first log file refers to the log file recording the memory row fault information, which can be the BMC log. Therefore, the first logical address information of the faulty row memory can be extracted from the first log file. The logical address is the abstract address used by the operating system or application program when accessing the memory.

[0057] After determining the first logical address, the logical address can be mapped to a physical address through the address translation function provided by the memory controller or BIOS, so as to determine the exact location of the fault, that is, to obtain the above-mentioned first physical address information. For example, mapping to a physical cell, row, and bank through the memory controller. This step is similar to accurately positioning the fault point on a map to clarify which physical area of the memory the error occurs in.

[0058] After determining the above-mentioned first physical address, determine whether the current isolation policy is a page isolation policy. If the isolation policy is a page isolation policy, determine the target memory page corresponding to the first physical address, and determine the physical address corresponding to the target memory page as the target physical address. Page Offline is a common strategy in fault isolation. When a row-level fault is detected in a certain memory page, the system will remove the page from the available memory resources to prevent it from being further used.

[0059] Through the above steps, memory faults can be effectively handled, avoiding performance degradation or system instability caused by row-level errors, and at the same time protecting data from being damaged or lost.

[0060] Optionally, in the memory fault handling method provided in the embodiments of the present application, triggering an isolation operation on the target physical address includes: isolating the target memory page; prohibiting the use of the memory area corresponding to the target memory page.

[0061] In an optional embodiment, triggering an isolation operation on the target physical address under the page isolation policy includes: implementing isolation error reporting for the target memory page through the operating system OS. First, the operating system receives the fault information transmitted from the BMC. The fault information at least includes the information of the target memory page to be isolated. The operating system will update the page table (the page table is a mapping table that converts virtual addresses to physical addresses) and mark the target memory page as unavailable or inaccessible.

[0062] In addition to isolating the specific fault page, it is also necessary to further ensure that the entire memory area where the target page is located is no longer used by the system to prevent potential fault spread. Therefore, it is necessary to prohibit the use of the memory area corresponding to the target memory page. For example, when attempting to access the memory of this page, the operating system will intercept these requests to prevent data from being incorrectly written or read.

[0063] Through the isolation of the faulty memory page and the disabling of the area where it is located coordinated by the operating system, the fault signal can be quickly responded to, the faulty area can be immediately isolated, and more serious system problems caused by fault spread can be avoided.

[0064] Optionally, in the memory fault handling method provided in the embodiments of the present application, after isolating the target memory page, the method further includes: determining second physical address information; migrating the service data located in the target memory page to the second physical address information.

[0065] In an optional embodiment, after a memory row fault is detected and isolated, the operating system OS checks the currently available healthy memory area to obtain the above-mentioned second physical address information. Copy the important service data in the faulty memory page to a temporary buffer to ensure the integrity and security of the original data during the migration process. Then, write the data from the temporary buffer to the memory area identified by the second physical address information. This process needs to ensure the accurate transmission of data and avoid introducing new errors during the migration process. Further, the operating system updates the page table to remap the virtual address of the original faulty memory page to the second physical address information, ensuring that the application or system process can automatically switch to the new, healthy memory area when accessing the same virtual address. After the migration is completed, the operating system can perform a data consistency check to verify whether the data in the new address is exactly the same as before the fault, thereby ensuring business continuity.

[0066] By querying and allocating resources of the healthy memory, the second physical address information is determined, and then the affected service data is safely migrated to this new address. At the same time, the page table in the operating system is updated to ensure that all applications and processes can access the correct and healthy memory area. This series of operations not only avoids the impact of the faulty memory page on the business, but also ensures the security of the data and the stable operation of the system, enhancing the fault tolerance and business continuity of the server.

[0067] In an optional embodiment, page isolation operations can be implemented through the Figure 4 schematic diagram shown as follows. Specifically, in the system startup phase, the BIOS reads the memory SPD chip information to complete the initialization of basic parameters such as capacity and timing. After entering the running phase, the BIOS synchronizes real-time parameters with the BMC. At the same time, the BMC actively polls sensor data based on the PECI protocol, combines the real-time configuration parameters obtained from the BIOS, and realizes the data collection of the memory health status through multi-source data fusion. According to the collected status data, it is determined whether there is a memory row fault. When a memory row fault is determined and the page isolation strategy is sampled, the BMC sends a fault message to the OS, and the OS updates the page table (the page table is a mapping table that converts virtual addresses to physical addresses), marks the target memory page as unavailable or inaccessible, and then the OS checks the currently available healthy memory area and migrates the data of the target memory page to this healthy memory area to achieve the isolation of the faulty row.

[0068] Optionally, in the memory fault handling method provided by the embodiments of the present application, after triggering the isolation operation on the target physical address, the method further includes: obtaining the target register snapshot information corresponding to the memory line fault; obtaining the performance data information of the memory at the time of the memory line fault; obtaining the memory operation information during the execution of the isolation operation; and generating a second log file based on the target register snapshot information, the performance data information, and the memory operation information.

[0069] In an alternative embodiment, when a memory line fault is detected, the snapshot information of the target register is obtained. The target register refers to the hardware register related to memory error detection and reporting. The register snapshot contains the status information in the register when the error occurs, such as error type, error frequency, error address, etc. In addition to the register information, the performance data of the memory before and after the error occurrence can also be obtained, for example, memory temperature, voltage, memory bandwidth utilization, memory access latency, etc. And all relevant memory operation information recorded during the execution of the isolation operation can also be obtained, including but not limited to the isolated physical address range, the switched new physical address, etc. Finally, a second log file is generated through the above data information.

[0070] The second log file not only provides comprehensive fault handling details for the operation and maintenance personnel, but also facilitates the fault analysis team to deeply understand the cause of the fault, the handling effect, and the potential improvement direction, so as to continuously optimize the memory management and fault recovery strategy of the server.

[0071] Optionally, in the memory fault handling method provided by the embodiments of the present application, after generating the second log file based on the target register snapshot information, the performance data information, and the memory operation information, the method further includes: analyzing the memory based on the second log file to obtain an analysis result; and determining the maintenance strategy for the memory according to the analysis result.

[0072] In an alternative embodiment, the second log file contains the whole process information of memory fault handling, including but not limited to the snapshot of the target register, system performance data, and isolation operation information. Therefore, the memory can be analyzed based on the second log file. For example, through the target register snapshot, the type and pattern of the error can be identified, such as whether it is a sudden error or a cumulative fault, which helps to judge whether the fault is temporary or persistent. Using the performance data information, evaluate the specific impact of the memory line fault on the server performance, including the changes of key indicators such as memory utilization, CPU load, network latency, etc., so as to quantify the necessity and effectiveness of fault isolation. Analyze the memory operation information to determine the isolation effect, memory usage, etc., and then obtain the above analysis result.

[0073] For example, the analysis results may include: Fault feature description: Describe in detail the type of the fault, such as Single-bit Error, Multi-bit Error, Recurring Error, etc., as well as the specific memory row and location where the fault occurs; Fault frequency and pattern: Analyze the frequency of the fault occurrence to determine whether it is an occasional noise error or a persistent hardware fault. At the same time, identify the fault pattern, such as whether it is a local fault (specific row or column) or a global fault (affecting the entire memory module); Performance impact assessment: Quantify the impact of the fault on system performance, including specific values such as decreased memory bandwidth, increased CPU load, and extended response time; Isolation operation effectiveness: Evaluate the actual effect of the isolation operation, including isolation time, resource consumption, and fault isolation success rate; Fault prediction and trend: Based on historical fault data, predict the possible future fault trends to provide a basis for preventive maintenance; Health score: Quantify the health score of the memory to help users or maintenance personnel determine whether to replace the memory module or take other actions.

[0074] After obtaining the analysis results, the maintenance strategy for the memory can be determined according to the analysis results. For example, Short-term emergency measures: If the fault has a small impact and is temporary, the spare resource pool can be temporarily used, such as the reserved healthy memory rows, to ensure the normal operation of the business. Long-term repair solutions: For continuous or severe memory faults, it is recommended to replace the faulty memory module or conduct deeper hardware inspections and upgrades. System optimization and prevention: By analyzing the fault pattern and performance data, potential weaknesses in the system design can be identified, such as optimizing the memory management algorithm, reasonably configuring redundant resources, and adjusting the fault warning threshold, to prevent the recurrence of similar faults. Software strategy adjustment: Based on the software operation feedback during fault handling, such as the resource management efficiency of the OS kernel, it is possible to consider optimizing the operating system configuration or the software process for memory fault handling to improve the real-time performance and accuracy of fault isolation.

[0075] For example, if the analysis results show that multiple rows of a certain memory module frequently have faults and the performance impact assessment shows that these faults have significantly degraded the system performance, the maintenance strategy may include immediately replacing the memory module and checking and adjusting the usage environment of the memory module (such as temperature, voltage) to prevent the recurrence of similar faults. On the other hand, if the analysis results indicate that the fault in a specific row is caused by improper software operation, the maintenance strategy may focus on software optimization, such as updating the memory access strategy of the application program or optimizing the BIOS configuration to reduce the number of accesses to sensitive memory rows.

[0076] Through a detailed analysis of the second log file, comprehensive and intelligent maintenance management can be provided for the server, which not only solves the current fault problems but also proactively prevents possible future faults, enhancing the overall reliability and operation and maintenance efficiency of the system.

[0077] Optionally, in the memory fault handling method provided in the embodiments of the present application, the method further includes: setting the memory row fault isolation mechanism in the baseboard management controller to the trigger state; setting the configuration items of the isolation policy of the basic input / output system through the baseboard management controller, and sending the configured signal to the basic input / output system through the target protocol.

[0078] In an optional embodiment, first, a new tab is added in the BMC for the memory row fault isolation mechanism. When this tab is set to the Enable state (i.e., setting the memory row fault isolation mechanism in the baseboard management controller to the trigger state as described above), it indicates that hierarchical isolation processing of memory row faults will be enabled.

[0079] When the memory row fault isolation mechanism is in the Enable state, the Configure file of the BIOS will be further modified, that is, the configuration items of the isolation policy of the basic input / output system will be set. For example, set the "Memory Faults Isolation" option in the Configure file of the BIOS option to the Enable state and configure the isolation policy. For example, change both RuntimePPR (Runtime Post Package Repair, runtime isolation policy) and Page Offline (page isolation policy) to the Enable state. It should be noted that only Runtime PPR or only Page Offline can be set to the Enable state. After the setting is completed, the BMC sends the configured signal to the BIOS through Redfish (i.e., the above-mentioned target protocol), and it will take effect after restart.

[0080] In an optional embodiment, under the OS, it can be sent to the BMC through IPMI commands, set the "memory row fault isolation mechanism" option to Enable, then the BMC modifies the Configure and sends it to the BIOS through redfish, and finally it will take effect after restart.

[0081] By adding specific memory row fault isolation settings in the BMC and collaborating with the BIOS through the Redfish protocol, flexible configuration of the isolation policy is achieved.

[0082] Optionally, in the memory fault handling method provided by the embodiments of this application, triggering an isolation operation on a target physical address includes: when the isolation policy is a runtime isolation policy, determining the physical address corresponding to the faulty row as the target physical address, and determining whether there is a redundant row in the memory; if there is a redundant row in the memory, marking the target physical address as unavailable, and migrating the service data in the faulty row to the redundant row.

[0083] In an alternative embodiment, if the runtime isolation policy (Runtime PPR) is configured, the physical address of the faulty row is directly determined as the target physical address. The BMC sends a fault message to the BIOS, and the BIOS checks whether the memory has sufficient redundant row resources. The redundant rows are a part of additional redundant rows pre-allocated during manufacturing. If there are indeed available redundant rows in the memory, the BIOS marks the physical address corresponding to the faulty row as unavailable, preventing any further system read and write access. Then, the service data in the faulty row is safely migrated to the redundant row, and the physical address of the faulty row is redirected to the physical address of the redundant row, ensuring that the application or system process can access the correct location without being affected by the faulty row.

[0084] Through Runtime PPR, the physical address remapping of the faulty row can be completed in real time during system operation without relying on the intervention of the operating system or system restart, minimizing service interruption to the greatest extent.

[0085] Optionally, in the memory fault handling method provided by the embodiments of this application, after determining whether there is a redundant row in the memory, the method further includes: when there is no redundant row in the memory and the basic input / output system has configured a page isolation policy, performing an isolation operation on the target physical address according to the page isolation policy.

[0086] In an alternative embodiment, if all redundant rows have been occupied or the number of faulty rows exceeds the coverage of the redundant rows, then the runtime isolation policy will no longer be able to repair with redundant rows, and it is necessary to further determine whether the page isolation policy is configured. If the BIOS has been configured to enable the page isolation policy, then the system will automatically trigger this policy. Page isolation is usually implemented at the operating system kernel level, and it uses a data structure called "Dirty Page" to track and isolate faulty memory. That is, the OS marks the page where the faulty memory row is located as unavailable and updates the page table structure to ensure that the application program points to a new and healthy memory page. At the same time, the BIOS and BMC will cooperate to update the memory mapping information to ensure system-level consistency and integrity.

[0087] Through data migration and page remapping at the operating system level, the page isolation strategy minimizes the impact of faults on business operations. Although it may cause a certain performance degradation, it generally ensures the continuity of system services.

[0088] In an optional embodiment, the row error isolation operation can be implemented through a schematic diagram as Figure 5 shown. Specifically, in the case where both the Runtime PPR and Page Offline isolation strategies are configured, the BMC determines whether a memory row fault is triggered. If there is a memory row fault, a fault error message is sent to the BIOS to perform the error isolation operation according to the Runtime PPR. If the BIOS isolation fails (for example, the redundant rows have been exhausted), the BMC sends the fault error message to the OS, and the Page Offline is executed by the OS to achieve row error isolation.

[0089] In an optional embodiment, the row error isolation operation can be implemented through a flowchart as Figure 6 shown. Specifically, a hierarchical isolation mechanism is constructed to optimize the isolation efficiency and reliability. The preferred solution is the Runtime PPR, which performs real-time physical address remapping of the faulty row during system operation through the memory controller, without relying on the intervention of the operating system or system restart, minimizing business interruption to the greatest extent. The secondary solution is page isolation (Page Offline). When the spare redundant row resources of the Runtime PPR are exhausted, a firmware-level soft isolation mechanism is automatically triggered to mark the faulty memory row as permanently unavailable and remove it from the system's available address space to ensure complete isolation of the faulty area. As Figure 6 shown, if it is determined to be a memory row fault, it is judged whether the spare redundant row resources are available. If available, the Runtime PPR is executed. If not available, the Page Offline is executed, and the resource status is updated in the case of successful isolation, and it continues to monitor whether there is a memory row fault. This hybrid solution combines high timeliness (fast response of the Runtime PPR), resource efficiency (seamless switching to soft isolation when the hard isolation resources are insufficient to avoid system downtime caused by resource exhaustion), and business non-damage. Through the deep cooperation between the hardware redundant resources and the operating system memory management module, it can improve the continuous operation ability and fault tolerance level of the critical business system.

[0090] Optionally, in the memory fault handling method provided in the embodiments of the present application, after determining whether there is a redundant row in the memory, the method further includes: when there is no redundant row in the memory and the basic input / output system is not configured with a page isolation strategy, triggering a first signal to the warning platform to instruct the target object to handle the memory row fault.

[0091] In an optional embodiment, when checking the memory resources, it is found that there are no available redundant rows, and the BIOS is not configured with a page isolation policy either. This means that the system cannot automatically isolate and repair faults through hardware or firmware-level means. Therefore, a first signal is sent to the warning platform through the BMC (Baseboard Management Controller). This signal contains detailed information about the faulty memory row, such as the physical address, error type, frequency of occurrence of the error, etc., as well as the current resource status and configuration information of the system. The warning platform is a system specifically used to monitor and manage the health status of the server. It can receive signals from the BMC and analyze and make decisions based on this information. The purpose of the first signal is to indicate that the memory row at the target physical address has a fault and the current system cannot automatically repair it, and the target object (i.e., the operation and maintenance personnel) needs to intervene in the memory row fault. It should be noted that in the case of failed fault isolation, the first signal will also be triggered to the warning platform to prompt the operation and maintenance personnel to handle the error in a timely manner.

[0092] Through real-time warning, the operation and maintenance team can quickly learn about the fault situation, accelerate the fault handling process, avoid the degradation of system performance or downtime caused by the failure to handle the fault in a timely manner, and thus reduce business losses and maintenance costs.

[0093] Optionally, in the memory fault handling method provided in the embodiment of the present application, determining whether there is a memory row fault based on the running state information includes: determining the target number of correctable errors that occur in each memory row in the memory during the current detection period according to the running state information; and determining whether there is a memory row fault based on the target number of times information.

[0094] In an optional embodiment, the BMC, as the core management unit, monitors and collects the running state data from the memory in real time, including but not limited to multi-dimensional information such as temperature, voltage, error logs, etc., and extracts the number of correctable errors (CE) that occur in each memory row during the current detection period, that is, the target number of times information. It should be noted that the detection period can be a fixed time interval or dynamically adjusted according to the system load. For example, 1s.

[0095] After obtaining the target number of times information, determine whether there is a memory row fault according to the target number of times information. For example, if the target number of times information of a certain memory row exceeds the set threshold, it is considered that the memory row has a row fault.

[0096] By counting the number of memory row faults, high-risk memory rows can be effectively screened out, ensuring that resources are reasonably allocated to the fault areas that really need attention and handling, and improving the accuracy and efficiency of fault detection.

[0097] Optionally, in the memory fault handling method provided in the embodiments of the present application, determining whether there is a memory row fault based on the target number of times information includes: obtaining the historical number of times information of correctable errors that occurred in each memory row during a historical time period; calculating the ratio between the historical number of times information and the target number of times information; determining the magnitude relationship between the ratio and a preset threshold to determine whether there is a memory row fault. Determining the magnitude relationship between the ratio and the preset threshold to determine whether there is a memory row fault includes: if the ratio is greater than or equal to the preset threshold, it is determined that a memory row fault has occurred; if the ratio is less than the preset threshold, it is determined that no memory row fault has occurred.

[0098] In an alternative embodiment, to improve the accuracy of determining memory row faults, it may further include: First, the BMC collects and collates the historical error counts of each memory row over a past period of time, which are typically data collected within a fixed or dynamically adjusted detection cycle. For the error count of each row, calculate the ratio between it and the total error count during the same period, that is, the row error density. For example, assume the error count of a certain row is 48 times, and the total error count of this row during the historical period is 50 times, then the error density of this row is 96%.

[0099] After calculating the error densities of all rows, the BMC will compare these error densities with a preset fault determination threshold one by one. The preset threshold can be set according to factors such as the type of server, operating load, error tolerance, etc., and is generally defaulted to 80%. In some high-load scenarios, to improve the sensitivity of fault detection, the threshold can be set at 70%. If the error density of a certain row reaches or exceeds the preset threshold, this indicates that the error occurrence frequency of this row is extremely high and may be a sign of a row-level fault. In this case, the BMC will mark this row as a faulty row and trigger a fault isolation process to prevent this row from causing more serious uncorrectable errors (UCE) and affecting the stability of the server and the integrity of the data. On the contrary, if the row error density is lower than the preset threshold, this indicates that the error occurrence frequency of this row is within a controllable range and does not constitute a row-level fault. Therefore, the BMC will consider this row to be in a normal working state and does not need to immediately take isolation or other fault handling measures.

[0100] By introducing the mechanism of comparing the row error density with the threshold, row-level faults can be identified at an early stage of the fault, and preventive measures can be taken in a timely manner to avoid the further expansion of the fault.

[0101] Optionally, in the memory fault handling method provided in the embodiments of the present application, after triggering the isolation operation on the target physical address, the method further includes: generating a third log file according to the timestamp information when the memory row fault occurs, the physical address information corresponding to the faulty row, and the remaining redundant row number corresponding to the memory; generating a fourth log file according to the target register snapshot information, performance data information when the memory row fault occurs, and memory operation information during the execution of the isolation operation, where the third log file and the fourth log file are used to evaluate the lifespan of the memory.

[0102] In an optional embodiment, when the BMC detects a memory fault and performs the isolation operation, it will automatically generate a maintenance log file, that is, the third log file. The third log file contains important information such as the timestamp at the fault moment, the specific physical address where the fault occurs, and the parameters when performing Runtime PPR (runtime repair) or Page Offline. These data are stored in a dedicated partition of NVRAM (Non-Volatile Random Access Memory) in a structured JSON format, facilitating real-time access and parsing by the early warning platform. The functions of the maintenance log include fault operation record: providing detailed information about the isolation operation for subsequent analysis, including when, where, and how the operation was performed. Resource status monitoring: recording the remaining redundant row number of the memory for evaluating future fault handling capabilities.

[0103] The BMC will also generate a one-key log, that is, the fourth log file, which contains the register snapshot at the time of fault trigger, performance monitoring data, and memory mapping information after the isolation operation. These information are integrated and compressed into a.tar file through the BMC Web interface. The functions of the maintenance log include retaining the hardware state at the time of fault through the register snapshot, facilitating backtracking and analysis. The memory mapping table after isolation can confirm whether the faulty row has been correctly isolated and the impact of the isolation operation on system resource allocation.

[0104] The two types of log information can be transmitted to the early warning platform through the Redfish interface, and the early warning platform uses these detailed data for in-depth analysis. Through the maintenance log, the early warning platform can understand the immediate handling situation of the fault and the resource status; while the one-key log provides a comprehensive perspective on the fault context and long-term performance trends. The early warning platform can perform cross-analysis based on these two types of logs to evaluate the remaining lifespan of the memory.

[0105] Through the dual guarantee mechanism of the maintenance log and the one-key log, each fault isolation operation can be clearly traced, and the detailed reasons and handling processes of the faults can be understood. Based on the lifespan assessment through log analysis, the operation and maintenance team can plan hardware replacement in advance, reducing the need for emergency maintenance and ensuring the stable operation of the system and business continuity.

[0106] In the traditional solution, after the faulty line is isolated, system events will be misreported as "normal". To solve the above problems, in the memory fault handling method provided in the embodiments of the present application, the method further includes: when a memory line fault is detected, the global state of the target system is changed from the first state to the second state; triggering the system event log to record the state change information of the target system, wherein, based on the state change information, the current state information of the target system is prompted to the target object. When it is detected that the isolation of the faulty line is successful, the global state of the target system is changed from the second state to the first state; locking and marking the memory area within the preset range corresponding to the faulty line as the target state, wherein the state information of the memory area is sent to the warning platform to prompt the target object to process.

[0107] In an optional embodiment, a state machine constraint is introduced. When the BMC (Baseboard Management Controller) discovers a memory line-level error through real-time monitoring, the system state immediately changes from "normal" (i.e., the above-mentioned first state) to "degraded" (i.e., the above-mentioned second state), and triggers the system event log (SEL, System Event Log) to record the state change information of the target system. At the same time, the access permission of the area where the faulty memory line is located is restricted to prevent potential errors from spreading. In addition, the change of state can also be notified to the user through the system event log in the form of a notice, prompting that the current system is in a degraded state.

[0108] Once the faulty line is successfully isolated, that is, the isolation operation is completed by means such as Runtime PPR or Page Offline, the global state of the system will change back from "degraded" to "normal" again. However, different from simple state restoration, the isolated memory area (i.e., the memory area corresponding to the preset range, for example, a memory page, or the corresponding memory module) is marked as the "sub-healthy" state (i.e., the above-mentioned target state) and locked. Although the system generally returns to the normal operation mode, the access to some memory resources is still restricted to maintain the stability and data integrity of the system. Synchronizing the "sub-healthy" state information to the warning platform can timely notify the operation and maintenance personnel and prompt them to take further measures, such as replacing the faulty memory module. The advantage of doing this is that, on the one hand, users can continuously obtain accurate system state feedback, and on the other hand, the operation and maintenance team can carry out targeted hardware maintenance to prevent the recurrence of faults and improve the overall reliability of the server.

[0109] As an information carrier of state changes, the SEL log not only records the state transition process of the system from "normal" to "degraded" and then back to "normal", but also includes specific details of fault line isolation, such as isolation time, physical address of the isolated line, etc. Users can understand this transformation through the SEL log, avoiding the problems of lagging state information or false alarms that may exist in traditional solutions.

[0110] In an optional embodiment, the state record can be implemented through a schematic diagram as Figure 7 shown, which specifically includes: in the form of a state machine, it shows the conversion logic of the server system state during the fault detection and isolation process, including normal state, degraded state, sub-healthy state, and emergency state. Normal state: This is the default state of the server, indicating that all hardware components, including memory, are operating normally and no faults are detected. Degraded state: When the BMC detects a memory row-level error, the system state will immediately transition to the degraded state. In this state, the system will trigger a SEL event record and at the same time restrict access to the area where the faulty memory row is located to prevent the spread of errors. Sub-healthy state: Once the faulty row is successfully isolated, the system state will recover from the degraded state to the normal state, but the associated memory area will be marked as sub-healthy and locked. The emergency state is the state of isolation failure, which requires manual intervention to make the system return to the normal state. It should be noted that if the memory area is manually repaired, it will also switch to the normal state.

[0111] Optionally, in the memory fault handling method provided in the embodiment of the present application, after generating the fourth log file, the method further includes: evaluating the remaining life of the memory based on the third log file and the fourth log file to obtain an evaluation result; determining a maintenance strategy for the memory based on the evaluation result.

[0112] In an optional embodiment, after fault isolation, the early warning platform can perform in-depth cross-analysis by combining the fault handling records in the third log file with the real-time system state data in the fourth log file to evaluate the remaining life of the memory. For example, the early warning platform will comprehensively consider the number of faults, types, success rate of isolation operations in the maintenance log (i.e., the third log file) and the decline trend of performance data in the one-key log (i.e., the fourth log file).

[0113] For example, if the fault line isolation operation is frequent and the number of remaining redundant lines gradually decreases, this may be a signal of accelerated memory aging; or, if the system performance indicators continue to decline after fault isolation, this also implies a decrease in the overall health of the memory and a possible short remaining life.

[0114] For example, machine learning models or pre-set evaluation rules can be used to analyze the health of memory based on the collected data, predict the probability of failure in a specific time period in the future, and thus obtain the evaluation result of the remaining life of the memory. The higher the failure probability, the shorter the corresponding life. Finally, based on the evaluation results, the intelligent early warning platform can generate targeted maintenance strategies.

[0115] For example, when the remaining life of the memory is predicted to be insufficient to support business needs, the system will recommend that users replace the memory sticks preventively within a specified time to avoid unexpected downtime. If the remaining life is acceptable but there is a risk trend, the early warning platform may recommend optimizing the allocation strategy of memory resources to reduce dependence on aging memory and ensure the operational stability of key businesses. For memory whose health status is still unclear, the early warning platform may recommend strengthening performance monitoring and regular inspections to capture any possible signs of degradation in a timely manner.

[0116] The remaining life assessment mechanism based on log analysis provides the operation and maintenance team with accurate maintenance strategy guidance, ensuring that the server can respond quickly and plan ahead when facing memory row-level failures, achieving the dual goals of short-term fault isolation and long-term hardware health management.

[0117] Optionally, in the memory fault handling method provided in the embodiment of the present application, after obtaining the evaluation result, the method also includes: determining the alarm level of the memory based on the evaluation result, and generating an alarm card based on the alarm level and maintenance strategy; and pushing the alarm card in the target interface of the early warning platform.

[0118] In an optional embodiment, the early warning platform evaluates the remaining life of the memory based on the comprehensive analysis of the third log file and the fourth log file, and gives a specific quantitative result. This quantitative result reflects the current health status of the memory and the expected continuous service capability. According to the evaluation results, the early warning platform will automatically divide the memory into different alarm levels, usually using color coding to indicate the severity: red represents fatal level failures, orange is a warning level, and blue represents routine reminders. For example, if the evaluation results show that the remaining life of the memory is extremely low, which may cause the system to crash at any time, it will be marked as a red alarm; if the failure trend is obvious but does not affect the system operation temporarily, it may be classified as an orange warning; and for mild or initially detected failures, the system will give a blue reminder.

[0119] According to the alarm level and maintenance strategy, the early warning platform will automatically generate an alarm card. The alarm card may include the severity level: visually displaying the severity of the fault through color icons, fault location: accurate to the component silk screen code and physical slot, facilitating the quick identification of the location of the faulty memory. Maintenance strategy: Based on the nature of the fault and the remaining life assessment results, specific maintenance suggestions are put forward, such as suggesting to replace the memory module within 72 hours. The alarm card will be pushed in real time on the target interface of the early warning platform and can also be pushed in real time in the alarm center module of the BMC Web interface to ensure that the operation and maintenance personnel can immediately receive the fault information and maintenance suggestions.

[0120] In an optional embodiment, after receiving the alarm, the user can complete three operations on the Web interface: view the scrolling alarm summary in the home page status bar; click on the alarm card to enter the details page to confirm the diagnostic report; and export the compressed log or submit a maintenance work order with one click through the "Log Management" module.

[0121] In an optional embodiment, the faulty component can also be highlighted in the topology diagram to visually emphasize the fault location, facilitating the operation and maintenance personnel to locate the faulty hardware on the physical server, and generating an RMA (Return Material Authorization) electronic label for tracking the spare part replacement process to ensure that the faulty memory module is effectively replaced within the specified time, forming a closed-loop management process from fault alarm to actual maintenance operation.

[0122] By pushing the alarm card in real time, it is ensured that the operation and maintenance team is notified at the first time of the fault occurrence, shortening the fault response time. Based on the assessment of the remaining life and the generation of the RMA electronic label, it can assist the operation and maintenance team to reasonably arrange the spare part inventory, perform preventive memory replacement, reduce the risk of unplanned downtime, and extend the service life of the server.

[0123] In an optional embodiment, it can also be as Figure 8The flowchart shown implements row fault handling: data collection, detecting the memory operation status, and recording memory error information; fault feature extraction and classification, analyzing the memory error information, and determining the fault type (i.e., whether it is a row fault); executing isolation decision-making, automatically triggering isolation operations based on the fault type; status synchronization and logging, updating the component status and system log information; verification and recovery, verifying the isolation effect, and if successful, resuming business operations. If it fails, an alarm is issued and manual intervention is required. For example, the integrity and execution result of the isolation action can be verified by parsing the detailed operation information recorded in the maintenance log. For soft faults, the operating system kernel automatically migrates the affected business processes to a healthy memory area to ensure a seamless service switch; for hard faults that cause physical damage at the row level, the spare part replacement process can be triggered through the early warning platform to achieve rapid replacement of the faulty component and restoration of the integrity of the business environment, ultimately achieving dual guarantees of system high availability and business continuity.

[0124] In the memory fault handling method provided by the embodiments of the present application, through the collaborative work of the BMC and the early warning system, the collection of fault information, fault type classification, fault isolation, and early warning are realized. The Runtime PPR and Page Offline technologies are used to achieve the isolation of memory row faults, a hierarchical isolation mechanism is constructed, and the isolation efficiency and reliability are optimized. After the isolation is completed, the state machine constraint and dual-log collaborative mechanism are used to synchronously update the system state to "normal" and record detailed log information, effectively solving the problem of inconsistent states and logs in the traditional solution, thus significantly improving the reliability of the server. This intelligent fault diagnosis and isolation method not only reduces the operation and maintenance costs but also is of great significance for improving the overall performance and reliability of the server.

[0125] In the memory fault handling method provided by the embodiments of the present application, when there is a memory row fault in the memory of the target system, the isolation process is started for the target physical address in a timely manner, avoiding the performance degradation and data error accumulation caused by the faulty row continuing to participate in data reading and writing. Therefore, the technical problem in the related art that the memory row fault cannot be quickly processed, resulting in system downtime, can be solved, and the technical effect of reducing the risk of system downtime can be achieved.

[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0127] The embodiments of the present application also provide a memory fault handling device. Figure 9 For the schematic diagram of the memory fault handling device provided by the embodiments of the present application, as Figure 9 shown, the schematic diagram of the memory fault handling device includes:

[0128] The first trigger unit 901 is configured to trigger an isolation operation on a target physical address when there is a memory line fault in the memory of the target system, where the target physical address at least includes the physical address of the faulty line with the memory line fault in the memory.

[0129] In the memory fault handling device provided in the embodiments of the present application, when there is a memory line fault in the memory of the target system, the first trigger unit 901 triggers an isolation operation on the target physical address, where the target physical address at least includes the physical address of the faulty line with the memory line fault in the memory, solving the technical problem in the related art that the memory line fault cannot be quickly processed, resulting in system downtime. When there is a memory line fault in the memory of the target system, the isolation process is started for the target physical address in a timely manner, avoiding the performance degradation and data error accumulation caused by the faulty line continuing to participate in data reading and writing, and thus achieving the technical effect of reducing the risk of system downtime.

[0130] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: a first acquisition unit configured to acquire the running state information of the memory; a judgment unit configured to judge whether there is a memory line fault according to the running state information.

[0131] Optionally, in the memory fault handling device provided in the embodiments of the present application, the acquisition unit includes: a first acquisition module configured to acquire the temperature information during the running of the memory through a temperature sensor in the memory; a second acquisition module configured to acquire the voltage information during the running of the memory through a voltage sensor in the memory; a third acquisition module configured to acquire the memory access information through a register in the memory; a first determination module configured to obtain the running state information according to the temperature information, voltage information, and memory access information.

[0132] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: a second acquisition unit configured to acquire a first log file in the target system before triggering an isolation operation on the target physical address, and extract first logical address information with a memory line fault from the first log file; a first determination unit configured to determine first physical address information corresponding to the first logical address information; a second determination unit configured to, when the isolation policy is a page isolation policy, determine the physical address of the target memory page corresponding to the first physical address information as the target physical address.

[0133] Optionally, in the memory fault handling device provided in the embodiments of the present application, the first trigger unit includes: an isolation module configured to isolate the target memory page; a disabling module configured to prohibit the use of the memory area corresponding to the target memory page.

[0134] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: a third determination unit, configured to determine second physical address information after isolating a target memory page; a migration unit, configured to migrate service data located in the target memory page to the second physical address information.

[0135] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: a third acquisition unit, configured to acquire target register snapshot information corresponding to a memory row fault after triggering an isolation operation on a target physical address; a fourth acquisition unit, configured to acquire performance data information of the memory during a memory row fault; a fifth acquisition unit, configured to acquire memory operation information during the execution of the isolation operation; a first generation unit, configured to generate a second log file according to the target register snapshot information, the performance data information, and the memory operation information.

[0136] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: an analysis unit, configured to analyze the memory based on the second log file after generating the second log file according to the target register snapshot information, the performance data information, and the memory operation information, to obtain an analysis result; a fourth determination unit, configured to determine a maintenance policy for the memory according to the analysis result.

[0137] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: a first setting unit, configured to set the memory row fault isolation mechanism in the baseboard management controller to a triggered state; a second setting unit, configured to set configuration items of the isolation policy of the basic input / output system through the baseboard management controller, and send a configured signal to the basic input / output system through a target protocol.

[0138] Optionally, in the memory fault handling device provided in the embodiments of the present application, for the first trigger unit, it includes: a second determination module, configured to, when the isolation policy is a runtime isolation policy, determine the physical address corresponding to the faulty row as the target physical address, and determine whether there is a redundant row in the memory; a marking module, configured to, when there is a redundant row in the memory, mark the target physical address as unavailable, and migrate the service data in the faulty row to the redundant row.

[0139] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: an isolation unit, configured to, after determining whether there is a redundant row in the memory, perform an isolation operation on the target physical address according to the page isolation policy when there is no redundant row in the memory and the basic input / output system has configured the page isolation policy.

[0140] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: a first trigger unit, configured to, after determining whether there is a redundant row in the memory, trigger a first signal to a warning platform when there is no redundant row in the memory and the basic input / output system is not configured with a page isolation policy, so as to instruct a target object to handle a memory row fault.

[0141] Optionally, in the memory fault handling device provided in the embodiments of the present application, the determining unit includes: a third determining module, configured to determine, according to running state information, target number information of correctable errors that occur in each memory row in the memory during a current detection period; and a judging module, configured to judge whether there is a memory row fault according to the target number information.

[0142] Optionally, in the memory fault handling device provided in the embodiments of the present application, the judging module includes: an obtaining sub-module, configured to obtain historical number information of correctable errors that occur in each memory row during a historical time period; a calculating sub-module, configured to calculate a ratio between the historical number information and the target number information; and a judging sub-module, configured to judge a magnitude relationship between the ratio and a preset threshold value, so as to judge whether there is a memory row fault.

[0143] Optionally, in the memory fault handling device provided in the embodiments of the present application, the judging sub-module includes: a first secondary determining sub-module, configured to determine that a memory row fault occurs if the ratio is greater than or equal to the preset threshold value; and a second secondary determining sub-module, configured to determine that no memory row fault occurs if the ratio is less than the preset threshold value.

[0144] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: a second generating unit, configured to generate a third log file according to timestamp information when a memory row fault occurs, physical address information corresponding to a faulty row, and the number of remaining redundant rows corresponding to the memory after triggering an isolation operation on a target physical address; and a third generating unit, configured to generate a fourth log file according to target register snapshot information when a memory row fault occurs, performance data information, and memory operation information during an isolation operation, where the third log file and the fourth log file are used to evaluate the lifespan of the memory.

[0145] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: a first conversion unit, configured to, when a memory row fault is detected, convert a global state of a target system from a first state to a second state; and a second trigger unit, configured to trigger a system event log to record state change information of the target system, where current state information of the target system is prompted to a target object based on the state change information.

[0146] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: a second conversion unit, configured to convert the global state of the target system from the second state to the first state when it is detected that the isolation of the faulty row is successful; a locking unit, configured to lock and mark the memory area within a preset range corresponding to the faulty row as the target state, where the status information of the memory area is sent to the warning platform to prompt the target object to process.

[0147] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: an evaluation unit, configured to evaluate the remaining life of the memory based on the third log file and the fourth log file after generating the fourth log file, to obtain an evaluation result; a fifth determination unit, configured to determine a maintenance policy for the memory according to the evaluation result.

[0148] Optionally, in the memory fault handling device provided in the embodiments of the present application, the device further includes: a sixth determination unit, configured to determine the alarm level of the memory according to the evaluation result after obtaining the evaluation result, and generate an alarm card based on the alarm level and the maintenance policy; a push unit, configured to push the alarm card in the target interface of the warning platform.

[0149] For the description of the features in the corresponding embodiments of the memory fault handling device, reference can be made to the relevant descriptions in the corresponding embodiments of the memory fault handling method, which will not be elaborated here one by one.

[0150] The embodiments of the present application further provide an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the memory fault handling method.

[0151] The embodiments of the present application further provide a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the steps in any of the above embodiments of the memory fault handling method when running.

[0152] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.

[0153] The embodiments of the present application further provide a computer program product, where the computer program product includes a computer program, and the computer program, when executed by a processor, implements the steps in any of the above embodiments of the memory fault handling method.

[0154] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps in any of the above-described embodiments of the memory fault handling method.

[0155] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0156] The above has introduced in detail a memory fault handling method provided by the present application. Specific examples are used herein to illustrate the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for handling memory faults, characterized in that, including: When there is a memory line fault in the memory of the target system, an isolation operation is triggered for the target physical address, where the target physical address at least includes the physical address of the faulty line with the memory line fault in the memory.

2. The memory fault handling method according to claim 1, wherein The method further includes: Obtaining the running state information of the memory; Judging whether there is a memory line fault according to the running state information.

3. The memory fault handling method according to claim 2, wherein The obtaining the running state information of the memory includes: Obtaining the temperature information during the running of the memory through a temperature sensor in the memory; Obtaining the voltage information during the running of the memory through a voltage sensor in the memory; Obtaining memory access information through a register in the memory; Obtaining the running state information according to the temperature information, the voltage information and the memory access information.

4. The memory fault handling method according to claim 1, wherein Before triggering the isolation operation for the target physical address, the method further includes: Obtaining a first log file in the target system and extracting first logical address information with a memory line fault from the first log file; Determining first physical address information corresponding to the first logical address information; When the isolation policy is a page isolation policy, determining the physical address of the target memory page corresponding to the first physical address information as the target physical address.

5. The memory fault handling method according to claim 4, wherein The triggering the isolation operation for the target physical address includes: Isolating the target memory page; Forbidding the use of the memory area corresponding to the target memory page.

6. The memory fault handling method according to claim 5, wherein, After isolating the target memory page, the method further includes: Determining second physical address information; Migrating the service data located in the target memory page to the second physical address information.

7. The memory fault handling method according to claim 1, wherein, After triggering the isolation operation for the target physical address, the method further includes: Obtaining target register snapshot information corresponding to the memory line fault; Obtaining performance data information of the memory at the time of the memory line fault; Obtaining memory operation information during the execution of the isolation operation; Generating a second log file according to the target register snapshot information, the performance data information and the memory operation information.

8. The memory fault handling method according to claim 7, wherein After generating the second log file according to the target register snapshot information, the performance data information and the memory operation information, the method further includes: Analyzing the memory based on the second log file to obtain an analysis result; Determining a maintenance strategy for the memory according to the analysis result.

9. The memory fault handling method according to claim 1, characterized in that The method further includes: Setting the memory line fault isolation mechanism in the baseboard management controller to a triggered state; Setting the configuration item of the isolation policy of the basic input / output system through the baseboard management controller and sending a signal indicating the completion of the configuration to the basic input / output system through a target protocol.

10. The memory fault handling method according to claim 1, wherein, The triggering the isolation operation for the target physical address includes: When the isolation policy is a runtime isolation policy, determining the physical address corresponding to the faulty line as the target physical address and judging whether there is a redundant line in the memory; When there is a redundant line in the memory, marking the target physical address as unavailable and migrating the service data in the faulty line to the redundant line.

11. The memory fault handling method according to claim 10, wherein, After judging whether there is a redundant line in the memory, the method further includes: When there is no redundant row in the memory and the basic input / output system has configured a page isolation policy, perform an isolation operation on the target physical address according to the page isolation policy.

12. The memory fault handling method according to claim 10, wherein After determining whether there is a redundant row in the memory, the method further includes: When there is no redundant row in the memory and the basic input / output system has not configured a page isolation policy, trigger a first signal to the warning platform to instruct the target object to handle the memory row failure.

13. The memory fault handling method according to claim 2, wherein The determining whether there is a memory row failure according to the operating status information includes: According to the operating status information, determine the target number of times information of correctable errors occurred by each memory row in the memory during the current detection period; According to the target number of times information, determine whether there is a memory row failure.

14. The memory fault handling method according to claim 13, wherein The determining whether there is a memory row failure according to the target number of times information includes: Obtain the historical number of times information of correctable errors occurred by each memory row during the historical time period; Calculate the ratio between the historical number of times information and the target number of times information; Judge the size relationship between the ratio and a preset threshold value to determine whether there is a memory row failure.

15. The memory fault handling method according to claim 14, wherein The judging the size relationship between the ratio and a preset threshold value to determine whether there is a memory row failure includes: If the ratio is greater than or equal to the preset threshold value, determine that a memory row failure has occurred; If the ratio is less than the preset threshold value, determine that no memory row failure has occurred.

16. The memory fault handling method according to claim 1, wherein After triggering the isolation operation on the target physical address, the method further includes: Generate a third log file according to the timestamp information when the memory row failure occurs, the physical address information corresponding to the failed row, and the remaining redundant row quantity corresponding to the memory; Generate a fourth log file according to the target register snapshot information when the memory row failure occurs, the performance data information, and the memory operation information during the execution of the isolation operation, wherein the third log file and the fourth log file are used to evaluate the lifespan of the memory.

17. The memory fault handling method according to claim 1, wherein The method further includes: When a memory row failure is detected, the global state of the target system changes from the first state to the second state; Trigger a system event log to record the state change information of the target system, wherein based on the state change information, prompt the target object of the current state information of the target system.

18. The memory fault handling method according to claim 17, wherein The method further includes: When it is detected that the isolation of the failed row is successful, the global state of the target system changes from the second state to the first state; Lock and mark the memory area within a preset range corresponding to the failed row as the target state, and send the state information of the memory area to the warning platform to prompt the target object to process it.

19. The memory fault handling method according to claim 16, wherein After generating the fourth log file, the method further includes: Evaluate the remaining lifespan of the memory according to the third log file and the fourth log file to obtain an evaluation result; Determine the maintenance strategy for the memory according to the evaluation result.

20. The memory fault handling method according to claim 19, characterized in that, After obtaining the evaluation result, the method further includes: Determine the alarm level of the memory according to the evaluation result, and generate an alarm card based on the alarm level and the maintenance strategy; Push the alarm card in the target interface of the warning platform.

21. A memory fault handling device, characterized in that, Comprising: A first trigger unit, configured to trigger an isolation operation on a target physical address when there is a memory line fault in the memory of the target system, where the target physical address at least includes the physical address of the faulty line with the memory line fault in the memory.

22. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the memory fault handling method according to any one of claims 1 to 20 when executing the computer program.

23. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program implements the steps of the memory fault handling method according to any one of claims 1 to 20 when executed by a processor.

24. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the memory fault handling method according to any one of claims 1 to 20 when executed by a processor.

Citation Information

Patent Citations

  • DRAM memory row disturbance error solving method

    CN110706733A

  • Memory fault early warning method and device, electronic equipment and readable medium

    CN115629905A

  • Fault detection method and device of DRAM (Dynamic Random Access Memory), storage medium and electronic equipment

    CN119065879A

Cited By

  • Integrated circuit memory fault management method and system based on NPU processor architecture

    CN121743098A

  • Memory fault processing method, electronic equipment, storage medium and program product

    CN122261900A