Memory fault repairing method and device, equipment, medium and computer program product

By obtaining memory error information and updating the memory address mapping table, the physical address of the faulty memory page is mapped to the physical address of the spare memory page, which solves the problem of high latency in memory failure processing, and realizes efficient memory failure repair and reduces system downtime.

CN120353632AInactive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510828619.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, memory fault handling depends on the system management interrupt (SMI) mechanism, resulting in high latency and affecting the server's real-time processing capabilities and overall performance.

Method used

By obtaining the memory error information sent by the memory controller, locate the fault target memory page, obtain the spare memory page from the spare memory pool, and update the memory address mapping table, map the physical address of the fault memory page to the physical address of the spare memory page, realizing memory repair during the system operation stage.

Benefits of technology

Reduces system downtime caused by memory failure repair, improves memory failure repair efficiency, and avoids the system's long-term waiting state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353632A_ABST
    Figure CN120353632A_ABST
Patent Text Reader

Abstract

The invention discloses a memory fault repairing method and device, equipment, a medium and a computer program product, and relates to the technical field of computers.The memory fault repairing method includes the steps that memory error information sent by a memory controller is obtained, a fault target memory page can be positioned, and a standby memory page is obtained from a standby memory pool; the memory address mapping table is updated, the physical address of the target memory page with the fault is mapped to the physical address of the standby memory page, the repairing process does not depend on triggering of a system management interrupt mechanism, memory repairing in the system running stage is achieved, the system downtime caused by memory fault repairing is shortened, and therefore the memory repairing efficiency is improved. The problem of high delay in memory fault processing can be solved, and the technical effect of improving the memory fault repairing efficiency is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a method, apparatus, device, medium, and computer program product for memory fault repair. Background Art

[0002] With the rapid development of technologies such as cloud computing and big data, the application scenarios of servers have been continuously expanded, and customers have higher and higher requirements for the stability and performance of servers. As a core component for data storage and processing in servers, memory plays an important role in server stability and performance.

[0003] In related technologies, memory fault handling mainly relies on the System Management Interrupt (SMI) mechanism. However, the triggering of SMI will introduce a relatively high latency, causing the system to be in a waiting state for a long time, affecting the real-time processing ability and overall performance of the server. Summary of the Invention

[0004] This application provides a method, apparatus, device, medium, and computer program product for memory fault repair, so as to at least solve the problem of high latency in memory fault handling in related technologies.

[0005] This application provides a method for memory fault repair, including: Obtaining memory error information sent by a memory controller; the memory error information includes the location where the error occurs; Locating a target memory page with a fault according to the memory error information; Obtaining a spare memory page from a spare memory pool; Mapping the physical address of the target memory page in a memory address mapping table to the physical address of the spare memory page; the memory address mapping table is used to record the mapping relationship between the virtual address and the physical address of the memory page.

[0006] This application also provides a memory fault repair system, including: a memory controller, a repair module, a baseboard management controller, and an operating system; The memory controller is configured to detect a memory error and send the memory error information to the baseboard management controller based on the detected memory error; The baseboard management controller is configured to locate a target memory page with a fault according to the memory error information, obtain a spare memory page allocated by the operating system from the spare memory pool, generate a repair instruction based on the physical address of the target memory page and the physical address of the spare memory page, and send the repair instruction to the repair module; The repair module is used to map the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page based on the repair instruction; the memory address mapping table is used to record the mapping relationship between the virtual address and the physical address of the memory page.

[0007] This application also provides a memory fault repair device, including: The first acquisition module is used to acquire the memory error information sent by the memory controller; the memory error information includes the location where the error occurs. The positioning module is used to locate the target memory page with a fault according to the memory error information. The second acquisition module is used to acquire a spare memory page from the spare memory pool. The mapping module is used to map the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page; the memory address mapping table is used to record the mapping relationship between the virtual address and the physical address of the memory page.

[0008] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above memory fault repair methods when executing the computer program.

[0009] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above memory fault repair methods are implemented.

[0010] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any of the above memory fault repair methods are implemented.

[0011] Through this application, by acquiring the memory error information sent by the memory controller, the target memory page with a fault can be located, a spare memory page can be acquired from the spare memory pool, and the pre-prepared spare resources are utilized. By updating the memory address mapping table, the physical address of the faulty target memory page is mapped to the physical address of the spare memory page. The repair process does not depend on the triggering of the system management interrupt mechanism, realizes memory repair during the system operation stage, reduces the system downtime caused by memory fault repair, and therefore, can solve the problem of high latency in memory fault handling and achieve the technical effect of improving the memory fault repair efficiency. Description of the Drawings

[0012] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0013] Figure 1 is a schematic flowchart of the memory fault repair method provided by the embodiments of the present application; Figure 2 is a schematic architecture diagram of the memory fault repair system provided by the embodiments of the present application; Figure 3 is a schematic flowchart of the work process in the scenario example provided by the embodiments of the present application; Figure 4 is a schematic structural diagram of the memory fault repair device provided by the embodiments of the present application; Figure 5 is a schematic structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0015] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0016] In the related art, memory fault handling mainly relies on the SMI mechanism. When the memory controller detects a fault, it triggers the SMI. The Central Processing Unit (CPU) transfers the fault information to the Baseboard Management Controller (BMC), and then the Basic Input / Output System (BIOS) performs Predictive Prediction and Repair (PPR) operations to repair the memory. The triggering of the SMI introduces a relatively high latency, causing the system to wait for a long time and affecting the real-time processing ability and overall performance of the server.

[0017] The BMC is an embedded controller that operates independently on the server motherboard. It has the ability to work independently of the server CPU and operating system, does not depend on other hardware in the system, nor on the BIOS or operating system. Even if the server loses power or the system crashes, it can still continuously monitor the hardware status.

[0018] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0019] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the memory fault repair method depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0020] The embodiments of the present application provide a memory fault repair method. The method will be described in detail in combination with the execution process of the memory fault repair method.

[0021] Specifically, Figure 1 FIG. is a flowchart of a memory fault repair method provided according to an embodiment of the present application.

[0022] As Figure 1 shown, the memory fault repair method includes: step 110, step 120, step 130, and step 140.

[0023] Step 110: Obtain the memory error information sent by the memory controller; the memory error information includes the location where the error occurs.

[0024] The memory of the server usually consists of multiple storage units. For the convenience of management and to improve storage efficiency, the memory space can be divided into memory pages of a fixed size. These memory pages are similar to independent "small warehouses" for storing different data segments.

[0025] The memory controller is responsible for managing the read and write operations of the memory and coordinating the data transfer between the CPU and the memory.

[0026] When a memory failure occurs, the memory controller will detect the abnormal situation and generate a memory error message. The memory error message includes the specific location where the error occurs and other contents, so as to locate and repair the fault subsequently.

[0027] In this step, the memory controller can, based on the internal detection mechanism, such as parity check, ECC (Error Correction Code) check, etc., monitor the read and write operations of the memory in real time. When abnormal situations such as data transfer errors and storage data verification failures are found, it will record the location where the error occurs and encapsulate this information into a memory error message and send it out.

[0028] In this step, communication can be carried out with the memory controller through the communication interface to obtain the memory error message sent by the memory controller.

[0029] Step 120: Locate the target memory page where the fault occurs according to the memory error message.

[0030] In this step, each memory page in the memory has a corresponding address. After obtaining the memory error message, the location of the error in the memory error message can be compared with the address of the memory page to determine the target memory page where the error occurs.

[0031] Step 130: Obtain a spare memory page from the spare memory pool.

[0032] In this step, the spare memory pool is an area pre-divided in the server memory, which stores at least one spare memory page. The spare memory page is in a standby state and is used to replace the main memory page in time when a failure occurs. The spare memory page is basically the same as the memory page used normally in terms of function and structure. It is also a fixed-size storage unit divided from the memory space and can undertake the tasks of data storage and processing. However, under normal circumstances, the spare memory page will not be allocated and used and remains idle.

[0033] In some embodiments, a free spare memory page can be randomly obtained from the spare memory pool. Of course, the spare memory pool can also be managed by the operating system of the server. When a spare memory page needs to be obtained, a request can be sent to the operating system. After receiving the request, the operating system allocates an available backup memory page through the spare memory pool allocation logic. For example, the operating system can maintain a memory page list, which records the free spare memory pages in the spare memory pool. When allocating a spare memory page, the operating system will select a free spare memory page from this list.

[0034] In some embodiments, obtaining a spare memory page from a spare memory pool includes: Determining the physical distance between each spare memory page in the spare memory pool and the central processing unit according to the physical address of each spare memory page; Obtaining the spare memory page with the shortest physical distance to the central processing unit.

[0035] In a computer hardware architecture, the memory is not closely connected to the central processing unit. Instead, data is transmitted through complex circuits and buses. There are actual physical distance differences between memory pages at different locations and the central processing unit, and this physical distance directly affects the efficiency of data transmission.

[0036] Each spare memory page in the spare memory pool corresponds to a unique physical address. By parsing these physical addresses, the specific location information of the spare memory page on the motherboard can be obtained. Combining the location of the central processing unit on the motherboard and using the hardware layout information, the physical distance between each spare memory page and the central processing unit can be calculated.

[0037] Since the shorter the physical distance, the shorter the path that data travels between the memory page and the central processing unit, the smaller the signal attenuation and transmission delay, and the faster the data transmission speed. Therefore, the spare memory page with the shortest physical distance to the central processing unit can be preferentially selected.

[0038] In this embodiment, by determining the physical distance between the spare memory page and the central processing unit and preferentially selecting the spare memory page with the shortest distance, the performance after memory repair is fundamentally optimized, the delay in data transmission between the memory and the central processing unit is reduced. After completing the memory page replacement, the central processing unit can access the repaired data faster, improving the data transmission efficiency after memory repair.

[0039] Step 130: Mapping the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page; the memory address mapping table is used to record the mapping relationship between the virtual address and the physical address of the memory page.

[0040] In this step, the memory address mapping table is a data structure used to record the mapping relationship between the virtual address and the physical address of the memory page. The virtual address is the logical address space allocated by the operating system for each process. It provides a continuous address range for the program, enabling the program to conveniently access the memory. The physical address is the address where data is actually stored in the memory hardware, corresponding to the specific location in the physical memory module. Since the physical memory may be shared by multiple processes and the layout of the physical memory may not be continuous, it is necessary to map the virtual address to the actual physical address through the memory address mapping table to achieve effective management and allocation of the memory.

[0041] After obtaining the spare memory page, in order to repair the memory fault and enable the program to continue to normally access the data originally stored in the target memory page, it is necessary to map the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page.

[0042] Specifically, first find the corresponding entry of the virtual address and physical address of the target memory page recorded in the memory address mapping table, and then modify the physical address originally pointing to the target memory page in this entry to the physical address of the newly obtained spare memory page. After the modification is completed, when the program accesses this virtual address again, the memory management unit will convert the virtual address to the physical address of the spare memory page according to the updated memory address mapping table, so that the program can successfully access the data stored on the spare memory page.

[0043] Through this application, since the memory error information sent by the memory controller can be obtained to locate the faulty target memory page, obtain the spare memory page from the spare memory pool, utilize the pre-prepared spare resources, and update the memory address mapping table to map the physical address of the faulty target memory page to the physical address of the spare memory page, the repair process does not rely on the triggering of the system management interrupt mechanism, realizes the memory repair during the system operation stage, reduces the system downtime caused by memory fault repair. Therefore, the problem of high latency in memory fault handling can be solved, and the technical effect of improving the memory fault repair efficiency can be achieved.

[0044] In some embodiments, before obtaining the memory error information sent by the memory controller, it includes: Detect the memory through the memory controller; In the case of detecting a memory error, identify the memory error type; In the case where the memory error type is a multi-bit error, generate a memory error information.

[0045] In this embodiment, the memory controller can monitor the state of the memory. For example, the memory can be detected by means of ECC technology or periodic scanning of the memory. Among them, ECC technology automatically generates a check code when data is written into the memory and stores it together with the data, and detects whether the data has an error by comparing the check code when reading. During the process of the memory controller scanning the memory, the existing memory errors can be found by writing known data and verifying the consistency of the reading results.

[0046] A bit is the smallest unit for a computer to store and process information and is also the basic data unit in the memory. Each bit can only represent two states: 0 or 1. In a memory chip, bits are arranged in an array and data is stored through the high and low levels of electrical signals. For example, a combination of 8 bits can represent a byte.

[0047] In this embodiment, memory error types can be classified into single-bit errors and multi-bit errors. A single-bit error refers to a situation where only one bit flips (e.g., from 0 to 1 or from 1 to 0) during data transmission or storage. Such errors are usually caused by transient factors such as random electromagnetic interference. Multi-bit errors involve multiple bits going wrong simultaneously, which may be due to physical damage to the memory chip or circuit.

[0048] In some embodiments, the memory error type can be determined as a single-bit error or a multi-bit error based on the number of bit errors. For example, by analyzing the error pattern of the ECC check result to distinguish the error type. When the check result shows that only one bit does not match, it is determined as a single-bit error; if there are multiple non-matching bits or the checksum is completely inconsistent, it is identified as a multi-bit error.

[0049] When the memory error type is determined to be a multi-bit error, the memory controller generates a memory error message. The memory error message can include error location information and error description information.

[0050] The error location information is used to record the location where the error occurs, that is, the physical address. For example, it can be presented in the form of a page frame number and an offset. Since memory is managed in pages, each memory page contains a fixed number of bytes, and each byte consists of 8 bits. Therefore, the error location information can be mapped from the bit level upwards to bytes, memory pages, and locations. For example, if it is detected that the 3rd and 5th bits in a certain byte are both in error, the memory controller can determine the page frame where the byte is located, and then combine the in-page offset to finally determine the complete physical address, facilitating subsequent rapid positioning to the target memory page.

[0051] The error description information can include metadata such as the timestamp when the error occurred, the error type identifier, and the size of the affected data, providing a reference for fault analysis. For example, the error type identifier can adopt a specific coding scheme, where different coding values represent different types of multi-bit errors, such as adjacent bit errors, randomly distributed multi-bit errors, etc.

[0052] In this embodiment, through the detection function of the memory controller, errors in the memory can be detected in a timely manner, the memory error types are identified to distinguish single-bit errors and multi-bit errors, and corresponding processing measures are taken for multi-bit errors, concentrating resources on handling multi-bit errors that need to be repaired, reducing unnecessary system overhead.

[0053] In some embodiments, the detection of the memory includes: Detecting the mapping relationship in the memory address mapping table; When there is an error in the mapping relationship, it is determined that a memory error has been detected.

[0054] The memory address mapping table is a key hub in the server memory management system, recording the correspondence between the virtual addresses and physical addresses of memory pages, enabling the operating system and application programs to accurately access data in physical memory through virtual addresses. During the operation of the server, the memory address mapping table is continuously read and updated to meet the dynamic allocation and release requirements of memory by programs.

[0055] In this embodiment, the memory controller can scan the memory address mapping table regularly or when triggered by a specific event, and check the mapping entries recorded in the table through a verification algorithm. For example, it can verify whether each virtual address can correspond to a reasonable range of physical addresses, and whether the physical address actually exists and is in an available state. It can also check the consistency of the mapping relationship, verifying whether there are contradictions such as one virtual address corresponding to multiple physical addresses, or one physical address being repeatedly mapped. When errors are found in the mapping relationship during the detection process of the memory address mapping table, it is determined that a memory error has been detected.

[0056] In this embodiment, by checking the mapping relationship in the memory address mapping table, memory mapping problems caused by hardware failures or software errors can be discovered in a timely manner, enhancing the detection ability and processing efficiency of memory errors.

[0057] In some embodiments, the detection of memory includes: Reading a data block and the corresponding storage checksum from the memory; the storage checksum is the checksum generated for storing the data block; Calculating the read checksum of the data block using a target algorithm; Determining that a memory error has been detected when the storage checksum and the read checksum are inconsistent.

[0058] In the server memory system, data is not stored chaotically, but is organized and managed in units of data blocks. Each data block can contain several bytes of data. To ensure the accuracy of data during storage and reading, a corresponding storage checksum may be generated for each data block. The storage checksum is similar to the fingerprint of the data, used to uniquely identify the corresponding data block. The storage checksum is calculated by a specific algorithm for the content of the data block when the data is written into the memory.

[0059] When the memory controller performs the detection task, it can read the data block and the corresponding storage checksum from the memory in a certain order. During the reading process, the memory control also records the memory address where the data block is located, so that the position of the problem data can be located when an error is detected.

[0060] After obtaining the data block, the memory controller will recalculate the checksum of the data block, that is, read the checksum, so as to determine whether the data has changed during storage. The target algorithm is the same algorithm as that for generating the storage checksum, such as parity check, cyclic redundancy check, error correction code, etc.

[0061] After generating the read checksum, the read checksum can be compared with the storage checksum previously read from the memory. The storage checksum represents the original state of the data block when it is written into the memory, while the read checksum reflects the current actual state of the data block. If the storage checksum and the read checksum are consistent, it means that the data block has not changed during storage and reading, and the data is accurate. If they are inconsistent, it indicates that there is an error in the data block, such as physical damage to the memory hardware, bit flipping caused by electromagnetic interference, etc.

[0062] When the storage checksum and the read checksum are consistent, it can be determined that a memory error has been detected, and information such as the memory address where the data block is located, the time when the error occurred, and the specific values of the storage checksum and the read checksum is recorded.

[0063] In this embodiment, by reading the data block and the storage checksum, and calculating the read checksum using the target algorithm, the integrity and consistency of the data can be verified. By comparing the storage checksum and the read checksum, problems in the memory can be discovered in a timely manner, improving the detection efficiency.

[0064] In some embodiments, the method further includes: When the memory error type is a single-bit error, locate the target position where the single-bit error occurs according to the storage checksum and the read checksum; Correct the bit value at the target position.

[0065] In this embodiment, the storage checksum and the read checksum can be compared bit by bit to find all the mismatched check bits. These mismatched check bits form a binary code, and the position in the data block corresponding to this code is the target position where the single-bit error occurs. For example, if the 3rd, 5th, and 7th check bits are in error, the binary numbers corresponding to these three positions (such as 011, 101, 111) are combined and calculated to determine the specific offset position of the error bit in the data block.

[0066] Since a single-bit error only involves the state flip of one bit, that is, from 0 to 1 or from 1 to 0, the corresponding data block address can be accessed according to the target position information obtained by localization, and the "bit flip" operation can be performed through a hardware circuit. For example, if an error is detected in the 12th bit of the data block, a level signal can be sent to the memory cell where the bit is located to invert the current state of the bit. For example, if the current state of the bit is 0, it becomes 1 after inversion; if the current state of the bit is 1, it becomes 0 after inversion. In this way, the correction of the single-bit error is completed, and the correction operation is executed at high speed by a dedicated circuit at the hardware level, usually within nanoseconds, and hardly affects the system performance.

[0067] In some embodiments, after the correction is completed, the storage checksum of the data block can be recalculated and the new storage checksum can be updated to the memory, so that the subsequent data reading can pass the check verification.

[0068] In this embodiment, the error location is accurately located through the storage checksum and the read checksum, and the bit value at the target position is corrected. The hardware-level correction operation can automatically repair the error without interrupting the service operation, improving the repair efficiency.

[0069] In some embodiments, obtaining a spare memory page from the spare memory pool includes: Determining the damage degree of the target memory page according to the memory error information; Obtaining a spare memory page from the spare memory pool when the damage degree is less than the preset degree.

[0070] In this embodiment, the memory error information may also include the location where the error occurs, the type of memory error, the frequency of error occurrence, etc. If it is analyzed that the target memory page frequently has multi-bit errors, or the bits involved in the error change multiple times in a short time, it is possible that the target memory page has relatively serious physical damage.

[0071] In some embodiments, the damage degree of the target memory page can be comprehensively evaluated by combining information such as the type of memory error, the frequency of error occurrence, and the duration. For example, the damage degree of the target memory page can be quantified by a score. For example, the more the number of error bits and the higher the frequency of error occurrence, the higher the score, and the higher the score, the higher the damage degree.

[0072] When the score is less than the preset value, it means that the damage degree is less than the preset degree, which means that although there are certain problems with the target memory page, the severity of the error is within an acceptable range, and a spare memory page can be obtained from the spare memory pool to replace the damaged target memory page.

[0073] In this embodiment, by evaluating the degree of damage, different processing strategies can be adopted according to the severity of the error, thereby reducing unnecessary resource waste and improving resource utilization.

[0074] In some embodiments, determining the degree of damage of the target memory page according to the memory error information includes: Determining the location where the error occurs and the number of bits in error according to the memory error information; Determining the degree of damage of the target memory page according to the location where the error occurs and the number of bits in error.

[0075] In this embodiment, the memory error information may further include the number of bits involved in the error and the location of the error.

[0076] The more the number of bits in error, the wider the range of the error involved in the target memory page and the higher the degree of damage.

[0077] And the location of the error also affects the degree of damage. For example, when multiple error locations are concentrated in a certain area, the degree of damage is higher than that of a scattered distribution; when the location of the error is in the area where important data or system instructions are stored in the target memory page, the degree of damage is higher than that in other areas.

[0078] In this embodiment, by combining the location of the error and the number of bits in error, and evaluating the degree of damage based on multi-dimensional data analysis, the accuracy of the degree of damage evaluation can be improved.

[0079] In some embodiments, the method further includes: Generating an error log according to the memory error information when the degree of damage is greater than a preset degree; Feeding back the error log to the user.

[0080] In this embodiment, when it is evaluated that the degree of damage of the target memory page is greater than the preset degree, it indicates that the current memory error may be relatively serious and exceeds the scope of automatic repair or replacement of the spare memory page. In this case, the user can be notified for further processing.

[0081] Specifically, a detailed error log can be generated according to the memory error information. For example, the memory error information can be deeply analyzed and structurally integrated, and the core elements in the memory error information, such as the time, location, number of bits involved, and the evaluation result of the degree of damage, can be extracted, and the possible error causes can also be analyzed. The generated error log can be stored in a standardized JSON format.

[0082] After generating the error log, the error log can be fed back to the user. Specifically, the error log can be sent to the user in various ways, such as through a log interface, email, instant messaging, or a dedicated monitoring tool, etc. The content of the feedback can include not only the error log itself, but also some preliminary processing suggestions or tips, such as suggesting the user to check the hardware, replace the memory module, or contact technical support, etc. By promptly feeding back the error log, the user can quickly understand the severity of the memory error and take appropriate actions based on the provided information, thereby reducing the impact of the failure on the system operation.

[0083] In this embodiment, by generating a detailed error log and feeding it back to the user, rich fault information is provided to the user, enabling the user to locate the root cause of the problem and take targeted repair measures, thereby improving the efficiency of fault troubleshooting.

[0084] In some embodiments, before mapping the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page, it further includes: Identifying a target process related to the target memory page according to the memory address mapping table and the memory mapping information of the process; the memory mapping information stores the virtual addresses of the memory pages used by the process; Detecting the state of the target process, and suspending the operation of the target process when the state of the target process is the running state.

[0085] In this embodiment, in the server system, the memory page is the basic unit of memory management, and each process will have its own memory mapping information, which records the virtual addresses of the memory pages used by the process. When it is detected that the target memory page has an error and it is ready to replace the target memory page with the spare memory page, in order to reduce the problem of data inconsistency caused by the process accessing the memory during the repair process, it is necessary to suspend the process related to the target memory page.

[0086] Specifically, the virtual address corresponding to the physical address of the target memory page can be queried from the memory address mapping table, and the memory mapping information of all processes can be traversed to compare this virtual address with the virtual addresses of the memory pages used by each process. Since the memory mapping information of the process is efficiently organized in the form of a data structure (such as a red-black tree or a hash table), this comparison process can be completed quickly. For example, it may be possible to find out those processes that use the same virtual address as the target memory page from hundreds or thousands of processes within milliseconds. The process that uses the same virtual address as the target memory page is the target process related to the target memory page.

[0087] After identifying the target processes, it is necessary to detect the current states of these target processes. The process states can include various states such as running state, ready state, blocked state, etc. Since the processes in the running state are executing instructions on the CPU and directly accessing memory resources, it may affect the repair process. Specific signals can be sent to the target processes to suspend the running of the target processes, thereby reducing data inconsistency or system instability caused by the read and write operations of the processes.

[0088] In this embodiment, through the memory address mapping table and the memory mapping information of the process, the target processes related to the target memory page can be quickly identified, and the target processes in the running state can be suspended, reducing the problem of data inconsistency or system instability caused by the conflict between the repair process and process access, and improving the stability of the memory fault repair process.

[0089] In some embodiments, after mapping the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page, it further includes: Detecting the mapping relationship in the memory address mapping table; When there is no error in the mapping relationship, resuming the running of the target process.

[0090] In this embodiment, after completing the mapping of the physical address of the target memory page to the physical address of the spare memory page, the accuracy of the memory address mapping table can be verified to determine whether the repair is successful. For example, it can start from traversing all entries of the memory address mapping table and check one by one whether the physical address corresponding to each virtual address is consistent with the actual storage location. It can also verify whether the physical address is within the range of valid memory space. It can also check whether there are conflicts in the mapping relationship, such as whether a virtual address is wrongly mapped to multiple physical addresses, or whether multiple virtual addresses wrongly point to the same physical address.

[0091] When it is determined that there is no error in the mapping relationship in the memory address mapping table, a resume signal can be sent to the target process to inform the target process that it can continue to execute the unfinished tasks.

[0092] In some embodiments, the results and time of the repair process can also be recorded to generate a repair log for subsequent optimization of the repair strategy.

[0093] In this embodiment, by detecting the correctness of the memory address mapping table, it can accurately identify whether the memory fault is successfully repaired. Resuming the running of the target process can seamlessly resume from the suspended state to the normal running state, thereby reducing the impact on system performance and user experience.

[0094] Such as Figure 2As shown in the figure, the present application also provides a memory fault repair system, including: a memory controller 211, a repair module 212, a baseboard management controller 220, and an operating system 230; The memory controller 211 is configured to detect a memory error and send memory error information to the baseboard management controller 220 based on the detected memory error; The baseboard management controller 220 is configured to locate a target memory page with a fault according to the memory error information, obtain a spare memory page allocated by the operating system 230 from a spare memory pool, generate a repair instruction based on the physical address of the target memory page and the physical address of the spare memory page, and send the repair instruction to the repair module; The repair module 212 is configured to map the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page based on the repair instruction; the memory address mapping table is used to record the mapping relationship between the virtual address and the physical address of the memory page.

[0095] In this embodiment, the memory controller 211 and the repair module 212 are hardware modules 210, which are responsible for detecting memory errors and performing repair operations on memory errors.

[0096] Specifically, the memory controller 211 can be integrated in the CPU or an independent chip. The method for the memory controller 211 to detect memory errors can refer to the above-mentioned embodiment of the memory fault repair method. The advantage of the memory controller 2112 is that it can detect memory errors in real time and reduce the risk of system crashes or data loss.

[0097] The baseboard management controller 220 is responsible for coordinating and managing the memory fault repair process in the system, and can communicate with the operating system 230 and the hardware module 210 through IPMI (Intelligent Platform Management Interface) or other communication protocols.

[0098] The baseboard management controller 220 can receive memory error information through a hardware interface and formulate a repair strategy according to the error type and severity. For example, for a target memory page with a high degree of damage, it can be determined as an irreparable error, and an error date is generated and sent to the user; for a target memory page with a low degree of damage, it can be determined as a repairable error, and the repair strategy can be determined as selecting a spare memory page and updating the memory address mapping table to complete the repair.

[0099] The baseboard management controller 220 can send a repair instruction to the repair module 212 through a hardware interface and notify the operating system 230 to coordinate. After receiving the repair instruction, the repair module 212 can map the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page based on the repair instruction.

[0100] In some embodiments, the operating system 230 may communicate with the hardware module 210 and the baseboard management controller 220 through kernel modules.

[0101] The operating system 230 may be responsible for managing memory resources to ensure the availability of spare memory pages in the spare memory pool. When the baseboard management controller 220 needs a spare memory page, it sends a request to the operating system 230. After receiving the request, the operating system 230 may allocate an available spare memory page from the spare memory pool to the baseboard management controller 220.

[0102] In some embodiments, the operating system 230 may also recycle damaged target memory pages through memory resource recovery logic, thereby optimizing the resources in the spare memory pool.

[0103] In some embodiments, to reduce the impact during the repair process, the operating system 230 may identify the target process associated with the target memory page according to the memory address mapping table and the memory mapping information of the process, and pause the operation of the target process when the state of the target process is the running state. When the repair is completed, the operation of the target process is resumed.

[0104] Through this application, since the repair process does not rely on the trigger of SMI, system interruptions are reduced, and the system performance of the server is improved. Through the collaborative work of the baseboard management controller hardware and the operating system, memory errors can be detected and repaired in a timely manner, reducing problems such as system crashes or data loss, maintaining the stability of the server, and improving the reliability of the system.

[0105] The working process of the memory fault repair system of this application is introduced below through a scenario example.

[0106] As Figure 3 shown, the working process starts from a memory error. When the memory controller detects a memory error, it triggers the subsequent processing mechanism.

[0107] Facing a memory error, first, the error type is determined to distinguish whether it is a single-bit error or a multi-bit error. A single-bit error is relatively minor. The memory controller can automatically detect and repair it with its own error checking and correction function without complex external intervention, quickly restoring the correctness of the memory data and ensuring the basic operation of the system.

[0108] If it is determined to be a multi-bit error, the baseboard management controller (BMC) intervenes. The BMC first obtains the error, collects detailed information such as the physical address and error code where the error occurs, and analyzes the error. It can combine the memory hardware status, historical error records, etc. to judge the root cause and scope of the failure, such as determining whether it is a local damage of the memory chip or a serious deviation in the address mapping.

[0109] After the analysis is completed, the BMC will determine whether the memory can be repaired. If it is determined that the memory cannot be repaired, it means that the memory failure has exceeded the system's automatic repair ability. The BMC can record the error log and notify the system administrator, write the details of the failure into the log for retention, and send an alarm to the administrator to remind the administrator to intervene manually.

[0110] If it is determined that the memory can be repaired, the BMC will start the repair preparation, select a spare memory page from the spare memory pool, and send a repair instruction to the hardware module. The hardware module performs the repair. According to the BMC repair instruction, it replaces the physical address of the faulty target memory page in the memory address mapping table with the physical address of the spare memory page to complete the repair action at the underlying hardware level.

[0111] After the repair action is completed, the repair result is fed back to the BMC and the operating system so that the BMC and the operating system can know the repair status. When the operating system receives the feedback, it will resume the relevant processes and restart the processes that were paused or affected due to the memory failure, so that the business program can run normally again.

[0112] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0113] The memory failure repair method provided by the embodiments of the present application may be executed by a memory failure repair device. In the embodiments of the present application, taking the memory failure repair device executing the memory failure repair method as an example, the memory failure repair device provided by the embodiments of the present application is described.

[0114] As Figure 4 shown, the memory failure repair device includes: A first acquisition module 410, configured to acquire memory error information sent by a memory controller; the memory error information includes the location where the error occurs; A positioning module 420, configured to locate a target memory page with a failure according to the memory error information; A second acquisition module 430, configured to acquire a spare memory page from a spare memory pool; A mapping module 440, configured to map the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page; the memory address mapping table is used to record the mapping relationship between the virtual address and the physical address of the memory page.

[0115] Through this application, by obtaining the memory error information sent by the memory controller, the faulty target memory page can be located, and a spare memory page can be obtained from the spare memory pool. By utilizing the pre-prepared spare resources and updating the memory address mapping table, the physical address of the faulty target memory page is mapped to the physical address of the spare memory page. The repair process does not rely on the triggering of the system management interrupt mechanism, achieving memory repair during the system operation phase and reducing the system downtime caused by memory fault repair. Therefore, the problem of high latency in memory fault handling can be solved, achieving the technical effect of improving the memory fault repair efficiency.

[0116] In some embodiments, the memory fault repair device may further include: A detection module, configured to detect the memory through the memory controller; in the case of detecting a memory error, identify the memory error type; and in the case where the memory error type is a multi-bit error, generate memory error information.

[0117] In some embodiments, the detection module may further be configured to: Detect the mapping relationship in the memory address mapping table; In the case where the mapping relationship is incorrect, determine that a memory error has been detected.

[0118] In some embodiments, the detection module may further be configured to: Read a data block and the corresponding storage check code from the memory; the storage check code is a check code generated for storing the data block; Calculate the read check code of the data block using a target algorithm; In the case where the storage check code and the read check code do not match, determine that a memory error has been detected.

[0119] In some embodiments, the detection module may further be configured to: In the case where the memory error type is a single-bit error, locate the target position where the single-bit error occurs based on the storage check code and the read check code; Correct the bit value at the target position.

[0120] In some embodiments, the second acquisition module 430 may further be configured to: Determine the degree of damage of the target memory page according to the memory error information; In the case where the degree of damage is less than a preset degree, obtain a spare memory page from the spare memory pool.

[0121] In some embodiments, the second acquisition module 430 may further be configured to: Determine the position where the error occurs and the number of error bits according to the memory error information; Determine the degree of damage to the target memory page according to the location where the error occurs and the number of bits in error.

[0122] In some embodiments, the second acquisition module 430 may further be configured to: Generate an error log according to the memory error information when the degree of damage is greater than a preset degree; Feed back the error log to the user.

[0123] In some embodiments, the mapping module 440 may further be configured to: Identify a target process associated with the target memory page according to the memory address mapping table and the memory mapping information of the process; the memory mapping information stores the virtual addresses of the memory pages used by the process; Detect the state of the target process, and pause the operation of the target process when the state of the target process is the running state.

[0124] In some embodiments, the mapping module 440 may further be configured to: Detect the mapping relationship in the memory address mapping table; Resume the operation of the target process when the mapping relationship has no error.

[0125] For the description of the features in the corresponding embodiments of the memory fault repair device, reference may be made to the relevant description of the corresponding embodiments of the memory fault repair method, which will not be elaborated here one by one.

[0126] An embodiment of the present application further provides an electronic device, as Figure 5 shown, including a memory 502 and a processor 501. The memory 502 stores a computer program, and the processor 501 is configured to run the computer program to execute the steps in any one of the above embodiments of the memory fault repair method.

[0127] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any one of the above embodiments of the memory fault repair method when running.

[0128] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.

[0129] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the steps in any of the above-described embodiments of the memory fault repair method.

[0130] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program which, when executed by a processor, implements the steps in any of the above-described embodiments of the memory fault repair method.

[0131] Those skilled in the art can further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0132] The above has introduced in detail a memory fault repair method, device, equipment, medium, and computer program product provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for repairing memory faults, characterized in that, Including: Obtain memory error information sent by a memory controller; The memory error information includes the location where the error occurs; Locate a target memory page with a fault according to the memory error information; Obtain a spare memory page from a spare memory pool; Map the physical address of the target memory page in a memory address mapping table to the physical address of the spare memory page; The memory address mapping table is used to record the mapping relationship between the virtual address and the physical address of a memory page.

2. The method according to claim 1, characterized in that, Before obtaining the memory error information sent by the memory controller, including: Detect the memory through the memory controller; When a memory error is detected, identify the memory error type; When the memory error type is a multi-bit error, generate memory error information.

3. The method according to claim 2, wherein The detecting of the memory includes: Detect the mapping relationship in the memory address mapping table; When there is an error in the mapping relationship, determine that a memory error is detected.

4. The method according to claim 2, wherein The detecting of the memory includes: Read a data block and a stored check code corresponding to the data block from the memory; the stored check code is a check code generated for storing the data block; Calculate a read check code of the data block by using a target algorithm; When the stored check code and the read check code are inconsistent, determine that a memory error is detected.

5. The method according to claim 4, wherein The method further includes: When the memory error type is a single-bit error, locate a target position where the single-bit error occurs according to the stored check code and the read check code; Correct the bit value at the target position.

6. The method according to claim 1, characterized in that The obtaining of the spare memory page from the spare memory pool includes: Determine the degree of damage of the target memory page according to the memory error information; When the degree of damage is less than a preset degree, obtain a spare memory page from the spare memory pool.

7. The method according to claim 6, characterized in that The determining of the degree of damage of the target memory page according to the memory error information includes: Determine the location where the error occurs and the number of error bits according to the memory error information; Determine the degree of damage of the target memory page according to the location where the error occurs and the number of error bits.

8. The method according to claim 6, wherein The method further includes: When the degree of damage is greater than a preset degree, generate an error log according to the memory error information; Feed back the error log to a user.

9. The method according to claim 1, wherein Before mapping the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page, further including: Identify a target process related to the target memory page according to the memory address mapping table and memory mapping information of a process; the memory mapping information stores the virtual address of a memory page used by the process; Detect the state of the target process, and when the state of the target process is a running state, suspend the running of the target process.

10. The method according to claim 9, characterized in that, After mapping the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page, further including: Detect the mapping relationship in the memory address mapping table; When there is no error in the mapping relationship, resume the running of the target process.

11. A memory failure repair system, characterized in that, Including: A memory controller, a repair module, a baseboard management controller, and an operating system; The memory controller is configured to detect memory errors and send memory error information to the baseboard management controller based on the detected memory errors; The baseboard management controller is configured to locate a target memory page with a fault according to the memory error information, obtain a spare memory page allocated by the operating system from a spare memory pool, generate a repair instruction based on the physical address of the target memory page and the physical address of the spare memory page, and send the repair instruction to the repair module; The repair module is configured to map the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page based on the repair instruction; The memory address mapping table is used to record the mapping relationship between the virtual address and the physical address of a memory page.

12. A memory fault repair device, characterized in that, Comprising: A first acquisition module, configured to acquire the memory error information sent by the memory controller; The memory error information includes the location where the error occurs; A location module, configured to locate a target memory page with a fault according to the memory error information; A second acquisition module, configured to acquire a spare memory page from the spare memory pool; A mapping module, configured to map the physical address of the target memory page in the memory address mapping table to the physical address of the spare memory page; The memory address mapping table is used to record the mapping relationship between the virtual address and the physical address of a memory page.

13. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the memory fault repair method according to any one of claims 1 to 10 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the memory fault repair method according to any one of claims 1 to 10 when executed by a processor.

15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the memory fault repair method according to any one of claims 1 to 10 when executed by a processor.

Citation Information

Patent Citations

  • Method and system for managing secure memory, device and storage medium

    CN111984374A

  • Memory fault information determination method and device

    CN114860432A

  • Memory repair method and device

    CN115129461A

Cited By

  • SPD fault processing method and SPD fault processing system

    CN120596305A

  • SPD fault processing method and SPD fault processing system

    CN120596305B

  • Fault detection system and method, electronic device, storage medium and program product

    CN120821605A

  • Method for recovering memory data in high-temperature environment

    CN121116711A

  • A memory data recovery method in a high temperature environment

    CN121116711B