Memory error processing method and device, equipment and storage medium

By using hardware-defined flip rules in the memory system to select co-memory blocks and record and save fault information, the system stability problem caused by the failure to handle memory block errors in a timely manner is solved, and the effect of quickly positioning errors and reducing the risk of underreport is achieved.

CN120492196APending Publication Date: 2025-08-15INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510517856.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, correctable errors of memory blocks are not processed in time, resulting in system failure and data loss, and the global counter cannot record error information in detail, increasing the difficulty of system maintenance.

Method used

Select the co-memory blocks logically associated with the main memory block through the hardware-defined flip rules, record the fault information and trigger the first correction error reporting mechanism, save the error record data to the lock register of the co-memory block, and avoid dynamically adjusting the main memory block mapping relationship.

Benefits of technology

It effectively reduces the risk of errors in main memory blocks, improves system stability, and helps engineers quickly locate the root cause of the problem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492196A_ABST
    Figure CN120492196A_ABST
Patent Text Reader

Abstract

The invention discloses a memory error processing method and device, equipment and a storage medium, and relates to the technical field of computers, when a main memory block triggers a preset fault event, a collaborative memory block logically associated with the main memory block is determined, fault information data of the main memory block is written into the collaborative memory block, and the fault information data of the main memory block is written into the collaborative memory block; the cooperative memory block triggers a mechanism for correcting an error report for the first time according to the fault information data, the error record data is locked after the error record data is stored in a locking register corresponding to the cooperative memory block, and the mapping relation between the main memory block and the cooperative memory block is based on an association rule of hardware. The mapping relation of the main memory block does not need to be dynamically adjusted by introducing extra delay, and the multiple collaborative memory blocks can record multiple correctable errors of the main memory block, so that an engineer can quickly position the root of a problem, the error report omission risk of the main memory block is effectively reduced, and the stability of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a memory error handling method, apparatus, device, and storage medium. Background Art

[0002] During the operation of a computer system, memory is prone to correctable errors due to factors such as electrical interference. If these errors are not handled in a timely manner, they may evolve into uncorrectable errors, resulting in system failure and data loss. Therefore, timely handling of correctable errors can ensure the stable operation of the system.

[0003] In the related art, when a correctable error is detected in a memory block for the first time, a lock register is used to record detailed information about the error. However, when a correctable error occurs in the memory block again, it is not recorded in the lock register, but the number of errors in the memory block is recorded in a global counter. Since the global counter cannot record error information in detail, it increases the difficulty of system maintenance. Summary of the Invention

[0004] The present application provides a memory error handling method, apparatus, device and storage medium to at least solve the problem of difficult system maintenance in related technologies.

[0005] The present application provides a memory error handling method, including: continuously monitoring the operating status of a main memory block, and judging whether the main memory block triggers a preset fault event based on the operating status, wherein the main memory block is a memory block that is reading and writing data in multiple storage modules, and one storage module includes multiple memory blocks; if it is determined that the main memory block triggers a preset fault event, determining a collaborative memory block logically associated with the main memory block; obtaining fault information data of the main memory block, and writing the fault information data to the collaborative memory block; triggering a first correction error reporting mechanism in the collaborative memory block based on the fault information data, and generating error record data; saving the error record data to a lock register corresponding to the collaborative memory block, and locking the lock register.

[0006] The present application also provides a memory error handling device, comprising:

[0007] The operation status monitoring module is used to continuously monitor the operation status of the main memory block and determine whether the main memory block triggers a preset fault event based on the operation status, where the main memory block is a memory block that is reading and writing data in multiple storage modules, and one storage module includes multiple memory blocks.

[0008] The collaborative memory block acquisition module is used to determine the collaborative memory block logically associated with the main memory block if it is determined that the main memory block triggers a preset fault event.

[0009] The fault information sending module is used to obtain the fault information data of the main memory block and write the fault information data into the coordinated memory block.

[0010] The error record generation module is used to trigger the first correction error reporting mechanism in the collaborative memory block according to the fault information data and generate error record data.

[0011] The memory error locking module is used to save the error record data to the locking register corresponding to the collaborative memory block and lock the locking register.

[0012] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any one of the above-mentioned memory error handling methods when executing the computer program.

[0013] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned memory error handling methods are implemented.

[0014] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned memory error handling methods when executed by a processor.

[0015] Through the memory error handling method, device, equipment and storage medium of the present application, when the main memory block triggers a preset fault event, the collaborative memory block logically associated with the main memory block is determined according to the pre-defined flipping rules, and the fault information data of the main memory block is written into the collaborative memory block. The collaborative memory block triggers the first correction error reporting mechanism according to the fault information data, and locks the error record data after saving the error record data to the lock register corresponding to the collaborative memory block. The mapping relationship between the main memory block and the collaborative memory block is based on the hardware flipping rules, and there is no need to introduce additional delays to dynamically adjust the mapping relationship of the main memory block. Multiple collaborative memory blocks can record multiple correctable errors of the main memory block, which helps engineers quickly locate the root cause of the problem, effectively reduces the risk of missed errors in the main memory block, and improves the stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 A schematic diagram of a memory error handling method according to an embodiment of the present invention;

[0018] Figure 2 A flowchart of a memory error handling method provided in an embodiment of the present application;

[0019] Figure 3 A flowchart of a method for determining a collaborative memory block provided in an embodiment of the present application;

[0020] Figure 4 A schematic diagram of the structure of a memory error handling device provided in an embodiment of the present application;

[0021] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0024] In order to clearly understand the technical solution of the present application, the solution of the prior art is first introduced in detail. During the operation of a computer system, the memory is susceptible to correctable errors due to factors such as electrical interference. If such errors are not handled in a timely manner, they may evolve into uncorrectable errors, resulting in system failure and data loss. Therefore, timely handling of correctable errors can ensure the stable operation of the system. In the related art, when a correctable error of a memory block is detected for the first time, a lock register is used to record the detailed information of the error. However, when a correctable error occurs again in the memory block, it will not be recorded in the lock register, but the number of errors of the memory block will be recorded in a global counter. Since the global counter cannot record error information in detail, it increases the difficulty of system maintenance.

[0025] To address the aforementioned technical issues, the inventors devised a method to utilize hardware-defined flipping rules to select a logically associated collaborative memory block as a collaborative unit for the main memory block experiencing a correctable error. This eliminates the need to dynamically adjust the mapping relationship of the main memory block and reduces the introduction of additional latency. First, the correctable error is transferred to the collaborative memory block of the main memory block. The first-correction error reporting mechanism is then triggered in the collaborative memory block to lock the lock register corresponding to the correctable error recorded in the collaborative memory block. This allows the error information data for multiple correctable errors in the main memory block to be recorded, helping engineers quickly locate the root cause of the problem, effectively reducing the risk of missed errors in the main memory block, and improving system stability.

[0026] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0027] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the memory error handling method depends, the specific application environment architecture or specific hardware architecture is described here. Figure 1 , Figure 1 Schematic diagram of a memory error handling method provided in an embodiment of the present application. Figure 1 As shown, the scenario includes: an electronic device 101, a main memory block 102, a cooperative memory block 103 and a locking register 104 corresponding to the cooperative memory block.

[0028] Specifically, the electronic device 101 continuously monitors the operating status of the main memory block 102, and determines whether the main memory block 102 has triggered a preset fault event based on the operating status of the main memory block 102; if it is determined that the main memory block 102 has triggered a preset fault event, the collaborative memory block 103 logically associated with the main memory block 102 is determined, and the electronic device 101 obtains the fault information data of the main memory block 102, and writes the fault information data to the collaborative memory block 103, triggers the first correction error reporting mechanism in the collaborative memory block 103 based on the fault information data, generates error record data, and saves the error record data to the lock register 104 corresponding to the collaborative memory block 103, and locks the lock register 104 at the same time.

[0029] Figure 2 This is a flow chart of a memory error handling method provided in an embodiment of the present application. Figure 2 As shown, an embodiment of the present application provides a memory error handling method, which is described in detail as follows:

[0030] S201: Continuously monitor the operating status of a main memory block, and determine whether the main memory block triggers a preset fault event based on the operating status, wherein the main memory block is a memory block that is reading and writing data in multiple storage modules, and one storage module includes multiple memory blocks.

[0031] The preset fault event may be an error condition that can be automatically corrected occurring in the main memory block, or when an action related to reliability, availability, and maintainability needs to be performed.

[0032] Specifically, during system operation, the operating status of the main memory block needs to be monitored in real time and continuously. For example, when the response time of memory read and write operations exceeds the normal range, or when the memory data verification shows an error, it is determined that the main memory block has triggered a preset fault event.

[0033] S202: If it is determined that the main memory block triggers a preset fault event, a coordinated memory block logically associated with the main memory block is determined.

[0034] Specifically, when it is determined that the main memory block has triggered a preset fault event, the pre-defined association rules are used to determine the collaborative memory block that is logically associated with it. The pre-defined association rules are a mapping relationship set based on the logical association of multiple memory blocks. When the main memory block fails, according to the pre-defined association rules, it can be determined which collaborative memory block to select to assist in processing.

[0035] S203: Acquire the fault information data of the main memory block, and write the fault information data into the cooperative memory block.

[0036] The fault information data can reflect the specific circumstances of the main memory block fault, including the event at which the fault occurred, the type of fault, such as parity error, correctable error, and the memory address where the fault occurred.

[0037] Specifically, after the cooperative memory block is determined, the fault information data of the main memory block is obtained by reading the status register of the main memory block or other related hardware information, and then the fault information data is written into the cooperative memory block.

[0038] S204: triggering a first correction error reporting mechanism in the collaborative memory block according to the fault information data, and generating error record data.

[0039] Specifically, after the fault information data is written to the collaborative memory block, the collaborative memory block triggers the first-correction error reporting mechanism based on this fault information data, corrects the preset fault event, and generates detailed error log data. The first-correction error reporting mechanism includes various processing methods, such as performing a parity check on the fault data and using a correctable error algorithm for error correction. During the processing, detailed information about the fault is recorded, including the fault handling result and whether the correction was successful, ultimately generating the error log data.

[0040] Specifically, the process of triggering the first correction error reporting mechanism in the collaborative memory block according to the fault information data includes:

[0041] Sa1: Control the memory controller to read fault information data from the cooperative memory block, where the fault information data includes memory data and a check code.

[0042] Among them, the memory controller is the component in the computer system responsible for managing memory access and coordinating data read and write operations.

[0043] Memory data refers to the actual data stored in the main memory block. A checksum is additional information generated by a specific algorithm when data is written to memory. It is used to detect and correct errors that may occur during data transmission or storage. Common checksum generation algorithms include parity check and cyclic redundancy check.

[0044] Specifically, the memory controller sends a read request to the cooperative memory block through the memory bus according to the address information of the cooperative memory block. After receiving the request, the cooperative memory block returns the stored fault information data to the memory controller through the memory bus.

[0045] Sa2: Use a preset Hamming code algorithm to process the fault information data to determine the error bits in the memory data and correct the memory data according to the check code.

[0046] Hamming code is a coding method with error correction capabilities. It implements error detection and correction by inserting additional check bits into the original data. When an error occurs in the data, the location of the error can be determined by calculating and comparing the check bits.

[0047] Specifically, the memory controller passes the read memory data and checksum to a processing module that uses a Hamming code algorithm. Based on the rules of the Hamming code, the processing module recalculates the checksum corresponding to the current memory data. The calculated checksum is then compared with the checksum read from the cooperating memory block. If the two differ, an error exists in the memory data. By comparing the checksum differences, the specific bit location of the error can be determined. After determining the location of the erroneous bit, the processing module modifies the memory data accordingly. This modification corrects the memory data to its correct state, thereby correcting the memory fault.

[0048] S205: Save the error record data to the lock register corresponding to the cooperative memory block, and lock the lock register.

[0049] The lock register is a special register used to store important information, and is locked after the error record data is written to ensure the integrity and security of the error record data.

[0050] Specifically, after the error record data is locked in the lock register, the error is stored in the error report. The storage method of the error report is divided into hardware-level reporting and software-level reporting. The hardware-level report triggers an exception through a designated pin of the processor, such as the MC# pin, and transmits the error to the baseboard management controller, and records it through the system event log or the integrated device list log. The software-level report mainly reads the exception information in the memory controller status register by the basic input and output system, then records the error in the advanced configuration and power interface table, and updates the memory mapping table to disable the faulty collaborative memory block. In addition, a notification is sent to the operating system through the advanced configuration and power interface. For example, the record is in the / sys / devices / system / edac / mc / mc0 / ce_count path under the Linux system directory.

[0051] In summary, when the main memory block triggers a preset fault event, the collaborative memory block logically associated with the main memory block is determined, and the fault information data of the main memory block is written into the collaborative memory block. The collaborative memory block triggers the first correction error reporting mechanism based on the fault information data, and locks the error record data after saving the error record data to the lock register corresponding to the collaborative memory block. The mapping relationship between the main memory block and the collaborative memory block is based on the hardware association rules. There is no need to introduce additional delays to dynamically adjust the mapping relationship of the main memory block. Multiple collaborative memory blocks can record multiple correctable errors of the main memory block, which helps engineers quickly locate the root cause of the problem, effectively reduces the risk of missed errors in the main memory block, and improves the stability of the system.

[0052] The method for determining a collaborative memory block logically associated with a main memory block provided in an embodiment of the present application includes:

[0053] S2021: Obtain a preset flag bit of the main memory block, where the preset flag bit includes a first flag bit and a second flag bit.

[0054] Each flag bit is used to point to a corresponding memory block, and each flag bit has a one-to-one correspondence with the corresponding memory block. The preset flag bits include a first flag bit and a second flag bit, wherein the first flag bit is used to point to the storage module corresponding to the memory block, and the second flag bit is used to point to the location of the memory block in the storage module.

[0055] Specifically, before obtaining the preset flag bit of the main memory, the following is also included:

[0056] Sb1: Determine the number of the plurality of storage modules, and set the first flag bit of each memory block according to the number of the plurality of storage modules.

[0057] Sb2: Determine the number of the multiple memory blocks, and set the second flag bit of each memory block according to the number of the multiple memory blocks.

[0058] S2022: Flip the first flag bit to generate a first flip flag bit.

[0059] Illustratively, flipping the first flag bit may include adding 1 to the first flag bit to indicate a storage module where another memory block is located.

[0060] S2023: Combine the start flag bit and the first flip flag bit to obtain a new flag bit.

[0061] The start flag is in the state of all 0 bits.

[0062] Specifically, the first flip flag is concatenated with the start flag to obtain a new flag.

[0063] S2024: Determine the target memory block pointed to by the new flag.

[0064] S2025: Obtain lock status information of the lock register corresponding to the target memory block.

[0065] Specifically, the lock status information of the lock register corresponding to the target memory block is read to determine whether the lock register is in a locked state or an unlocked state.

[0066] S2026: If the lock register corresponding to the target memory block is in an unlocked state, the target memory block is determined as a collaborative memory block logically associated with the main memory block.

[0067] Specifically, after finding the collaborative memory block, if the main memory block occupies multiple memory sub-channels at the same time, the collaborative memory block will completely follow this data transmission mode when receiving the fault information data from the main memory and in the subsequent data interaction process. When the main memory block sends the fault information data to the collaborative memory block, it will be transmitted according to the previously used memory sub-channel combination and transmission method. The collaborative memory block will also receive the data in the same way to ensure that the data can be accurately transferred from the main memory block to the collaborative memory block.

[0068] S2027: If the locking register corresponding to the target memory block is in a locked state, iteratively execute the steps of "flipping the start flag to generate a second flip flag; combining the second flip flag with the first flip flag to obtain a new flag; determining the target memory block pointed to by the new flag; obtaining the locking state information of the locking register corresponding to the target memory block". If the locking register corresponding to the target memory block is determined to be in an unlocked state before the second flip flag becomes the start flag again, the target memory block is determined to be a collaborative memory block logically associated with the main memory block.

[0069] The second flip flag is used to indicate the position of the target memory block in the current storage module.

[0070] Specifically, if the lock register corresponding to the target memory block is in a locked state, the start flag is incremented to determine the second flip flag, the first flip flag and the second flip flag are concatenated to obtain a new flag indicating the target memory block, and the lock state information of the lock register corresponding to the target memory block is obtained. If the lock register corresponding to the target memory block is not locked, the target memory block is determined to be a collaborative memory block logically associated with the main memory block. If the lock register corresponding to the target memory block is locked, the above process is iteratively executed, and the second flip flag is continuously incremented and flipped until the target memory block is found, or the second flip flag is incremented and flipped until it becomes the start flag, and the iterative process ends.

[0071] Specifically, if the second flip flag is flipped incrementally and becomes the start flag, and the cooperative memory block is still not found, the cooperative memory block will be re-determined according to steps Sc1 to Sc6:

[0072] Sc1: If the second flip flag becomes the start flag again, the first flip flag is flipped to generate the third flip flag.

[0073] Specifically, when the second flip flag becomes the starting flag again, the first flip flag is flipped to obtain the third flip flag, indicating that no usable collaborative memory block is found in the storage module corresponding to the first flip flag, and the first flip flag is flipped incrementally, indicating that a collaborative memory block is selected in the next storage module.

[0074] Sc2: Combine the start flag and the third flip flag to obtain a new flag.

[0075] Specifically, in the storage module corresponding to the third flip flag, judgment is started from the first memory block, the third flip flag and the start flag are concatenated to determine a new flag.

[0076] Sc3: Determine the target memory block pointed to by the new flag.

[0077] Sc4: Get the lock status information of the lock register corresponding to the target memory block.

[0078] Specifically, the lock status information of the lock register corresponding to the target memory block is read to determine whether the lock register is in a locked state or an unlocked state.

[0079] Sc5: If the lock register corresponding to the target memory block is in an unlocked state, the target memory block is determined as a collaborative memory block logically associated with the main memory block.

[0080] Sc6: If the locking register corresponding to the target memory block is in a locked state, iteratively execute the steps of "flipping the start flag to generate a second flip flag; combining the second flip flag with the third flip flag to obtain a new flag; determining the target memory block pointed to by the new flag; obtaining the locking state information of the locking register corresponding to the target memory block". If the locking register corresponding to the target memory block is determined to be in an unlocked state before the second flip flag becomes the start flag again, the target memory block is determined to be a collaborative memory block logically associated with the main memory block.

[0081] Specifically, if the lock register corresponding to the target memory block is in a locked state, then continue to search for available memory blocks in the storage module corresponding to the third flip flag, flip the start flag, and generate a second flip flag; combine the second flip flag with the third flip flag to obtain a new flag; determine the target memory block pointed to by the new flag; obtain the lock status information of the lock register corresponding to the target memory block, and if the lock register corresponding to the target memory block is not locked, then determine the target memory block as a collaborative memory block logically associated with the main memory. If the target memory is locked, then iterate the above process and continue to flip the second flip flag until the target memory block is found, or the second flip flag is flipped incrementally and becomes the start flag, and the iterative process ends.

[0082] Specifically, if the second flip flag is flipped incrementally and becomes the start flag, and the collaborative memory block is still not found, the collaborative memory block will be re-determined according to the steps Sd1 to Sd2:

[0083] Sd1: If the second flip flag becomes the starting flag again, iteratively flip the third flip flag to generate a new third flip flag; combine the starting flag with the new third flip flag to obtain a new flag, and then perform the following steps.

[0084] Specifically, if the second flip flag becomes the start flag again, the search continues in other storage modules, and the process of iteratively querying in each storage module whether the lock register of each memory block is in the locked state is executed.

[0085] Sd2: If the lock register corresponding to the target memory block is determined to be in an unlocked state before the new third flip flag bit becomes the first flag bit again, the target memory block is determined to be a collaborative memory block logically associated with the main memory block.

[0086] Specifically, during the iterative execution of querying in each storage module whether the lock register of each memory block is in a locked state, if before the new third flip flag bit returns to the first flag bit, the lock register corresponding to the target memory block is determined to be in an unlocked state, the target memory is determined to be a collaborative memory block logically associated with the main memory block. If the new third flip flag bit returns to the first flag bit, it indicates that no usable memory block has been found in other storage modules, and the search continues in the storage module where the main memory block is located, executing steps Se1 to Se4:

[0087] Se1: If the new third flip flag becomes the first flag again, continue to determine whether the second flag is consistent with the start flag.

[0088] Specifically, if the new third flip flag becomes the first flag again, it means that no available memory block is found in other storage modules. The search will continue in the storage module where the main memory block is located based on the first flag to determine whether the second flag is consistent with the starting flag, that is, to determine whether the position of the main memory block in the storage module is the first memory block.

[0089] Se2: If the second flag is consistent with the start flag, iteratively execute the steps of "flipping the second flag to generate a fourth flip flag; combining the fourth flip flag with the new third flip flag to obtain a new flag; determining the target memory block pointed to by the new flag; and obtaining the lock status information of the lock register corresponding to the target memory block."

[0090] Se3: If the lock register corresponding to the target memory block is determined to be in an unlocked state before the fourth flip flag becomes the start flag again, the target memory block is determined to be a collaborative memory block logically associated with the main memory block.

[0091] Specifically, if the second flag bit is consistent with the start flag bit, that is, if the main memory block is the first memory block in the storage module, then the second flag bit is flipped to generate a fourth flip flag bit, and the new third flip flag bit is combined with the fourth flip flag bit to determine a new flag bit and the target memory block pointed to by the new flag bit. If the lock register corresponding to the target memory block is not locked, then the target memory block is determined to be a collaborative memory block logically associated with the main memory block. If the lock register corresponding to the target memory block is locked, the fourth flip flag bit is iteratively flipped, and the search for the target memory block pointed to by the new flag bit continues.

[0092] Se4: If the fourth flip flag becomes the start flag again, the main memory block is determined to be a collaborative memory block.

[0093] Specifically, if the fourth flip flag becomes the start flag again, it means that there is no usable collaborative memory block logically associated with the main memory block among the multiple memory blocks. The main memory block itself is set as a collaborative memory block logically associated with the main memory block, and the correctable fault information data that occurs is recorded in the lock register corresponding to the main memory block.

[0094] Sf1: If the second flag is inconsistent with the start flag, combine the start flag with the new third flip flag to obtain a new flag.

[0095] Sf2: Determine the target memory block pointed to by the new flag.

[0096] Sf3: Get the lock status information of the lock register corresponding to the target memory block.

[0097] Sf4: If the lock register corresponding to the target memory block is in an unlocked state, the target memory block is determined as a collaborative memory block logically associated with the main memory block.

[0098] Specifically, if the second flag bit is inconsistent with the start flag bit, that is, the main memory block is not the first memory block in the memory blocks, it is determined from the first memory block whether the corresponding lock register is locked.

[0099] Sf5: If the lock register corresponding to the target memory block is in a locked state, iteratively execute the steps of "flipping the start flag to generate a second flip flag; combining the second flip flag with the new third flip flag to obtain a new flag; determining the target memory block pointed to by the new flag; and obtaining the lock state information of the lock register corresponding to the target memory block."

[0100] Sf6: If the lock register corresponding to the target memory block is determined to be in an unlocked state before the second flip flag bit becomes the second flag bit, the target memory block is determined to be a collaborative memory block logically associated with the main memory block.

[0101] Specifically, if the first memory block is unusable, the judgment continues from the lock register state corresponding to the second memory block of the current storage module until the flag bit corresponding to the main memory block is flipped.

[0102] Sf7: If the second flip flag becomes the second flag, continue to iterate the steps of "flipping the second flip flag to generate the fifth flip flag; combining the fifth flip flag with the new third flip flag to obtain a new flag; determining the target memory block pointed to by the new flag; obtaining the lock status information of the lock register corresponding to the target memory block".

[0103] Sf8: If the lock register corresponding to the target memory block is determined to be in an unlocked state before the fifth flip flag becomes the start flag again, the target memory block is determined to be a collaborative memory block logically associated with the main memory block.

[0104] Specifically, if the flag bit corresponding to the main memory block is flipped, the main memory block is skipped, and the search continues in the memory blocks after the main memory block to determine whether there is a memory block that is not locked by the lock register.

[0105] Sf9: If the fifth flip flag becomes the start flag again, the main memory block is determined to be a collaborative memory block.

[0106] Specifically, if all memory blocks except the main memory block cannot be used, the main memory block itself is set as a collaborative memory block logically associated with the main memory block, and the correctable fault information data that occurs is recorded in the lock register corresponding to the main memory block.

[0107] In summary, setting flags for storage modules and memory blocks provides a clear identification and indexing basis for subsequent searches. When a main memory block triggers a fault, the available collaborative memory blocks are determined based on pre-defined association rules and flag combinations. If the collaborative memory block is unavailable, the search strategy is gradually adjusted. By iteratively flipping the first and second flags, the search continues across different combinations of storage modules and memory blocks. This multi-level, multi-strategy collaborative memory block search mechanism greatly increases the probability of finding available collaborative memory blocks in the event of a memory fault, ensuring the smooth progress of the memory error handling process and enhancing the reliability and stability of the entire memory system.

[0108] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0109] Figure 3 Schematic diagram of the process of determining the collaborative memory block provided in the embodiment of the present application. Figure 3 As shown:

[0110] S301: Obtaining a preset flag bit of the main memory, including a first flag bit and a second flag bit.

[0111] The first flip flag is obtained by flipping the first flag, and step S302 is performed: the first flip flag and the start flag are concatenated to obtain a new flag.

[0112] S303: Determine the target memory block according to the new flag bit, and obtain the locking status information of the locking register corresponding to the target memory block.

[0113] S304: If the lock register corresponding to the target memory block is not locked, the target memory block is determined as a cooperative memory block.

[0114] S305: If the lock register corresponding to the target memory block is locked, flip the start flag to obtain a second flip flag.

[0115] S306: Concatenate the first flip flag bit and the second flip flag bit to obtain a new flag bit.

[0116] S307: Determine the target memory block according to the new flag bit, and obtain the locking status information of the locking register corresponding to the target memory block.

[0117] S304: If the lock register corresponding to the target memory block is not locked, the target memory block is determined as a cooperative memory block.

[0118] S308: If the lock register corresponding to the target memory block is locked, determine whether the second flip flag is the start flag.

[0119] If the second flip flag is not the start flag, the second flip flag is flipped and steps S306 to S308 are executed cyclically.

[0120] If the second flip flag is the start flag, step S309 is executed to determine whether the first flip flag is the first flag.

[0121] If the first flip flag is not the first flag, steps S302 to S309 are executed in a loop.

[0122] If the first flip flag is the first flag, step S401 is executed: determining whether the second flag is the start flag.

[0123] If the second flag is the start flag, a third flip flag is obtained by flipping the second flag, and step S402 is performed: the first flip flag and the third flip flag are concatenated to obtain a new flag.

[0124] S403: Determine the target memory block according to the new flag bit, and obtain the locking status information of the locking register corresponding to the target memory block.

[0125] S404: If the lock register corresponding to the target memory block is not locked, the target memory block is determined as a cooperative memory block.

[0126] S405: If the lock register corresponding to the target memory block is locked, determine whether the third flip flag is a start flag.

[0127] If the third flip flag is not the start flag, steps S402 to S405 are executed in a loop.

[0128] S406: If the third flip flag is the start flag, the main memory block is determined as the collaborative memory block.

[0129] If the second flag is not the start flag, step S407 is executed: the first flip flag and the start flag are concatenated to obtain a new flag.

[0130] S408: Determine the target memory block according to the new flag bit, and obtain the locking status information of the locking register corresponding to the target memory block.

[0131] If the lock register corresponding to the target memory block is not locked, step S404 is executed: the target memory block is determined as a cooperative memory block.

[0132] If the lock register corresponding to the target memory block is locked, a fourth flip flag is obtained by flipping the start flag, and step S409 is executed: the first flip flag and the fourth flip flag are concatenated to obtain a new flag.

[0133] S410: Determine the target memory block according to the new flag bit, and obtain the locking status information of the locking register corresponding to the target memory block.

[0134] If the lock register corresponding to the target memory block is not locked, step S404 is executed: the target memory block is determined as a cooperative memory block.

[0135] S411: If the lock register corresponding to the target memory block is locked, determine whether the fourth flip flag is the second flag.

[0136] If the fourth flip flag is not the second flag, the fourth flip flag is flipped and steps S409 to S411 are executed in a loop.

[0137] If the fourth flip flag is the second flag, the fifth flip flag is obtained by flipping the fourth flip flag.

[0138] S412: Concatenate the first flip flag bit and the fifth flip flag bit to obtain a new flag bit.

[0139] S413: Determine the target memory block according to the new flag bit, and obtain the locking status information of the locking register corresponding to the target memory block.

[0140] If the lock register corresponding to the target memory block is not locked, step S404 is executed: the target memory block is determined as a coordinated memory block.

[0141] S414: If the lock register corresponding to the target memory block is locked, determine whether the fifth flip flag is the start flag.

[0142] If the fifth flip flag is not the start flag, the fifth flip flag is flipped and steps S412 to S414 are executed in a loop.

[0143] S415: If the fifth flip flag is the start flag, the main memory block is determined as the collaborative memory block.

[0144] In summary, setting flags for storage modules and memory blocks provides a clear identification and indexing basis for subsequent searches. When a main memory block triggers a fault, the available collaborative memory blocks are determined based on pre-defined association rules and flag combinations. If the collaborative memory block is unavailable, the search strategy is gradually adjusted. By iteratively flipping the first and second flags, the search continues across different combinations of storage modules and memory blocks. This multi-level, multi-strategy collaborative memory block search mechanism greatly increases the probability of finding available collaborative memory blocks in the event of a memory fault, ensuring the smooth progress of the memory error handling process and enhancing the reliability and stability of the entire memory system.

[0145] Figure 4 This is a schematic diagram of the structure of the memory error handling device provided in the embodiment of the present application. Figure 4 As shown, an embodiment of the present application also provides a memory error handling device, which includes an operation status monitoring module 401, a collaborative memory block acquisition module 402, a fault information sending module 403, an error record generation module 404 and a memory error locking module 405.

[0146] The operation status monitoring module 401 is used to continuously monitor the operation status of the main memory block and determine whether the main memory block triggers a preset fault event based on the operation status, where the main memory block is a memory block that is reading and writing data in multiple storage modules, and one storage module includes multiple memory blocks.

[0147] The coordinated memory block acquisition module 402 is configured to determine a coordinated memory block that is logically associated with the main memory block if it is determined that the main memory block triggers a preset fault event.

[0148] The fault information sending module 403 is used to obtain the fault information data of the main memory block and write the fault information data into the coordinated memory block.

[0149] The error record generating module 404 is configured to trigger a first correction error reporting mechanism in the collaborative memory block according to the fault information data, and generate error record data.

[0150] The memory error locking module 405 is used to save the error record data to the locking register corresponding to the cooperative memory block and lock the locking register.

[0151] In one possible embodiment, the collaborative memory block acquisition module 402 is specifically used to obtain a preset flag bit of the main memory block, wherein the preset flag bit includes a first flag bit and a second flag bit; flip the first flag bit to generate a first flip flag bit; combine the start flag bit with the first flip flag bit to obtain a new flag bit; determine the target memory block pointed to by the new flag bit; obtain the locking status information of the locking register corresponding to the target memory block; if the locking register corresponding to the target memory block is in an unlocked state, the target memory block is determined to be a collaborative memory block logically associated with the main memory block; if the locking register corresponding to the target memory block is in a locked state, iteratively execute the steps of "flipping the start flag bit to generate a second flip flag bit; combining the second flip flag bit with the first flip flag bit to obtain a new flag bit; determining the target memory block pointed to by the new flag bit; obtaining the locking status information of the locking register corresponding to the target memory block". If the second flip flag bit is determined to be in an unlocked state before it becomes the start flag bit again, the target memory block is determined to be a collaborative memory block logically associated with the main memory block.

[0152] In a possible embodiment, the collaborative memory block acquisition module 402 is also specifically used to flip the first flip flag to generate a third flip flag if the second flip flag becomes the starting flag again; combine the starting flag with the third flip flag to obtain a new flag; determine the target memory block pointed to by the new flag; obtain the locking status information of the locking register corresponding to the target memory block; if the locking register corresponding to the target memory block is in an unlocked state, determine the target memory block as a collaborative memory block logically associated with the main memory block; if the locking register corresponding to the target memory block is in a locked state, iteratively execute the steps of "flipping the starting flag to generate a second flip flag; combining the second flip flag with the third flip flag to obtain a new flag; determining the target memory block pointed to by the new flag; obtaining the locking status information of the locking register corresponding to the target memory block", if the second flip flag is determined to be in an unlocked state before becoming the starting flag again, then the target memory block is determined to be a collaborative memory block logically associated with the main memory block.

[0153] In one possible embodiment, the collaborative memory block acquisition module 402 is also specifically used to iteratively flip the third flip flag to generate a new third flip flag if the second flip flag becomes the starting flag again; combine the starting flag with the new third flip flag to obtain a new flag, and subsequent steps; if before the new third flip flag becomes the first flag again, it is determined that the lock register corresponding to the target memory block is in an unlocked state, then the target memory block is determined to be a collaborative memory block logically associated with the main memory block.

[0154] In a possible embodiment, the collaborative memory block acquisition module 402 is also specifically used to continue to determine whether the second flag bit is consistent with the start flag bit if the new third flip flag bit becomes the first flag bit again; if the second flag bit is consistent with the start flag bit, iteratively execute the steps of "flipping the second flag bit to generate a fourth flip flag bit; combining the fourth flip flag bit with the new third flip flag bit to obtain a new flag bit; determining the target memory block pointed to by the new flag bit; obtaining the locking status information of the locking register corresponding to the target memory block"; if before the fourth flip flag bit becomes the start flag bit again, it is determined that the locking register corresponding to the target memory block is in an unlocked state, then the target memory block is determined to be a collaborative memory block logically associated with the main memory block; if the fourth flip flag bit becomes the start flag bit again, the main memory block is determined to be a collaborative memory block.

[0155] In a possible embodiment, the collaborative memory block acquisition module 402 is further specifically used to, if the second flag bit is inconsistent with the start flag bit, combine the start flag bit with the new third flip flag bit to obtain a new flag bit; determine the target memory block pointed to by the new flag bit; obtain the locking status information of the locking register corresponding to the target memory block; if the locking register corresponding to the target memory block is in an unlocked state, determine the target memory block as a collaborative memory block logically associated with the main memory block; if the locking register corresponding to the target memory block is in a locked state, iteratively execute the steps of "flipping the start flag bit to generate a second flip flag bit; combining the second flip flag bit with the new third flip flag bit to obtain a new flag bit; determining the target memory block pointed to by the new flag bit; obtaining the locking status information of the locking register corresponding to the target memory block"; if the third flag bit is inconsistent with the start flag bit, the collaborative memory block acquisition module 402 is further specifically used to, if the second flag bit is inconsistent with the start flag bit, combine the start flag bit with the new third flip flag bit to obtain a new flag bit; determine the target memory block pointed to by the new flag bit; obtain the locking status information of the locking register corresponding to the target memory block; if the third flag bit is inconsistent with the start flag bit, the collaborative memory block acquisition module 402 is further specifically used to, Before the second flip flag bit becomes the second flag bit, it is determined that the locking register corresponding to the target memory block is in an unlocked state, and the target memory block is determined to be a collaborative memory block logically associated with the main memory block; if the second flip flag bit becomes the second flag bit, then continue to iterate the steps of "flipping the second flip flag bit to generate the fifth flip flag bit; combining the fifth flip flag bit with the new third flip flag bit to obtain a new flag bit; determining the target memory block pointed to by the new flag bit; obtaining the locking state information of the locking register corresponding to the target memory block"; if before the fifth flip flag bit becomes the starting flag bit again, it is determined that the locking register corresponding to the target memory block is in an unlocked state, the target memory block is determined to be a collaborative memory block logically associated with the main memory block; if the fifth flip flag bit becomes the starting flag bit again, the main memory block is determined to be a collaborative memory block.

[0156] In a possible embodiment, the collaborative memory block acquisition module 402 is also specifically used to determine the number of multiple storage modules and set the first flag bit of each memory block according to the number of multiple storage modules; determine the number of multiple memory blocks and set the second flag bit of each memory block according to the number of multiple memory blocks.

[0157] For the description of the features in the embodiment corresponding to the memory error handling device, reference can be made to the relevant description of the embodiment corresponding to the memory error handling method, which will not be repeated here.

[0158] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 5 As shown, the electronic device provided by this embodiment includes: at least one processor 501 and a memory 502. Optionally, the electronic device further includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus.

[0159] During the specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502, so that the at least one processor 501 executes the above-mentioned memory error handling method embodiment.

[0160] The specific implementation process of the processor 501 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0161] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0162] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.

[0163] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0164] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above-mentioned memory error handling method embodiments when running.

[0165] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0166] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned memory error handling method embodiments are implemented.

[0167] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned memory error handling method embodiments are implemented.

[0168] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0169] The above is a detailed introduction to a memory error handling method, apparatus, device and storage medium provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A memory error handling method, characterized in that: include: Continuously monitoring the operating status of a main memory block, and determining whether the main memory block triggers a preset fault event based on the operating status, wherein the main memory block is a memory block that is reading and writing data in a plurality of storage modules, and one storage module includes a plurality of memory blocks; If it is determined that the main memory block triggers a preset fault event, determining a collaborative memory block logically associated with the main memory block; Acquiring fault information data of the main memory block, and writing the fault information data into the collaborative memory block; triggering a first correction error reporting mechanism in the collaborative memory block according to the fault information data, generating error record data; The error record data is saved in a lock register corresponding to the cooperative memory block, and the lock register is locked.

2. The memory error handling method according to claim 1, wherein: The determining of a coordinated memory block logically associated with the main memory block comprises: Acquire a preset flag bit of the main memory block, wherein the preset flag bit includes a first flag bit and a second flag bit; Flipping the first flag bit to generate a first flip flag bit; Combining the start flag bit with the first flip flag bit to obtain a new flag bit; Determine the target memory block pointed to by the new flag bit; Obtaining lock status information of a lock register corresponding to the target memory block; If the lock register corresponding to the target memory block is in an unlocked state, determining the target memory block as a collaborative memory block logically associated with the main memory block; If the locking register corresponding to the target memory block is in a locked state, the steps of "flipping the starting flag to generate a second flip flag; combining the second flip flag with the first flip flag to obtain a new flag; determining the target memory block pointed to by the new flag; and obtaining the locking state information of the locking register corresponding to the target memory block" are iteratively executed. If the locking register corresponding to the target memory block is determined to be in an unlocked state before the second flip flag becomes the starting flag again, the target memory block is determined to be a collaborative memory block logically associated with the main memory block.

3. The memory error handling method according to claim 2, wherein: Also includes: If the second flip flag becomes the start flag again, flip the first flip flag to generate a third flip flag; Combining the start flag bit with the third flip flag bit to obtain a new flag bit; Determine the target memory block pointed to by the new flag bit; Obtaining lock status information of a lock register corresponding to the target memory block; If the lock register corresponding to the target memory block is in an unlocked state, determining the target memory block as a collaborative memory block logically associated with the main memory block; If the lock register corresponding to the target memory block is in a locked state, iteratively executing "flipping the start flag to generate a second flip flag; combining the second flip flag with the third flip flag to obtain a new flag; Determine the target memory block pointed to by the new flag bit; In the step of "obtaining locking status information of the locking register corresponding to the target memory block", if it is determined that the locking register corresponding to the target memory block is in an unlocked state before the second flip flag bit is changed back to the start flag bit, the target memory block is determined as a collaborative memory block logically associated with the main memory block.

4. The memory error handling method according to claim 3, wherein: Also includes: If the second flip flag becomes the starting flag again, iteratively flipping the third flip flag to generate a new third flip flag; combining the starting flag with the new third flip flag to obtain a new flag, and the following steps; If the lock register corresponding to the target memory block is determined to be in an unlocked state before the new third flip flag bit becomes the first flag bit again, the target memory block is determined to be a collaborative memory block logically associated with the main memory block.

5. The memory error handling method according to claim 4, characterized in that: Also includes: If the new third flip flag becomes the first flag again, continue to determine whether the second flag is consistent with the start flag; If the second flag bit is consistent with the start flag bit, iteratively executing "flipping the second flag bit to generate a fourth flipped flag bit; combining the fourth flipped flag bit with the new third flipped flag bit to obtain a new flag bit; Determine the target memory block pointed to by the new flag bit; The step of "obtaining locking status information of the locking register corresponding to the target memory block"; If, before the fourth flip flag bit becomes the start flag bit again, it is determined that the lock register corresponding to the target memory block is in an unlocked state, the target memory block is determined as a collaborative memory block logically associated with the main memory block; If the fourth flip flag becomes the start flag again, it is determined that the main memory block is the collaborative memory block.

6. The memory error handling method according to claim 5, characterized in that: After continuing to determine whether the second flag is consistent with the start flag, the method further includes: If the second flag bit is inconsistent with the start flag bit, combining the start flag bit with the new third flip flag bit to obtain a new flag bit; Determine the target memory block pointed to by the new flag bit; Obtaining lock status information of a lock register corresponding to the target memory block; If the lock register corresponding to the target memory block is in an unlocked state, determining the target memory block as a collaborative memory block logically associated with the main memory block; If the lock register corresponding to the target memory block is in a locked state, iteratively executing the steps of "flipping the start flag to generate a second flip flag; combining the second flip flag with the new third flip flag to obtain a new flag; determining the target memory block pointed to by the new flag; and obtaining the lock state information of the lock register corresponding to the target memory block." If, before the second flip flag bit becomes the second flag bit, it is determined that the lock register corresponding to the target memory block is in an unlocked state, the target memory block is determined as a collaborative memory block logically associated with the main memory block; If the second flip flag bit becomes the second flag bit, continue iterating the steps of "flipping the second flip flag bit to generate a fifth flip flag bit; combining the fifth flip flag bit with the new third flip flag bit to obtain a new flag bit; determining the target memory block pointed to by the new flag bit; and obtaining the lock status information of the lock register corresponding to the target memory block"; If, before the fifth flip flag bit becomes the start flag bit again, it is determined that the lock register corresponding to the target memory block is in an unlocked state, the target memory block is determined as a collaborative memory block logically associated with the main memory block; If the fifth flip flag becomes the start flag again, it is determined that the main memory block is the collaborative memory block.

7. The memory error handling method according to claim 2, wherein: Before obtaining the preset flag bit of the main memory block, the method further includes: Determining the number of the plurality of storage modules, and setting a first flag bit of each memory block according to the number of the plurality of storage modules; Determine the number of the multiple memory blocks, and set a second flag bit of each memory block according to the number of the multiple memory blocks.

8. A memory error handling device, characterized in that: include: an operating status monitoring module, configured to continuously monitor the operating status of a main memory block and determine, based on the operating status, whether the main memory block has triggered a preset fault event, wherein the main memory block is a memory block in the plurality of storage modules that is currently reading or writing data, and wherein one storage module includes a plurality of memory blocks; a collaborative memory block acquisition module, configured to determine a collaborative memory block logically associated with the main memory block if it is determined that the main memory block triggers a preset fault event; a fault information sending module, configured to obtain the fault information data of the main memory block and write the fault information data into the collaborative memory block; An error record generating module, configured to trigger a first correction error reporting mechanism in the collaborative memory block according to the fault information data, and generate error record data; A memory error locking module is used to save the error record data to a locking register corresponding to the collaborative memory block and lock the locking register.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the memory error handling method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the memory error handling method according to any one of claims 1 to 7 are implemented.