Memory error rate monitoring and repairing method and system
By associating data importance levels with physical addresses during memory allocation and selecting differentiated strategies when errors occur, the problem of failing to consider data importance in existing technologies is solved, and intelligent and differentiated processing of memory errors is achieved, thereby improving system reliability and resource utilization efficiency.
Patent Information
- Application Number
- CN202510828216.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Existing memory error response strategies fail to fully consider data importance, resulting in imprecise handling of data of different values, which may cause system crashes, data inconsistency or resource waste.
By associating data importance levels with physical addresses during memory allocation, the importance of error addresses is obtained when monitoring memory errors, and differentiated repair or mitigation strategies are selected based on error type and data importance, including building a memory allocation metadata table and a differentiated strategy set to optimize the memory error handling process.
It improves the protection level of high-value data, avoids unnecessary processing of low-value data, improves system reliability and resource utilization efficiency, and ensures business continuity.
Smart Images

Figure CN120670207A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data storage technology, and in particular to a memory error rate monitoring and repair method and system. Background Art
[0002] During the continuous operation of a computing system, memory cells are responsible for storing and accessing data. Error events such as data bit state changes may occur in memory cells. Memory error monitoring is typically implemented within the system to continuously monitor the status of memory cells and catch error events.
[0003] When memory error monitoring detects an error, traditional response strategies might be based primarily on the error type, frequency, or location. For uncorrectable errors, the memory page containing the address might be marked as bad, or the entire memory area might be taken offline to prevent further use or propagation of the erroneous data. This response is relatively straightforward and easy to implement.
[0004] However, this response strategy, based solely on the nature and physical location of the error, fails to fully consider the importance of the data currently stored at that memory address. If the address where the error occurred stores the status information of a critical system lock or the core data of an ongoing high-value transaction, simply taking the memory page offline could lead to system crashes, data inconsistencies, severe business losses, or even security vulnerabilities.
[0005] Furthermore, the distribution of data in system memory is highly dynamic. As tasks start, execute, and terminate, memory allocation and deallocation occur frequently. The same physical address may store data of varying types and importance at different times. If the system lacks effective means to quickly and accurately determine and assess the real-time importance of the affected data when an error occurs, its memory error response strategy will be neither refined nor intelligent.
[0006] Therefore, in the computing system memory, facing the challenge of stored data with different importance levels and dynamic changes in their physical locations, when a memory error is detected, how can we quickly and accurately perceive the real-time importance of the affected data, and based on this importance information and error type, intelligently and differentiatedly formulate and execute memory error repair or mitigation strategies to ensure the integrity and business continuity of high-value data, while avoiding excessive resource consumption or unnecessary interruptions caused by erroneous handling of low-value data.
[0007] In view of the above problems, the existing technology needs to be improved urgently. Summary of the Invention
[0008] In view of the above-mentioned shortcomings of the existing technology, the present application provides a memory error rate monitoring and repair method and system, which has the advantages of being able to differentiate and select and execute memory error repair or mitigation strategies according to the real-time importance of the data stored in the memory and the error type, thereby improving the protection level of high-value data, avoiding unnecessary processing of low-value data, and improving system reliability and resource utilization efficiency.
[0009] In a first aspect, a memory error rate monitoring and repair method is provided, the method comprising the steps of: S1: Get the data importance level identifier in the memory allocation request; S2: In response to the memory allocation request, allocate a physical address range from the physical memory pool, associate the data importance level identifier with the allocated physical address range, and store the data in a memory allocation metadata table; S3: When a memory error is detected, obtaining an error physical address and an error type, and determining the data importance level identifier corresponding to the memory error from the memory allocation metadata table according to the error physical address; S4: selecting a corresponding memory error repair or mitigation strategy from a preset differentiated strategy set according to the error type, the error physical address, and the data importance level identifier corresponding to the error memory; S5: Execute the selected memory error repair or mitigation strategy.
[0010] A memory error rate monitoring and repair method proposed in this application solves the problem of the prior art that the importance of data is not taken into account, resulting in imprecise strategies, by associating data importance with physical addresses during memory allocation, obtaining the data importance corresponding to the error address when an error occurs, and selecting differentiated repair strategies based on the error type, error address, and data importance. It has the advantage of being able to differentially select and execute memory error repair or mitigation strategies based on the real-time importance of data stored in the memory and the error type, thereby improving the level of protection for high-value data, avoiding unnecessary processing of low-value data, and improving system reliability and resource utilization efficiency.
[0011] Furthermore, step S1 includes: S11: Receive a memory allocation request, wherein the memory allocation request includes a data importance level identifier and a requested memory size; S12: Verify the validity of the data importance level identifier; if the data importance level identifier is not within a preset level range, replace it with a preset default data importance level identifier, and record the replacement information in the system log; S13: Sending the verified or replaced data importance level identifier and the requested memory size to the memory management module.
[0012] A memory error rate monitoring and repair method proposed in this application adds a validity verification and processing mechanism for the data importance level identifier obtained in the memory allocation request, thereby ensuring that subsequent steps can be operated based on the valid data importance level, solving potential problems caused by invalid level identifiers, and improving the robustness and reliability of the entire method.
[0013] Furthermore, step S2 includes: S21: In response to a memory allocation request, select an idle physical address range that meets the requested size from a physical memory pool and generate a page table entry, where the page table entry includes at least a start address and an end address of the physical address range; S22: Determine the corresponding metadata storage structure type according to the data importance level identifier in the memory allocation request; S23: Based on the determined metadata storage structure type, an entry is created in the memory allocation metadata table, with the starting address of the physical address range as the key value, and the data corresponding to the key value including the data importance level identifier and the metadata information of the physical address range is stored in the memory allocation metadata table.
[0014] This application proposes a memory error rate monitoring and repair method, which improves the organization efficiency and subsequent search efficiency of the memory allocation metadata table by proposing a method for determining different metadata storage structure types based on data importance level identification and storing metadata in the structure, thereby solving the problem of rapid positioning in a large amount of dynamically changing metadata.
[0015] Furthermore, step S22 includes: S221: Establishing a storage correspondence between the data importance level and the metadata storage structure, wherein the storage correspondence includes: a high importance level corresponds to a hash table, a medium importance level corresponds to a B+ tree, and a low importance level corresponds to a linked list; S222: querying the storage correspondence according to the data importance level identifier in the memory allocation request to determine the metadata storage structure type; S223: If the query fails, the preset default metadata storage structure type is used.
[0016] A memory error rate monitoring and repair method proposed in this application, by specifying the process of determining the metadata storage structure type based on the data importance level identifier, provides a technical solution for differentially selecting metadata storage structures based on the data importance level, thereby improving the management efficiency and subsequent search efficiency of the memory allocation metadata table, and solving the problem of low metadata management efficiency of data with different importance levels.
[0017] Furthermore, step S3 includes: S31: When a memory error is detected, the error physical address and the error type are obtained, and the error type is encoded into an error type enumeration value; S32: Searching for a physical address range entry containing the erroneous physical address in the memory allocation metadata table; if the search fails, determining that it is an illegal memory access, generating an error report containing an illegal access flag, and terminating; S33: If the search is successful, extract the data importance level identifier from the found physical address range entry, and combine it with the error type enumeration value to generate a memory error report containing the data importance level identifier and the error type enumeration value.
[0018] Furthermore, step S32 includes: S321: Constructing a sorting index table, which stores the starting addresses of physical address range entries and the offsets of corresponding entries in the memory allocation metadata table, and the physical address range entries are sorted in ascending order according to the starting addresses; S322: searching, according to the erroneous physical address, in the sorted index table by a binary search method for the maximum entry offset whose starting address is less than or equal to the erroneous physical address; S323: If the entry offset is found, read the corresponding physical address range entry from the memory allocation metadata table according to the offset, and determine whether the erroneous physical address falls within the physical address range; S324: If it is not within the range, it is determined to be an illegal memory access, an error report containing an illegal access flag is generated, and the process is terminated; if the entry offset is not found, it is determined to be an illegal memory access, an error report containing an illegal access flag is generated, and the process is terminated.
[0019] Furthermore, step S4 includes: S41: Constructing a policy query key value, where the key value includes the error type enumeration value, the data importance level identifier, and the physical address range to which the error physical address belongs; S42: searching for a matching memory error repair or mitigation strategy in a preset differentiated strategy set according to the strategy query key value; S43: If a matching memory error repair or mitigation strategy is found, the strategy is output; if no matching memory error repair or mitigation strategy is found, a preset default memory error repair or mitigation strategy corresponding to the data importance level identifier is selected and output.
[0020] Furthermore, step S42 includes: S421: According to the policy query key value, in a preset differentiated policy set, determining whether the error type enumeration value indicates a tolerable error; if so, selecting a preset default memory error mitigation policy corresponding to the data importance level identifier; S422: If an intolerable error is indicated, the error type enumeration value and the data importance level identifier are used as key values, and a Bloom filter is used to filter the preset differentiated strategy set, excluding the repair strategies that do not contain the key value, and then the remaining repair strategies are matched to determine a matching memory error repair strategy.
[0021] Furthermore, step S5 includes: S51: Establish a priority queue for executing memory error repair or mitigation strategies; S52: Selecting a policy with the highest priority from the memory error repair or mitigation policy execution priority queue, determining whether current computing system resources meet the requirements for executing the policy, and if so, executing the policy; S53: If not satisfied, the policy is added back to the memory error repair or mitigation policy execution priority queue, the priority of the policy is lowered, and S52 is re-executed; if the policy execution fails, the policy with the next highest priority is selected for execution until the policy execution succeeds or the policy execution priority queue is empty.
[0022] In a second aspect, a memory error rate monitoring and repair system is provided, which is applied to the steps of any of the above methods, and the system includes: The first acquisition module: obtains the data importance level identifier in the memory allocation request; A storage module: in response to a memory allocation request, allocates a physical address range from a physical memory pool, associates the data importance level identifier with the allocated physical address range, and stores the data in a memory allocation metadata table; The second acquisition module is configured to acquire an error physical address and an error type when a memory error is detected, and determine the data importance level identifier corresponding to the memory error from the memory allocation metadata table according to the error physical address; A strategy selection module: selects a corresponding memory error repair or mitigation strategy from a preset differentiated strategy set according to the error type, the error physical address, and the data importance level identifier corresponding to the error memory; Policy Execution Module: Executes the selected memory error repair or mitigation strategy.
[0023] Beneficial effect: The memory error rate monitoring and repair method and system proposed in the present application solves the problem of the prior art that the importance of data is not considered, resulting in imprecise strategies, by associating data importance with physical addresses during memory allocation, obtaining the data importance corresponding to the error address when an error occurs, and selecting differentiated repair strategies based on the error type, error address, and data importance. It has the advantage of being able to differentially select and execute memory error repair or mitigation strategies based on the real-time importance of data stored in the memory and the error type, thereby improving the protection level of high-value data, avoiding unnecessary processing of low-value data, and improving system reliability and resource utilization efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a flowchart of a memory error rate monitoring and repair method proposed in this application.
[0025] Figure 2 This is a structural diagram of a memory error rate monitoring and repair system proposed in this application.
[0026] Figure 3 This is an architectural diagram of a memory error rate monitoring and repair system proposed in this application.
[0027] Description of reference numerals: 201, first acquisition module; 202, storage module; 203, second acquisition module; 204, policy selection module; 205, policy execution module. DETAILED DESCRIPTION
[0028] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and marked in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.
[0029] It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0030] Please refer to Figure 1 , a memory error rate monitoring and repair method, the method comprising the steps of: S1: Get the data importance level identifier in the memory allocation request; S2: In response to the memory allocation request, allocate a physical address range from the physical memory pool, associate the data importance level identifier with the allocated physical address range, and store the data in the memory allocation metadata table; S3: When a memory error is detected, the error physical address and error type are obtained, and the data importance level identifier corresponding to the memory error is determined from the memory allocation metadata table based on the error physical address; S4: Selecting a corresponding memory error repair or mitigation strategy from a preset differentiated strategy set based on the error type, the error physical address, and the importance level of the data corresponding to the error memory; S5: Execute the selected memory error repair or mitigation strategy.
[0031] Among them, the data importance level identification refers to a mark used to distinguish the importance of different data in the system or business. It can be implemented using a numerical value. For example, the integer value 1 represents high importance, the integer value 2 represents medium importance, and the integer value 3 represents low importance. It is mainly used to record the important attributes of the data when allocating memory to provide a basis for subsequent error handling.
[0032] The memory allocation metadata table refers to a structure used to store information related to memory allocation. It can be implemented using a hash table, such as a set of key-value pairs, where the key is the starting address of the physical address range and the value contains information about the data importance level and the physical address range. Its main purpose is to establish and maintain the association between the physical address range and the data importance level so that it can be quickly queried when an error occurs.
[0033] A preset differentiated strategy set refers to a collection of multiple predefined processing methods for different memory error situations. It can be organized in the form of lookup tables, rule engines, decision trees, etc. For example, a mapping relationship, the input is the error type and data importance level, and the output is the corresponding repair or mitigation operation sequence. Its main purpose is to select the most appropriate processing solution based on the error attributes and data importance, and to achieve refined error response.
[0034] As a preferred embodiment, the solution of this application is specifically implemented as follows: When a memory allocation request is received, it includes an integer value indicating the importance of the data, for example, 1 for high importance, 2 for medium importance, and 3 for low importance. The system allocates a physical memory region from the physical memory pool, for example, starting at address 0x10000000 and 4KB in size. The system associates this integer value with the physical address range (0x10000000 - 0x10000FFF) and stores this association in a hash table, using the starting address of the physical address range, 0x10000000, as the key and a structure containing the importance level and physical address range as the value.
[0035] When a memory error is detected at physical address 0x10000100, the system obtains the error type, such as a single-bit error. Using the error physical address 0x10000100, the system searches the hash table for the physical address range containing that address, finds the entry with the key 0x10000000, and extracts the data importance level identifier, such as the integer value 1 (high importance). Based on the error type (single-bit error) and the data importance level (high importance), the system searches a preset policy lookup table for the corresponding policy. For example, the lookup table indicates that for single-bit errors occurring on high-importance data, ECC correction should be performed and detailed logging should be recorded. The system then performs the ECC correction and logging operations.
[0036] Through the above scheme, this application solves the problem that traditional memory error handling methods cannot perceive the real-time importance of data. By establishing an association between physical addresses and data importance during memory allocation, and using this association information when an error occurs, combined with the error type, the real-time importance of the affected data can be quickly and accurately determined. Based on the importance of data and the type of error, the most appropriate repair or mitigation strategy is selected and executed from a preset set of differentiated strategies, thereby achieving intelligent and differentiated handling of memory errors. This avoids system crashes or data loss caused by insufficient processing of high-value data, and also avoids waste of resources or unnecessary interruptions caused by excessive processing of low-value data, thereby improving system stability and business continuity.
[0037] Furthermore, step S1 includes: S11: Receive a memory allocation request, where the memory allocation request includes a data importance level identifier and a requested memory size; S12: Verify the validity of the data importance level identifier. If the data importance level identifier is not within the preset level range, replace it with a preset default data importance level identifier and record the replacement information in the system log; S13: Sending the verified or replaced data importance level identifier and the requested memory size to the memory management module.
[0038] This scheme constructs a robust data importance information acquisition mechanism by refining the steps of obtaining data importance level identification.
[0039] Step S11 is the starting point for obtaining original information.
[0040] Step S12 verifies the validity of the extracted data importance level identifier to determine whether it meets the system's preset specifications. This verification step is key, as it can identify and isolate data identifiers that do not meet the requirements. If the verification finds that the identifier is invalid, the system will not directly use the invalid information, but will replace it with a preset default data importance level identifier. This replacement mechanism ensures that even if the request source provides an incorrect or non-standard identifier, the system can obtain a valid and usable importance level information, avoiding subsequent processing interruptions or errors caused by invalid data.
[0041] Step S13 records this replacement behavior in the system log, providing a basis for subsequent system analysis, problem troubleshooting or strategy optimization. Finally, the importance level of the data processed after validity processing is passed to the memory management module together with the original requested memory size.
[0042] The memory management module relies on validated data importance level information when performing subsequent memory allocations and associating importance levels with allocated physical address ranges. This ensures that subsequent differentiated metadata management, error monitoring, error type determination, and most importantly, the selection and execution of differentiated repair strategies based on data importance levels can be performed based on accurate and reliable importance information. This improves the reliability and effectiveness of the entire memory error rate monitoring and repair method, enabling truly refined and intelligent processing based on data importance levels, and resolving the issue of the entire differentiated processing chain failing due to invalid original identifiers.
[0043] Furthermore, step S2 includes: S21: In response to a memory allocation request, select an idle physical address range that meets the requested size from a physical memory pool and generate a page table entry, where the page table entry includes at least a start address and an end address of the physical address range; S22: Determine the corresponding metadata storage structure type according to the data importance level identifier in the memory allocation request; S23: Based on the determined metadata storage structure type, an entry is created in the memory allocation metadata table, with the starting address of the physical address range as the key value, and the data corresponding to the key value including the data importance level identifier and the metadata information of the physical address range is stored in the memory allocation metadata table.
[0044] The method proposed in this application achieves efficient organization of memory allocation metadata tables by refining the metadata storage method during the memory allocation process. Specifically, after receiving a memory allocation request and determining the physical address range to be allocated, the method further obtains the data importance level associated with the request. Based on this importance level, the system intelligently determines a storage structure type suitable for storing the data metadata.
[0045] For example, for high-importance data, a hash table with high search efficiency can be selected; for medium-importance data, a B+ tree with range search support can be selected; and for low-importance data, a linked list with high space efficiency can be selected. After determining the storage structure type, the system creates or updates an entry in the memory allocation metadata table using the determined structure type. This entry uses the starting address of the allocated physical address range as the index key value and stores metadata information including the data importance level and physical address range.
[0046] In this way, metadata for data of different importance levels is stored or organized separately in structures tailored to their access characteristics. When a memory error subsequently occurs and the faulty physical address is obtained, the system can use this faulty physical address, particularly the start address of its physical address range, to quickly locate the corresponding metadata entry in the optimized metadata storage structure, thereby obtaining the real-time importance level of the data affected by the error.
[0047] This method of organizing and storing metadata based on the level of data importance improves the search efficiency of the memory allocation metadata table, especially when the number of metadata entries is large and dynamically changing. It can quickly and accurately obtain the required information, providing efficient data support for subsequent differentiated error handling based on importance level.
[0048] As a specific implementation method, when a memory allocation request is received, for example, a request to allocate a memory area for storing high-importance data, the system first finds a free physical address range that meets the size requirement from the physical memory pool, assuming that the range starts from address A and ends at address B. The system generates a page table entry to record address A and address B. Then, the system recognizes that the data importance level corresponding to the request is "high". According to the preset mapping relationship, for example, the "high" importance level corresponds to the hash table storage structure, the system determines that the metadata should be stored in the hash structure for high-importance data in the memory allocation metadata table.
[0049] The system then creates a new entry in the hash table, using address A as the key value, and stores metadata containing information such as the data importance level "high" and address range AB as the data corresponding to the key value. If another request is to allocate a memory area for storing low-importance data, with addresses ranging from C to D and an importance level of "low", the system will store the entry with address C as the key value and the corresponding metadata information in the linked list structure for low-importance data in the memory allocation metadata table based on the preset mapping relationship (for example, the "low" importance level corresponds to a linked list).
[0050] Furthermore, step S22 includes: S221: Establishing a storage correspondence between data importance levels and metadata storage structures, wherein the storage correspondence includes: a high importance level corresponds to a hash table, a medium importance level corresponds to a B+ tree, and a low importance level corresponds to a linked list; S222: Query the storage correspondence according to the data importance level identifier in the memory allocation request to determine the metadata storage structure type; S223: If the query fails, the preset default metadata storage structure type is used.
[0051] This solution establishes a storage correspondence between data importance levels and metadata storage structures, clarifying the metadata storage structure that should be used for data of different importance levels. For example, high-importance data requires extremely fast search speeds, so a hash table is selected; medium-importance data requires both search and range queries, so a B+ tree is selected; low-importance data is large in quantity and has a relatively low search frequency, so a linked list is selected. When a memory allocation request is received, the system queries the preset storage correspondence based on the data importance level identifier in the request to determine the metadata storage structure type that is most suitable for data of that level. If an exception occurs during the query process, such as an invalid or undefined identifier, the preset default storage structure is used as an alternative, ensuring the robustness of the metadata storage process.
[0052] Through this differentiated storage structure selection, this solution enables the memory allocation metadata table to be optimized and managed according to the actual importance of the data. This is combined with the steps in the aforementioned method of allocating a physical address range according to the memory allocation request and associating the data importance level identifier with the physical address range and storing it in the memory allocation metadata table, ensuring that the metadata is organized in a structure that is conducive to subsequent fast search when it is stored. When a memory error is detected and the corresponding metadata needs to be found based on the erroneous physical address, the search efficiency is significantly improved because the metadata has been stored in different, optimized structures according to the importance level. In particular, for high-importance data, its importance identifier can be obtained more quickly, laying the foundation for the rapid selection and execution of subsequent differentiated error repair or mitigation strategies.
[0053] Furthermore, step S3 includes: S31: When a memory error is detected, the error physical address and error type are obtained, and the error type is encoded as an error type enumeration value; S32: Searching for a physical address range entry containing an erroneous physical address in the memory allocation metadata table; if the search fails, determining that it is an illegal memory access, generating an error report containing an illegal access flag, and terminating; S33: If the search is successful, extract the data importance level identifier from the found physical address range entry, and combine it with the error type enumeration value to generate a memory error report containing the data importance level identifier and the error type enumeration value.
[0054] Among them, the memory allocation metadata table refers to a structure that stores physical memory allocation information, which records metadata information such as the allocated physical address range and the associated data importance level. It can be organized in the form of key-value pairs, where the key is the physical address range and the value is the corresponding metadata information. The underlying storage structure can be selected according to needs, such as a hash table, tree structure or linked list.
[0055] The error type enumeration value refers to the numerical representation of different types of memory errors after unified encoding. It can use predetermined integers or symbolic constants to represent specific error types, such as single-bit errors, multi-bit errors, parity errors, etc.
[0056] Among them, illegal memory access refers to read and write operations on physical memory addresses that are not legally allocated by the system, which can be manifested as the wrong physical address being outside any recorded physical address range in the memory allocation metadata table.
[0057] In one embodiment, the above steps can be implemented as follows: When the memory monitoring hardware detects a single-bit ECC error at the physical address 0x12345678, the system interrupt handler is triggered and step S31 is executed.
[0058] At this point, the acquired error physical address is 0x12345678, and the error type is "single-bit ECC error." The system encodes "single-bit ECC error" as a predetermined error type enumeration value, such as the value 1. Next, proceeding to step S32, the system uses the address 0x12345678 to search the memory allocation metadata table. This table may store entries such as {Start Address: 0x12340000, End Address: 0x1234FFFF, Data Importance Level: High}, {Start Address: 0x20000000, End Address: 0x2000FFFF, Data Importance Level: Low}, and so on.
[0059] The system searches for a range containing address 0x12345678. If the search finds that address 0x12345678 does not fall within any recorded physical address range, for example, if the error occurs at address 0xFFFFFFFF, which is unallocated, the system determines that the memory access is illegal, generates an error report containing an illegal access flag, and terminates the process. In this embodiment, assuming that address 0x12345678 falls within the range [0x12340000, 0x1234FFFF], the search is successful.
[0060] The system proceeds to step S33, extracting the associated data importance level indicator (i.e., "High") from the entry. The extracted "High" importance level indicator is then combined with the error type enumeration value 1 obtained in step S31 to generate a memory error report. This report may include information such as the error address, error type, error type enumeration value 1, and the data importance level "High." This report can then be passed to the subsequent policy selection module.
[0061] Furthermore, step S32 includes: S321: Constructing a sorting index table, which stores the starting addresses of physical address range entries and the offsets of corresponding entries in the memory allocation metadata table, and the physical address range entries are sorted in ascending order according to the starting addresses; S322: searching the sorted index table for the maximum entry offset whose starting address is less than or equal to the error physical address by a binary search method according to the error physical address; S323: If the entry offset is found, then read the corresponding physical address range entry from the memory allocation metadata table according to the offset, and determine whether the erroneous physical address falls within the physical address range; S324: If it is not within the range, it is determined to be an illegal memory access, an error report containing an illegal access flag is generated, and the process is terminated; if the entry offset is not found, it is determined to be an illegal memory access, an error report containing an illegal access flag is generated, and the process is terminated.
[0062] The search method of the present application converts the search operation of the original memory allocation metadata table into a search on an optimized index table by introducing an auxiliary sorting index table.
[0063] Specifically, a sorted index table is constructed that is arranged in ascending order according to the starting address of the physical address range. The index table only stores the starting address of each physical address range entry and the position offset of the entry in the original memory allocation metadata table.
[0064] This structure centralizes and sorts the key information for lookups (starting address), while retaining a reference to the original complete data through an offset.
[0065] When searching for an entry containing a specific erroneous physical address, a binary search algorithm is used to quickly locate it in the sorted index table. This algorithm exploits the ordered nature of the index table and can quickly find the position of the largest entry in the index table whose starting address is less than or equal to the erroneous physical address in logarithmic time, thereby obtaining its corresponding offset.
[0066] After obtaining the offset, the corresponding physical address range entry can be accurately read directly from the original memory allocation metadata table using the offset.
[0067] The read entry is then finally verified to determine whether the erroneous physical address is indeed within the physical address range. If the erroneous physical address is not within the searched range, or if no suitable entry offset is found in the index table, the access is deemed an illegal memory access, and a corresponding error report is generated, terminating subsequent processing.
[0068] By constructing a sorted index table and adopting binary search, the method of the present application avoids inefficient linear scanning of the huge memory allocation metadata table, significantly improves the search efficiency, ensures that the affected memory range can be located quickly and accurately when a memory error is detected, and provides an efficient data foundation for subsequent differentiated error handling based on the data importance level.
[0069] This method, combined with steps such as obtaining the error physical address and extracting the data importance level identifier from the found entry, constitutes a complete process for quickly responding to memory errors and determining data importance, improving the real-time and accuracy of the entire memory error handling process.
[0070] Furthermore, step S4 includes: S41: Constructing a policy query key value, where the key value includes an error type enumeration value, a data importance level identifier, and a physical address range to which the error physical address belongs; S42: searching for a matching memory error repair or mitigation strategy in a preset differentiated strategy set according to the strategy query key value; S43: If a matching memory error repair or mitigation strategy is found, the strategy is output; if no matching memory error repair or mitigation strategy is found, a preset default memory error repair or mitigation strategy corresponding to the data importance level identifier is selected and output.
[0071] The error type enumeration value refers to a numerical value or symbolic representation of different memory error types after encoding, which can be implemented using an integer, an enumeration type, or a string.
[0072] After receiving memory error information, this solution first integrates three key dimensions of information into a single query condition: error type, data importance level, and the physical address range to which the error's physical address belongs. This integration is the foundation of this solution's refined policy selection, as it enables subsequent policy searches to simultaneously consider the nature of the error, the value of the affected data, and the specific memory region where the error occurred. This combination of multi-dimensional information allows the system to form a more discriminating query condition, enabling efficient and accurate searches within a pre-set set of differentiated policies.
[0073] Based on the constructed policy query key, the system searches for an exact match for a memory error repair or mitigation policy within the pre-set set of differentiated policies. This composite key-based search method quickly locates pre-customized solutions for specific error scenarios, improving the efficiency and accuracy of policy selection. If a matching policy is found, the system directly outputs the customized policy, ensuring that the optimal solution is executed first.
[0074] If an exact matching policy cannot be found, the system does not interrupt processing. Instead, it falls back to selecting and outputting a preset default memory error repair or mitigation policy corresponding to the data importance level. This mechanism ensures that even without precise preset policies for all possible error combinations, the system can still provide a reasonable default handling solution based on the importance of the affected data, enhancing system robustness.
[0075] By combining this with the data importance level identification obtained in the previous steps, this solution can break through the limitations of traditional strategy selection based only on error type or location, and truly achieve differentiated responses based on data value, thereby optimizing resource utilization while ensuring the security of high-value data.
[0076] Furthermore, step S42 includes: S421: According to the policy query key value, in the preset differentiated policy set, it is determined whether the error type enumeration value indicates a tolerable error. If it indicates a tolerable error, a preset default memory error mitigation policy corresponding to the data importance level identifier is selected; S422: If an intolerable error is indicated, the error type enumeration value and the data importance level identifier are used as key values, and a Bloom filter is used to filter the preset differentiated strategy set, excluding the repair strategies that do not contain the key value, and then the remaining repair strategies are matched to determine the matching memory error repair strategy.
[0077] Among them, tolerable errors refer to error types that have little impact on system functions or data integrity and can be handled by lightweight mitigation measures or even ignored, such as certain single-bit correctable errors; intolerable errors refer to error types that have a serious impact on system functions or data integrity and require repair measures to restore normal status or prevent further damage, such as uncorrectable errors or errors occurring in critical data areas.
[0078] A Bloom filter is a space-efficient probabilistic data structure used to test whether an element is a member of a set. It may produce false positive matches but never false negative matches. That is, if a Bloom filter indicates that an element is not in a set, then it is definitely not in the set. If it indicates that an element may be in a set, then it may or may not be in the set. In this solution, a Bloom filter is used to quickly eliminate repair strategies that are definitely not applicable to the current combination of error type and data importance.
[0079] This solution addresses the efficiency and pertinence issues of searching for matching strategies in a preset differentiated strategy set, and proposes a technical means to distinguish error types and adopt different search strategies.
[0080] Specifically, the error type enumeration value in the policy query key is used to determine whether the error is tolerable. This determination is made because different types of errors have different impacts on system stability and data integrity, requiring different responses.
[0081] If the error is determined to be tolerable, the default memory error mitigation strategy corresponding to the data importance level is directly selected. This approach has the advantage of fast response, avoiding the complex search process in the entire policy set, reducing system overhead, and combining the data importance level to ensure differentiated and lightweight handling of tolerable errors.
[0082] If an error is determined to be intolerable, a more precise and effective repair strategy is required. At this point, a Bloom filter is used to perform a preliminary screening of the preset differentiated strategy set, using the error type enumeration value and the data importance level identifier as key values. The Bloom filter can quickly eliminate repair strategies that are definitely inappropriate for the current error type and data importance combination, significantly narrowing the range of strategies requiring further consideration and improving search efficiency. After the Bloom filter screening, the remaining repair strategies are more precisely matched to ultimately determine the most appropriate repair strategy for the current intolerable error.
[0083] This hierarchical search method, which combines error type judgment, Bloom filter pre-screening, and subsequent exact matching, improves the efficiency and accuracy of finding the exact repair strategy when facing intolerable errors.
[0084] This strategy selection process is combined with the overall process of obtaining data importance levels, monitoring errors, determining the physical address of errors, and subsequently executing strategies. This enables the entire memory error handling method to quickly and accurately select and execute the most appropriate response strategy based on the specific nature of the error and the value of the affected data, thereby improving the system's robustness and data protection capabilities in the face of memory errors.
[0085] Furthermore, step S5 includes: S51: Establish a priority queue for executing memory error repair or mitigation strategies; S52: Selecting a policy with the highest priority from the memory error repair or mitigation policy execution priority queue, determining whether current computing system resources meet the requirements for executing the policy, and if so, executing the policy; S53: If not satisfied, the policy is added back to the memory error repair or mitigation policy execution priority queue, the priority of the policy is lowered, and S52 is re-executed; if the policy execution fails, the policy with the next highest priority is selected for execution until the policy execution succeeds or the policy execution priority queue is empty.
[0086] Establishing a priority queue for executing memory error repair or mitigation policies involves constructing an ordered data structure for the set of policies to be executed. This can be achieved using a heap-based priority queue, where each policy is associated with a priority value, with smaller values indicating higher priority. The initial priority of a policy can be determined based on the policy type, error type, and the real-time importance of the affected data. For example, repair policies targeting high-importance data can be assigned a higher initial priority.
[0087] Current computing system resources refer to the system resources required to execute the policy, such as CPU time, memory, I / O bandwidth, etc.
[0088] As a preferred embodiment, the solution of this application is specifically implemented as follows: Assume the system detects a memory error affecting a data area marked as "high importance," and the error type is a correctable single-bit error. Based on the preceding steps, the system selects a set of strategies, including: Strategy A: Hardware ECC correction + data rewrite, the highest priority; Strategy B: Software data recovery + migration to a spare area, the second highest priority; Strategy C: Marking the error area as read-only and logging, the lowest priority.
[0089] Step S51: Create a priority queue and add policies A, B, and C to the queue.
[0090] Step S52: Take out policy A from the queue. Check the system resources and find that the current CPU utilization and memory bandwidth are both at high levels, which do not meet the resource conditions required for executing policy A.
[0091] Step S53: Policy A is re-enqueued and its priority is lowered to 1.5. Policy B currently has the highest priority in the queue. The system re-executes step S52 and retrieves Policy B. System resources are checked and the resource requirements for Policy B are met. Policy B is executed. Assuming Policy B executes successfully, the data is restored and migrated. At this point, the error is mitigated, and the execution process terminates.
[0092] In another scenario, assume that after Policy A is retrieved in step S52, resources are sufficient and Policy A is executed. However, Policy A fails to execute, for example, because the hardware correction circuit reports that the error cannot be corrected. The system then proceeds to the failure handling branch in step S53. The next highest priority policy, Policy B, is selected from the queue. Policy B is executed. If Policy B succeeds, the process terminates. If Policy B also fails, Policy C is selected for execution. If Policy C succeeds, the process terminates. If Policy C also fails and the queue is empty, the error was not effectively handled, and the system may trigger a higher-level error handling mechanism, such as a system restart or an alarm.
[0093] Please refer to Figure 1 、 Figure 3 A memory error rate monitoring and repair system is applied to the steps of any of the above methods, and the system includes: The first acquisition module 201: acquires the data importance level identifier in the memory allocation request; Storage module 202: In response to a memory allocation request, allocates a physical address range from a physical memory pool, associates the data importance level identifier with the allocated physical address range, and stores the data in a memory allocation metadata table; The second acquisition module 203: when a memory error is detected, acquires the error physical address and error type of the error, and determines the data importance level identifier corresponding to the memory error from the memory allocation metadata table according to the error physical address; Strategy selection module 204: selects a corresponding memory error repair or mitigation strategy from a preset differentiated strategy set based on the error type, the error physical address, and the importance level of the data corresponding to the error memory; Policy execution module 205: executes the selected memory error repair or mitigation policy.
[0094] Among them, the first acquisition module 201 refers to the component responsible for receiving or parsing memory allocation requests from the outside and extracting data importance level information therefrom. The storage module 202 refers to the component responsible for managing the allocation of physical memory resources and establishing a mapping relationship between physical addresses and data importance levels. The second acquisition module 203 refers to the component responsible for receiving or monitoring memory error events and querying the mapping relationship established by the storage module based on the error information to obtain the data importance level. The policy selection module 204 refers to the component responsible for determining appropriate countermeasures from a variety of preset policies based on error-related information (including data importance level). The policy execution module 205 refers to the component responsible for specifically executing the countermeasures determined by the policy selection module.
[0095] This system provides a set of functional modules for implementing a memory error monitoring and repair method based on data importance levels, thereby solving the problem of lacking a specific implementation architecture when applying this method to actual computing systems. Specifically, when the system receives a memory allocation request, the first acquisition module 201 first obtains a data importance level identifier from the request. Subsequently, the storage module 202 responds to this request, allocates a physical address range from the physical memory pool, and associates the data importance level identifier obtained by the first acquisition module 201 with this allocated physical address range, and records this association information in the memory allocation metadata table. Through this process, the system establishes a corresponding relationship between data importance and physical location during the memory allocation stage. When the system detects a memory error, the second acquisition module 203 obtains the physical address and error type where the error occurred. Using the error physical address, the second acquisition module 203 queries the memory allocation metadata table maintained by the storage module 202 to quickly determine the data importance level identifier corresponding to the data affected by the error. Next, the policy selection module 204 receives the error type, the error physical address, and the data importance level identifier corresponding to the error memory, and based on this, selects the most appropriate memory error repair or mitigation policy from a preset set of differentiated repair policies. Finally, the policy execution module 205 is responsible for executing the selected policy.
[0096] Through the collaborative work of these modules, the system can correlate physical memory error events with logical data importance and take differentiated countermeasures based on this information, effectively implementing a memory error monitoring and repair method based on data importance level. This system architecture enables the reliable and efficient implementation and coordination of each step in the method, providing support for its effective implementation and operation.
[0097] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.
[0098] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Persons skilled in the art will readily appreciate that the present application may be modified and altered in various ways. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A memory error rate monitoring and repair method, characterized in that: The method comprises the steps of: S1: Get the data importance level identifier in the memory allocation request; S2: In response to the memory allocation request, allocate a physical address range from the physical memory pool, associate the data importance level identifier with the allocated physical address range, and store the data in a memory allocation metadata table; S3: When a memory error is detected, obtaining an error physical address and an error type, and determining the data importance level identifier corresponding to the memory error from the memory allocation metadata table according to the error physical address; S4: selecting a corresponding memory error repair or mitigation strategy from a preset differentiated strategy set according to the error type, the error physical address, and the data importance level identifier corresponding to the error memory; S5: Execute the selected memory error repair or mitigation strategy.
2. A memory error rate monitoring and repair method according to claim 1, characterized in that: Step S1 includes: S11: Receive a memory allocation request, wherein the memory allocation request includes a data importance level identifier and a requested memory size; S12: Verify the validity of the data importance level identifier; if the data importance level identifier is not within a preset level range, replace it with a preset default data importance level identifier, and record the replacement information in the system log; S13: Sending the verified or replaced data importance level identifier and the requested memory size to the memory management module.
3. The memory error rate monitoring and repair method according to claim 1, characterized in that: Step S2 includes: S21: In response to a memory allocation request, select an idle physical address range that meets the requested size from a physical memory pool and generate a page table entry, where the page table entry includes at least a start address and an end address of the physical address range; S22: Determine the corresponding metadata storage structure type according to the data importance level identifier in the memory allocation request; S23: Based on the determined metadata storage structure type, an entry is created in the memory allocation metadata table, with the starting address of the physical address range as the key value, and the data corresponding to the key value including the data importance level identifier and the metadata information of the physical address range is stored in the memory allocation metadata table.
4. A memory error rate monitoring and repair method according to claim 3, characterized in that: Step S22 includes: S221: Establishing a storage correspondence between the data importance level and the metadata storage structure, wherein the storage correspondence includes: a high importance level corresponds to a hash table, a medium importance level corresponds to a B+ tree, and a low importance level corresponds to a linked list; S222: querying the storage correspondence according to the data importance level identifier in the memory allocation request to determine the metadata storage structure type; S223: If the query fails, the preset default metadata storage structure type is used.
5. The memory error rate monitoring and repair method according to claim 1, characterized in that: Step S3 includes: S31: When a memory error is detected, the error physical address and the error type are obtained, and the error type is encoded into an error type enumeration value; S32: Searching for a physical address range entry containing the erroneous physical address in the memory allocation metadata table; if the search fails, determining that it is an illegal memory access, generating an error report containing an illegal access flag, and terminating; S33: If the search is successful, extract the data importance level identifier from the found physical address range entry, and combine it with the error type enumeration value to generate a memory error report containing the data importance level identifier and the error type enumeration value.
6. A memory error rate monitoring and repair method according to claim 5, characterized in that: Step S32 includes: S321: Constructing a sorting index table, which stores the starting addresses of physical address range entries and the offsets of corresponding entries in the memory allocation metadata table, and the physical address range entries are sorted in ascending order according to the starting addresses; S322: searching, according to the erroneous physical address, in the sorted index table by a binary search method for the maximum entry offset whose starting address is less than or equal to the erroneous physical address; S323: If the entry offset is found, read the corresponding physical address range entry from the memory allocation metadata table according to the offset, and determine whether the erroneous physical address falls within the physical address range; S324: If it is not within the range, it is determined to be an illegal memory access, an error report containing an illegal access flag is generated, and the process is terminated; if the entry offset is not found, it is determined to be an illegal memory access, an error report containing an illegal access flag is generated, and the process is terminated.
7. The memory error rate monitoring and repair method according to claim 5, characterized in that: Step S4 includes: S41: Constructing a policy query key value, where the key value includes the error type enumeration value, the data importance level identifier, and the physical address range to which the error physical address belongs; S42: searching for a matching memory error repair or mitigation strategy in a preset differentiated strategy set according to the strategy query key value; S43: If a matching memory error repair or mitigation strategy is found, the strategy is output; if no matching memory error repair or mitigation strategy is found, a preset default memory error repair or mitigation strategy corresponding to the data importance level identifier is selected and output.
8. The memory error rate monitoring and repair method according to claim 7, characterized in that: Step S42 includes: S421: According to the policy query key value, in a preset differentiated policy set, determining whether the error type enumeration value indicates a tolerable error; if so, selecting a preset default memory error mitigation policy corresponding to the data importance level identifier; S422: If an intolerable error is indicated, the error type enumeration value and the data importance level identifier are used as key values, and a Bloom filter is used to filter the preset differentiated strategy set, excluding the repair strategies that do not contain the key value, and then the remaining repair strategies are matched to determine a matching memory error repair strategy.
9. The memory error rate monitoring and repair method according to claim 1, characterized in that: Step S5 includes: S51: Establish a priority queue for executing memory error repair or mitigation strategies; S52: Selecting a policy with the highest priority from the memory error repair or mitigation policy execution priority queue, determining whether current computing system resources meet the requirements for executing the policy, and if so, executing the policy; S53: If not satisfied, the policy is added back to the memory error repair or mitigation policy execution priority queue, the priority of the policy is lowered, and S52 is re-executed; if the policy execution fails, the policy with the next highest priority is selected for execution until the policy execution succeeds or the policy execution priority queue is empty.
10. A memory error rate monitoring and repair system, characterized in that: In the steps of the method according to any one of claims 1 to 9, the system comprises: The first acquisition module: obtains the data importance level identifier in the memory allocation request; A storage module: in response to a memory allocation request, allocates a physical address range from a physical memory pool, associates the data importance level identifier with the allocated physical address range, and stores the data in a memory allocation metadata table; The second acquisition module is configured to acquire an error physical address and an error type when a memory error is detected, and determine the data importance level identifier corresponding to the memory error from the memory allocation metadata table according to the error physical address; A strategy selection module: selects a corresponding memory error repair or mitigation strategy from a preset differentiated strategy set according to the error type, the error physical address, and the data importance level identifier corresponding to the error memory; Policy Execution Module: Executes the selected memory error repair or mitigation strategy.
Citation Information
Patent Citations
Method and system for optimizing system management interrupt processing hardware error time
CN113076213A
SSD storage data error processing method and device, equipment and storage medium
CN114385404A
Memory fault prediction method and system, central processing unit and computing equipment
CN115640174A
Remote replication method and device, equipment and storage medium
CN118152179A
Photoelectromagnetic storage method and system based on hierarchical management
CN119620934A