Memory correctable error handling method and apparatus
Patent Information
- Application Number
- CN202610746822.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-05-27
AI Technical Summary
[0003]本申请提供了一种内存可纠正错误处理方法及设备,以至少解决相关技术中采用固定内存可纠正错误阈值上报策略,导致在错误风暴下频繁触发系统管理中断占用处理器资源,同时因数据类型不同容易导致系统性能损耗与数据安全风险并存的问题
[0005]This application embodiment traverses the processor and its memory list during the server startup phase to obtain the data type and historical access frequency of storage units. It then determines the data sensitivity level based on the data type, determines the correctable error threshold based on the data sensitivity level and historical access frequency, and sets differentiated hierarchical processing strategies for different storage units in conjunction with the data sensitivity level. This achieves adaptive and refined memory error management, proactively distinguishing between critical and non-critical data. It implements stricter monitoring and more timely reporting of highly sensitive data to prevent potential risks, while relaxing the tolerance limits for low-sensitivity data to significantly reduce unnecessary interruptions and performance overhead. Ultimately, it improves overall operating efficiency while ensuring the reliability of critical system data, thereby enhancing the server's reliability, availability, and serviceability.
Smart Images

Figure CN122285364B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of server technology, and in particular to methods and devices for correcting memory errors. Background Technology
[0002] In related technologies, a correctable memory error triggers a system management interrupt, forcing the processor to enter a high-privilege system management mode to execute firmware processing programs, thereby suspending normal computing tasks. However, when a correctable memory error storm occurs, a fixed correctable error threshold reporting strategy can lead to frequent system management interrupts that consume significant processor resources, significantly reducing server performance. Furthermore, applying a uniform fault tolerance threshold and reporting strategy to non-critical and critical data can cause correctable errors in non-critical data to trigger frequent interrupts and reports, generating unnecessary performance overhead and reducing overall server efficiency. On the other hand, it may fail to trigger timely warnings about the reliability degradation trend of critical data, thereby increasing the potential risk of data corruption and system downtime. Summary of the Invention
[0003] This application provides a memory correctable error handling method and device to at least solve the problems in related technologies where the fixed memory correctable error threshold reporting strategy leads to frequent system management interrupts that consume processor resources during error storms, and where different data types can easily cause both system performance loss and data security risks.
[0004] This application provides a method for handling correctable memory errors. The method is applied to the boot system of a server. The boot system communicates with at least one processor, and the processor is connected to at least one memory. A counter is set in the storage unit of the memory. The method includes the following steps: during the server boot process, traversing the processors in a first list of the server and a second list of the processors, and obtaining the stored data of the memory storage units in the second list; identifying the data type and historical access frequency of the stored data, determining the corresponding data sensitivity level according to the data type, and determining the corresponding correctable memory error threshold according to the data sensitivity level and historical access frequency; reading the correctable error count value of the storage unit in the counter, and when the correctable error count value reaches the correctable error threshold, processing the correctable error of the storage unit according to the data sensitivity level. This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the memory-correctable error handling methods described above when executing the computer program.
[0005] This application embodiment traverses the processor and its memory list during the server startup phase to obtain the data type and historical access frequency of storage units. It then determines the data sensitivity level based on the data type, determines the correctable error threshold based on the data sensitivity level and historical access frequency, and sets differentiated hierarchical processing strategies for different storage units in conjunction with the data sensitivity level. This achieves adaptive and refined memory error management, proactively distinguishing between critical and non-critical data. It implements stricter monitoring and more timely reporting of highly sensitive data to prevent potential risks, while relaxing the tolerance limits for low-sensitivity data to significantly reduce unnecessary interruptions and performance overhead. Ultimately, it improves overall operating efficiency while ensuring the reliability of critical system data, thereby enhancing the server's reliability, availability, and serviceability. Attached Figure Description
[0006] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0007] Figure 1 A flowchart illustrating a memory error correction handling method provided in this application embodiment; Figure 2 A schematic diagram of the server startup system interaction interface provided in this application embodiment; Figure 3 A schematic diagram illustrating the option settings provided in the embodiments of this application; Figure 4 This is a physical schematic diagram of a server motherboard provided in an embodiment of this application; Figure 5 This is a schematic diagram of a storage unit provided in an embodiment of this application; Figure 6 An overall logic diagram of memory-correctable error handling provided in the embodiments of this application; Figure 7 A flowchart illustrating the overall process of memory-correctable error handling provided in this application embodiment; Figure 8 This is a block diagram of a memory error correction processing device provided in an embodiment of this application. Detailed Implementation
[0008] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0009] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0010] Most server components currently include several basic peripheral devices, such as CPU (Central Processing Unit) components, memory components, PCI (Peripheral Component Interconnect) device components, storage components, and communication components. These components are the basic component types that any server architecture must include.
[0011] Given the aforementioned basic components, the server can support CPU power-on. During power-on, the CPU's on-chip cache initializes various internal components of the CPU and the server's components. However, due to the limited capacity of the CPU's on-chip cache, it can only initialize some basic and essential core components. After memory initialization is complete, once the memory parameters are confirmed through training, they are written back to the SPD (Serial Presence Detect) data structure of each memory module via the I2C (Inter-Integrated Circuit) bus. At this point, the memory can provide a stable memory address range. Each device informs the BIOS (Basic Input / Output BIOS) of the memory address range, memory start address, or memory capacity required for stable operation or initialization through memory address mapping. The BIOS (Basic Input / Output System) allocates memory address space to each device within a specific address range. Once the server has initialized the memory, it can provide a stable address space range. During the PCI enumeration phase, each device actively reports the memory size range it requires to the BIOS. The requested memory range and space can only be used by the specific device that requested it, and other devices cannot request it again. After the request and allocation are completed, the corresponding devices can use their respective requested memory space to perform read and write operations, i.e., data read and write actions.
[0012] Because memory addresses are contiguous, with a default address length of 64 bits, plus an additional 8 bits for ECC (Error Correcting Code) correction, totaling 72 bits, the memory address mapping of each server component will be mapped to a 64-bit memory address. This 64-bit memory address is composed of multiple memory chips. Assuming one memory chip has x8 bandwidth and a memory module consists of 16 memory chips, then it will have 128 bits of memory bandwidth, or two x64-bit memory bandwidths, known as dual RANK (memory unit). If it consists of 8 memory chips with x64-bit memory bandwidth, it is called single RANK memory. However, in practice, the maximum number of RANKs that can be supported is only 4, regardless of whether the memory chip bandwidth is x2, x4, x8, etc. If the bandwidth of memory chips exceeds x16 or x32 in the future, the number of RANKs will also be upgraded accordingly.
[0013] As explained above, all peripherals perform read and write operations through the memory address space. Because the memory address spaces of each component are segmented and cannot overlap, only the base address of the memory and the required memory size are available. Therefore, each component can only operate within its own memory space. However, since memory addresses are contiguous, it is impossible to pinpoint exactly which specific chip caused the error. But it is possible to pinpoint the RANK level where the error occurred. Because memory addresses are contiguous, memory errors can only be located at the RANK level. Data can be corrupted whenever read or write operations are performed on memory, as the operating frequency of memory and CPU are both at high frequencies. If correctable errors occur in memory and cannot be repaired in time, or if the number of errors exceeds the number of correctable errors, it will lead to a correctable error storm in memory or frequent system management interrupts for error reporting, causing the server to become unusable. Alternatively, correctable memory errors may turn into uncorrectable errors, causing even more fatal errors.
[0014] Therefore, any server component needs to be allocated memory addresses upon startup. During the initialization process of each component, resources such as memory address ranges and base addresses are allocated. Especially during stable server operation, each component performs read and write operations within a fixed allocated memory address range. Although memory modules possess uncorrectable memory error handling capabilities, their ability to handle such errors is limited by various factors. Once the server boots into the operating system, the allocated memory address space for each component is unlocked, enabling read and write operations. Any read or write operation will lead to data errors within the firmware space. Correctable memory errors have a threshold; exceeding this threshold without appropriate measures will result in unforeseen consequences such as system crashes, reboots, uncorrectable memory errors, and performance degradation.
[0015] The essence of correctable errors in memory is single-bit or multi-bit correctable errors occurring at the physical level of dynamic random access memory. The fundamental causes can be attributed to the following physical and electrical mechanisms: (1) Physical disturbances inherent in random access memory: Bit flips: mainly caused by high-energy charged particles (such as alpha particles and thermal neutrons in cosmic rays) bombarding silicon wafers, generating electron-hole pairs and disturbing the charge of storage nodes; Read and write disturbances: frequent or specific read and write operation modes may lead to capacitive coupling or charge leakage between adjacent cells, interfering with non-target cell data; due to capacitor leakage, the stored charge will decay over time. If the charge loss exceeds the tolerance during the refresh cycle, or leakage is aggravated at high temperature, it will lead to data loss. (2) Signal integrity, electrical parameters and aging: In high-speed data transmission, problems such as clock jitter, crosstalk, and reflection may cause the signal timing or amplitude to be at the tolerance edge, increasing the risk of misjudgment. Transistors generate negative bias temperature instability and hot carrier injection effects with continuous electrical stress, leading to device threshold voltage drift and performance degradation. Increased temperature exacerbates leakage current and noise; unstable voltage (ripple, voltage drop) directly affects the stability and read / write reliability of storage cells.
[0016] Therefore, when any one or more of the above factors cause an error in the stored bit state, the error checking and correction circuit built into the memory subsystem or processor will detect and automatically correct the single bit error in real time. After correction, the system will record the event in the hardware error status register and optionally report it through interrupts or other means.
[0017] The historical access frequency of storage units directly affects the rate at which correctable errors are detected and recorded. In areas with frequent read and write operations, the random access memory is activated, sensed, and refreshed more often. Therefore, the probability of errors induced by access-related mechanisms such as read and write disturbances is relatively higher. The correctable error counter in high-access-frequency areas accumulates faster and may reach the preset alarm threshold more quickly, thus affecting the judgment of error severity and resource scheduling strategies.
[0018] Meanwhile, a rapid, concentrated outbreak (storm) of correctable errors is not only a performance issue but also a serious warning sign of system reliability and data integrity, posing the following evolutionary risks: A sustained storm of correctable errors often manifests as an escalation of physical memory failures (such as chip defects, row / column failures) or systemic electrical problems (such as power supply degradation, clock instability). A single-bit error in the same or adjacent memory cells may develop into a multi-bit error. Once this exceeds the error correction capability of the error check code, it leads to uncorrectable errors, triggering system-level errors (such as machine check anomalies), typically resulting in data corruption or system crashes. Furthermore, when high-risk (i.e., highly sensitive, high-value) data is corrupted due to hardware failures, improperly corrected errors, or silent data errors (SDC), it can cause system startup failures (crashes) or fatal runtime errors (blue screens / crashes), resulting in a complete interruption of business services for the entire server or cluster and potentially irreparable losses.
[0019] Therefore, this application can identify the data type and historical access frequency of each memory storage unit, determine the data sensitivity level based on the data type of the storage unit, determine the correctable error threshold based on the data sensitivity level and historical access frequency, and process the correctable errors of the storage unit according to the data sensitivity level when the correctable error count reaches the correctable error threshold. This allows for setting differentiated fault tolerance thresholds for storage units with different sensitivity levels and executing corresponding hierarchical processing strategies. While ensuring the reliability of highly sensitive data, it significantly reduces unnecessary error processing overhead for low-sensitivity data, optimizes the allocation of system reliability resources, and achieves a synchronous improvement in overall operating performance.
[0020] This application's embodiments relate to the specific application environment architecture or hardware architecture upon which the execution of the memory-correctable error handling method depends. The specific application environment architecture or hardware architecture is described herein, such as... Figure 4 As shown.
[0021] The server's system architecture includes a CPU (Central Processing Unit), a Dual In-line Memory Module (DIM), a BIOS (Basic Input / Output System), a BMC (Baseboard Management Controller), and a VGA (Video Graphics Array).
[0022] The CPU connects to the memory module via the I2C bus to read memory parameter information and complete memory initialization. It communicates with the BIOS system via the SPI (Serial Peripheral Interface) or LPC (Low Pin Count) bus.
[0023] Each BIOS system communicates with multiple CPUs via LPC / SPI / QSPI (Quad Serial Peripheral Interface) buses to receive boot commands and execute initialization processes. It interacts with the BMC via IPMI (Intelligent Platform Management Interface) / SM Bus (System Management Bus) to transmit system status information. The BIOS runs first after the server boots up, completing hardware self-tests, loading memory parameters, initializing devices, and finally transferring control to the operating system. It is also responsible for configuring memory correctable error thresholds and handling mechanisms. These settings are read during the boot process and applied to runtime monitoring. Among these settings, timers are configured in the BIOS.
[0024] like Figure 5 As shown, the memory module contains multiple DRAM (Dynamic Random-Access Memory) chips, divided into one or more RANKs. Each RANK provides a 64-bit (or 72-bit including error correction codes) data width and is one of the basic units for memory access and management. It should be noted that each RANK in the memory module contains a counter used to count correctable memory errors.
[0025] The BMC communicates with the CPU via LPC / SPI / QSPI to obtain system event logs and hardware status, exchanges information with the BIOS via IPMI / SM Bus, and connects to an external monitor or remote management terminal via a VGA interface or network interface. The VGA is used to connect to the monitor and allows the BMC to output debugging information or system status interface.
[0026] In this embodiment, the BIOS is responsible for providing and executing a monitoring policy for correctable memory errors. Specifically, during the startup phase, the BIOS provides an interactive interface to the user, displaying a first setting button and a second setting button. The first setting button guides the user in configuring a count threshold for correctable memory errors; the second setting button guides the user in configuring the target handling mechanism (such as disabling reporting, single reporting, or frequent reporting) when the error count exceeds the threshold. The configuration completed by the user is saved by the BIOS and used to make decisions and execute strategies during server runtime.
[0027] The BMC also serves as a key recorder and alarm terminal for error information. When the BIOS decides to report an error based on the above policy, it will send a log containing specific error RANK location information (such as CPU, channel, memory module, RANK identifier) to the BMC through interfaces such as IPMI (Intelligent Platform Management Interface). The BMC will persistently store this log and can trigger remote alarms based on it, thereby improving fault location and maintenance efficiency.
[0028] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] Specifically, Figure 1 This application provides a flowchart illustrating a memory error correction handling method according to an embodiment of the present application.
[0030] like Figure 1 As shown, this memory-correctable error handling method is applied to the boot system of a server. The boot system communicates with at least one processor, which is connected to at least one memory. A counter is set in the storage unit of the memory. The method includes the following steps: In step S101, during the server startup process, the processors in the first list of the server and the processors in the second list are traversed to obtain the storage data of the memory storage units in the second list.
[0031] The first list of processors can be a collection of all processors in the current server; the second list can be a collection of memory that each processor is directly connected to or can access, without specific limitations.
[0032] It is understood that the embodiments of this application can scan the storage contents of all processors and associated memory units during the server startup process, thereby achieving a global awareness of the memory data layout and status of the entire system, so as to facilitate the subsequent determination of the corresponding correctable error thresholds and processing strategies.
[0033] It should be noted that the boot system in this embodiment is BIOS, and no specific limitation is made.
[0034] Specifically, during BIOS execution, a callback function is used for callback. The function uses a system management interrupt to check whether the memory in the memory channel memory module slot under the CPU identifier of the server is present. If it is not present, it continues to poll the next memory slot. If it is present, it records the specific location information of this memory slot, including but not limited to the processor identifier, memory channel identifier, memory module identifier, and storage unit identifier. At the same time, it obtains the stored data of the memory storage unit so that the data type can be determined later based on the stored data.
[0035] In step S102, the data type and historical access frequency of the stored data are identified, the corresponding data sensitivity level is determined according to the data type, and the corresponding memory correctable error threshold is determined according to the data sensitivity level and historical access frequency.
[0036] It is understood that the embodiments of this application can determine the corresponding data sensitivity level by the data type of the stored data, and determine the corresponding memory correctable error threshold based on the data sensitivity level and historical access frequency. Thus, data with high sensitivity or high frequency of access is assigned a lower threshold to achieve stricter error monitoring and earlier fault warning, while data with low sensitivity or low frequency of access adopts a relatively lenient threshold to reduce unnecessary interruption overhead. This achieves deep coupling between memory error management strategy and business data characteristics, and optimizes the overall system performance and resource utilization efficiency while ensuring the integrity of critical data.
[0037] In this embodiment of the application, determining the corresponding data sensitivity level based on the data type includes: obtaining a pre-set second relation mapping table, wherein the second relation mapping table is a mapping relationship table between data types and data sensitivity levels; querying the second relation mapping table using the data type as an index to determine the corresponding data sensitivity level.
[0038] It is understood that the embodiments of this application can establish a correspondence between data types and data sensitivity levels by obtaining a pre-set second relation mapping table, thereby using the identified data type as an index to query the mapping table, quickly and accurately determining its corresponding data sensitivity level, making the sensitivity determination process standardized, configurable and easy to expand.
[0039] It should be noted that the second relationship mapping table can be set according to actual needs. The mapping table is a predefined configuration rule in the policy engine, which is configured by the administrator in the BIOS or BMC interface. It supports customization according to business needs. Data types can be obtained through memory attribute tags passed by the operating system, virtualization layer metadata, or application annotations. The sensitivity level can be understood as a comprehensive assessment result of the importance, confidentiality, integrity, and business impact of the data. It is usually divided into three levels: high, medium, and low. It can also be extended to a numerical score (such as 1-10 points) to facilitate quantitative decision-making without specific limitations.
[0040] Taking enterprise data as an example, the second mapping relationship table is shown in Table 1 below.
[0041] Table 1
[0042] As shown in Table 1, the running log is an internal log used for system debugging and does not contain user or business data, so its data sensitivity is relatively low. On the other hand, financial transaction records contain core financial data such as user payments and transfers, so their data sensitivity is the highest.
[0043] Specifically, the storage unit has a built-in configurable sensitivity level mapping table, which administrators can customize according to security policies. For example, data involving confidentiality or security, such as "source code", "financial statements", and "user credentials", have the highest sensitivity level, such as a numerical score of 9; data such as "project documents" and "internal emails" have a medium sensitivity level; and data such as "public information" and "operation logs" have the lowest sensitivity level, such as a numerical score of 2, without specific restrictions.
[0044] In this embodiment of the application, determining the corresponding correctable memory error threshold based on the data sensitivity level and historical access frequency includes: obtaining a pre-set first relational mapping table, wherein the first relational mapping table is a mapping relationship table between the data sensitivity level, historical access frequency and the correctable memory error threshold; querying the first relational mapping table using the data sensitivity level and historical access frequency as indexes to determine the corresponding correctable memory error threshold.
[0045] It is understood that the embodiments of this application can directly map the data sensitivity level and historical access frequency as a combined index to the corresponding memory correctable error threshold through a preset first relation mapping table, thereby achieving fast, deterministic, and configurable adaptive threshold setting: assigning lower thresholds to highly sensitive and frequently accessed areas to trigger detection and handling earlier; assigning higher thresholds to low-sensitivity and sparsely accessed areas to reduce unnecessary reporting and interruptions, avoiding performance loss and alarm noise caused by a one-size-fits-all strategy, suppressing error accumulation and storm risk, and the mapping table facilitates unified management and policy updates, enabling fine-grained differentiated control without complex calculations, improving the accuracy and timeliness of error detection, and enhancing the reliability and stability of the system.
[0046] It should be noted that the first relationship mapping table can be set according to actual needs. The mapping table is a predefined configuration rule in the policy engine, which is configured by the administrator in the BIOS or BMC interface. It supports customization according to business needs and is not specifically limited.
[0047] For example, the first mapping relationship table is shown in Table 2 below, which can be set according to actual needs.
[0048] Table 2
[0049] Specifically, as shown in Table 2 above, if the memory block is identified as storing an encryption key, and its data sensitivity level is found to be 9 points through the second relational mapping table, and if the memory area is detected to have been accessed approximately 1200 times / second in the past hour, it is determined to be high-frequency access. The pre-set first relational mapping table is retrieved, and the table is queried using the composite index. If the correctable error threshold is found to be 30 times / 24 hours, then stricter correctable error monitoring is enabled for the memory unit. When the correctable error count reaches the correctable error threshold, a system management interrupt is immediately triggered, and the correctable error data is reported to the BMC.
[0050] In step S103, the correctable error count value of the storage cell in the counter is read. When the correctable error count value reaches the correctable error threshold, the correctable error of the storage cell is processed according to the data sensitivity level.
[0051] It is understood that the embodiments of this application can read the correctable error count value of the storage unit in the counter, and when the count value reaches the correctable error threshold dynamically set based on the data sensitivity level and historical access frequency, trigger a processing mechanism that matches the data sensitivity level. The system can implement stricter error response strategies for highly sensitive or critical business data, such as immediate reporting or access restriction, while adopting lenient processing for low-sensitivity data to reduce performance overhead. This realizes the transformation of memory error management from a static unified strategy to dynamic differentiated protection, which optimizes the overall system operating efficiency and resource utilization while ensuring the integrity of core data.
[0052] It should be noted that before reading the correctable error count value of the storage cell in the counter, it is necessary to identify the current state of the storage cell in memory. Only when the storage cell is in the pre-set in-place state will the correctable error count value be read from its internal counter and counted. This ensures that error monitoring is accurately applied to the actual installed and usable memory resources, avoids invalid access to empty slots or faulty offline cells, improves the accuracy of error collection and system initialization efficiency. The pre-set in-place state is a state that is physically present, electrically connected, and can be recognized and accessed by the system, without specific limitations.
[0053] In this embodiment, processing correctable errors of the storage unit according to the data sensitivity level includes: reading a pre-set first sensitivity level threshold; if the data sensitivity level is less than or equal to the first sensitivity level threshold, disabling the reporting function of correctable errors in the memory, and not triggering an interrupt operation when the correctable errors in the memory are detected to reach the correctable error threshold; if the data sensitivity level is greater than the first sensitivity level threshold, enabling the reporting function of correctable errors in the memory, triggering an interrupt operation when the correctable errors in the memory are detected to reach the correctable error threshold, and reporting the correctable errors of the storage unit according to the reporting function.
[0054] The threshold for the first sensitivity level can be set according to actual needs, without specific limitations.
[0055] It is understood that the embodiments of this application can intelligently determine whether to enable the reporting function of memory-correctable errors by reading a pre-set first sensitivity level threshold and comparing the actual data sensitivity level of the storage unit with the first sensitivity level threshold. For data with a sensitivity level less than or equal to the threshold, the reporting function is turned off and no interruption is triggered when the error reaches the threshold, effectively avoiding unnecessary performance overhead caused by low-risk data. For data with a sensitivity level higher than the threshold, the reporting function is turned on, and an interruption is triggered and error information is reported when the error count reaches the threshold, ensuring that abnormal behavior of critical data is captured and processed in a timely manner, thereby improving the overall operating efficiency of the system while ensuring the reliability of highly sensitive data.
[0056] For example, the data sensitivity level of the public static resources cached in the first storage unit is low. When the correctable error count of the storage unit reaches the set correctable error threshold, no system management interrupt is triggered, and no correctable error data is reported to the BMC. However, the database index blocks stored in the second storage unit have a data sensitivity level of medium. When the correctable error in the memory is detected to reach the correctable error threshold, an interrupt operation is triggered, and the correctable error of the storage unit is reported according to the reporting function.
[0057] Specifically, if the data sensitivity level is less than or equal to the first sensitivity level threshold, then it is equivalent to... Figure 3 The first processing mechanism (shutdown policy) is as follows: when the memory correctable error count reaches the memory correctable error threshold, no processing is performed and no system management interrupt is triggered for reporting. The server continues to operate normally. If the data sensitivity level exceeds the first sensitivity level threshold, it is equivalent to... Figure 3 The second processing mechanism (one-time strategy) is as follows: when the memory correctable error counter value reaches the memory correctable error threshold, only one system management interrupt is triggered, and the memory correctable error is reported to the BMC terminal for recording, and then the memory correctable error counter value statistics are turned off.
[0058] In this embodiment of the application, reporting correctable errors of the storage unit according to the reporting function includes: obtaining correctable error events of memory and corresponding location identifiers, wherein the location identifiers include target processor identifier, target memory channel identifier, target storage module slot identifier and target storage unit identifier; generating target correctable error data according to the correctable error events of memory and corresponding location identifiers, and reporting the target correctable error data to the baseboard management controller.
[0059] It is understood that the embodiments of this application can obtain correctable memory error events and their corresponding location identifiers, including target processor identifier, target memory channel identifier, target storage module slot identifier, and target storage unit identifier, and generate target correctable error data containing complete information and report it to the baseboard management controller. This achieves accurate location of memory errors, enabling maintenance personnel to quickly identify the processor, channel, physical slot, and logic unit where the correctable error is located without having to troubleshoot step by step, greatly shortening fault diagnosis time and improving remote maintenance efficiency.
[0060] In this embodiment of the application, after reporting correctable errors of the storage unit according to the reporting function, the process includes: obtaining the number of times correctable errors of the memory have been reported; reading a pre-set second sensitivity level threshold, wherein the second sensitivity level threshold is greater than the first sensitivity level threshold; and disabling the reporting function of correctable errors of the memory according to the number of reports and the second sensitivity level threshold.
[0061] It is understood that the embodiments of this application can determine whether to close the reporting function by introducing a joint judgment of the number of reports and the second sensitivity level threshold (higher than the first threshold) after a report is completed. This achieves hierarchical throttling based on business importance, ensuring timely alarms on the first time, and adaptively converging subsequent events to avoid the impact of error storms on performance and stability, while taking into account both the reliability of key data and the efficiency of system operation.
[0062] In this embodiment of the application, disabling the reporting function of correctable errors in memory based on the number of reports and a second sensitivity level threshold includes: if the data sensitivity level is less than or equal to the second sensitivity level threshold, then disabling the reporting function of correctable errors in memory when the number of reports exceeds a preset first threshold; if the data sensitivity level is greater than the second sensitivity level threshold, then disabling the reporting function of correctable errors in memory when the number of reports exceeds a preset second threshold, wherein the second threshold is greater than the first threshold.
[0063] The thresholds for the second sensitivity level, the first count, and the second count can all be set according to actual needs without specific limitations.
[0064] It is understood that the embodiments of this application can achieve adaptive throttling based on the priority of business importance by linking the reporting shutdown condition with data sensitivity and setting graded reporting number thresholds for different sensitivities. In low-sensitivity areas, reporting is shut down when the number of reporting times exceeds a certain threshold, quickly suppressing interruptions and log noise caused by non-critical errors. In high-sensitivity areas, reporting is maintained until a higher number of reporting times threshold is reached, ensuring the continuous visibility and timely handling of critical errors. While ensuring the intensity of critical data monitoring, the impact of error storms on performance and stability is avoided, thereby improving the overall reliability and operating efficiency of the system.
[0065] In this embodiment of the application, before disabling the reporting function of correctable errors in memory when the number of reports exceeds a preset second threshold, the method further includes: obtaining the reporting period of memory; if the number of reports within the reporting period is less than or equal to the second threshold, then restarting the timing and continuing to enable the reporting function of correctable errors in memory; if the number of reports within the reporting period exceeds the second threshold, then disabling the reporting function of correctable errors in memory.
[0066] The reporting period can be set according to actual needs, such as 24 hours, without any specific limit.
[0067] It is understood that the embodiments of this application can introduce a reporting period as a time window before shutting down reporting, and make decisions by comparing the number of reports within the period with the second threshold. When the number of reports within the period does not exceed the threshold, the timer is reset and reporting continues to be enabled, ensuring the visibility of intermittent or sporadic errors. When the number of reports within the period exceeds the threshold, reporting is shut down to quickly suppress persistent error storms. This can distinguish between short-lived and continuously accumulated errors, avoid accidental shutdown due to occasional jitter, and achieve rapid convergence when a real storm occurs. It dynamically balances visibility and overhead, significantly reduces system management interruptions and log load, and improves system performance and stability.
[0068] For example, when the number of reports exceeds a pre-set second threshold, it is equivalent to... Figure 3 The third processing mechanism (frequent strategy) is as follows: When the memory correctable error counter reaches the memory correctable error threshold, a system management interrupt is triggered, and the memory correctable error is reported to the BMC for recording. At the same time, the memory correctable error count continues, and the timer keeps counting. It is then determined whether the memory correctable error threshold is reached again within a time period (e.g., 24 hours). If the threshold is reached again, the memory correctable error is reported again via a system management interrupt and transmitted to the BMC for recording. If the target number (e.g., 10 times) of the memory correctable error threshold is not exceeded within the time period, only the memory correctable error is reported. If 10 memory correctable errors are reached within 24 hours, a system management interrupt is triggered, and the timer restarts. If more than 10 memory correctable error reports are received within 24 hours, the memory correctable error counter is closed after the memory correctable error is reported via a system management interrupt, and the system can then operate normally.
[0069] In this embodiment of the application, after processing the correctable errors of the storage unit according to the data sensitivity level, the method includes: responding to a user-triggered memory correctable error threshold change instruction, updating the memory correctable error threshold according to the memory correctable error threshold change instruction; and processing the memory correctable errors according to the updated memory correctable error threshold.
[0070] It is understood that the embodiments of this application can be used to implement a response and immediate effect mechanism for user threshold change instructions, so that the memory correctable error threshold can be dynamically adjusted as needed during runtime, avoiding reliance on restarts or firmware upgrades to complete policy reconstruction, thereby keeping the monitoring intensity consistent with business priorities, load changes and compliance requirements in real time, and improving the controllability of operation and maintenance and the efficiency of configuration management.
[0071] It should be noted that this application can not only modify correctable memory errors, but also modify memory correctable error handling mechanisms triggered by users, without making specific limitations.
[0072] Specifically, such as Figure 2 As shown, in response to the user's first interactive action, the system enters the interactive interface for starting the system. When the first setting button is triggered, a first setting sub-interface is displayed on the interactive interface. Responding to the user's interactive action, the correctable memory error threshold is set within the first setting sub-interface. Therefore, the user can change the correctable memory error threshold within the first setting sub-interface. In addition to fixed options, the setting of the correctable memory error threshold also provides custom options, facilitating modification according to the user's intentions. The interactive action can be a specific operation performed by the user in the first setting sub-interface to set the correctable memory error threshold. This can include: using the keyboard arrow keys or mouse to select a numerical input box and inputting a specific error count threshold using the number keys; or adjusting a preset threshold slider or option (such as low, medium, high, or a specific numerical level) using the up and down arrow buttons or scroll wheel; or selecting a system-predefined threshold strategy template (such as conservative mode, performance priority, custom, etc.) from a drop-down menu, without specific limitations.
[0073] When the second setting button is triggered, the second setting sub-interface is displayed on the interactive interface to respond to the user's interaction. Within the second setting sub-interface, the target processing mechanism for when the correctable error count value of memory reaches the correctable error threshold is set. The interaction can be a specific operation that the user actively performs in the second setting sub-interface to configure the target processing mechanism after the correctable error count value of memory reaches the threshold. Specifically, it can include selecting from multiple preset processing strategy options (such as turning off reporting, single reporting, frequency-limited reporting, or triggering system alarms) via keyboard, mouse, or touch, or adjusting parameter sliders or input boxes related to the processing mechanism (such as setting the maximum number of reporting times, reporting time interval, etc.), or customizing the target processing mechanism options, etc., without specific limitations.
[0074] In this embodiment of the application, before processing the correctable errors of the storage unit according to the data sensitivity level, the method includes: identifying whether the data type of the stored data has changed; if the data type has changed, updating the data sensitivity level according to the changed data type; and processing the correctable errors of the storage unit according to the updated data sensitivity level.
[0075] It is understood that the embodiments of this application can continuously identify changes in the data type of stored data before correctable error handling and re-implement correctable error policies according to the new sensitivity level. This allows monitoring and threshold settings to be aligned with data risks in real time, avoiding policy mismatch and outdated configurations caused by page reuse, load migration, or memory reallocation. When data type changes increase sensitivity, thresholds are tightened and reporting is strengthened in a timely manner to capture critical errors in advance. When sensitivity decreases, thresholds are appropriately relaxed and interruptions and log noise are suppressed to reduce performance overhead. This runtime adaptive capability improves the timeliness and accuracy of error governance, reduces the risk of missed and over-reported errors, and maintains system performance, stability, and reliability without requiring a restart.
[0076] According to the memory correctable error handling method of this application embodiment, the processor and its memory list are traversed during the server startup phase to obtain the data type and historical access frequency of the storage unit. The data sensitivity level is determined according to the data type, and the correctable error threshold is determined based on the data sensitivity level and historical access frequency. Then, a differentiated hierarchical processing strategy is set for different storage units in combination with the data sensitivity level. This achieves adaptive and refined memory error management, actively distinguishes between critical and non-critical data, implements stricter monitoring and more timely reporting of highly sensitive data to prevent potential risks, and relaxes the fault tolerance limit for low-sensitivity data to significantly reduce unnecessary interruptions and performance overhead. Ultimately, while ensuring the reliability of critical system data, the overall operating efficiency is improved, and the reliability, availability, and serviceability of the server are enhanced.
[0077] The following will combine Figure 6-7 This application provides a detailed description of the memory correctable error handling method. The method primarily utilizes a memory module, a BIOS module, and a BMC module. The memory module comprises multiple logical units, each with an independent correctable error counter to record the number of correctable memory errors and to generate and record correctable error events in memory. The BIOS module includes the following sub-functions: receiving correctable error thresholds and various processing mechanisms configured by the user through the BIOS interface; executing corresponding strategies based on the user-selected processing mechanism when the correctable error count reaches the correctable error threshold, such as disabling the reporting function, triggering a system management interrupt, and encapsulating the correctable error data before reporting it to the BMC module. The BMC module receives the correctable error data reported from the BIOS and parses the target processor identifier, target memory channel identifier, target storage module slot identifier, and target storage unit identifier to generate an event log. The specific steps are as follows: Step 1: Start the server.
[0078] After the server is powered on, it begins the boot process. During this stage, the BIOS controls the system to enter the initialization state, preparing to load the configuration and start the operating system.
[0079] Step 2, BIOS interactive interface settings.
[0080] The system boots up by first loading the BIOS. Based on user commands, the user can choose whether to enter the setup interface, which can be accessed via a key to configure parameters. At this stage, the system provides two key configuration entry points: a first setting button for setting the correctable memory error threshold; and a second setting button for setting the target handling mechanism after exceeding the threshold. Users can configure or change each memory module through interactive actions. Different memory modules may exhibit different error behaviors due to manufacturing differences, installation location, workload, or aging. By independently configuring correctable error thresholds and handling mechanisms for each memory module, the system can enhance monitoring and alarms for high-risk modules and reduce interrupt overhead for stable modules, thereby optimizing the overall balance between reliability and performance. These settings will be saved to non-volatile memory for later use.
[0081] Step 3: Traverse the CPU's memory channels.
[0082] The system starts with the first CPU and iterates through each memory channel in turn, determining whether the currently processed memory channel belongs to the memory domain of the first CPU. If it does, it proceeds to the next step; otherwise, it jumps to the next memory channel to continue checking.
[0083] Step 4: Check if the memory module of the memory channel is in place.
[0084] For the current memory channel, the system checks whether the connected memory module is physically installed and electrically normal. It determines whether the memory module is in place by reading SPD information or detecting the chip select signal response. If it is not in place (e.g., the slot is empty or faulty), the system skips the memory channel and continues to process the next memory channel. If it is in place, the system obtains the memory information (including address, capacity, timing, etc.), data type, and historical access frequency of the current memory module. The system determines the corresponding data sensitivity level based on the data type, as shown in Table 1 above. The system determines the corresponding memory correctable error threshold based on the data sensitivity level and historical access frequency, as shown in Table 2 above.
[0085] Step 5: Obtain the memory error event and its location identifier.
[0086] The system reads correctable error counter data from each storage unit of the in-situ memory module, and simultaneously collects complete location identification information, including: target processor identifier, target memory channel identifier, target storage module slot identifier, and target storage unit identifier; and packages the above information with correctable error events to generate target correctable error data.
[0087] Step 6: Determine whether the memory correctable count has reached the threshold.
[0088] The correctable error count of the read storage unit is compared with the corresponding correctable error threshold. If the threshold is not reached, no action is triggered, and the next memory channel or memory module is checked. If the threshold is reached or exceeded, the target processing mechanism is entered, that is, the corresponding strategy is executed according to the data sensitivity level of the current storage unit.
[0089] Step 7: Determine the response strategy based on the target processing mechanism.
[0090] The system invokes the target processing mechanism selected by the user in the second settings sub-interface, specifically including three modes: First processing mechanism: disables the reporting function and no longer triggers system management interrupts. Second processing mechanism: triggers a system management interrupt and reports correctable memory error data to the BMC, but automatically disables the reporting function after the number of reports exceeds the first threshold. Third processing mechanism: reports periodically, but disables the function when the number of reports exceeds the second threshold within a preset time window.
[0091] If the data sensitivity level is less than or equal to the first sensitivity level threshold, and if the first processing mechanism is enabled, then when the number of correctable errors in memory reaches the correctable error threshold, reporting will be directly turned off and the current processing will end.
[0092] If the data sensitivity level is greater than the first sensitivity level threshold, the reporting function for correctable errors in memory is enabled. When the number of correctable errors in memory reaches the correctable error threshold, an interrupt operation is triggered, and the correctable errors of the storage unit are reported according to the reporting function. If the data sensitivity level is less than or equal to the second sensitivity level threshold, when the number of reports exceeds the preset first number threshold, the second processing mechanism is enabled. When the value of the correctable error counter in memory reaches the correctable error threshold in memory, only one system management interrupt is triggered, and the correctable error in memory is reported to the BMC. Then, the statistics of the correctable error counter value are turned off, and the system operates normally at this time.
[0093] If the data sensitivity level is greater than the second sensitivity level threshold, then when the number of reports exceeds the preset second threshold, a third processing mechanism is activated. Specifically, when the memory correctable error counter reaches the memory correctable error threshold, a system management interrupt is triggered to report memory correctable errors to the BMC for recording. Simultaneously, the memory correctable error count continues, and a timer keeps track of the time. If the memory correctable error threshold is reached again within 24 hours, a system management interrupt is triggered to report memory correctable errors to the BMC for recording. If the threshold is not exceeded 10 times within 24 hours, only memory correctable error reporting is required. If 10 memory correctable error reports are received within 24 hours, the timer restarts. If more than 10 memory correctable error reports are received within 24 hours, the memory correctable error counter is closed after the system management interrupt reporting is completed, and the system can then operate normally.
[0094] After a report is triggered, the correctable error counter is reset based on the selected processing mechanism: if the third processing mechanism is used and the number of reports has not exceeded the second threshold, the count will restart after the report is completed; if the second processing mechanism is used and the maximum number of reports has been reached, subsequent reports will be turned off; if the first processing mechanism is used, the count will remain unchanged and only reporting will be stopped.
[0095] Step 8: Check if this is the last memory channel.
[0096] The system determines whether all memory channels have been traversed. If not, it returns to process the next memory channel. If all memory channels have been traversed, it enters the final processing stage.
[0097] Step 9: Enter the operating system and run.
[0098] After all memory error correction monitoring and handling processes are completed, the BIOS continues the boot process, the server enters the operating system, the operating system takes over control, and normal operation begins.
[0099] In summary, this application leverages the inherent characteristics of memory by adding two BIOS options to the server's BIOS management firmware. One BIOS option sets the correctable error threshold for all memory in the server. The other BIOS option, based on user needs, determines whether to trigger a system management interrupt to report fault location information when correctable memory errors reach the threshold. This option can be set to disabled (no reporting), single reporting, or frequent reporting. In frequent reporting, the default maximum number of reports is 10 within 24 hours. If the number of reports does not exceed 10 within 24 hours, the timer is reset and the timing process restarts. The reporting mechanism for the correctable memory error threshold... It can accurately locate all CPU location information, memory channel information, and memory module information in the server, precisely pinpoint the memory with correctable error thresholds, and report it to the BMC (Browser Management Center). The BMC can clearly record and display which memory module on the server has generated a correctable memory error exceeding the correctable error threshold, allowing for timely replacement to avoid affecting normal user operation or causing server downtime. At the same time, it limits the number of correctable memory error threshold reports to prevent excessively frequent triggering of system management interrupts, reducing the core technical problem of server performance degradation caused by frequent system management interrupts, and improving server security, stability, and reliability.
[0100] Embodiments of this application also provide a memory-correctable error handling apparatus.
[0101] like Figure 8 As shown, the memory error correction processing device 10 is applied to the boot system of a server. The boot system communicates with at least one processor, the processor is connected to at least one memory, and a counter is provided in the storage unit of the memory. The memory error correction processing device 10 includes: an acquisition module 100, an identification module 200, and a reading module 300.
[0102] The acquisition module 100 is used to traverse the processors in the first list and the second list of processors in the server during the server startup process, and acquire the stored data of the memory storage units in the second list; the identification module 200 is used to identify the data type and historical access frequency of the stored data, determine the corresponding data sensitivity level according to the data type, and determine the corresponding memory correctable error threshold according to the data sensitivity level and historical access frequency; the reading module 300 is used to read the correctable error count value of the storage unit in the counter, and when the correctable error count value reaches the correctable error threshold, process the correctable error of the storage unit according to the data sensitivity level.
[0103] In this embodiment, the reading module 300 is further configured to: read a pre-set first sensitivity level threshold; if the data sensitivity level is less than or equal to the first sensitivity level threshold, disable the reporting function of correctable errors in memory, and do not trigger an interrupt operation when the correctable errors in memory are detected to reach the correctable error threshold; if the data sensitivity level is greater than the first sensitivity level threshold, enable the reporting function of correctable errors in memory, trigger an interrupt operation when the correctable errors in memory are detected to reach the correctable error threshold, and report the correctable errors of the storage unit according to the reporting function.
[0104] In this embodiment of the application, the reading module 300 is further configured to: obtain the number of correctable errors reported in the memory; read a pre-set second sensitivity level threshold, wherein the second sensitivity level threshold is greater than the first sensitivity level threshold; and disable the reporting function of correctable errors in the memory based on the number of reports and the second sensitivity level threshold.
[0105] In this embodiment of the application, the reading module 300 is further configured to: disable the reporting function of correctable errors in memory based on the number of reports and the second sensitivity level threshold, including: if the data sensitivity level is less than or equal to the second sensitivity level threshold, then disable the reporting function of correctable errors in memory when the number of reports is greater than a preset first threshold; if the data sensitivity level is greater than the second sensitivity level threshold, then disable the reporting function of correctable errors in memory when the number of reports is greater than a preset second threshold, wherein the second threshold is greater than the first threshold.
[0106] In this embodiment of the application, it further includes: a processing module, used to obtain the memory reporting cycle before closing the memory correctable error reporting function when the number of reports exceeds a preset second number threshold; if the number of reports within the reporting cycle is less than or equal to the second number threshold, then the timing is restarted and the memory correctable error reporting function is continued to be enabled; if the number of reports within the reporting cycle exceeds the second number threshold, then the memory correctable error reporting function is disabled.
[0107] In this embodiment of the application, the identification module 200 is further configured to: obtain a pre-set first relational mapping table, wherein the first relational mapping table is a mapping relationship table between data sensitivity level and historical access frequency and memory correctable error threshold; and query the first relational mapping table using data sensitivity level and historical access frequency as indexes to determine the corresponding memory correctable error threshold.
[0108] In this embodiment of the application, the identification module 200 is further configured to: obtain a pre-set second relation mapping table, wherein the second relation mapping table is a mapping relationship table between data types and data sensitivity levels; and query the second relation mapping table using data types as indexes to determine the corresponding data sensitivity levels.
[0109] In this embodiment of the application, it further includes: a response module, configured to respond to a user-triggered memory correctable error threshold change instruction, update the memory correctable error threshold according to the memory correctable error threshold change instruction, and process memory correctable errors according to the updated memory correctable error threshold.
[0110] In this embodiment of the application, it further includes: an update module, used to identify changes in the data type of the stored data; and to process correctable errors in the storage unit according to the data sensitivity level corresponding to the changed data type.
[0111] It should be noted that the specific implementation of the memory correctable error handling device in this embodiment of the invention is similar to the specific implementation of the memory correctable error handling method. To reduce redundancy, it will not be described in detail here.
[0112] The memory error correction handling device according to the embodiments of this application traverses the processor and its memory list during the server startup phase to obtain the data type and historical access frequency of the storage unit, determines the data sensitivity level based on the data type, determines the correctable error threshold based on the data sensitivity level and historical access frequency, and then sets differentiated hierarchical processing strategies for different storage units in combination with the data sensitivity level. This achieves adaptive and refined memory error management, actively distinguishes between critical and non-critical data, implements stricter monitoring and more timely reporting of highly sensitive data to prevent potential risks, and relaxes the fault tolerance limit for low-sensitivity data to significantly reduce unnecessary interruptions and performance overhead. Ultimately, it improves the overall operating efficiency while ensuring the reliability of critical system data, and enhances the reliability, availability, and serviceability of the server.
[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0114] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the memory-correctable error handling method.
[0115] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the memory-correctable error handling method when running.
[0116] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0117] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the memory correctable error handling method.
[0118] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the memory-correctable error handling method.
[0119] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0120] The foregoing has provided a detailed description of a memory error correction handling method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A memory correctable error handling method, characterized by, The method is applied to a server boot system, which communicates with at least one processor, which is connected to at least one memory, and the memory's storage units contain counters. The method includes the following steps: During server startup, the processors in the first list of the server and the second list of the processors are traversed to obtain the storage data of the memory storage units in the second list; Identify the data type and historical access frequency of the stored data, determine the corresponding data sensitivity level based on the data type, and determine the corresponding memory correctable error threshold based on the data sensitivity level and the historical access frequency; Read the correctable error count value of the storage cell in the counter, and when the correctable error count value reaches the correctable error threshold, process the correctable error of the storage cell according to the data sensitivity level; The step of processing the correctable errors of the storage unit according to the data sensitivity level includes: reading a pre-set first sensitivity level threshold; if the data sensitivity level is greater than the first sensitivity level threshold, then enabling the reporting function of the correctable errors of the memory; triggering an interrupt operation when the correctable errors of the memory reach the correctable error threshold; and reporting the correctable errors of the storage unit according to the reporting function. After reporting the correctable errors of the storage unit according to the reporting function, the process includes: obtaining the number of times the correctable errors of the memory have been reported; reading a pre-set second sensitivity level threshold, wherein the second sensitivity level threshold is greater than the first sensitivity level threshold; and disabling the reporting function of the correctable errors of the memory according to the number of reports and the second sensitivity level threshold. The step of disabling the reporting function of correctable errors in the memory based on the number of reports and the second sensitivity level threshold includes: if the data sensitivity level is less than or equal to the second sensitivity level threshold, then disabling the reporting function of correctable errors in the memory when the number of reports is greater than a preset first number threshold; if the data sensitivity level is greater than the second sensitivity level threshold, then disabling the reporting function of correctable errors in the memory when the number of reports is greater than a preset second number threshold, wherein the second number threshold is greater than the first number threshold; Before disabling the reporting function of correctable errors in the memory when the number of reports exceeds a preset second threshold, the method further includes: obtaining the reporting period of the memory; if the number of reports within the reporting period is less than or equal to the second threshold, restarting the timing and continuing to enable the reporting function of correctable errors in the memory; if the number of reports within the reporting period exceeds the second threshold, disabling the reporting function of correctable errors in the memory.
2. The memory correctable error handling method of claim 1, wherein, The process of handling correctable errors in the storage unit according to the data sensitivity level includes: If the data sensitivity level is less than or equal to the first sensitivity level threshold, the reporting function of correctable errors in the memory is turned off, and no interruption operation is triggered when the correctable errors in the memory are detected to reach the correctable error threshold.
3. The memory correctable error handling method of claim 1, wherein, The step of determining the corresponding correctable memory error threshold based on the data sensitivity level and the historical access frequency includes: Obtain a pre-set first relation mapping table, wherein the first relation mapping table is a mapping relationship table between the data sensitivity level and the historical access frequency and the memory correctable error threshold; Using the data sensitivity level and the historical access frequency as indexes, query the first relation mapping table to determine the corresponding memory correctable error threshold.
4. The memory correctable error handling method of claim 1, wherein, The step of determining the corresponding data sensitivity level based on the data type includes: Obtain a pre-set second relation mapping table, wherein the second relation mapping table is a mapping relationship table between the data type and the data sensitivity level; Using the data type as an index, query the second relational mapping table to determine the corresponding data sensitivity level.
5. The memory correctable error handling method of claim 1, wherein, After processing the correctable errors of the storage unit according to the data sensitivity level, the process includes: In response to a user-triggered memory correctable error threshold change instruction, the memory correctable error threshold is updated according to the memory correctable error threshold change instruction; Correctable errors in the memory are processed based on the updated correctable error threshold for the memory.
6. The memory correctable error handling method of claim 1, wherein, Before processing the correctable errors of the storage unit according to the data sensitivity level, the process includes: Identify whether the data type of the stored data has changed; If a change in the data type is detected, the data sensitivity level is updated according to the changed data type. Correctable errors in the storage unit are handled based on the updated data sensitivity level.
7. An electronic device, comprising: include: Memory, used to store computer programs; A processor, configured to implement the steps of the memory-correctable error handling method as described in any one of claims 1 to 6 when executing the computer program.
Citation Information
Patent Citations
Memory fault-tolerant method and device capable of correcting error storm and medium
CN116954986A
Data retention in memory devices
US20220137842A1