Methods and devices for repairing server memory, storage media and electronic devices

By dynamically adjusting the server memory CE threshold and utilizing the first resource information and error information of the repaired resources, the problem of low resource utilization in existing technologies is solved, and efficient resource management and system stability of the server memory repair method are achieved.

CN120723525BActive Publication Date: 2025-12-02LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511208465.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-12-02
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Among existing server memory repair methods, the CE threshold setting method is static and inflexible, resulting in low resource utilization and inability to adapt to changes in the server operating environment. In particular, when memory resources are scarce, it is impossible to repair critical errors in a timely manner, which may lead to system instability or shutdown.

Method used

By determining the primary resource information of the repair resources in the server, searching for target rules based on the rule lookup table, dynamically adjusting the CE threshold, and using the repair resources to repair memory under preset conditions, including isolating, replacing, resetting or backing up faulty areas, and using repair mechanisms such as spare memory modules, ECC, hot-swapping and system recovery strategies.

Benefits of technology

This method improves resource utilization of server memory repair, ensuring timely repair when resources are sufficient and avoiding unnecessary repair operations when resources are scarce, thus maintaining system stability and performance and adapting to complex operating environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723525B_ABST
    Figure CN120723525B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, storage medium, and electronic device for repairing server memory, relating to the field of artificial intelligence technology. The method includes: determining first resource information corresponding to repair resources in the server, the first resource information indicating the state of resources in the server that can be used to repair memory errors; searching for a target rule in a rule lookup table based on the first resource information and first error information of the server memory, the first error information indicating errors occurring in the server memory within a preset period, the rule lookup table indicating the correspondence between the first resource information, the first error information, and reference rules, the reference rules including the target rule; determining a target threshold based on the target rule and the first error information; and repairing the server memory using repair resources when the first error information and the target threshold meet preset repair conditions. This solves the technical problem of low resource utilization in related server memory repair methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a method and apparatus for repairing server memory, a storage medium, and an electronic device. Background Technology

[0002] In servers, correctable errors (CEs) are an unavoidable part of memory operation. They refer to data errors that occasionally occur during memory read and write operations, but these errors can be automatically detected and corrected by the system to a certain extent. For server stability, when the number of CEs exceeds a CE threshold, the server will implement error handling policies. Obviously, the setting of the CE threshold affects the server's performance and stability.

[0003] Existing methods for setting CE thresholds are mostly based on static, empirical indicators, maintaining a constant CE threshold regardless of whether the server is under light or heavy load. This approach often struggles to adapt to changes in the server's operating environment, such as sudden traffic spikes or large-scale data processing, leading to delayed error handling or wasted resources. Furthermore, when server memory resources are scarce, existing CE threshold adjustments often fail to adequately consider the availability of repair resources. For example, continuing to use a high threshold strategy when memory repair resources are about to run out may result in critical CEs not being corrected in time, ultimately leading to system instability or even system failure. In short, existing server memory repair methods suffer from low resource utilization. Summary of the Invention

[0004] This application provides a method and apparatus for repairing server memory, a storage medium and an electronic device, to at least solve the problem of low resource utilization in related technologies for repairing server memory.

[0005] This application provides a method for repairing server memory, including: determining first resource information corresponding to the repair resource in the server, wherein the first resource information is used to indicate the resource status in the server that can be used to repair memory errors;

[0006] The target rule is obtained by looking up the first resource information and the first error information of the server memory in the rule lookup table. The first error information is used to indicate the error that occurs in the server memory within a preset period. The rule lookup table is used to indicate the correspondence between the first resource information, the first error information and the reference rule. The reference rule includes the target rule.

[0007] Determine the target threshold based on the target rules and the first error message;

[0008] If the first error message and the target threshold meet the preset repair conditions, repair resources are used to repair the server memory.

[0009] This application also provides a server memory repair device, including: an information determination module, used to determine first resource information corresponding to the repair resources in the server, wherein the first resource information is used to indicate the resource status in the server that can be used to repair memory errors;

[0010] The rule lookup module is used to find the target rule in the rule lookup table based on the first resource information and the first error information of the server memory. The first error information is used to indicate the error that occurred in the server memory within a preset period. The rule lookup table is used to indicate the correspondence between the first resource information, the first error information and the reference rule, which includes the target rule.

[0011] The threshold determination module is used to determine the target threshold based on the target rules and the first error information.

[0012] The repair module is used to repair server memory using repair resources when the first error message and target threshold meet the preset repair conditions.

[0013] This application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described server memory repair methods when executing the computer program.

[0014] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described server memory repair methods.

[0015] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described server memory repair methods.

[0016] This application identifies first resource information corresponding to repairable resources in a server. This first resource information indicates the status of resources available for repairing memory errors in the server. Based on the first resource information and first error information in the server memory, a target rule is retrieved from a rule lookup table. The first error information indicates errors occurring in the server memory within a preset period. The rule lookup table indicates the correspondence between the first resource information, the first error information, and reference rules, including the target rule. A target threshold is determined based on the target rule and the first error information. When the first error information and the target threshold meet preset repair conditions, the repairable resources are used to repair the server memory. By real-time monitoring of the first resource information of repairable resources in the server (such as spare memory modules, memory repair mechanisms, system recovery strategies, etc.), the system can dynamically understand the status of available repairable resources. This allows the server to intelligently adjust its strategy based on current resource availability, avoiding unnecessary repair operations when resources are scarce, or performing timely repairs when resources are plentiful, thereby maximizing resource utilization efficiency and overall system performance. Based on the first resource information and the first error information (i.e., details of errors occurring in the server memory within a preset period), the system can retrieve the most suitable reference rule, i.e., the target rule, from the rule lookup table. This process allows error handling strategies to adaptively adjust based on the server's current operating status and resource availability. It ensures that threshold settings are neither too lenient, leading to potential faults going undetected, nor too stringent, causing unnecessary resource consumption and performance impact. Therefore, it addresses the low resource utilization problem of server memory repair methods in related technologies. Attached Figure Description

[0017] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the hardware environment for an optional server memory repair method according to an embodiment of this application;

[0019] Figure 2 This is a flowchart of an optional server memory repair method according to an embodiment of this application;

[0020] Figure 3 This is a schematic diagram of an optional server memory repair method according to an embodiment of this application;

[0021] Figure 4 This is a schematic diagram of an optional reference rule according to an embodiment of this application;

[0022] Figure 5 This is a schematic diagram of another optional reference rule according to an embodiment of this application;

[0023] Figure 6 This is a schematic diagram of another optional reference rule according to an embodiment of this application;

[0024] Figure 7 This is a schematic diagram of another optional method for repairing server memory according to an embodiment of this application;

[0025] Figure 8 This is a schematic diagram of another optional server memory repair method according to an embodiment of this application;

[0026] Figure 9 This is a structural block diagram of an optional server memory repair device according to an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0028] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0029] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] According to one aspect of the embodiments of this application, a method for repairing server memory is provided. As an optional implementation, the above-described method for repairing server memory can be applied, but is not limited to, to applications such as... Figure 1 The hardware environment shown includes a server memory repair system. This server memory repair system may include, but is not limited to, server 102, network 104, and processor 106.

[0031] Optionally, in this embodiment, the aforementioned network may include, but is not limited to, wired networks and wireless networks. The wired network includes local area networks (LANs), metropolitan area networks (MANs), and wide area networks (WANs). The wireless network includes Bluetooth, Wi-Fi, and other networks that enable wireless communication. The aforementioned server may be a single server, a server cluster consisting of multiple servers, or a cloud server. The processor may be an internal processor within the server capable of implementing a method for repairing server memory, or it may be an external system. The above is merely an example, and no limitations are imposed in this embodiment.

[0032] In an optional implementation, server 102 executes step S102, sending first resource information to processor 106 via network 104. Then, processor 106 executes steps S104-S108, searching for a target rule in a rule lookup table based on the first resource information and the first error information of the server memory. The first error information indicates an error that occurs in the server memory within a preset period, and the rule lookup table indicates the correspondence between the first resource information, the first error information, and reference rules, including the target rule. A target threshold is determined based on the target rule and the first error information. If the first error information and the target threshold meet preset repair conditions, step S110 is executed to repair the server memory using repair resources.

[0033] Embodiments of this application provide a method for repairing server memory. Figure 2 This is a flowchart of an optional server memory repair method according to an embodiment of this application; as follows: Figure 2 As shown, the methods for repairing the server's memory include:

[0034] Step S202: Determine the first resource information corresponding to the repair resource in the server, wherein the first resource information is used to indicate the resource status in the server that can be used to repair memory errors;

[0035] It's important to note that in a computing environment, a server is a high-performance computer designed to handle large amounts of data and service requests. It is typically deployed in data centers or network environments to provide various IT services, such as data storage, network services, and computing resources.

[0036] Repair resources refer to the resources or mechanisms within a server used to detect and repair memory errors. This can include spare or replaceable memory modules, memory error self-healing functions, system recovery strategies, fault isolation zones, or any system components and capabilities that can be used to handle and resolve memory error correction (CE). Primary resource information refers to data describing the status of repair resources within the server; it indicates the type, quantity, location, and repair capabilities of currently available repair resources. Through primary resource information, the system can understand the real-time status of repair resources, including their sufficiency, availability for immediate memory error repair, and the potential impact of subsequent repair operations on system performance and stability.

[0037] In an optional implementation, all resources on the server that can be used to detect and repair memory errors can be examined and quantified, including but not limited to spare memory units, the availability of self-healing functions, and the readiness of system recovery strategies. By acquiring this initial resource information, the system can understand the availability of repair resources in real time, which is crucial for dynamically adjusting the memory error correction (CE) threshold based on resource status in subsequent steps.

[0038] In server memory memory error correction (CE) management, the rational use of repair resources is crucial for maintaining system stability and improving service availability. While memory CE events may not cause system crashes in the short term, their accumulation can gradually erode system performance and security. Therefore, real-time monitoring and understanding of the status of repair resources are essential for timely and effective handling of memory CEs. In step S202, by collecting first resource information, the system can accurately determine the currently available repair resources (such as spare memory and the readiness status of the repair mechanism) and their efficiency, providing foundational data for intelligent decision-making in subsequent steps. For example, if the first resource information indicates a shortage of repair resources, the system may choose a conservative memory CE threshold to avoid frequent memory repair operations, thus saving resources and avoiding unnecessary system interruptions. Conversely, if repair resources are sufficient, the system may adopt a more aggressive strategy, lowering the CE threshold to handle memory errors more promptly, improving system stability and response speed. Therefore, step S202 is a key step in building a dynamic adjustment mechanism for the intelligent memory CE threshold. It ensures that the server can make the optimal threshold setting decision based on the actual repair resource status, thereby achieving a good balance between resource management and system performance, and enhancing the server's adaptability and overall operating efficiency in complex dynamic environments.

[0039] It should be noted that repair resources are an important component of server maintenance and management mechanisms, used to detect, prevent, and repair various errors and anomalies in server hardware or software. In the context of server memory management, repair resources are primarily used to handle correctable errors (CE) and uncorrectable errors (UE) in memory, aiming to ensure stable server operation and data integrity.

[0040] In an optional implementation, server memory is repaired using repair resources, including one of the following:

[0041] 1) Isolate the faulty memory region and replace it with a spare storage region;

[0042] 2) Reset the memory fault region, and then start the memory fault region after the memory fault region has been reset;

[0043] 3) Back up the memory error area to the redundant memory area and start the corresponding redundant memory area.

[0044] Repair resources can exist in various forms. For example:

[0045] 1. Backup or Redundant Memory Modules: Servers are typically equipped with backup memory modules to provide an alternative in case of main memory failure. These modules are idle when not in use, and the system automatically switches to the backup module if an uncorrectable error is detected in a main memory module to avoid data loss or system crash.

[0046] 2. Error Correction Coding (ECC): ECC is a technology embedded in memory used to detect and correct errors during data transmission. When a Error Correction (CE) occurs, ECC can automatically correct the error without downtime or manual intervention, thus maintaining data integrity and system performance.

[0047] 3. Hot-swappable memory: Hot-swappable technology allows users to replace faulty memory modules without shutting down the server, reducing downtime required to fix memory errors and improving server availability.

[0048] 4. System recovery strategy: The server can be configured to immediately start a recovery mechanism after a memory error is detected, such as restarting, memory mirroring, or data migration to a healthy module, to reduce the impact of the error on the system.

[0049] 5. Fault Isolation Area (FAR): In some server architectures, memory can be divided into multiple independent areas. When a part of the memory fails, the system can isolate it without affecting the normal operation of other parts until the failed memory is replaced or repaired.

[0050] 6. Memory repair instruction set: The server's hardware or firmware may contain instruction sets used to diagnose memory status and attempt to repair specific types of errors. These repair instruction sets are usually low-level implementations and are transparent to the user.

[0051] While the existence of repair resources greatly improves the fault tolerance and maintenance efficiency of servers, they are also subject to certain limitations in their use, namely:

[0052] 1. Resource capacity limitation: The number and capacity of spare memory modules are limited. Once they are exhausted, the server will lose the ability to automatically repair memory errors. At this time, manual intervention is required to replace the memory modules.

[0053] 2. Limitations of ECC correction capabilities: Although ECC can automatically correct CE, it also has certain error correction boundaries. Once these boundaries are exceeded (e.g., by the UE), ECC will become ineffective and will need to rely on other repair resources.

[0054] 3. Physical limitations of hot-swapping: Hot-swapping requires hardware design support and proper cooling conditions. If the server's physical design does not support hot-swapping, or if the memory area is overheated, the hot-swapping function may not work properly.

[0055] 4. Complexity of system recovery strategies: Automatic recovery strategies may introduce additional system overhead, especially during rapid repair and data migration, which may lead to temporary performance degradation or delays.

[0056] 5. Efficiency of the isolation area: Although FAR can isolate faulty memory, it also means that even if only a small part of the memory fails, the entire isolation area will be abandoned, reducing the effective utilization of memory resources.

[0057] 6. Scope and efficiency of instruction set: The memory repair instruction set is effective for specific types of errors, but it is not a panacea. Its execution efficiency is also limited by hardware. For large-scale memory errors, its repair capability may be insufficient.

[0058] Therefore, understanding the diversity and limitations of repair resources is crucial for designing effective dynamic adjustment strategies for memory error correction (CE) thresholds. Servers must comprehensively consider the availability, performance, and limitations of various resources to formulate the most appropriate error handling and resource allocation strategies, thereby maximizing resource utilization efficiency and system performance while ensuring system stability and data integrity.

[0059] Step S204: Based on the first resource information and the first error information of the server memory, the target rule is found in the rule lookup table. The first error information is used to indicate the error that occurs in the server memory within a preset period. The rule lookup table is used to indicate the correspondence between the first resource information, the first error information and the reference rule. The reference rule includes the target rule.

[0060] It's important to note that a rule lookup table is a data structure or database that contains a series of rules and their correspondence with specific conditions (such as first resource information and first error information). The purpose of the rule lookup table is to find and determine the most suitable reference rule, or target rule, based on the server's current resource repair status (first resource information) and memory error status (first error information). The target rule is the best rule selected from the rule lookup table based on the current status, used to guide subsequent operations, such as determining a dynamic adjustment strategy for the memory error correction (CE) threshold.

[0061] In an optional implementation, the server system will take the first resource information (i.e., the availability of repaired resources) and the first error information (referring to the memory error state within a preset period, including the frequency and severity of CE events) as input, and query a rule lookup table to identify the best target rule. This rule will directly affect the setting of the subsequent memory CE threshold and the selection of error handling strategies.

[0062] In an optional implementation, first resource information (the status of the repaired resource) and first error information (the frequency and nature of errors occurring in server memory within a period) are obtained to form a parameter set for a query rule lookup table. Based on these parameters, the corresponding rule entry is searched in the rule lookup table. The rule lookup table contains optimal reference rules corresponding to different combinations of repaired resource status and error information. The rules are designed to cover different strategies, from conservative to aggressive, to adapt to the needs of different scenarios. The target rule most suitable for the current situation is selected from the rule lookup table. This step requires the system to be able to quickly and accurately match the input parameters and preset rules, ensuring that the selected rule can effectively balance the use of repaired resources and the processing requirements of memory error correction (CE).

[0063] The Baseboard Management Controller (BMC) is a key component in modern servers used for monitoring and managing hardware status. It can operate independently of the main operating system, enabling remote monitoring and control of server hardware. The BMC can be used to obtain initial error information (i.e., server memory error messages such as CE).

[0064] Server hardware and firmware (such as memory controllers, CPUs, BIOS, etc.) need to have error detection and logging capabilities. When a memory error (CE) occurs, these components generate corresponding error log entries and record them in the hardware log buffer. Log entries typically contain information such as the error type, occurrence time, error location (e.g., specific memory address), and error severity level. The BMC (Browser Control Center) communicates with the server hardware via an interface (such as IPMI, Intelligent Platform Management Interface) and periodically checks the contents of the log buffer. Once the BMC detects a new log entry, it reads these entries and extracts the error information related to the memory CE. The BMC may also upload these log entries to a remote management system for remote monitoring and analysis by maintenance personnel. After reading the logs, the BMC parses the information to identify the specific details of the CE error. This includes the frequency of the error, the type of error (e.g., single-bit error, multi-bit error), the timing pattern of the error (e.g., whether it occurs frequently under specific loads), and whether the error has been corrected through mechanisms such as ECC (Error Correcting Code). To further analyze the potential impact of CE errors on server performance and stability, BMC will count the number and type of CE errors occurring within a preset period (such as 1 hour, 24 hours, etc.) and summarize this information into the first error information. The first error information may include the total number of CEs, the average CE frequency per unit time, the time period when CEs occur most frequently, and whether CEs have led to a decline in system performance or instability.

[0065] Step S206: Determine the target threshold based on the target rule and the first error information;

[0066] It should be noted that the target threshold is a detection threshold for CE events determined at a specific moment based on the first resource information of the repaired resource and the first error information of the memory error, using selected target rules. It is the standard value for the system to determine whether to trigger the memory CE repair mechanism.

[0067] In an optional implementation, after obtaining the target rule, the system applies the logic and parameters within that rule to process the first error message. For example, if the target rule is based on a conservative strategy, it might suggest setting a higher CE threshold under conditions of resource scarcity and frequent CE events to avoid the impact of frequent repair operations on system performance. Following the guidance in the target rule, the system calculates parameters related to the first error message to determine the target threshold. This may involve quantitative analysis of the frequency, severity, and trends of CE events, as well as assessment of the availability and consumption rate of repair resources. Taking into account the results of these analyses, the system determines a target threshold that reflects the system's tolerance limit for memory CE events in the current state. For example, if the rule suggests setting the CE threshold to 1500 when the number of CE events is medium to high, the system will set the threshold accordingly to guide subsequent error detection and repair operations. The determined target threshold is not static; it dynamically adjusts with changes in system state (including repair resource status and CE event conditions). This allows the system to always maintain the optimal CE detection and response strategy, regardless of whether resources are abundant or scarce.

[0068] Step S208: If the first error message and the target threshold meet the preset repair conditions, repair resources are used to repair the server memory.

[0069] It should be noted that the preset repair conditions are a series of conditions set by the system to determine when to initiate the server memory repair process. These conditions may be based on the frequency and severity of errors, resource availability, and the current operating status of the system. The system will only trigger the repair operation when both the initial error message (memory error details) and the target threshold are met.

[0070] In an optional implementation, after detecting a memory error correction (CE) event and dynamically adjusting the CE threshold based on the current repair resource status and error information, the system checks whether the memory error has reached the preset repair conditions. If the first error information (i.e., the number and nature of errors occurring within a preset period) exceeds the dynamically set target threshold, and the system has sufficient repair resources available, the system will automatically initiate the repair process, using the preset repair resources to process and correct the errors in memory.

[0071] In an optional implementation, the system needs to monitor the first error information in real time, namely the number and status of Error Detection (CE) events in the server memory within a preset period, and simultaneously confirm the current CE threshold setting. When the actual frequency or cumulative number of CE events exceeds the target threshold dynamically set based on the first resource information and the first error information, the system enters the stage of determining whether the preset repair conditions are met. The system also needs to check whether sufficient repair resources are available. This includes assessing the number of spare memory modules, the remaining correction capability of the ECC mechanism, the availability of hot-swappable memory slots, and the readiness of the system recovery strategy. If resources are sufficient, the system will prepare to perform repair operations. Once it is confirmed that the repair conditions are met, the system will use the corresponding repair resources to process the memory CE events. This may include replacing faulty memory modules, enabling spare memory regions, performing ECC error correction, isolating faulty memory units, or applying other recovery strategies, depending on the nature of the CE events and the type of repair resources. After the repair operation is completed, the system will evaluate the repair effect and feed the results back to the BMC or other management systems for subsequent threshold adjustments and resource management decisions. This helps the system form a closed-loop control and continuously optimize the error handling strategy.

[0072] Example 1:

[0073] Suppose that the server detects frequent memory termination events (CE events) within a certain period, with 5 CE interrupts, while the current dynamically adjusted target CE threshold is 4 (calculated based on the threshold in step S206). Meanwhile, the BMC determines that the proportion of remaining spare memory modules in the repair resources is 50%, the ECC mechanism still has a high correction capability, and the server's current operating load is at a moderate level.

[0074] First, the system checks if the first error message exceeds the target threshold. In this example, the number of CE interrupts reached 5, exceeding the dynamic CE threshold by 4, thus meeting the preset repair conditions. The BMC further checks if there are sufficient resources for repair. Since the proportion of remaining spare memory modules is high and the ECC mechanism still has sufficient corrective capacity, this indicates that the system has the resources to repair memory errors. Based on these conditions, the system will initiate the repair process, potentially using the ECC mechanism to automatically correct CE, or replacing the faulty unit with a spare memory module if necessary. In this example, due to the high number of CE interrupts, the system may choose a more aggressive repair strategy, such as replacing the faulty memory module, to reduce potential future interrupt events. After the repair operation is completed, the BMC records the repair effect and adjusts the CE threshold and repair strategy as needed. For example, if the repair operation successfully reduces CE events, the BMC may appropriately increase the CE threshold to reduce future resource consumption; conversely, if the frequency of CE events remains high, the BMC may further decrease the threshold to ensure the problem is addressed more promptly.

[0075] This application identifies first resource information corresponding to repairable resources in a server. This first resource information indicates the status of resources available for repairing memory errors in the server. Based on the first resource information and first error information in the server memory, a target rule is retrieved from a rule lookup table. The first error information indicates errors occurring in the server memory within a preset period. The rule lookup table indicates the correspondence between the first resource information, the first error information, and reference rules, including the target rule. A target threshold is determined based on the target rule and the first error information. When the first error information and the target threshold meet preset repair conditions, the repairable resources are used to repair the server memory. By real-time monitoring of the first resource information of repairable resources in the server (such as spare memory modules, memory repair mechanisms, system recovery strategies, etc.), the system can dynamically understand the status of available repairable resources. This allows the server to intelligently adjust its strategy based on current resource availability, avoiding unnecessary repair operations when resources are scarce, or performing timely repairs when resources are plentiful, thereby maximizing resource utilization efficiency and overall system performance. Based on the first resource information and the first error information (i.e., details of errors occurring in the server memory within a preset period), the system can retrieve the most suitable reference rule, i.e., the target rule, from the rule lookup table. This process allows error handling strategies to adaptively adjust based on the server's current operating status and resource availability. It ensures that threshold settings are neither too lenient, leading to potential faults going undetected, nor too stringent, causing unnecessary resource consumption and performance impact. Therefore, it addresses the low resource utilization problem of server memory repair methods in related technologies.

[0076] In an optional implementation, the process of finding a target rule in a rule lookup table based on the first resource information and the first error information of the server memory includes: determining second error information based on the first error information, wherein the second error information is used to indicate correctable errors that occur in the server memory within a preset period; determining the number of correctable errors corresponding to each time period within the preset period based on the second error information; determining third error information based on the number of correctable errors corresponding to each time period, wherein the third error information is used to indicate the change in the number of correctable errors within the preset period; and finding a target rule in a rule lookup table based on the first resource information and the third error information, wherein the rule lookup table is used to indicate the correspondence between the first resource information, the third error information, and reference rules.

[0077] It should be noted that the second error message indicates the specific details of server memory CE errors within the preset period, including information such as error frequency, type, and distribution. It is data further refined from the first error message, used to more accurately assess the memory error status. The third error message reflects the trend of CE error quantity changes within the preset period, including upward, downward, or stable trends. It is derived by comparing the number of CE errors across different time periods, used to identify the development direction of memory health and help the system predict future error risks. The preset period is defined by the system for monitoring and evaluating CE errors; it is a fixed time period, such as 1 hour or 24 hours. The setting of the preset period should comprehensively consider system characteristics, application requirements, and resource management efficiency.

[0078] In an optional implementation, the current resource status (first resource information) and CE error details (first error information) are analyzed and further refined into second error information (error details) and third error information (error trend). Then, a rule lookup table is consulted to determine the most suitable strategy (target rule) to adapt to the current memory health status and resource availability.

[0079] In an optional implementation, the system first parses the first error information, extracting details of all CE events within a preset period, including the type, timestamp, and severity of each event. Based on this, it generates the second error information, providing specific data for subsequent error quantity trend analysis. The system subdivides the preset period into multiple time periods, counting the number of CE events in each time period. This helps identify time patterns of error occurrence, such as concentrated outbreaks during peak hours or sporadic occurrences during off-peak hours. Based on the number of CE events in each time period, the system analyzes the overall trend of change, determining whether it is a stable state, an upward trend, or a downward trend, thus forming the third error information. This information is crucial for judging the long-term development trend of the system's memory status. Finally, the system uses the first resource information (such as the proportion of remaining spare memory) and the third error information (CE trend) as query keywords to find the most suitable target rule from a rule lookup table. This rule will guide subsequent CE threshold setting and resource allocation decisions.

[0080] The above-described implementation method of this application involves starting from fuzzy overall information about CE events (first error information), gradually refining it to a detailed list of events (second error information), then to trends revealed by time series analysis (third error information), and finally matching it with strategies (target rules) in a rule lookup table. This method not only improves the server's perception and response speed to memory health status, but also optimizes the allocation and use of repair resources through an intelligent decision-making mechanism, ensuring that the server can maintain stable performance and reliable operation in complex and ever-changing operating environments.

[0081] In an optional implementation, the target rule is obtained by searching a rule lookup table based on the first resource information and the third error information, including: determining multiple reference rules in the rule lookup table; searching multiple reference values ​​corresponding to the first resource information and the third error information in the rule lookup table, wherein the reference value corresponds one-to-one with the reference rule; and determining the reference rule corresponding to the largest reference value as the target rule.

[0082] It should be noted that reference rules are a series of preset rules in the rule lookup table. Each rule is designed based on different resource states and CE trends, and is used to guide the setting of CE thresholds and the execution of remediation strategies. The reference value can be the Q-value corresponding to each reference rule, reflecting the expected return or value of selecting that rule as the target rule in a specific state. The larger the Q-value, the higher the probability that the rule will be selected by the agent (such as a BMC or server management system) in the current state, because it can bring better system performance or resource utilization efficiency.

[0083] In an optional implementation, during the dynamic adjustment of the CE threshold, the system (such as BMC) will search and determine the target rule most suitable for the current situation from a rule lookup table based on the current repair resource status (first resource information) and the trend of memory CE (third error information). This process involves evaluating and selecting multiple reference rules, using the highest reference value (Q value) as guidance, to ensure the intelligence and adaptability of the system when handling memory CE.

[0084] In an optional implementation, the system identifies all potentially applicable reference rules from a rule lookup table. These rules cover different resource management and CE handling strategies, such as conservative, balanced, and aggressive approaches. For each reference rule, the system looks up the Q-value, or reference value, corresponding to the first resource information and the third error information in the rule lookup table. The Q-value is obtained through multiple rounds of training and optimization using a reinforcement learning algorithm (such as Q-learning), reflecting the expected performance and resource utilization efficiency of selecting this rule as the target rule under a specific resource state and CE trend. After collecting the Q-values ​​of all reference rules, the system selects the rule with the highest reference value as the target rule. This means that the selected rule will bring the best CE management effect to the system under the current state, whether it is increasing the threshold to reduce resource consumption or decreasing the threshold to enhance the sensitivity of CE detection.

[0085] Example 2:

[0086] Assuming that within a preset period, the server detects that the remaining percentage of repaired resources is 50%, and the number of memory error codes (CEs) is trending upwards (third error message). Based on this information (first resource information and third error message), the system performs the following operations:

[0087] The system identifies three strategy rules from the rule lookup table: conservative rule (high resource consumption, low CE threshold), balanced rule (medium resource consumption, moderate CE threshold), and aggressive rule (low resource consumption, high CE threshold). The system queries the Q table for the Q value corresponding to the current remaining repair resource percentage (50%) and the upward trend in CE. Assume the obtained Q values ​​are: 120 for the conservative rule, 150 for the balanced rule, and 130 for the aggressive rule. After comparing these three Q values, the system determines the balanced rule (Q=150) with the highest Q value as the target rule. This indicates that under the current resource status and CE trend, the balanced repair strategy delivers the best overall performance, effectively addressing the rising CE without excessively consuming repair resources.

[0088] As can be seen from the above-described embodiments of this application, the process of "finding the target rule in the rule lookup table based on the first resource information and the third error information" is actually a dynamic decision-making process in which the system intelligently selects the optimal memory CE management strategy using reinforcement learning algorithms and fuzzy control technology. The core of this process lies in enabling the system to make timely CE threshold adjustments and resource allocation decisions in a complex and ever-changing operating environment through continuous learning and optimization, thereby achieving the goals of stable operation and efficient resource management.

[0089] In an optional implementation, searching for multiple reference values ​​corresponding to the first resource information and the third error information in a rule lookup table includes: determining the first sub-state to which the first resource information belongs, and determining the second sub-state to which the third error information belongs, wherein the first sub-state is used to indicate the degree of the first resource information, and the second sub-state is used to indicate the trend of correctable errors; calculating a table index based on the first sub-state and the second sub-state, wherein the table index is used to indicate a specific row or column in the rule lookup table; and finding the corresponding multiple reference values ​​in the rule lookup table based on the table index.

[0090] It should be noted that the first sub-state represents the status classification of the first resource information. For example, the remaining proportion of repaired resources is divided into multiple levels (such as low, medium, and high), and each level represents a sub-state of the repaired resources. The second sub-state reflects the status classification of the third error information, that is, the status classification of the CE trend, such as rising, stable, and falling, and each state represents a sub-state of the CE trend. The table index can be an index value used to locate a specific row or column in the rule lookup table. It is calculated based on the first and second sub-states and is used to identify the set of rules in the rule lookup table that match the current resource status and CE trend.

[0091] In an optional implementation, the system analyzes the first resource information and categorizes it into a predefined first sub-state (e.g., high / low resource remaining ratio classification). Next, based on the third error information (CE trend), it determines its corresponding second sub-state (e.g., CE quantity change trend classification). Then, the system calculates a table index based on these two sub-states to quickly locate relevant rules in a rule lookup table. Finally, the system uses the table index to look up and obtain multiple reference values ​​(e.g., Q-values) to prepare for the next step of selecting the optimal rule.

[0092] In an optional implementation, during the dynamic adjustment of the server memory CE threshold, precise state classification and rule index calculation allow for efficient searching and matching of optimal rules in the rule lookup table. The system categorizes the remaining percentage of repaired resources into a preset first sub-state; for example, a remaining percentage less than 20% is classified as a "low" sub-state, 20%-80% as a "medium" sub-state, and greater than 80% as a "high" sub-state. The system analyzes the third error information, namely the trend of CE quantity changes within a preset period, and categorizes it into a second sub-state. A continuous increase in CE quantity is marked as an "upward" trend, stability as a "stable" trend, and a decrease as a "downward" trend. Based on the combination of the first and second sub-states, the system calculates a unique table index corresponding to a specific position in the rule lookup table, used to find rules matching the current state and their reference value. In the rule lookup table, the table index is used to locate the set of rules matching the first and second sub-states, and the Q-values ​​(reference values) of these rules are obtained, providing a basis for subsequent selection of the optimal rule.

[0093] Example 3:

[0094] Assume the server currently has 65% of its repairable resources remaining, and CE (Cost Efficiency) has been trending upwards in recent periods. The 65% remaining repairable resources fall within the "Medium" sub-state range of the rule lookup table; the upward trend in CE is marked as an "Upward" sub-state. Based on the first sub-state ("Medium" resource state) and the second sub-state ("Upward" CE trend), the system calculates a table index that is likely 5 (assuming the rule lookup table has columns for "Low," "Medium," and "High" resource states and rows for "Declining," "Stable," and "Upward" CE trends, and the table index is calculated by cross-referencing rows and columns). In the rule lookup table, the system uses table index 5 to find rules that match the "Medium" resource state and the "Upward" CE trend. Assuming the rule lookup table stores Q-values ​​calculated based on Q-learning, for index 5, the Q-values ​​found by the system are: Conservative rule Q-value = 130, Balanced rule Q-value = 147, and Aggressive rule Q-value = 120.

[0095] Through the above-described implementation methods of this application, the system can not only quickly locate the rule set that matches the current resource status and CE trend, but also select the optimal rule for threshold adjustment by comparing Q values ​​(reference value), ensuring that the server memory management strategy can both cope with resource constraints and adapt to changes in CE trend, thereby achieving efficient system operation and effective resource utilization.

[0096] In an optional implementation, before finding the target rule in the rule lookup table based on the first resource information and the third error information, the method includes: cross-combining at least one first sub-state in the first sub-state set and at least one second sub-state in the second sub-state set to obtain multiple first states in the first state set; constructing a first rule lookup table with the multiple first states and multiple reference rules as rows and columns respectively, and setting the reference value in the rule lookup table to a preset value, wherein the reference value corresponds to the first state and the reference rule;

[0097] The rule lookup table is updated M times. The i-th update of the rule lookup table yields the (i+1)-th rule lookup table, which includes: setting the repaired resources to conform to the second resource information and setting the number of correctable errors to the fourth error information, where M is an integer greater than 0 and i is an integer greater than 0 and less than or equal to M; updating the i-th rule lookup table according to the second resource information and the fourth error information to obtain the (i+1)-th rule lookup table; and determining the (M+1)-th rule lookup table as the rule lookup table if i is greater than or equal to M.

[0098] It should be noted that the first sub-state set contains information related to repair resources, such as the proportion of remaining resources and the rate of resource consumption. These states describe the availability and changes of repair resources from a macroscopic perspective. The second sub-state set contains information related to CE event trends, such as the rate of change of the number of CEs and their upward or downward trends. These states describe the dynamic development characteristics of memory CE events from a microscopic perspective. The first state set is composed of the intersection of the first and second sub-state sets, representing the state space of the system under different combinations of repair resources and CE trends. This forms the basis for constructing the rule lookup table.

[0099] The first rule lookup table can be an initial rule lookup table, containing all possible pairings of first states with reference rules, and their corresponding initial reference values ​​(Q-values). This rule lookup table is the basic framework formed based on expert experience and initial settings. M updates represent the number of iterative updates used to progressively optimize the Q-values ​​in the rule lookup table, making them more consistent with the actual system state and behavior patterns. M is a positive integer representing the scale of the training and optimization loops.

[0100] In an optional implementation, the first sub-state set and the second sub-state set are cross-combined to form a comprehensive first state set. Then, an initial rule lookup table is constructed and an initial Q value is set. Through multiple iterations of the learning process, the Q value in the rule lookup table is continuously updated according to the actual system state and CE events until the preset number of training times M is reached, and finally the optimized rule lookup table is determined.

[0101] In an optional implementation, the system cross-combines a first sub-state set (repair resource state) with a second sub-state set (CE trend) to generate a first state set, covering all possible combinations of resource state and CE trend. Based on the first state set and preset reference rules, the system constructs an initial rule lookup table. Each state-rule pair corresponds to an initial Q-value in the table, which is set to a preset value (e.g., 0) as the starting point of the learning process. The system evaluates the application effect of each reference rule through simulation or actual operation, based on second resource information (e.g., current repair resource state) and fourth error information (actual CE count). In each iteration (denoted as the i-th update), the system executes the reference rule, observes system state changes and repair results, obtains immediate rewards (e.g., improvements in system stability, resource utilization, etc.), and updates the Q-value in the rule lookup table accordingly. The update follows the Q-learning algorithm formula, taking into account immediate rewards and possible future rewards, and is calculated using parameters such as the learning rate α and discount factor γ. When the number of iterations reaches a predetermined M, the training process stops, and the rule lookup table after the last update is considered the optimized version. This optimized rule lookup table contains more accurate Q values, reflecting the execution effect and expected value of different rules under various resource states and CE trends, providing more solid data support for subsequent search of target rules.

[0102] Example 4:

[0103] Assume the server memory repair resource states are categorized as low (L), medium (M), and high (H), and the CE trend is categorized as decreasing (D), stable (S), and increasing (U). The cross-combination of the first sub-state set and the second sub-state set generates 9 possible first states: LD, LS, LU, MD, MS, MU, HD, HS, and HU.

[0104] The system creates a rule lookup table that contains combinations of these 9 states and 3 preset reference rules (conservative, balanced, and aggressive). The Q-value of each state-rule pair is initially set to 0.

[0105] The system underwent M simulation training iterations. In each iteration, a state from the first state set was randomly selected, and corresponding second resource information (such as the remaining resource ratio) and fourth error information (such as the number of memory errors) were set. For example, in the 5th iteration, state MU was selected, where the number of memory errors increased and the repair resources were at a moderate level. The system executed aggressive rules, observed the repair effect, and obtained immediate rewards (such as the system not crashing and resource consumption being moderate). Based on the observation results, the system updated the rules by comparing the Q-values ​​of state MU and the aggressive rules in the rule comparison table. Assuming the Q-value before the update was 5, the immediate reward was 10, and the estimated maximum possible reward in the future was 15 (based on a discount factor γ=0.5), the learning rate was α=0.2.

[0106] Once all M updates are completed, the system stops training and considers the rule lookup table updated in the last update as the final version. This rule lookup table contains the optimal Q-value for each state-rule pair, providing crucial information for finding the target rule in the next step.

[0107] Here, we will explain in more detail the process of defining the state space and action space. The remaining repair resource level (first resource information) can be divided into three levels:

[0108] Low: Resource levels are between 0% and 33%, indicating that repair resources are nearly exhausted and the system's ability to handle memory errors is weak.

[0109] Medium: Resource levels are between 34% and 66%, resource repair is in a moderate state, and the system has a certain error handling capability.

[0110] High: Resource levels are between 67% and 100%, with sufficient resources for repair, and the system can efficiently handle memory CE.

[0111] CE trend (third error message) is divided into:

[0112] Decrease: This indicates that the number of CEs is decreasing, and the risk of system memory errors is reduced.

[0113] Stable: The number of CEs has not changed significantly, and system memory errors are under control.

[0114] Rising: The number of CEs is increasing, exacerbating potential memory health issues.

[0115] Combining the two dimensions mentioned above, the state space has a total of 9 possible states, each of which represents the server's operation under different levels of repair resources and CE trends.

[0116] The action space defines the policy options an agent can take in reinforcement learning, used to adjust the memory completion threshold. The action space can include three different sets of fuzzy rules:

[0117] Conservative: Corresponds to "1" in the action space. It is suitable for repairing situations where resource levels are low or CE trends are rising. The strategy tends to set a higher CE threshold to reduce resource consumption, even if it means potentially delaying error handling.

[0118] Balanced: Corresponding to "2" in the action space, it is the default or general strategy, which aims to balance system stability and resource utilization efficiency, with a moderate CE threshold setting.

[0119] Aggressive: Corresponding to "3" in the action space, this strategy is suitable for situations where resources are sufficient and CE (Error Detection) trends are declining. It tends to set a lower CE threshold to improve the sensitivity of error detection and handle memory errors promptly. The specific settings for the fuzzy rule set will be explained in detail later and will not be repeated here.

[0120] Through the above-described implementation methods of this application, an adaptive optimization of the CE threshold adjustment strategy is achieved by constructing a state set, initializing a rule lookup table, iteratively updating the Q-value, and finally determining the optimized rule lookup table. This process fully utilizes the trial-and-error learning capability and adaptive adjustment characteristics of reinforcement learning, improving the system's intelligent management and response capabilities to memory CE problems. This allows the system to make optimal repair decisions under different operating environments, balancing system performance and reliability.

[0121] In an optional implementation, the i-th rule lookup table is updated based on the second resource information and the fourth error information to obtain the (i+1)-th rule lookup table. This includes: when the current time reaches a preset period, searching for multiple reference values ​​corresponding to the second resource information and the fourth error information in the rule lookup table, and determining the reference rule corresponding to the highest reference value as the target rule; determining a target threshold based on the target rule and the fourth error information, and repairing the server memory using repair resources when the fourth error information and the target threshold meet preset repair conditions; obtaining information on correctable errors and repair resource information in the server memory, and updating the second resource information, the fourth error information, and the reference values ​​corresponding to the second resource information and the fourth error information; and obtaining the (i+1)-th rule lookup table when the second resource information reaches a preset resource amount or the number of correctable errors is greater than the preset number of errors.

[0122] It should be noted that the second resource information refers to the real-time status of server memory repair resources, such as the percentage of remaining spare memory and the total capacity of repair resources, reflecting the current resource availability. The fourth error information records detailed information about current CE events in server memory, including the type, number, frequency of occurrence, and impact on system performance of CEs, used to assess the current CE trend and severity. Preset repair conditions can be a set of rules or conditions used to determine whether to immediately use repair resources for CE repair, such as the number of CEs exceeding a certain threshold or the impact on system performance reaching a certain level.

[0123] In an optional implementation, based on the current resource status (second resource information) and the real-time status of CE events (fourth error information), and according to the results of the previous learning round (the i-th rule lookup table), real-time feedback and decisions are made to update the rule lookup table and obtain a more optimized i+1-th rule lookup table. Specifically, this includes the steps of finding the rule corresponding to the maximum reference value, determining the target threshold, performing repair operations, providing feedback updates, and continuing iteration after meeting specific conditions.

[0124] In an optional implementation, updating the rule lookup table is a crucial step in the agent's strategy of dynamically adjusting the memory CE threshold based on Q-learning. This step, through a real-time feedback mechanism, continuously optimizes the agent's decision-making strategy, ensuring the system achieves an optimal balance between resource management and CE handling. The agent, based on the current second resource information (e.g., 65% remaining repair resources) and fourth error information (e.g., the increasing trend of CE count), searches the rule lookup table for the corresponding Q-value (reference value) and determines the reference rule with the maximum Q-value, i.e., the target rule. Based on the target rule and the fourth error information, the agent calculates the most suitable CE threshold (target threshold) for the current state. When preset repair conditions are met (e.g., the number of CEs exceeds the threshold, system performance degrades), the agent performs memory repair operations using the repair resources. After the repair operation, the agent needs to re-acquire the server memory's CE information (updated fourth error information) and the repair resource status (updated second resource information). Simultaneously, based on the repair results and the current state, the agent updates the Q-values ​​of the relevant state-rule pairs in the rule lookup table. When certain conditions are met (such as the second resource information reaching a preset resource amount or the number of CEs exceeding a preset error number), the agent uses the feedback results of the current state and actions (such as successful repair or improved system performance) to update the i-th rule lookup table, generating a more optimized i+1-th rule lookup table to adapt to future state changes.

[0125] Example 5:

[0126] Suppose that the server detects the second resource information (the remaining repaired resource ratio is 65%) and the fourth error information (the number of CEs is increasing, and the current number of CEs is 800) within a certain preset period. The agent looks up the relevant Q value in the rule lookup table according to the current state. Suppose that the Q values ​​of the conservative, balanced and aggressive strategies are 130, 155 and 120 respectively. Therefore, the balanced strategy has the largest Q value and is determined as the target rule.

[0127] The agent calculates the target CE threshold to be 700 based on the target rules and the fourth error information. At this point, the number of CEs (800) exceeds the target threshold of 700, and the remaining repair resources are at a moderate level, meeting the preset repair conditions (number of CEs exceeds the threshold). The agent then uses the repair resources to process the CE events, successfully repairing some of them. After the repair, the number of CEs drops to 600, and the remaining repair resources decrease to 60%.

[0128] Subsequently, the agent obtains the updated second resource information (repaired resource remaining ratio 60%) and fourth error information (CE count decreased to 600), and re-evaluates the Q-value based on the repair results, updating the rule lookup table. Assuming that under the new state of 60% remaining resource and CE count decreased to 600, the Q-values ​​of the conservative, balanced, and aggressive strategies are re-evaluated to 140, 160, and 130, respectively. The balanced strategy still has the highest Q-value, but due to the decrease in CE count, the agent may choose the conservative strategy in the next round to reduce resource consumption and achieve efficient resource utilization.

[0129] Through the above-described implementation methods of this application, a continuous feedback and decision-making process continuously optimizes the rule lookup table to adapt to the dynamic needs of server memory CE threshold adjustment. This Q-learning-based dynamic rule selection mechanism not only improves the intelligence and adaptability of memory CE threshold adjustment but also effectively balances system stability and resource utilization efficiency, representing a key innovation in modern server memory management.

[0130] In an optional implementation, updating the reference value corresponding to the second resource information and the fourth error information includes: determining the current reward value based on the number of correctable errors in the server memory and the repair resource information within a preset time period; estimating the maximum value of the reference value in the next preset period as the predicted reference value; and determining the updated reference value as the weighted sum of the first parameter and the reference value corresponding to the second resource information and the fourth error information, and the second parameter and the current reward value and the predicted reference value.

[0131] It should be noted that the current reward value is an immediate evaluation value calculated based on the current second resource information and fourth error information, reflecting the immediate impact of the agent's actions (such as adjusting the CE threshold) on system performance and resource management in the current state. The predicted reference value is the optimal reference value that may appear after executing a certain rule (a conservative, balanced, or aggressive fuzzy rule set) in the next preset period, based on the current state and historical data; that is, the expected Q value, used to evaluate long-term gains. The first and second parameters represent the learning rate α and discount factor γ in the reinforcement learning update formula, respectively, and are used to adjust the weights of the current reward value and the predicted reference value when updating the reference value.

[0132] In an optional implementation, the system first assesses the immediate impact of executing a specific rule (such as lowering the CE threshold) on the system under the current state, based on the current number of CEs, the number of CE processing interruptions, and the remaining amount of repair resources. If the number of CEs decreases and resource consumption is reasonable, the reward value may be higher, and vice versa. The reward value can be calculated based on an expert-designed function or determined through experiments and simulations. Based on the current system state and historical data, the system attempts to predict the potential Q value in the next preset period after executing a certain rule, which involves predicting future resource states and CE trends. The predicted Q value can be obtained using time series analysis, machine learning regression models, or simple statistical inference methods. The system updates the reference value (Q value) in the rule lookup table according to the reinforcement learning update formula, combining the current reward value and the predicted Q value (predicted reference value). The Q value update is adjusted by the learning rate α and the discount factor γ. The learning rate determines the proportion of immediate reward in updating the reference value, while the discount factor reflects the importance attached to future rewards.

[0133] Through the above-described embodiments of this application, it can be seen how the agent dynamically adjusts the Q-value in the rule lookup table to reflect the immediate feedback after actual operation (current reward value) and the prediction of the likelihood of future policy execution (prediction reference value). This update mechanism ensures that the agent can continuously learn and adapt to environmental changes, ultimately formulating the most effective CE threshold adjustment strategy and optimizing server memory management.

[0134] In an optional implementation, the current reward value is determined based on the number of correctable errors in the server memory and the repair resource information within a preset time period, including: determining the number of correctable errors in the server memory within the preset time period as a first reward value; determining the change in the server's repair resources within the preset time period as a second reward value; determining the preset reward value as a third reward value if the server stops running within the preset time period; and determining the sum of the first reward value, the second reward value, and the third reward value as the current reward value if a third reward value exists.

[0135] It should be noted that the first reward value corresponds to the feedback value of the number of correctable errors (CEs) in the server's memory within a preset time. The fewer the CEs, the higher the first reward value, indicating that the system's memory health is good and error handling is timely and effective. The second reward value corresponds to the feedback value of the amount of resource changes repaired within a preset time. The less resource consumption for repair, the higher the second reward value, indicating that the system's resource management is efficient and resource usage is economical. The third reward value is the penalty reward value directly applied if the server stops running within a preset time. The third reward value is usually a negative number, used to penalize system downtime events caused by improper handling of CEs, reinforcing the agent's emphasis on system stability.

[0136] In an optional implementation, at the end of each preset period, the change in the number of CEs in the server memory is statistically analyzed and converted into a first reward value. For example, if the number of CEs decreases by 10%, the first reward value can be set to a positive number, such as +5; if the number of CEs increases by 10%, the first reward value can be set to a negative number, such as -5, thereby incentivizing the agent to reduce CEs. Repair resource information is monitored, and the changes in the consumption or remaining amount of repair resources within the preset period are calculated and converted into a second reward value. For example, if the consumption of repair resources decreases by 5% within the period, the second reward value can be set to a positive number, such as +4; if the consumption increases by 5%, the second reward value can be set to a negative number, such as -4, to reward resource conservation and punish resource waste. A third reward value is a reward setting for extreme cases. When the server stops running within the preset period, the agent will be severely penalized. The third reward value is usually set to a large negative number, such as -100, to ensure that the agent avoids actions that cause system downtime in subsequent decisions. The first reward value, the second reward value, and the third reward value (if any) are summed to obtain the current reward value. This is an important part of the formula for updating the Q-value in the Q-learning algorithm, used to evaluate the overall effect of the actions taken by the agent in the current cycle.

[0137] Example 6:

[0138] Assuming the server runs for a certain period (e.g., 30 minutes), at the start of the preset period, the agent selects a balanced rule set to adjust the CE threshold. At the end of the period:

[0139] First reward value: The number of CEs in the server memory is reduced from 1000 to 800, a reduction of 20%, so the agent's first reward value is +10 (assuming that the first reward value is +5 for every 10% reduction in the number of CEs, here a 20% reduction gives double the reward).

[0140] Second reward value: The repair resources decreased from 70% to 65%, which consumed 5%, so the agent's second reward value is -3 (assuming that for every 5% consumed, the second reward value is reduced by 3).

[0141] Third reward value: During this period, the server did not stop running, so there is no third reward value; therefore, the third reward value is 0.

[0142] The current reward value is calculated to be +7. Based on the current reward value (+7), the agent updates the Q value in the Q table using the Q-learning algorithm, providing a data basis for decision-making in the next cycle.

[0143] Through the above-described embodiments of this application, the intelligent agent can continuously optimize its strategy based on the actual operation of the system and guided by the reward mechanism, so as to achieve the optimal adjustment of the server memory CE threshold, ensure the balance between resource management and fault handling, and enhance system stability and performance.

[0144] Example 7:

[0145] Figure 3 This is a schematic diagram of an optional server memory repair method according to an embodiment of this application; as shown. Figure 3 The training process is shown below:

[0146] S1 defines the state space and action space;

[0147] The state space consists of two dimensions: remaining repair resource level and CE trend. The remaining repair resource level is divided into three levels: low (0%-33%), medium (34%-66%), and high (67%-100%), while the CE trend is divided into three states: decreasing, stable, and increasing. Combining these two dimensions, nine possible state combinations can be obtained: low resources + decreasing trend, low resources + stable trend, low resources + increasing trend, medium resources + decreasing trend, medium resources + stable trend, medium resources + increasing trend, high resources + decreasing trend, high resources + stable trend, and high resources + increasing trend.

[0148] The action space contains three fuzzy rule sets, corresponding to conservative (action 1), balanced (action 2), and aggressive (action 3) CE threshold adjustment strategies, respectively.

[0149] S2, Initialize the Q-table and learning parameters;

[0150] The Q-table is a two-dimensional matrix where rows represent states in the state space and columns represent actions in the action space. Before training begins, each element of the Q-table is initialized to 0, indicating that the agent's initial evaluation of the state-action pair is unknown.

[0151] Set learning parameters:

[0152] Learning rate (alpha): Set to 0.2 to control how much importance the agent places on new experiences when updating its Q-value.

[0153] Discount factor (gamma): Set to 0.9, it measures how much the agent values ​​future rewards.

[0154] Exploration rate (epsilon): Set to 0.3 to balance the agent's exploration and exploitation behaviors during the learning process.

[0155] Training rounds (total_episodes): Set to 1000 rounds, which means the agent will go through 1000 complete learning cycles.

[0156] Maximum repair resources (max_resources): Set to 100, which represents the total capacity of repair resources in the system.

[0157] Percentage of resources consumed per repair (resource_consumption_rate): Set to 1%, meaning that each repair operation will consume 1% of the repair resources.

[0158] S3, training the model;

[0159] Before training begins, the initial remaining repair resource level and the number of memory errors (CEs) are randomly set to simulate different operating environments. At each time step, the agent queries the Q-table based on the current system state and selects an action according to an ε-greedy strategy. Specifically, it randomly selects an action (exploration) with a 30% probability and selects the action with the largest Q-value in the current state (exploitation) with a 70% probability. The agent executes the selected action, i.e., selects the corresponding fuzzy rule set, and calculates the target CE threshold through fuzzy inference based on the rule set and the current second resource information (remaining repair resource level) and fourth error information (number of CEs). Changes in the system environment include the occurrence of new CE events, repair events (resource consumption), and updates to CE trends. The agent needs to monitor and record these changes to evaluate the effect of the action execution. The reward value reflects the impact of the current action on the system, including stability rewards (proper CE handling leads to stable system operation), resource penalties (excessive resource consumption), and system crash penalties (excessive CEs lead to system instability). These reward values ​​are calculated based on changes in system performance indicators and resource management status. Based on environmental changes, the agent calculates the system state for the next time step. Based on the current state, the action performed, the reward obtained, and the new state, update the relevant Q values ​​in the Q table using the Q-learning update formula, as follows:

[0160] Q(s,a)←Q(s,a)+α[r+γ·max a 'Q(s',a')−Q(s,a)];

[0161] The agent updates the system state to the new state \(s'\) and continues training for the next time step. If the termination condition is met (such as the remaining repair resources falling below a preset minimum threshold or the number of CEs exceeding the maximum capacity of the system), the current round of training ends, the agent saves the learning results in the current state (the updated Q-table), and prepares for the next round of training.

[0162] In an optional implementation, after finding the target rule in the rule lookup table based on the first resource information and the third error information, the method includes: updating the reference value corresponding to the target rule according to the number of correctable errors in the server memory within a preset time and the repair resource information.

[0163] In an optional implementation, after the agent determines the target rule, it updates the reference value of the target rule based on the actual handling of CE events and the consumption of repair resources within a preset period. Within this preset period, the agent continuously monitors the number of CE events in server memory and the consumption of repair resources, recording whether CE events are effectively handled and the impact of handling these events on resources. Based on the effectiveness of CE event handling (e.g., reduced number of CE events, no system crash) and resource consumption (e.g., moderate resource consumption, not exhausted), the agent calculates the reward value after the target rule is executed. The reward value consists of stability rewards, resource penalties, and system crash penalties. The agent combines the reward value with the current Q-value of the target rule and uses the Q-learning update formula to update the Q-value of the target rule. The update formula considers the current reward, the maximum possible future reward (considered through a discount factor γ), and the learning rate α (which determines the degree of influence of new experience on the old Q-value). The updated Q-value is stored in a rule lookup table, corresponding to a combination of first resource information and third error information, providing a data basis for subsequently selecting the optimal rule under similar conditions.

[0164] Example 8:

[0165] Suppose that within a certain preset cycle, the agent, based on the first resource information (the remaining repaired resources are 60%) and the third error information (the number of CEs is 500 and is on the rise), searches for and selects a balanced set of rules as the target rule in the rule lookup table.

[0166] After applying the balanced rule set, the agent detects that the number of CEs has decreased to 400, and 5% of the repair resources have been consumed, leaving 55% remaining. Based on the effective reduction in the number of CEs, the agent calculates a stability reward of +20 (assuming an increase of 5 points for every 100 fewer CEs); simultaneously, due to the 5% resource consumption, the agent calculates a resource penalty of -10 (assuming a decrease of 2 points for every 1% resource consumption). Assuming the system does not crash within the preset period, the system crash penalty is 0. Assuming the current target rule (balanced) has a Q value of 10, a learning rate α = 0.2, and a discount factor γ = 0.9, the agent needs to update the Q value using the Q-learning formula. Assuming that in the next state (repair resources remaining at 55%, CEs at 400 and trending towards stability), among all possible actions, the balanced rule set has the highest Q value of 15, then the updated Q value is 12.7.

[0167] The updated Q value of 12.7 is stored in the rule lookup table, corresponding to the combination of the first resource information (remaining repaired resource ratio is 60%) and the third error information (the number of CEs is 500 and is on the rise), which provides a more accurate reference value for subsequent decisions under similar conditions.

[0168] Through the above-described embodiments of this application, based on real-time system state changes and CE event handling results, the Q-value information in the rule lookup table is dynamically updated, thereby gradually learning how to select the optimal CE threshold adjustment strategy under different resource and error trends, so as to achieve the optimal balance between system performance and resource management.

[0169] In an optional implementation, if the first error information and the target threshold meet the preset repair conditions, the server memory is repaired using repair resources, including: determining the number of correctable errors based on the first error information; and repairing the server memory using repair resources if the number of correctable errors is greater than or equal to the target threshold.

[0170] It should be noted that the preset repair conditions can be a set of conditions defined by an algorithm or system administrator to determine when to use repair resources to repair CEs. Conditions may include the number of CEs exceeding a threshold, a decline in system performance metrics, etc., ensuring that the repair operation can respond to memory errors promptly while avoiding resource waste. The number of correctable errors refers to the total number of CEs that the system has currently detected that need to be corrected, as shown in the first error message; it is one of the important criteria for determining whether to initiate the repair program.

[0171] In an optional implementation, the system continuously monitors the CE (Completely Executable) status in the server memory, including the number, type, and other relevant characteristics of CEs; this information constitutes the first error message. Once an abnormal increase in the number of CEs is detected, the system immediately enters emergency handling mode to prepare for assessing whether to initiate repair operations. The agent determines whether the total number of CEs in the first error message has reached or exceeded the calculated target CE threshold based on the current system state. If the number of CEs does exceed the threshold, the system proceeds to the next step of assessing the necessity of repair. Besides the excessive number of CEs, the system also needs to determine whether the current situation meets preset repair conditions, such as resource sufficiency and system operating status. If conditions permit, the agent will decide to initiate repair operations. Once the repair conditions are confirmed to be met, the agent will allocate the necessary resources from the repair resource pool to perform CE repair or replacement. After repair is completed, the system will update the level of remaining repair resources to prepare for possible subsequent repair needs.

[0172] Example 9:

[0173] Suppose the server detects 1200 memory collision events (CEs) within a certain period, while the target CE threshold calculated according to the fuzzy rule set selected by the current agent is 1100. Based on preset repair conditions, when the number of CEs exceeds the target threshold by more than 5%, a repair procedure should be initiated immediately to reduce potential system risks.

[0174] The agent detected 1200 CE (Critical Errors) events in the first error message, exceeding the target threshold of 1100, representing an excess rate of 9.09%, thus meeting the repair trigger condition. Simultaneously, it checked the server's current repair resource status. If the remaining resource pool percentage exceeded 50%, resources were considered sufficient for repair. The agent decided to utilize repair resources for CE repair. Assuming that repairing one CE event consumes approximately 0.1% of the total resource pool, repairing 1200 CE events would consume approximately 120% of the resources. However, considering the resource pool remaining above 50%, the agent prioritized handling the most severe CE events until resources were exhausted or the CE count dropped below the target threshold. After the repair operation, assuming 25% of the repair resources were consumed, most high-priority CE events were successfully handled, reducing the CE count to around 800, and the system status improved.

[0175] Through the above-described embodiments of this application, when performing CE (Cycles Exceeded) repair, it is necessary not only to judge based on the current absolute number of CEs and the dynamically adjusted target threshold, but also to comprehensively consider the remaining repair resources and the urgency of the repair, so as to utilize repair resources in an optimal way and reduce the negative impact of CEs on system performance. This repair strategy based on real-time monitoring and intelligent decision-making can significantly improve the stability and security of server memory, extend system life, and reduce maintenance costs.

[0176] In an optional implementation, the target rule is obtained by looking up the first resource information and the first error information of the server memory in a rule lookup table, including one of the following:

[0177] 1) When the first resource information indicates that the server memory is under resource pressure and the first error information indicates that the number of correctable errors in the server memory is low, the target rule is determined as the first reference rule.

[0178] 2) If the first resource information indicates that the server memory resources are sufficient, the target rule is determined as the second reference rule;

[0179] 3) If the first resource information indicates that the server memory is under resource pressure and the first error information indicates that the number of correctable errors in the server memory is high, the target rule is determined as the third reference rule.

[0180] It should be noted that the first reference rule, the second reference rule, and the third reference rule correspond to different fuzzy rule sets. The first reference rule is selected when resources are scarce and the number of CEs is low, tending to set a higher CE threshold to save resources; the second reference rule is applicable when resources are sufficient, and a more flexible CE threshold adjustment strategy can be adopted; the third reference rule is activated when resources are scarce and the number of CEs is high, tending to adopt a more aggressive strategy to reduce CE risk.

[0181] In an optional implementation, when the first resource information indicates resource scarcity and the first error information indicates a low number of CEs (Errors), the agent determines the target rule as a first reference rule (e.g., a conservative rule set), aiming to reduce resource consumption and maintain sustainable resource use by setting a higher CE threshold. When the first resource information indicates sufficient resources, regardless of the number of CEs, the agent tends to determine the target rule as a second reference rule (e.g., a balanced rule set), allowing for more flexible threshold adjustments, effectively handling CEs while avoiding excessive resource consumption. Faced with the challenge of resource scarcity and a high number of CEs, the agent determines the target rule as a third reference rule (e.g., an aggressive rule set), lowering the CE threshold to handle CE events promptly. Although this may increase resource consumption, it helps ensure system stability and security.

[0182] Figure 4 This is a schematic diagram of an optional reference rule according to an embodiment of this application; as shown... Figure 4As shown, the CE threshold is represented by the z-axis, the CE number by the x-axis, and the number of system interrupts by the y-axis. This indicates the correspondence between the CE threshold, the number of CEs, and the number of interrupts. For example, the first reference rule could be: [1 1 1;1 2 2;1 3 3;1 4 4;1 5 5;2 1 2;2 2 3;2 3 4;2 45;2 5 5;3 1 3;3 2 4;3 3 5;3 4 6;3 5 7;4 1 4;4 2 5;4 3 6;4 4 7;4 5 7;5 1 5;5 26;5 3 7;5 4 7;5 5 7;]; and the corresponding rule table could be as shown in Table 1:

[0183] Table 1 First Reference Rule

[0184]

[0185] Based on Table 1, the fuzzy conditional statements are as follows, totaling 25:

[0186] For example (only 4 examples are shown): If "Memory CE count = Very Low" and "Interrupt count = Very Low", Then "Memory CE threshold = Very Low";

[0187] If "Memory CE count = Very Low" and "Interrupt count = Low", then "Memory CE threshold = Low";

[0188] If "Memory CE count = Very High" and "Interrupt count = High", then "Memory CE threshold = Very High";

[0189] If "Memory CE count = Very High" and "Interrupt count = Very High", then "Memory CE threshold = Very High";

[0190] There are 5 fuzzy sets for the number of memory CEs and the number of system interrupts caused by memory CEs, and 7 fuzzy sets for the memory CE threshold. Among them, 5, 3, and 7 can represent "number of memory CEs = Very High" and "number of interrupts = Medium", then "memory CE threshold = Very High".

[0191] Similarly, Figure 5 This is a schematic diagram of another optional reference rule according to an embodiment of this application; such as Figure 5As shown, the CE threshold is on the z-axis, the CE number is on the x-axis, and the number of system interrupts is on the y-axis. This indicates the correspondence between the CE threshold, the number of CEs, and the number of interrupts. The second reference rule can be: [1 1 1;1 2 2;1 3 3;1 4 4;1 5 5;2 1 2;2 2 3;2 3 4;2 4 5;2 5 6;3 1 3;3 2 4;3 3 5;3 4 6;3 5 7;4 1 4;4 2 5;4 3 6;4 4 7;4 5 7;5 15;5 2 6;5 3 7;5 4 7;5 5 7;];

[0192] Figure 6 This is a schematic diagram of another optional reference rule according to an embodiment of this application; as shown... Figure 6 As shown, the CE threshold is the z-axis, the CE number is the x-axis, and the number of system interrupts is the y-axis. This indicates the correspondence between the CE threshold, the number of CEs, and the number of interrupts. The third reference rule can be: [1 1 1;1 2 1;1 3 2;1 4 3;1 5 4;2 1 2;2 2 3;2 3 4;2 4 5;2 5 6;3 1 3;3 2 4;3 3 5;3 4 6;3 5 7;4 1 4;4 2 5;4 3 6;4 4 7;4 5 7;5 1 6;5 27;5 3 7;5 4 7;5 5 7;].

[0193] Example 10:

[0194] Assume the current state of the server's memory is as follows:

[0195] First resource information: The remaining repair resource ratio is 20%, indicating that resources are relatively scarce.

[0196] The first error message: The number of error codes (CEs) is 800 and is increasing, indicating that the server memory is facing a high risk of errors.

[0197] Based on the rules defined above, the agent determines that the current state meets the condition of "resource scarcity and high number of repair operations (CEs)". Therefore, it selects the third reference rule (aggressive rule set) as the target rule to reduce the CE threshold to cope with the high number of CEs in memory, even if resource consumption may increase. After applying the aggressive rule set, the agent will continuously monitor changes in the repair resource status and the number of CEs to evaluate the effectiveness of the strategy, and in the next cycle, it will redetermine the target rule based on the updated first resource information and first error information. For example, during the cycle in which the aggressive rule set is applied, the number of CEs is effectively controlled, decreasing to 600, but due to frequent repair operations, the remaining repair resource ratio further decreases to 15%. In the next cycle, based on the updated first resource information (resource remaining ratio 15%) and first error information (number of CEs 600 and trending towards stability), the agent may choose a more conservative rule set to slow down resource consumption and achieve longer-term system stability and resource management.

[0198] Through the above embodiments of this application, it can be seen how an intelligent agent can dynamically select the optimal CE threshold adjustment strategy based on different resource conditions and CE error information, so as to achieve the optimal balance between system performance and resource management.

[0199] In an optional implementation, determining the target threshold based on the target rule and the first error information includes: determining the first number of correctable errors in the server memory within a preset period and the first number of memory interruptions based on the first error information; determining at least one first fuzzy set based on the first number of errors and at least one second fuzzy set based on the first number of interruptions; determining at least one third fuzzy set based on the target rule based on the at least one first fuzzy set and at least one second fuzzy set, and obtaining the target threshold based on the third fuzzy set, wherein the third fuzzy set is used to indicate the fuzzy set to which the target threshold belongs.

[0200] It should be noted that the first and second fuzzy sets are fuzzy sets partitioned according to the number of CEs (first error count) and the number of CE interruptions (first interruption count), used to describe the fuzzy linguistic values ​​of the input variables. For example, the first fuzzy set may include fuzzy linguistic values ​​such as "extremely low," "low," "medium," "high," and "extremely high," while the second fuzzy set may be partitioned according to the frequency of interruptions. The third fuzzy set is the output fuzzy set determined through a fuzzy inference process based on the target rule (conservative / balanced / aggressive) and the first and second fuzzy sets. It indicates the fuzzy interval in which the target CE threshold lies. The determination of the third fuzzy set is based on inference from the input fuzzy set and fuzzy rules, and is used to guide the adjustment of the CE threshold.

[0201] In an optional implementation, a first number of errors and a first number of interruptions are determined. Then, based on these numbers, the input fuzzy sets (the first fuzzy set and the second fuzzy set) are determined respectively. Finally, through fuzzy rule reasoning and in combination with the target rule, the output fuzzy set (the third fuzzy set) is determined, and the specific target threshold is calculated accordingly.

[0202] In an optional implementation, the specific number of CEs and the number of system interruptions caused by CEs within a preset period are collected as input variables for fuzzy inference. Based on the number of CEs and the number of interruptions, corresponding fuzzy linguistic values ​​are determined. For example, if the number of CEs is very low, it belongs to the first fuzzy set; if the number of interruptions is medium, it belongs to the second fuzzy set. Based on the selected target rule (conservative / equilibrium / aggressive) and the determined input fuzzy sets (first and second fuzzy sets), the agent determines the output fuzzy set (third fuzzy set) through a fuzzy rule inference process. This process may involve matching multiple fuzzy rules, ultimately determining an output fuzzy set that indicates the fuzzy range of the CE threshold. Based on the determined third fuzzy set, the agent converts the fuzzy linguistic values ​​into specific CE threshold values, i.e., the target threshold, through a defuzzification process. This process typically uses a membership function, such as a trigonometric function, to quantize the fuzzy set, and then uses defuzzification methods such as the centroid method or the mean method to calculate the specific output value.

[0203] Through the above-described embodiments of this application, a suitable CE threshold for the current system state can be determined through fuzzy reasoning based on specific first error information and target rules, thereby achieving precise adjustment of the server memory CE threshold. This method can provide a flexible and accurate CE threshold adjustment strategy in a real-time changing system environment, which is beneficial for balancing stable system operation and efficient resource utilization.

[0204] In an optional implementation, determining at least one third fuzzy set based on at least one first fuzzy set and at least one second fuzzy set according to a target rule includes: calculating the first membership degree corresponding to the first number of errors and each of the at least one first fuzzy set; calculating the second membership degree corresponding to the first number of interruptions and each of the at least one second fuzzy set; and determining a target threshold based on the at least one first membership degree and at least one second membership degree.

[0205] It should be noted that the first fuzzy set represents a fuzzy set used to describe the changes in the number of CEs in server memory, converting continuous CE counts into fuzzy linguistic descriptions (e.g., "Very Low", "Low", "Medium", "High", "Very High"). The second fuzzy set corresponds to another fuzzy set used to describe the number of system interrupts caused by CEs, similarly converting continuous interrupt counts into fuzzy linguistic descriptions (e.g., "Very Low", "Low", "Medium", "High", "Very High"). The third fuzzy set is a fuzzy set obtained through fuzzy inference based on the first and second fuzzy sets and the selected target rule, used to determine the fuzzy linguistic description of the CE threshold (e.g., "Very Low", "Low", "Medium", "High", "Very High").

[0206] The first membership degree and the second membership degree represent the degree of membership between the number of CEs and the number of interruptions and each fuzzy set, respectively. They are used to quantify the matching degree between the input variables and the fuzzy sets and are important parameters in fuzzy inference.

[0207] In an optional implementation, the observed number of CEs and the number of interruptions are fuzzified using a first fuzzy set and a second fuzzy set. This step converts continuous input variables into a fuzzy linguistic description, facilitating fuzzy inference. The membership degrees between the number of CEs and the number of interruptions and the first and second fuzzy sets are determined. The membership degree is calculated based on the membership function of each fuzzy set, reflecting the degree of matching between the input variables and the fuzzy set. According to the target rule, fuzzy logic (such as the Mamdani or Sugeno model) is applied to combine the first and second membership degrees to determine the fuzzy linguistic description of the CE threshold (the third fuzzy set). The fuzzy description of the CE threshold obtained from fuzzy inference is converted into a specific CE threshold for practical application. Various defuzzification methods exist, including the centroid method and the maximum membership method; a suitable method is selected to convert the fuzzy linguistic description into a deterministic numerical output.

[0208] Example 11:

[0209] Assume that there are 7500 instances of CE in the current server memory, and the number of system interrupts caused by CE is 1000.

[0210] The number of CEs is 7500, falling into the two fuzzy sets "High" and "Very High".

[0211] The number of interruptions is 1000, and the result falls into the two fuzzy sets "Medium" and "High".

[0212] The membership degree between the number of CEs and the "High" fuzzy set is calculated to be 0.4 (assuming that it is calculated according to the triangular membership function).

[0213] The membership degree between the number of CEs and the "Very High" fuzzy set is calculated to be 0.6.

[0214] The number of interruptions and the membership degree of the "Medium" fuzzy set are calculated to be 0.3.

[0215] The number of interruptions and the membership degree of the "High" fuzzy set are calculated to be 0.7.

[0216] Suppose the agent determines the target rule to be an aggressive rule set (fis_aggressive) based on the previous training round. Based on the current input membership degrees and rule set, the agent will determine a fuzzy description of the CE threshold through fuzzy inference. Using the "Mamdani" fuzzy inference model, the first and second membership degrees are combined with the target rule to calculate the fuzzy output of the CE threshold. This might be a fuzzy description like "Medium High," indicating that the CE threshold should be set at a moderately high level. Assuming the defuzzification method is the centroid method, the agent will calculate a specific CE threshold, such as 1423, based on the membership function of the "Medium High" fuzzy output.

[0217] The implementation methods described in this application not only consider the absolute values ​​of the number of CEs and the number of interruptions, but also incorporate the guidance of the target rule set. Through fuzzy reasoning, a CE threshold adapted to the current system state is derived, ensuring a balance between system stability and resource utilization efficiency. This dynamic adjustment mechanism helps the agent make optimal decisions in complex and ever-changing environments, avoiding the adaptability problems caused by fixed thresholds.

[0218] In an optional implementation, determining the target threshold based on at least one first membership degree and at least one second membership degree includes: cross-combining at least one first membership degree and at least one second membership degree to obtain at least one membership degree combination; calculating the corresponding reference threshold area in at least one third fuzzy set using the membership degree combination; superimposing the reference threshold areas to obtain the target threshold area; and determining the median value of the target threshold area as the target threshold.

[0219] It should be noted that membership combination refers to pairing the first membership degree (fuzzy set membership degree related to the number of CEs) and the second membership degree (fuzzy set membership degree related to the number of interruptions) to form a combined input for fuzzy inference. The reference threshold area is an intermediate result in the fuzzy inference process, representing the area of ​​the fuzzy output of the CE threshold corresponding to each membership degree combination, used to calculate the final CE threshold. The target threshold area is formed by superimposing or fusing all reference threshold areas according to certain rules (such as the maximum membership principle), resulting in a final fuzzy output that represents the CE threshold distribution under the current system state. The median value is a representative value selected from the target threshold area during the defuzzification process; that is, a specific CE threshold value is obtained by quantizing the fuzzy output. The median value is usually chosen because it reflects the central tendency and distribution characteristics of the fuzzy output.

[0220] In an optional implementation, the fuzzy set membership degree of the number of CEs (first membership degree) and the fuzzy set membership degree of the number of interruptions (second membership degree) are cross-paired to form different membership degree combinations. Each combination represents a description of the current system state and can be input into the fuzzy inference system to obtain the corresponding CE threshold fuzzy output. Using each membership degree combination, the corresponding reference threshold area is calculated in the third fuzzy set through fuzzy inference. Here, the reference threshold area actually refers to the area under the fuzzy output curve with membership applied, which is related to the final CE threshold. All calculated reference threshold areas are superimposed or merged to obtain the target threshold area. This process can be implemented in various ways, such as the maximum membership principle, weighted average, etc., to ensure that the final CE threshold can fully reflect the evaluation result of the current system state. The median value is calculated from the target threshold area as the determined CE threshold. The selection of the median value is based on the defuzzification principle, which represents the central tendency of the fuzzy output, can balance the setting of the CE threshold, avoid extreme cases of being too high or too low, and thus optimize the stability and performance of the system.

[0221] Example 12:

[0222] Assuming that after the steps of fuzzification and membership degree calculation, the following membership degree information is obtained:

[0223] The number of CEs is 7500, and the corresponding first membership degrees are 0.4 (High fuzzy set) and 0.6 (Very High fuzzy set).

[0224] The number of interruptions is 1200, and the corresponding second membership degrees are 0.2 (Medium fuzzy set) and 0.8 (High fuzzy set).

[0225] Two membership combinations are formed: (High, Medium) and (Very High, High). For the combination (High, Medium), the corresponding CE threshold fuzzy output is calculated in the third fuzzy set, assuming the result is a curve distributed across the threshold [1200, 1500]. For the combination (Very High, High), the calculated CE threshold fuzzy output is distributed across [1400, 1800]. The two fuzzy output curves are superimposed to form a target threshold area covering [1200, 1800]. The shape of this area will depend on the overlap of the two curves and their respective contributions. The median value is calculated from the superimposed target threshold area. Assuming the distribution center of the target threshold area is 1600, then 1600 is set as the CE threshold for the current period.

[0226] Through the above-described embodiments of this application, it can be seen that the intelligent agent uses fuzzy control theory to transform the fuzzy description of the number of CEs and interrupts into a specific CE threshold. This threshold fully considers the actual situation of the current system, such as resource availability and the urgency of CE processing, so as to achieve effective monitoring and processing of memory CE problems while balancing system stability and resource efficiency.

[0227] In an optional implementation, before determining at least one first fuzzy set based on the first number of errors and at least one second fuzzy set based on the first number of interruptions, at least one of the following is included:

[0228] 1) Set multiple first fuzzy sets, and set a range for the number of errors for each of the multiple first fuzzy sets;

[0229] 2) Set multiple second fuzzy sets, and set the range of interruption counts for each of the multiple second fuzzy sets;

[0230] 3) Set multiple third fuzzy sets and set the range of the target threshold for each of the multiple third fuzzy sets.

[0231] In an optional implementation, multiple first fuzzy sets are defined, each fuzzy set corresponding to a range of error counts. For example, the "very low" fuzzy set corresponds to a range of error counts of [0, 1000], the "low" fuzzy set corresponds to [1000, 3000], the "medium" fuzzy set corresponds to [3000, 5000], the "high" fuzzy set corresponds to [5000, 7000], and the "very high" fuzzy set corresponds to [7000, 10000].

[0232] Similar to the first fuzzy set, multiple second fuzzy sets are defined, each corresponding to a range of interruption counts. For example, the "extremely low" fuzzy set corresponds to the interruption count range of [0, 100], the "low" fuzzy set corresponds to [100, 300], the "medium" fuzzy set corresponds to [300, 500], the "high" fuzzy set corresponds to [500, 700], and the "extremely high" fuzzy set corresponds to [700, 1500].

[0233] Define multiple third fuzzy sets, each corresponding to a target CE threshold range. For example, the threshold range corresponding to the "extremely low threshold" fuzzy set is [100, 300], the "low threshold" fuzzy set is [300, 600], the "medium threshold" fuzzy set is [600, 1000], the "high threshold" fuzzy set is [1000, 1400], and the "extremely high threshold" fuzzy set is [1400, 2000].

[0234] Suppose the server detects 6500 instances of event cancellation (CE) and 450 system interruptions caused by CE. The agent first converts these numbers into fuzzy language: the 6500 CEs fall within the defined "high" fuzzy set, so the agent fuzzifies the CE count as "high." The 450 interruptions fall within the defined "medium" fuzzy set, so the agent fuzzifies the interruption count as "medium." Assume the agent has determined the target rule to be an aggressive rule set. According to the fuzzy inference rules, when the CE count is "high" and the interruption count is "medium," the output should be a "high threshold" fuzzy set, with a threshold range of [1000, 1400]. Next, the agent will calculate the specific target CE threshold based on the "high" CE count, the "medium" interruption count, and the aggressive rule set through fuzzy inference. For example, if the agent calculates the output fuzzy set as "high threshold" based on the fuzzy rules and membership functions, it might determine the target CE threshold to be 1200 based on a defuzzification method (such as the centroid method).

[0235] The embodiments described above in this application can convert the number of CEs and interrupts into fuzzy language descriptions, and then, through a fuzzy rule reasoning process, obtain a CE threshold that adapts to the current system state, thereby achieving dynamic adjustment of the server memory CE threshold. This method provides a more flexible and intelligent solution when dealing with fuzzy and uncertain system states.

[0236] Figure 7 This is a schematic diagram of another optional server memory repair method according to an embodiment of this application; as shown. Figure 7As shown, the memory CE information (first error information) is fuzzified and fuzzy inference is performed. This process may utilize a fuzzy knowledge base, which includes data (universe of discourse, fuzzy sets, membership functions, etc.) and rules (reference rules). The result obtained after fuzzy inference is defuzzified and sent to the control object to determine the memory CE threshold. After this process is complete, statistics can be collected and feedback can be provided.

[0237] Specifically, the input variables are the number of CEs (Continuous Errors) in the server memory and the number of CE interrupts within a preset period. The output variable is the CE threshold.

[0238] Membership functions define fuzzy sets for input and output variables and specify the membership function. For example, the number of CEs can be divided into "very low", "low", "medium", "high", and "very high". The membership function of each fuzzy set is represented by a triangular distribution to ensure a reasonable conversion from input values ​​to fuzzy language descriptions.

[0239] Based on expert experience or system requirements, three fuzzy rule sets are created: conservative, balanced, and aggressive. Each rule set contains a series of if-then conditions to transform different input fuzzy sets into output fuzzy sets. For example, "If the number of CEs is 'very high' and the number of interruptions is 'high', then the CE threshold is 'very high'."

[0240] The measured number of CEs and interruptions are compared with the membership functions of each fuzzy set to calculate the membership degree of each fuzzy set. If the input value is close to the boundary of a fuzzy set, the same input value may belong to multiple fuzzy sets, requiring cross-membership calculation to ensure that each input has a clear fuzzy language description. The membership degrees of the number of CEs and the number of interruptions are combined according to the rule set requirements to obtain multiple possible membership degree combinations. Based on the membership degree of each combination, the corresponding fuzzy rules are applied to calculate the membership degree of the output fuzzy set. The strength of the rule, i.e., the membership degree of the output fuzzy set, is determined by applying least squares or algebraic product operations.

[0241] The membership degree of the output fuzzy set is converted into a clear CE threshold value through a defuzzification algorithm. For example, if the output fuzzy set is "medium", the threshold obtained after defuzzification may be 900.

[0242] Figure 8 This is a schematic diagram of another optional server memory repair method according to an embodiment of this application; as shown. Figure 8As shown, the z-axis represents the CE threshold, the x-axis represents remaining repair resources, and the y-axis represents the number of memory CEs, illustrating the relationship between the number of memory CEs, remaining repair resources, and the CE threshold. In an optional implementation, the number of CEs and the number of interrupts are continuously read in real time. The Q-table is queried based on the current state to find the optimal fuzzy rule set for adjusting the CE threshold. The determined CE threshold is applied to the system to guide the CE processing strategy.

[0243] In optional implementations, deep reinforcement learning (DRL) methods, particularly deep Q-networks (DQN), can be introduced to further improve the intelligence level of memory CE threshold adjustment and the efficiency of resource scheduling. DQN can handle more complex environmental states and action spaces, extract state features through neural networks, optimize Q-value estimation, and thus improve the accuracy and robustness of policy selection.

[0244] In addition to the number of memory event cancellers (CEs) and the number of interrupts, the system also includes additional information such as server load status, time period, historical CE event processing success rate, and current threshold settings. This additional state information provides a more comprehensive system operating context, helping the agent make more accurate decisions. Besides adjusting the CE threshold, resource scheduling actions can be added, such as allocating repair resources to the highest priority CE events, batch processing low-priority CE events, and preloading spare memory resources. In this way, the agent can not only dynamically adjust the threshold but also actively schedule resources, achieving dual optimization. A multi-layer neural network is constructed using DQN instead of traditional Q-learning as an approximator for the Q-value function. The network input is a state representation vector, and the output is the Q-value of different actions (including threshold adjustment and resource scheduling actions). Through iterative training, the agent learns to select the optimal action sequence in complex environments. The reward function design considers three objectives: the timeliness of CE processing, system stability, and resource utilization efficiency. The weights of the reward function can be dynamically adjusted according to the system's current primary needs. For example, when resources are scarce, the weight of resource utilization efficiency can be increased; conversely, the timeliness of CE processing can be emphasized more. The agent learns online during system operation and can also learn offline using historical data, improving the model's generalization ability and prediction accuracy. Online learning ensures real-time policy optimization, while offline learning enhances the model's ability to handle unknown states. Long Short-Term Memory (LSTM) units are incorporated into the DQN architecture to capture the temporal dependencies of state information. LSTMs can remember important features from long-term series data, which is crucial for understanding the development trends of CE events and the long-term impact of resource scheduling. In addition to traditional reward signals, the agent also receives multimodal feedback from system logs, performance metrics, and hardware states as auxiliary information for decision-making. The fusion of multimodal data improves the comprehensiveness and reliability of decision-making, helping the agent make reasonable judgments in more complex scenarios.

[0245] Example 13:

[0246] Imagine a large data center server cluster facing frequent CE events and resource scheduling pressure. In this scenario, the following steps can be taken:

[0247] S1. State Information Fusion: The agent collects and fuses state information such as the number of CEs, the number of interrupts, the server CPU utilization, the memory usage, and the network traffic to form a comprehensive state representation vector.

[0248] S2. Action Space Design: In addition to the three CE threshold adjustment strategies—conservative, balanced, and aggressive—resource scheduling actions have been added, such as "prioritize repairing high-priority CEs," "batch process low-priority CEs," "preload spare memory resources," and "adjust CPU scheduling priority." These actions are designed to take into account the multi-dimensional needs of CE event handling.

[0249] S3. DQN Model Training: The DQN model is iteratively trained using historical operating data and a simulated environment. The agent experiences different states in the simulation, learning how to minimize the impact of CE events while maximizing resource utilization efficiency under resource-constrained conditions through optimal threshold adjustments and resource scheduling strategies.

[0250] S4. Online Decision Making: In a real-world operating environment, the agent makes online decisions based on real-time state information using the DQN model, dynamically adjusting the CE threshold and resource allocation strategy. For example, when high server load and frequent CE events are detected, the agent may choose to lower the CE threshold, prioritize resource allocation to repair high-priority CE events, and optimize CPU scheduling strategies to mitigate the impact of handling CE events on system performance.

[0251] S5. Performance Evaluation and Feedback Loop: The agent's decision-making effectiveness is evaluated through the actual performance of the system, including the timeliness of CE processing, system stability, and resource utilization efficiency. If performance does not meet expectations, the agent will adjust model parameters or modify action strategies and state representations to achieve continuous performance improvement.

[0252] This innovative extension scheme enables more precise dynamic adjustment of CE thresholds and intelligent resource scheduling to minimize the impact of CE events on server performance and stability. This approach is particularly suitable for handling large-scale, high-performance server clusters, significantly improving the overall operational efficiency and resource utilization of the data center.

[0253] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0254] Embodiments of this application also provide a server memory repair device. Figure 9 This is a structural block diagram of an optional server memory repair device according to an embodiment of this application, such as... Figure 9 As shown, the device includes:

[0255] The information determination module 902 is used to determine the first resource information corresponding to the repair resource in the server, wherein the first resource information is used to indicate the status of the resources in the server that can be used to repair memory errors;

[0256] The rule lookup module 904 is used to look up the target rule in the rule lookup table based on the first resource information and the first error information of the server memory. The first error information is used to indicate the error that occurred in the server memory within a preset period. The rule lookup table is used to indicate the correspondence between the first resource information, the first error information and the reference rule. The reference rule includes the target rule.

[0257] The threshold determination module 906 is used to determine the target threshold based on the target rule and the first error information;

[0258] Repair module 908 is used to repair server memory using repair resources when the first error message and target threshold meet preset repair conditions.

[0259] Optionally, the rule lookup module 904 is further configured to: determine second error information based on first error information, wherein the second error information indicates that there are correctable errors occurring within a preset period in the server; determine the number of correctable errors corresponding to each time period within the preset period based on the second error information; determine third error information based on the number of correctable errors corresponding to each time period, wherein the third error information indicates the change in the number of correctable errors within the preset period; and find the target rule in a rule lookup table based on the first resource information and the third error information, wherein the rule lookup table indicates the correspondence between the first resource information, the third error information, and the reference rule.

[0260] Optionally, the rule lookup module 904 is further configured to: determine multiple reference rules in the rule lookup table; search for multiple reference values ​​corresponding to the first resource information and the third error information in the rule lookup table, wherein the reference value corresponds one-to-one with the reference rule; and determine the reference rule corresponding to the highest reference value as the target rule.

[0261] Optionally, the rule lookup module 904 is further configured to: determine the first sub-state to which the first resource information belongs, and determine the second sub-state to which the third error information belongs, wherein the first sub-state is used to indicate the degree of the first resource information, and the second sub-state is used to indicate the trend of correctable errors; calculate a table index based on the first sub-state and the second sub-state, wherein the table index is used to indicate a specific row or column in the rule lookup table; and find multiple corresponding reference values ​​in the rule lookup table based on the table index.

[0262] Optionally, the rule lookup module 904 is further configured to: cross-combine at least one first sub-state in the first sub-state set and at least one second sub-state in the second sub-state set to obtain multiple first states in the first state set; construct a first rule lookup table by using the multiple first states and multiple reference rules as rows and columns respectively, and set the reference value in the rule lookup table to a preset value, wherein the reference value corresponds to the first state and the reference rule; update the rule lookup table M times, wherein the i-th update of the rule lookup table to obtain the (i+1)-th rule lookup table includes: setting the repair resource to conform to the second resource information and setting the number of correctable errors to the fourth error information, wherein M is an integer greater than 0, and i is an integer greater than 0 and less than or equal to M; update the i-th rule lookup table according to the second resource information and the fourth error information to obtain the (i+1)-th rule lookup table; and determine the (M+1)-th rule lookup table as the rule lookup table when i is greater than or equal to M.

[0263] Optionally, the rule lookup module 904 is further configured to: when the current time reaches a preset period, search for multiple reference values ​​corresponding to the second resource information and the fourth error information in the rule lookup table, and determine the reference rule corresponding to the highest reference value as the target rule; determine the target threshold based on the target rule and the fourth error information, and when the fourth error information and the target threshold meet the preset repair conditions, repair the server memory using repair resources; obtain information on correctable errors and repair resource information of the server memory, and update the second resource information, the fourth error information, and the reference values ​​corresponding to the second resource information and the fourth error information; when the second resource information reaches a preset resource amount or the number of correctable errors is greater than the preset number of errors, obtain the (i+1)th rule lookup table.

[0264] Optionally, the rule lookup module 904 is further configured to: determine the current reward value based on the number of correctable errors in the server memory and the repair resource information within a preset time period; estimate the maximum value of the reference value in the next preset period as the predicted reference value; and determine the updated reference value by weighting the reference value corresponding to the first parameter, the second resource information, and the fourth error information, and the second parameter, the current reward value, and the predicted reference value.

[0265] Optionally, the rule lookup module 904 is further configured to: determine the number of correctable errors in the server memory within a preset time as a first reward value; determine the change in the server's repair resources within a preset time as a second reward value; determine the preset reward value as a third reward value if the server stops operating within a preset time; and determine the sum of the first reward value, the second reward value, and the third reward value as the current reward value if a third reward value exists.

[0266] Optionally, the rule lookup module 904 is further configured to: update the reference value corresponding to the target rule based on the number of correctable errors in the server memory and the repair resource information within a preset time period.

[0267] Optionally, the repair module 908 is further configured to: determine the number of correctable errors based on the first error information; and repair the server memory using repair resources if the number of correctable errors is greater than or equal to a target threshold.

[0268] Optionally, the repair module 908 is further configured to: isolate the memory error region and replace the memory error region with a spare storage region; reset the memory error region and start the memory error region after the memory error region has been reset; back up the memory error region to a redundant memory region and start the corresponding redundant memory region.

[0269] Optionally, the rule lookup module 904 is further configured to: determine the target rule as a first reference rule when the first resource information indicates that the server memory is under-resourced and the first error information indicates that the number of correctable errors in the server memory is low; determine the target rule as a second reference rule when the first resource information indicates that the server memory is under-resourced; and determine the target rule as a third reference rule when the first resource information indicates that the server memory is under-resourced and the first error information indicates that the number of correctable errors in the server memory is high.

[0270] Optionally, the threshold determination module 906 is further configured to: determine the first number of correctable errors in the server memory and the first number of memory interruptions within a preset period based on the first error information; determine at least one first fuzzy set based on the first number of errors and at least one second fuzzy set based on the first number of interruptions; determine at least one third fuzzy set according to the target rule based on the at least one first fuzzy set and the at least one second fuzzy set, and obtain the target threshold based on the third fuzzy set, wherein the third fuzzy set is used to indicate the fuzzy set to which the target threshold belongs.

[0271] Optionally, the threshold determination module 906 is further configured to: calculate the first membership degree corresponding to the first number of errors and at least one first fuzzy set; calculate the second membership degree corresponding to the first number of interruptions and at least one second fuzzy set; and determine the target threshold based on at least one first membership degree and at least one second membership degree.

[0272] Optionally, the threshold determination module 906 is further configured to: cross-combine at least one first membership degree and at least one second membership degree to obtain at least one membership degree combination; calculate the corresponding reference threshold area in at least one third fuzzy set using the membership degree combination; superimpose the reference threshold areas to obtain the target threshold area, and determine the median value of the target threshold area as the target threshold.

[0273] Optionally, the threshold determination module 906 is further configured to: set multiple first fuzzy sets and set a range for the number of errors corresponding to each of the multiple first fuzzy sets; set multiple second fuzzy sets and set a range for the number of interruptions corresponding to each of the multiple second fuzzy sets; set multiple third fuzzy sets and set a range for the target threshold corresponding to each of the multiple third fuzzy sets.

[0274] For a description of the features in the embodiment corresponding to the server memory repair device, please refer to the relevant description in the embodiment corresponding to the server memory repair method, which will not be repeated here.

[0275] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the server memory repair method.

[0276] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described server memory repair method embodiments when running.

[0277] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0278] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described server memory repair method embodiments.

[0279] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described server memory repair method embodiments.

[0280] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0281] The foregoing has provided a detailed description of a server memory repair method, apparatus, storage medium, and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to aid in understanding the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for repairing server memory, characterized in that, include: Determine the first resource information corresponding to the repair resource in the server, wherein the first resource information is used to indicate the resource status in the server that can be used to repair memory errors; A second error message is determined based on the first error message, wherein the second error message is used to indicate a correctable error that occurs in the server memory within a preset period, and the first error message is used to indicate an error that occurs in the server memory within a preset period; The number of correctable errors corresponding to each time period within the preset period is determined based on the second error information; A third error information is determined based on the number of correctable errors corresponding to each time period, wherein the third error information is used to indicate the change in the number of correctable errors within the preset period; Multiple reference rules are determined in the rule lookup table; multiple reference values ​​corresponding to the first resource information and the third error information are found in the rule lookup table, wherein the reference value is the Q value corresponding to each reference rule, and the reference value is obtained through multiple rounds of training and optimization, and is used to indicate the expected performance and resource utilization efficiency of selecting the corresponding reference rule as the target rule under a specific resource state and the third error information; The reference rule corresponding to the highest reference value is determined as the target rule, wherein the rule lookup table is used to indicate the correspondence between the first resource information, the third error information and the reference rule, and the reference rule includes the target rule; Determine the target threshold based on the target rule and the first error information; If the first error message and the target threshold meet the preset repair conditions, the server memory is repaired using the repair resources.

2. The method according to claim 1, characterized in that, The step of searching for multiple reference values ​​corresponding to the first resource information and the third error information in the rule lookup table includes: Determine a first sub-state to which the first resource information belongs, and determine a second sub-state to which the third error information belongs, wherein the first sub-state is used to indicate the degree of the first resource information, and the second sub-state is used to indicate the changing trend of the correctable error; A table index is calculated based on the first sub-state and the second sub-state, wherein the table index is used to indicate a specific row or column in the rule lookup table; Based on the table index, the corresponding reference values ​​are found in the rule lookup table.

3. The method according to claim 1, characterized in that, Before the target rule is found in the rule lookup table based on the first resource information and the third error information, the process includes: By cross-combining at least one first substate in the first substate set and at least one second substate in the second substate set, multiple first states in the first state set are obtained; A first rule lookup table is constructed by using the plurality of first states and the plurality of reference rules as rows and columns, respectively, and the reference values ​​in the rule lookup table are set to preset values, wherein the reference values ​​correspond to the first states and the reference rules; The rule lookup table is updated M times, wherein the i-th update of the rule lookup table yields the (i+1)-th rule lookup table, including: The repair resource is set to conform to the second resource information, and the number of correctable errors is set to the fourth error information, where M is an integer greater than 0, and i is an integer greater than 0 and less than or equal to M; The i-th rule lookup table is updated based on the second resource information and the fourth error information to obtain the (i+1)-th rule lookup table; If i is greater than or equal to M, the (M+1)th rule lookup table is determined as the rule lookup table.

4. The method according to claim 3, characterized in that, The step of updating the i-th rule lookup table based on the second resource information and the fourth error information to obtain the (i+1)-th rule lookup table includes: When the current time reaches the preset period, multiple reference values ​​corresponding to the second resource information and the fourth error information are searched in the rule lookup table, and the reference rule corresponding to the largest reference value is determined as the target rule. A target threshold is determined based on the target rule and the fourth error information, and the server memory is repaired using the repair resources when the fourth error information and the target threshold meet the preset repair conditions. Obtain information on correctable errors and repairable resources in the server memory, and update the second resource information, the fourth error information, and the reference value corresponding to the second resource information and the fourth error information; When the second resource information reaches a preset resource quantity or the number of correctable errors is greater than the preset number of errors, the (i+1)th rule lookup table is obtained.

5. The method according to claim 4, characterized in that, The reference value of updating the second resource information and the fourth error information includes: The current reward value is determined based on the number of correctable errors in the server memory and the information on repairable resources within a preset time period; The maximum value estimated for the next preset period is the predicted reference value. The reference value is updated using a first parameter, a second parameter, the current reward value, and the predicted reference value, wherein the first parameter is used to indicate the learning rate, the second parameter is used to indicate the discount factor, and the first parameter and the second parameter are used to adjust the weights of the current reward value and the predicted reference value when updating the reference value; The step of determining the current reward value based on the number of correctable errors in the server memory and repair resource information within a preset time period includes: The number of correctable errors in the server memory within the preset time period is determined as the first reward value; The amount of change in the server's repair resources within the preset time period is determined as the second reward value; If the server stops operating within a preset time, the preset reward value will be determined as the third reward value; If a third reward value exists, the sum of the first reward value, the second reward value, and the third reward value is determined as the current reward value.

6. The method according to claim 1, characterized in that, After the target rule is obtained by searching the rule lookup table based on the first resource information and the third error information, the process includes: The reference value corresponding to the target rule is updated based on the number of correctable errors in the server memory and the repair resource information within a preset time period.

7. The method according to any one of claims 1 to 6, characterized in that, The step of repairing the server memory using the repair resources when the first error message and the target threshold meet preset repair conditions includes: The number of correctable errors is determined based on the first error information; If the number of correctable errors is greater than or equal to the target threshold, the server memory is repaired using the repair resources.

8. The method according to claim 7, characterized in that, The repair of the server memory using the repair resources includes one of the following: The faulty memory region is isolated and replaced with a spare storage region. The memory error region is reset, and once the memory error region has been reset, it is activated. The memory error area is backed up to the redundant memory area, and the corresponding redundant memory area is activated.

9. The method according to any one of claims 1 to 6, characterized in that, The method of finding the target rule in the rule lookup table based on the first resource information and the first error information of the server memory includes one of the following: When the first resource information indicates that the server memory is under resource constraints and the first error information indicates that the number of correctable errors in the server memory is low, the target rule is determined as the first reference rule, wherein the first reference rule is used to indicate a conservative rule set; If the first resource information indicates that the server memory has sufficient resources, the target rule is determined as the second reference rule, wherein the second reference rule is used to indicate a balanced rule set; When the first resource information indicates that the server memory is under resource constraints and the first error information indicates that the number of correctable errors in the server memory is high, the target rule is determined as a third reference rule, wherein the third reference rule is used to indicate an aggressive rule set.

10. The method according to claim 9, characterized in that, Determining the target threshold based on the target rule and the first error information includes: Based on the first error information, determine the first number of correctable errors in the server memory and the first number of memory interrupts within a preset period; At least one first fuzzy set is determined based on the number of errors, and at least one second fuzzy set is determined based on the number of interruptions. At least one third fuzzy set is determined based on the target rule according to the at least one first fuzzy set and the at least one second fuzzy set, and the target threshold is obtained based on the third fuzzy set, wherein the third fuzzy set is used to indicate the fuzzy set to which the target threshold belongs.

11. The method according to claim 10, characterized in that, The step of determining at least one third fuzzy set based on the at least one first fuzzy set and the at least one second fuzzy set according to the target rule includes: Calculate the first membership degree corresponding to the first number of errors and the at least one first fuzzy set; Calculate the second membership degree corresponding to the first interruption count and each of the at least one second fuzzy set; At least one third fuzzy set is determined based on at least one first membership degree and at least one second membership degree according to the target rule.

12. The method according to claim 11, characterized in that, The step of determining at least one third fuzzy set based on the at least one first fuzzy set and the at least one second fuzzy set according to the target rule, and obtaining the target threshold based on the third fuzzy set, includes: By cross-combining at least one first membership degree and at least one second membership degree, at least one membership degree combination is obtained; The corresponding reference threshold area is calculated in at least one third fuzzy set using the membership degree combination; The target threshold area is obtained by superimposing the reference threshold areas, and the median value of the target threshold area is determined as the target threshold.

13. The method according to claim 10, characterized in that, Before determining at least one first fuzzy set based on the first number of errors and at least one second fuzzy set based on the first number of interruptions, at least one of the following is included: Set multiple first fuzzy sets, and set a range for the number of errors corresponding to each of the multiple first fuzzy sets; Set multiple second fuzzy sets, and set a range of interruption counts for each of the multiple second fuzzy sets; Multiple third fuzzy sets are set, and the range of the corresponding target threshold is set for each of the multiple third fuzzy sets.

14. A server memory repair device, characterized in that, include: The resource determination module is used to determine the first resource information corresponding to the repair resource in the server, wherein the first resource information is used to indicate the resource status in the server that can be used to repair memory errors; A rule determination module is used to determine second error information based on first error information, wherein the second error information indicates that there are correctable errors occurring in the server within a preset period, and the first error information indicates that there are errors occurring in the server within the preset period; determine the number of correctable errors corresponding to each time period within the preset period based on the second error information; determine third error information based on the number of correctable errors corresponding to each time period, wherein the third error information indicates the change in the number of correctable errors within the preset period; determine multiple reference rules in a rule lookup table; search for multiple reference values ​​corresponding to the first resource information and the third error information in the rule lookup table, wherein the reference value is a Q-value corresponding to each reference rule, and the reference value is obtained through multiple rounds of training and optimization, used to indicate the expected performance and resource utilization efficiency of selecting the corresponding reference rule as the target rule under a specific resource state and the third error information; and determine the reference rule corresponding to the largest reference value as the target rule, wherein the rule lookup table indicates the correspondence between the first resource information, the third error information, and the reference rules, and the reference rules include the target rule; The threshold determination module is used to determine the target threshold based on the target rule and the first error information; The memory repair module is used to repair the server memory using the repair resources when the first error message and the target threshold meet preset repair conditions.

15. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the server memory repair method as described in any one of claims 1 to 13 when executing the computer program.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the server memory repair method as described in any one of claims 1 to 13.

17. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the server memory repair method as described in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Fault monitoring method and system for analog integrated circuit

    CN119471293A

  • Memory state detection method and apparatus, and communication device and storage medium

    WO2025025683A1