Memory isolation
Patent Information
- Application Number
- PCT/CN2026/080044
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-02-26
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026080044_01102026_PF_FP_ABST
Abstract
Description
Memory isolation Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to memory isolation. Background Technology
[0002] Memory is a crucial component of computing devices. Memory errors are among the most common hardware system failures and a leading cause of server downtime, significantly impacting server reliability, availability, and performance. Memory errors are primarily categorized into Correctable Errors (CE) and Uncorrectable Errors (UE). UEs can cause server crashes, while a large number of CEs can lead to impaired system performance or even unavailability. Therefore, effectively avoiding UEs and a large number of CEs is essential for server reliability, availability, and serviceability.
[0003] Memory page isolation strategies isolate faulty memory regions, preventing further error accumulation and system crashes. Traditional page isolation strategies accumulate the number of Corrected Errors (CEs) in memory. When the number of CEs reaches a threshold, the kernel in the operating system issues a command to trigger a low-level memory isolation replacement action built into the Central Processing Unit (CPU). This locally isolates the faulty region, preventing the use of that memory and thus avoiding further read / write operations, thereby reducing the risk of system crashes or performance issues. However, adjusting the memory page isolation strategy requires modifying the kernel code of the computing device. The kernel is a core component of the operating system, directly managing hardware resources and system processes. Modifying the kernel code may introduce new errors or vulnerabilities, leading to system instability or even crashes. Summary of the Invention
[0004] This disclosure provides a memory isolation method, device, system, storage medium, and program product to achieve centralized memory isolation, which helps improve the stability of the system on the service node side.
[0005] This disclosure provides a memory isolation method applicable to a central control node, comprising: acquiring memory error information of a target service node; determining memory error characteristics of the memory page corresponding to the memory error information based on the memory error information; using an abnormal page prediction model to determine a target memory page to be isolated from the memory pages corresponding to the memory error information based on the memory error characteristics; and controlling the target service node to isolate the target memory page.
[0006] This disclosure also provides an electronic device, including: a memory and a processor; wherein the memory is used to store a computer program; and the processor is coupled to the memory and used to execute the computer program to perform the steps in the memory isolation method described above.
[0007] This disclosure also provides a computer-readable storage medium storing computer instructions that, when executed by one or more processors, cause the one or more processors to perform the steps in the memory isolation method described above.
[0008] This disclosure also provides a computer program product, including a computer program that, when executed by one or more processors, causes the one or more processors to perform the steps in the memory isolation method described above.
[0009] In this embodiment, the logic for determining the target memory page that needs to be isolated is implemented at the central control node, thus realizing a centralized page isolation strategy. Specifically, the central control node can determine the memory error characteristics of the memory page that has malfunctioned based on the memory error information of the service nodes; and using the abnormal page prediction model set in the central control node, it determines the target memory page to be isolated from the malfunctioning memory pages based on the memory error characteristics, and controls the target service node where the target memory page is located to perform page isolation on the target memory page, thus realizing a centralized page isolation strategy. This page isolation method allows for flexible adjustment of the memory page isolation strategy through software updates or modifications to the abnormal page prediction model on the central side, without requiring modifications to the kernel code on the service node side, which helps improve the system stability on the service node side. Attached Figure Description
[0010] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of this disclosure and are used to explain this disclosure, but do not constitute an undue limitation of this disclosure.
[0011] Figure 1 is a schematic diagram of the structure of the memory isolation system provided in an embodiment of this disclosure.
[0012] Figure 2 is a schematic diagram of the internal structure of the memory provided in an embodiment of this disclosure.
[0013] Figure 3 is a flowchart illustrating the memory isolation method provided in an embodiment of this disclosure.
[0014] Figure 4 is a schematic diagram of the memory isolation process provided in an embodiment of this disclosure.
[0015] Figure 5 is a flowchart illustrating another memory isolation method provided in an embodiment of this disclosure.
[0016] Figure 6 is a schematic diagram of the structure of the electronic device provided in the embodiment of this disclosure. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0018] It should be noted that, in the cases involving user information in the embodiments of this disclosure, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0019] The terms and concepts involved in the embodiments of this disclosure will be explained below.
[0020] Memory Page: A memory page is 4KB in size. The operating system (OS) typically manages a contiguous logical address space of 4KB, while the corresponding physical page frames are distributed across contiguous regions on the same row across multiple different Dynamic Random Access Memory Chips (DRAM Chips). A 2MB large page consists of 512 4KB basic memory pages.
[0021] Page Offlining: When a memory page fails and meets the conditions for going offline, the OS calls an interface to copy the contents of the memory page to a new physical page frame, isolates the memory page, and updates the mapping table from virtual pages to physical page frames.
[0022] Centralization: Centralizing computing and data storage on one or more servers, and centrally managing and accessing various resources.
[0023] On-device: refers to applications and resources on client devices or specific physical nodes. Most of the computation and processing is done on the client devices, rather than on the server.
[0024] The following provides an exemplary description of a traditional page isolation scheme.
[0025] Error correction mechanisms within the central processing unit (CPU) of a computing device, such as error checking and correcting (ECC) technology, can identify corrected errors (CEs) in memory. The kernel in the operating system of the computing device counts the CEs and compares them with a set CE threshold. When the number of CEs reaches the CE threshold, the kernel issues a command to trigger the page isolation action of the CPU's built-in underlying memory, which locally isolates the faulty area, that is, it stops using the faulty memory area to avoid reading and writing to the faulty area again.
[0026] However, in traditional solutions, the operating system kernel executes a memory page isolation strategy, predicting that a memory page might experience a UE (User Error) based on the number of CEs (Error Detection) thresholds, and then isolating that memory page. Adjusting the memory page isolation strategy requires modifying the kernel code of the computing device. However, the kernel is a core component of the operating system, directly managing hardware resources and system processes; modifying the kernel code may introduce new errors or vulnerabilities, leading to system instability or even crashes. To address these technical problems, this disclosure proposes a centralized memory page isolation scheme in some embodiments. Specifically, a centralized page isolation strategy is implemented by implementing the logic for determining the target memory page that needs to be isolated at the central control node. Specifically, the central control node can determine the memory error characteristics of the erroneous memory page based on the memory error information from the service nodes; and using the abnormal page prediction model set in the central control node, it determines the target memory page to be isolated from the erroneous memory pages based on the memory error characteristics, and controls the target service node where the target memory page is located to perform page isolation on the target memory page, thus implementing a centralized page isolation strategy. This page isolation method allows for flexible adjustment of the memory page isolation strategy through software updates or modifications to the abnormal page prediction model on the central side, without requiring modifications to the kernel code on the service node side, which helps improve the system stability on the service node side.
[0027] The technical solutions provided by the embodiments of this disclosure are described in detail below with reference to the accompanying drawings.
[0028] It should be noted that the same reference numerals in the following figures and embodiments denote the same object or the same step. Therefore, once an object or step is defined in one figure or embodiment, it does not need to be discussed further in subsequent figures and embodiments.
[0029] Figure 1 is a schematic diagram of the structure of a memory isolation system provided in an embodiment of this disclosure. As shown in Figure 1, the system includes: a central management node 10 and service nodes 20. The number of service nodes 20 is one or more. "Multiple" means two or more (including two). Preferably, the number of service nodes 20 is multiple.
[0030] In this embodiment, the central management node 10 refers to a device, software module, or apparatus that performs memory exception management on the service node 20. The central management node 10 and the service node 20 can be on the same physical machine or separate physical machines. Figure 1 illustrates the central management node 10 and the service node 20 as separate physical machines, but this is not a limitation. For example, the central management node 10 can be implemented as a software module, virtual machine (VM), or container instance within the service node 20, or it can be implemented as an independent physical machine. When the central management node 10 and the service node 20 are located on the same physical machine, the central management node 10 can be deployed within the operating system of the device where the service node 20 resides, or deployed on the Base Board Management Controller (BMC) of the device where the service node 20 resides (not shown in the figures).
[0031] In this embodiment, service node 20 refers to a computing device with computing and communication functions, such as a server, mobile phone, computer, wearable device, etc. Service node 20 can be a service node in a distributed computing system or distributed storage system. For service node 20, memory anomalies are the cause of memory errors. Memory anomalies can easily lead to service node 20 crashing, affecting its reliability, availability, and serviceability (RAS).
[0032] In this embodiment, the service node 20 may include a CPU 201 and memory 202. In this embodiment, the specific implementation of the memory is not limited. Optionally, the memory may be Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), also known as DDR memory. DDR memory may be DDR4 memory or DDR5 memory. Of course, the memory may also be Phase Change Memory (PCM) or High Bandwidth Memory (HBM), etc.
[0033] Of course, the service node 20 may also include other storage media (not shown in Figure 1), such as random access memory (ROM), disk, etc. In this embodiment, the ROM may include: Basic Input Output System (BIOS) ROM (not shown in Figure 1). Among them, the BIOS ROM has a BIOS201b embedded in it.
[0034] In this embodiment, the error correction mechanism (such as ECC) within the CPU of the service node 20 can repair memory errors (CEs). For embodiments where the central management node 10 is deployed on the operating system (OS) 201a of a computing device or implemented as an independent physical machine, the OS 201a of the service node 20 can collect memory error information of the service node 20. This memory error information may include CE information and / or inspected UE information. Inspected UEs refer to memory UEs discovered through memory inspection. Inspection refers to automatically checking the memory status at set time intervals. The purpose is to proactively discover potential problems, rather than waiting for problems to occur before addressing them. Discovering UEs through the inspection mechanism (i.e., inspected UEs) will not directly cause system crashes or result in page isolation of the inspected UE's memory pages.
[0035] In some embodiments, service node 20 may further include an Error Detection and Correction Driver (EDAC Driver) (not shown in Figure 1). The EDAC Driver can collect memory error information from service node 20. The EDAC Driver is a software driver module corresponding to the operating system in service node 20, which can parse and print errors after receiving error events. Dedicated hardware for ECC is integrated within the memory controller. Correspondingly, the central control node can also acquire the memory error information collected by the EDAC Driver.
[0036] In embodiments where the central control node 10 is deployed on a computing device's BMC or implemented as an independent physical machine, the BMC can collect memory error information from the service node 20 (not shown in the accompanying drawings). The BMC can directly collect the memory error information from the service node 20, or it can collect the memory error information from the service node 20 through the BIOS 201b. The methods by which the central control node 10 obtains memory error information from the service node 20 shown in the foregoing embodiments are merely illustrative and do not constitute a limitation.
[0037] In this embodiment of the disclosure, memory error information refers to information describing memory errors (CE and / or inspection UE), which may include: memory location information where the error occurred and time information where the error occurred. The memory location information where the error occurred may include: the socket where the error occurred, the integrated memory controller (IMC), the channel, the dual in-line memory module (DIMM), the physical array (Rank), the chip selector, the memory chip, the logical array group (Bank Group), the logical array (Bank), and the row and column addresses, etc. Of course, memory error information may also include one or more of the following: the number of error bits in the memory page where the error occurred, the burst position corresponding to the error bit, and the identifier of the data queue (DQ) line corresponding to the error bit. "Multiple" refers to two or more types (including two).
[0038] The concepts in the memory location information where the error occurred are illustrated below with reference to Figure 2. As shown in Figure 2, a channel refers to the CPU data bus width. If the CPU has a 64-bit data bus and the memory chip is also 64-bit, then it's a single-channel memory; if the CPU has a 64-bit data bus, but the memory chip can support 128-bit, then the memory is dual-channel. The memory in Figure 2 is dual-channel memory. A DIMM (Digital Memory Modulated Module) is a memory module whose printed circuit board has gold fingers on both sides that contact the memory slots on the motherboard; this structure is called a DIMM.
[0039] A physical array (rank) refers to memory chips connected to the same chip selector (Chip Select, CS). A memory chip consists of several logical arrays (banks). A cell (memory unit) is the basic unit of memory storage, and a bank is a two-dimensional array of cells, with one cell used to store one bit of data.
[0040] A burst refers to the ability to quickly and continuously read or write a series of data from an activated memory page. It allows for the automatic transfer of data from multiple consecutive addresses after a single operation. The burst position corresponding to an error bit indicates which burst the error occurred in. DQ lines are physical lines connected to the data input / output (I / O) of the memory chip. They are responsible for the actual data transfer. For a single memory access, a memory chip generates multiple bursts in batches, with each burst typically carrying 4 or 8 bits of data simultaneously. Identical bit positions in different bursts share a set of DQ lines. For DRAM memory, due to the characteristics of DRAM I / O, bits associated with a burst or a DQ line are prone to error simultaneously. Therefore, ECC can completely correct single bursts and single DQs, but when multiple bursts and multiple DQ lines fail, UE (User Equipment) errors may occur. Therefore, memory-related burst and DQ line errors have a certain impact on UE crashes.
[0041] Since a CE (Error Detection) event may occur in the component of service node 20 in the future, a UE (User Equipment) may also occur, resulting in UE downtime. UE downtime refers to downtime caused by a UE-related memory error in service node 20. In this embodiment of the disclosure, in order to reduce the probability of UE downtime, service node 20 can send the collected memory error information to the central management node 10.
[0042] In this embodiment, the central control node 10 can perform abnormal memory page isolation management on the service node 20. The process of the central control node 10 performing abnormal page isolation management on the service node 20 is described below by way of example.
[0043] Figure 3 is a flowchart illustrating the memory isolation method provided in this embodiment. This method is applicable to central control nodes. As shown in Figure 3, the memory isolation method mainly includes the following steps.
[0044] 301. Obtain memory error information of the target service node.
[0045] 302. Based on the memory error information, determine the memory error characteristics of the memory page corresponding to the memory error information.
[0046] 303. Using the abnormal page prediction model, based on the characteristics of memory errors, determine the target memory page to be isolated from the memory pages corresponding to the memory error information.
[0047] 304. Control the target service node to isolate the target memory page.
[0048] In this embodiment, for the central control node, in step 301, memory error information of any one of the multiple service nodes (defined as the target service node) can be obtained. The method by which the central control node obtains memory error information can be found in the relevant content of the aforementioned system embodiment, and will not be repeated here. The memory error information obtained by the central control node can be the memory error information that the service node immediately sends to the central control node when it collects memory error information. Alternatively, it can be memory error information within any time period. In this embodiment, to improve the timeliness of memory exception handling, memory error information within a set time period sent by the service node can be obtained. The set time period can be the set time period closest to the current time, such as memory error information within 3 minutes, 5 minutes, or 10 minutes closest to the current time. In practice, the central control node can obtain memory error information from multiple service nodes; and can determine the memory error information of the target service node based on the identifier of the service node where the memory page corresponding to the memory error information is located.
[0049] Further, in step 302, the memory error characteristics of the memory page corresponding to the memory error information can be determined based on the memory error information. Here, the memory page corresponding to the memory error information refers to the memory page where the error occurred, as recorded in the memory error information. The memory error characteristics of the memory page refer to information related to the probability of determining a UE (User Error Detection) error in the memory page. In this embodiment, the specific implementation of the memory error characteristics of the memory page is not limited. The implementation of the memory error characteristics of each memory page is the same. The following example uses any memory page A among the memory pages corresponding to the memory error information to illustrate the implementation method of determining the memory error characteristics of the memory page.
[0050] Implementation Method 1: Memory page isolation isolates memory pages that may cause UE (User Equipment) crashes. Since a higher number of faulty bits in a memory page indicates a greater probability of the page causing a UE crash, the number of faulty bits can be implemented as a memory error characteristic. Therefore, for any memory page A within the memory pages corresponding to memory error information, the number of faulty bits in memory page A can be determined from its memory error information as its memory error characteristic.
[0051] Implementation Method 2: Due to the characteristics of DRAM I / O, bits associated with a single Burst or DQ line are prone to error simultaneously. Therefore, ECC can completely correct errors in a single Burst and DQ line, but when multiple Bursts and DQ lines fail, a UE (User Equipment) may be affected. Thus, memory errors in Bursts and DQ lines have a certain impact on UE downtime. Based on this, in some embodiments of this disclosure, for any memory page A, the burst position corresponding to the error bit can be obtained from the memory error information of memory page A. The burst position corresponding to the error bit is used to indicate in which burst transmission the error bit occurred. Error bits at the same burst position are concentrated in the same burst transmission. Furthermore, the burst error characteristics of memory page A can be determined based on the burst position corresponding to the error bit of any memory page A or the error time of the error bit of memory page A, serving as the memory error characteristics of memory page A.
[0052] In some embodiments, the burst error time of memory page A can be determined based on the error time of the error bits of memory page A, and used as the burst error characteristic of memory page A.
[0053] In other embodiments, UE refers to an error that exceeds the error correction capability of ECC. If the error bits are concentrated in the same burst transmission, ECC can usually correct these errors. However, if multiple burst transmissions fail simultaneously, the error correction capability of ECC may be insufficient to handle these errors, resulting in UE. Therefore, the number of bursts corresponding to the error bits can be implemented as a memory error characteristic. Based on this, for any memory page A, the burst position corresponding to the error bit can be obtained from the memory error information of memory page A. The burst position corresponding to the error bit is used to indicate in which burst transmission the error bit occurred. Error bits at the same burst position are concentrated in the same burst transmission. Therefore, the number of bursts in which memory page A has errors can be determined based on the burst position corresponding to the error bit, serving as the burst error characteristic of memory page A.
[0054] In some embodiments, since the number of error bits that ECC can correct is limited, typically only 1 bit of data can be corrected. If the number of error bits in the same burst transmission exceeds the number of error bits that ECC can correct, UE crashes may also occur. Based on this, for any memory page A, the number of error bits in the same burst can be determined according to the burst position corresponding to the error bits of memory page A, which serves as the burst error characteristic of memory page A.
[0055] It is worth noting that, in actual implementation, the characteristics of a burst error in memory page A may include: the time of the burst corresponding to the fault bit of memory page A, and / or, the number of bursts in memory page A that result in an error, and / or, the number of fault bits in the same burst.
[0056] Implementation Method 3: Referring to the analysis in Implementation Method 2 above, memory errors in the DQ line have a certain impact on UE crashes. Based on this, in some embodiments of this disclosure, for any memory page A, the identifier of the DQ line corresponding to the error bit can be obtained from the memory error information of memory page A. Further, the error characteristics of the DQ line of memory page A can be determined based on the identifier of the DQ line corresponding to the error bit of any memory page A or the error time of the error bit of any memory page A, and used as the memory error characteristics of memory page A.
[0057] In some embodiments, the error time of the DQ line in memory page A that malfunctions can be determined based on the error time of the error bit in memory page A, and used as the error characteristic of the DQ line in memory page A that malfunctions.
[0058] In some embodiments, since multiple Burst transmissions share the same set of DQ lines, the bits on these DQ lines are prone to error simultaneously. If the error occurs only on a single DQ line, the ECC can usually correct it. However, if errors occur simultaneously on multiple DQ lines, the ECC may be unable to correct them, leading to UE errors. Therefore, the number of DQ lines corresponding to the error bits can also be implemented as a memory error characteristic. Based on this, for any memory page A, the number of DQ lines in memory page A that have errors is determined according to the identifiers of the DQ lines corresponding to the error bits of memory page A, serving as the error characteristic of the DQ lines of memory page A.
[0059] In some embodiments, since the number of error bits that ECC can correct is limited, typically only 1 bit of data can be corrected. If the number of error bits in the same DQ line exceeds the number of error bits that ECC can correct, UE crashes may also occur. Based on this, for any memory page A, the number of error bits in the same DQ line can be determined according to the identifier of the DQ line corresponding to the error bits of memory page A, which serves as the error characteristic of the DQ line of memory page A.
[0060] It is worth noting that, in actual implementation, the characteristics of a burst error in memory page A may include: the identifier of the DQ line corresponding to the error bit of memory page A, and / or the number of DQ lines in memory page A where the error occurs, and / or the number of error bits in the same DQ line.
[0061] Implementation Method 4: Due to the locality of memory access (temporal and spatial locality), programs tend to frequently access certain memory regions or adjacent memory regions. This access pattern leads to certain memory cells being accessed frequently, making them more susceptible to interference or wear, thus causing errors. If consecutive addresses within a memory page are frequently accessed, these addresses may simultaneously encounter errors due to spatial locality, forming a CE storm. A CE storm refers to the frequent occurrence of CEs in a certain memory region (such as a memory page or consecutive memory cells within a memory page) over a period of time. CE storms are highly likely to cause UE (User Experience) issues. Therefore, the address information of the memory cells (cells) where errors occur within a memory page can also be implemented as a memory error feature. The address information of the memory cell can be represented using the memory row and column where the memory cell is located. Based on this, for any memory page A, the memory location information of the error bits in memory page A can be obtained from the memory error information of memory page A; and based on the memory location information of the error bits in memory page A, the address information of the memory cells in memory page A that have encountered errors can be determined as the memory error feature of memory page A.
[0062] The implementation methods of memory error features for memory page A shown in the foregoing embodiments are merely illustrative and do not constitute a limitation. In actual implementation, any one or more of the memory error features shown in implementation methods 1-4 can be selected as the memory error features of the memory page; of course, all the memory error features shown in implementation methods 1-4 can also be selected as the memory error features of the memory page. Preferably, all the memory error features obtained in implementation methods 1-4 are used as the memory error features in step 302. In this way, the subsequent use of multi-dimensional memory error features can provide more reference for predicting the possibility of a faulty memory page causing a UE, which helps to improve the accuracy of predicting the occurrence of a UE due to a memory page.
[0063] To identify the memory pages that truly need isolation from the memory pages corresponding to memory error information, in this embodiment, the central control node is pre-configured with an abnormal page prediction model. This abnormal page detection model is used to determine the memory pages from which a UE might be affected, i.e., to predict or determine the memory pages where a UE might be affected based on the memory error information. These memory pages are the ones identified by the abnormal page prediction model as needing isolation, i.e., the memory pages to be isolated. Since the abnormal page prediction model is set on the central control node side, the memory page isolation strategy can be adjusted by updating or modifying the abnormal page prediction model on the central side, without modifying the kernel code on the service node side, which helps improve the system stability on the service node side.
[0064] Based on the abnormal page detection model, in step 303, the abnormal page prediction model can be used to determine the memory pages to be isolated from the memory pages corresponding to the memory error information according to the memory error characteristics of each memory page. For ease of description, the memory pages to be isolated are defined as target memory pages.
[0065] In this embodiment, the specific implementation of the abnormal page prediction model is not limited. In some embodiments, the abnormal page prediction model includes multiple abnormal page prediction rules. These abnormal page prediction rules can be implemented in the form of conditional expressions, etc. If the memory error information of the memory page that has an error satisfies the abnormal page prediction rules, it indicates that the probability of the memory page causing a UE is relatively high, and the memory page can be identified as a memory page to be isolated. Based on this, for any memory page A among the memory pages corresponding to the memory error information, if the memory error characteristics of memory page A satisfy some or all of the abnormal page prediction rules among the multiple abnormal page prediction rules, it indicates that the probability of memory page A causing a UE is relatively high, and memory page A needs to be isolated. Therefore, if the memory error characteristics of memory page A satisfy some or all of the abnormal page prediction rules among the multiple abnormal page prediction rules, memory page A can be identified as a target memory page to be isolated. The more abnormal page prediction rules that the memory error characteristics of memory page A satisfy, the greater the probability of memory page A causing a UE, and the more necessary it is to isolate memory page A.
[0066] The following section, in conjunction with the memory error characteristics shown in the aforementioned implementation methods 1-4, provides an exemplary description of the implementation form of the abnormal page prediction rule and the specific implementation method of using the abnormal page prediction rule to determine the target memory page to be isolated.
[0067] Abnormal Page Prediction Rule 1: Since a higher number of erroneous bits in a memory page increases the probability of it causing a UE crash, the abnormal page prediction rule can be implemented as follows: if the number of erroneous bits in a memory page is greater than or equal to a set first threshold, then that memory page is identified as a page to be isolated. The first threshold is determined based on the error correction capability of ECC. Generally, ECC can correct 1 bit error; therefore, the first threshold can be 2.
[0068] Based on the abnormal page prediction rule 1, for any memory page A in the memory pages corresponding to the memory error information in step 301, if the number of error bits in memory page A is greater than or equal to the set first number threshold, then memory page A is determined to be the target memory page to be isolated.
[0069] Abnormal Page Prediction Rule 2: Based on the impact analysis of burst errors in memory pages in the aforementioned implementation method 2 of memory error characteristics, a quantity threshold (defined as the second quantity threshold) can be set based on the number of burst errors occurring in a memory page. The second quantity threshold can be 2. This is mainly because ECC can correct single-burst errors but cannot repair multiple burst errors. Based on this, Abnormal Page Prediction Rule 2 can be implemented as follows: If the number of burst errors occurring in a memory page is greater than or equal to the set second quantity threshold, then the memory page is determined to be the target memory page to be isolated. If the number of burst errors occurring in a memory page is less than the set second quantity threshold, but the number of error bits in the same burst is greater than or equal to the first quantity threshold, it indicates that ECC cannot repair the errors in the burst either, which may also lead to UE errors. Therefore, if the number of burst errors occurring in a memory page is less than the set second quantity threshold, but the number of error bits in the same burst is greater than or equal to the first quantity threshold, the memory page can also be determined to be the target memory page to be isolated.
[0070] Based on abnormal page prediction rule 2, for any memory page A in the memory pages corresponding to the memory error information in step 301, if the number of faulty bursts in memory page A is greater than or equal to a set second threshold, then memory page A is determined to be a target memory page to be isolated. Correspondingly, if the number of faulty bursts in a memory page is less than the set second threshold, but the number of faulty bits in the same burst is greater than or equal to a first threshold, then memory page A is determined to be a target memory page to be isolated.
[0071] Abnormal Page Prediction Rule 3: Based on the impact analysis of DQ line errors corresponding to memory pages in the aforementioned memory error feature implementation method 3, a quantity threshold (defined as the third quantity threshold) can be set based on the number of DQ lines corresponding to a memory page that have errors. The third quantity threshold is greater than or equal to 2. This is mainly because ECC can correct errors of a single DQ limiting, but cannot repair errors of multiple DQ limiting. Based on this, abnormal page prediction rule 3 can be implemented as follows: if the number of DQ lines with errors in a memory page is greater than or equal to the set third quantity threshold, then the memory page is determined to be the target memory page to be isolated. If the number of DQ lines with errors in a memory page is less than the set second quantity threshold, but the number of error bits in the same DQ line is greater than or equal to the first quantity threshold, it means that ECC cannot repair the errors in the DQ line either, which may also lead to UE errors. Therefore, if the number of DQ lines with errors in a memory page is less than the set third quantity threshold, but the number of error bits in the same DQ line is greater than or equal to the first quantity threshold, the memory page can also be determined to be the target memory page to be isolated.
[0072] Based on error page prediction rule 3, for any memory page A in the memory pages corresponding to the memory error information in step 301, if the number of DQ lines with errors in memory page A is greater than or equal to a set third threshold, then memory page A is determined to be a target memory page to be isolated. Correspondingly, if the number of DQ lines with errors in a memory page is less than the set third threshold, but the number of error bits in the same DQ line is greater than or equal to a first threshold, then the memory page can be determined to be a target memory page to be isolated.
[0073] Error Page Prediction Rule 4: Due to the characteristics of DRAM I / O, a burst and the bits associated with a DQ line are prone to fail together. When both the burst and the DQ line fail simultaneously, it indicates a memory hardware fault to a certain extent. If the number of faulty bits in the burst exceeds the ECC's repair capability (i.e., the set first threshold), and the number of faulty bits in the DQ line also exceeds the ECC's repair capability (i.e., the set first threshold), then a UE crash may occur, and this memory page can be identified as the target memory page.
[0074] Based on this, for any memory page A, if the error time of the burst in which memory page A has an error is the same as the error time of the data queue line in which memory page A has an error, and the error bits in the burst in which memory page A has an error are greater than or equal to a first quantity threshold, and the data queue line in which memory page A has an error is greater than or equal to a first quantity threshold, then memory page A is determined to be the target memory page.
[0075] Abnormal Page Prediction Rule 5: Based on the impact analysis of CE storms in the aforementioned implementation method 4 of memory error characteristics, abnormal page prediction rule 4 can be set based on the address information of the memory cells that have erroneous events in a memory page: if the address information of the memory cells that have erroneous events in a memory page is continuous, and the number of continuous memory cells is greater than or equal to a set fourth quantity threshold, then the memory page is determined as the target memory page to be isolated. The fourth quantity threshold can be greater than or equal to 2, and its specific value can be determined based on the offline data of the UE caused by historical memory error information.
[0076] Based on the abnormal page prediction rule 5, for any memory page A in the memory pages corresponding to the memory error information in step 301, if the address information of the memory units in memory page A that have errors is continuous, and the number of memory units with continuous addresses is greater than or equal to the set fourth quantity threshold, then memory page A is determined to be the target memory page to be isolated.
[0077] It is worth noting that the abnormal page prediction rules 1-5 shown in the foregoing embodiments are merely illustrative and do not constitute a limitation. In some embodiments, the aforementioned abnormal page prediction rules 1-5 can be implemented individually, or multiple abnormal page prediction rules 1-5 can be combined to form a single abnormal page prediction rule. For example, the aforementioned abnormal page prediction rules 4 and 5 can be combined to implement abnormal page prediction rule 6: If the address information of the memory cells in memory page A that have erroneous is continuous, and the number of continuous memory cells is greater than or equal to a set fourth quantity threshold, and the error time of the burst in memory page A that has erroneous is the same as the error time of the data queue line in memory page A that has erroneous, and the error bits in the burst in memory page A that has erroneous are greater than or equal to a first quantity threshold, and the data queue line in memory page A that has erroneous is greater than or equal to a first quantity threshold, then memory page A is determined to be the target memory page.
[0078] For example, the aforementioned abnormal page prediction rules 1 and 2 can be combined to implement abnormal page prediction rule 7: if the number of error bits in memory page A is greater than or equal to a set first quantity threshold, and the number of bursts in memory page A that cause errors is greater than or equal to a set second quantity threshold, then memory page A is determined to be the target memory page to be isolated; or, if the number of error bits in memory page A is greater than or equal to a set first quantity threshold, and the number of bursts in memory page A that cause errors is less than a set second quantity threshold, but the number of error bits in the same burst is greater than or equal to a first quantity threshold, then memory page A is determined to be the target memory page to be isolated.
[0079] For example, the aforementioned abnormal page prediction rules 1 and 3 can be combined to implement abnormal page prediction rule 8: If the number of error bits in memory page A is greater than or equal to a set first threshold, and the number of DQ lines in memory page A that have errors is greater than or equal to a set third threshold, then memory page A is determined to be a target memory page to be isolated. Correspondingly, if the number of error bits in memory page A is greater than or equal to the set first threshold, and the number of DQ lines in memory page A that have errors is less than the set third threshold, but the number of error bits in the same DQ line is greater than or equal to the first threshold, then the memory page can be determined to be a target memory page to be isolated. And so on.
[0080] It is worth noting that when multiple abnormal page prediction rules in the aforementioned abnormal page prediction rules 1-5 are implemented in combination, if the memory error characteristics of memory page A satisfy all the rules in the combined abnormal page prediction rules, then the memory page is determined to be the target memory page to be isolated.
[0081] In some embodiments, the risk level of a memory page can be determined based on the abnormal page prediction rules satisfied by the memory page. For ease of description and distinction, the abnormal page prediction rules satisfied by the memory page are defined as the first abnormal page prediction rules. The first abnormal page prediction rules can be some or all of the adopted abnormal page prediction rules. In some embodiments, the risk level of a memory page can be determined based on the number of first abnormal page prediction rules satisfied by the memory page. The risk level of a memory page is positively correlated with the number of first abnormal page prediction rules. That is, the more first abnormal page prediction rules a memory page satisfies, the greater the probability of a user experience (UE) occurring on that memory page, and the higher the risk level of the memory page.
[0082] In other embodiments, the risk level of a memory page can be determined based on the priority of the first abnormal page prediction rule satisfied by the memory page. The risk level of a memory page is positively correlated with the priority of the first abnormal page prediction rule. That is, the higher the priority of the first abnormal page prediction rule satisfied by a memory page, the greater the probability of a user experience (UE) occurring on that memory page, and the higher the risk level of the memory page. The priority of the abnormal page prediction rule is preset and determined based on the probability that the memory error characteristics in the abnormal page prediction rule lead to a UE. Specifically, the higher the probability that the memory error characteristics in the abnormal page prediction rule lead to a UE, the higher the priority of that abnormal page prediction rule. The probability that the memory error characteristics in the abnormal page prediction rule lead to a UE is obtained by the inventors of this disclosure through analysis of historical memory error information leading to UEs.
[0083] In some embodiments of this disclosure, considering that the static resource information of service nodes can also affect memory performance, the central control node can also obtain the static resource information of multiple service nodes. The static resource information of a service node refers to the static dimension information of the service node's resources. The static dimension information is fixed and unchanging, such as model and quantity.
[0084] The static resource information of a service node includes both hardware and software resources. Hardware resources include the CPU and memory. Software resources include the operating system and BIOS. CPU static information may include the model number. Different CPUs may use different ECC algorithms, and different ECC algorithms may have different CE processing capabilities; therefore, different CPUs may have varying degrees of impact on memory failures.
[0085] Static information about memory can include: memory model, number of modules, and insertion method. The memory model can indicate memory errors, while the number of modules and insertion method determine the interleaving pattern of memory accesses, both of which have a certain impact on memory failures. Static information about the operating system can include: the operating system model; static information about the BIOS can also include: the BIOS model. Of course, the server node model can also be considered part of the server node's static resource information. The operating system model, BIOS model, and server node model can all affect the server node's system behavior and load type, and the system behavior and load type of computing devices have a certain impact on memory failures. Therefore, the server node's static resource information can also be obtained.
[0086] Because different resource static information has varying degrees of impact on memory faults, the abnormal page prediction rules adapted to service nodes with different resource static information differ. The abnormal page prediction rules adapted to service nodes for various resource static information were obtained by the inventors of this disclosure through analysis of historical memory error information of service nodes for various resource static information leading to UE failures. The correspondence between resource static information and abnormal page prediction rules is pre-set on the central control node. In this correspondence, the abnormal page prediction rule corresponding to the resource static information is the abnormal page prediction rule adapted to that resource static information.
[0087] Based on the static resource information of service nodes, in some embodiments of this disclosure, a target abnormal page prediction rule adapted to the static resource information of any service node can be determined from a variety of abnormal page prediction rules. Then, the target abnormal page prediction rule can be used to determine the target memory page corresponding to the service node from the memory pages corresponding to the memory error information of that service node, based on the memory error characteristics of that service node. For a detailed implementation of determining the target memory page corresponding to the service node from the memory pages corresponding to the memory error information of that service node using the target abnormal page prediction rule based on the memory error characteristics of that service node, please refer to the aforementioned content on determining the target memory page to be isolated from the memory pages corresponding to the memory error information of that service node using the abnormal page prediction rule, which will not be repeated here. In this embodiment, both the memory error characteristics of the memory pages and the static resource information of the service node are considered. The abnormal page prediction rule adapted to the static resource information of the service node can be used to determine the memory pages that need to be isolated from the memory pages where errors occur at the service node, which helps to improve the accuracy of the determined memory pages that need to be isolated.
[0088] The foregoing embodiments exemplify how to determine the target memory page to be isolated from the memory pages that have malfunctioned when constructing an abnormal page prediction model using multiple abnormal page prediction rules, but this is not intended to be limiting. In particular, constructing an abnormal page prediction model using abnormal page prediction rules eliminates the need for model training compared to neural network models, thus saving sample training resources.
[0089] In other embodiments, the abnormal page prediction model may be a trained neural network model. This model predicts the probability of a memory page malfunctioning and triggering a UE (User Error), and determines the memory pages to be isolated based on this probability. In this embodiment, the abnormal page prediction model may be pre-trained. The training process of the abnormal page prediction model is illustrated below.
[0090] Specifically, historical memory error information of the sample device within a historical time period and the actual probability of a UE occurring on the sample device within a historical time period can be obtained. The sample device may include: the service node shown in Figure 1 and / or other computing devices.
[0091] Optionally, the sample devices can collect their own historical memory error information (including real-time memory CE information and UE information) and send it in batches to an offline database. UE information may include: UE address information and UE time information. UE address information refers to the memory address where the UE error occurred; see the relevant content regarding CE address information above. UE time information refers to the time when the UE error occurred.
[0092] The device used for model training can obtain historical memory error information of the sample device within a historical time period from an offline database; and determine the actual probability of a UE occurring on the sample device within the historical time period based on the UE information of the sample device. Specifically, if the sample device experiences a UE during the historical time period, the actual probability of the sample device experiencing a UE during the historical time period is 1; if the sample device does not experience a UE during the historical time period, the actual probability of the sample device experiencing a UE during the historical time period is 0.
[0093] Furthermore, the historical memory error characteristics of the faulty memory pages in the sample device can be determined based on the historical memory error information of the sample device. For a description of the historical memory error characteristics, please refer to the relevant memory information in the foregoing embodiments. For a detailed implementation of determining the historical memory error characteristics of the faulty memory pages in the sample device based on the historical memory error information of the sample device, please also refer to the specific content of step 302 above.
[0094] After identifying the historical memory error characteristics of memory pages that malfunctioned in the sample devices, the initial model corresponding to the abnormal page prediction model is trained using the historical memory error characteristics of memory pages that malfunctioned in the sample devices as the training objective, so as to obtain the abnormal page prediction model.
[0095] The loss function is determined by the error between the predicted probability of a UE occurring on a sample device during a historical time period, as output by the model training, and the actual probability of a UE occurring on the sample device during the same historical time period. This error between the predicted and actual probabilities can be represented by the mean squared error between the predicted and actual probabilities, or by the cross-entropy between them. The loss function expressed as the mean squared error between the predicted and actual probabilities can be represented as follows:
[0096] In equation (1) above, N represents the total number of sample devices. i represents the i-th sample device. P i f P represents the predicted probability of a UE occurring in the i-th sample device, as output by the model; i r This represents the actual probability of a UE occurring in the i-th sample device.
[0097] Based on the trained abnormal page prediction model, the memory error features of the memory pages obtained in step 302 can be input into the abnormal page prediction model. The abnormal page prediction model can determine the probability that these memory pages cause a UE based on their memory error features; and based on the probability that a memory page causes a UE, it can determine the target memory pages to be isolated from the memory pages corresponding to the memory error information. Specifically, the abnormal page prediction model can determine memory pages whose probability of causing a UE is greater than or equal to a set probability threshold from the memory pages corresponding to the memory error information, and use these as the target memory pages to be isolated.
[0098] The abnormal page prediction model is a neural network model, which can obtain a deeper correlation between memory error characteristics and UE. Therefore, using a neural network model to implement the abnormal page prediction model has high accuracy in identifying the target memory page to be isolated.
[0099] In other embodiments, since the static resource information of service nodes also affects memory performance, the static resource information of sample devices can be obtained during the training of the abnormal page prediction model. Using the loss function as the training objective, the abnormal page prediction model is trained using the static resource information and memory error characteristics of the sample devices to obtain the abnormal page prediction model. Based on the abnormal page prediction model trained in this embodiment, the static resource information of multiple service nodes can also be obtained. The static resource information of multiple service nodes, along with the memory error characteristics determined in step 302, are then input into the abnormal page prediction model. The abnormal page prediction model can determine the probability that these memory pages cause UE based on the memory error characteristics of the memory pages and the static resource information of multiple service nodes. Based on the probability that the memory pages cause UE, the model can determine the target memory pages to be isolated from the memory pages corresponding to the memory error information. Specifically, the abnormal page prediction model can determine memory pages whose probability of causing UE is greater than or equal to a set probability threshold from the memory pages corresponding to the memory error information, and use these as the target memory pages to be isolated. In this embodiment, the impact of memory page error characteristics and service node resource static information on memory failures is considered, which helps to improve the accuracy of identifying memory pages that need to be isolated.
[0100] For the central control node, after identifying the target memory page to be isolated, step 304 allows control to isolate the target memory page by the target service node. The target service node refers to the service node containing the target memory page among multiple service nodes. The target service node can respond to a request from the central control node and perform page isolation operations on the target memory page.
[0101] In this embodiment, the logic for determining the target memory page that needs to be isolated is implemented at the central control node, thus realizing a centralized page isolation strategy. Specifically, the central control node can determine the memory error characteristics of the memory page that has malfunctioned based on the memory error information of the service nodes; and using the abnormal page prediction model set in the central control node, it determines the target memory page to be isolated from the malfunctioning memory pages based on the memory error characteristics, and controls the target service node where the target memory page is located to perform page isolation on the target memory page, thus realizing a centralized page isolation strategy. This page isolation method allows for flexible adjustment of the memory page isolation strategy through software updates or modifications to the abnormal page prediction model on the central side, without requiring modifications to the kernel code on the service node side, which helps improve the system stability on the service node side.
[0102] Moreover, in large-scale cloud server scenarios, the memory page isolation strategy can be adjusted by updating the software or modifying the abnormal page prediction model on the central side. This allows for the adjustment of the memory page isolation strategy for the entire cloud service cluster without updating the kernel code of each service node in the cloud server cluster. This reduces the update cost of the memory page isolation strategy and prevents security and stability issues caused by large-scale updates to the kernel code of all service nodes in the cloud server cluster.
[0103] On the other hand, the embodiments of this disclosure utilize an abnormal page prediction model to identify the target memory page that truly needs to be isolated from the memory pages that have malfunctioned, and perform page isolation on the target memory page instead of page isolation on all memory pages that have malfunctioned. This can reduce the probability of insufficient memory capacity due to excessive isolation of memory pages, thereby causing service node failure.
[0104] In some embodiments of this disclosure, the memory address of the target memory page to be isolated may be a user-space address or a kernel-space address. In actual systems, user-space memory management is performed according to huge pages. A huge page is composed of multiple memory pages. Typically, for a 4KB memory page, a huge page is a 2MB huge page composed of 512 memory pages. User-space memory management according to huge pages reduces the number of page table entries for each process, which can reduce the Translation Lookaside Buffer (TLB) miss rate and reduce the number of page table traversals, thereby improving memory access efficiency.
[0105] Since user space manages memory in units of large pages, if one or more bits within a large page of a user-space memory address are corrupted, the entire large page will be isolated. Therefore, to reduce the number of I / O operations between the central control node and the target service node, when the memory address of the target memory page to be isolated is a user-space address, the central control node can determine the target large page to which the target memory page belongs and provide a page isolation request containing the memory address information of that target large page to the target service node. The target service node can respond to this page isolation request and isolate the target large page based on its memory address information, thereby achieving the isolation of the target memory page. This implementation requests page isolation from the target service node in units of large pages for the user-space target memory page to be isolated, reducing the number of page isolation requests sent to the target service node. Since frequent I / O can lead to system crashes, this implementation reduces the number of page isolation requests sent to the target service node, thus lowering the probability of system crashes at the service node.
[0106] Since the kernel manages memory on a page-by-page basis, when the target memory page to be isolated has a user-space address, the central control node can provide a page isolation request containing the target memory page's address information to the target service node. The target service node can then respond to this page isolation request and isolate the target memory page based on its address information.
[0107] Because the central control node employs different isolation strategies for kernel-mode and user-mode memory pages, it must first determine whether the target memory page's address is a user-mode or kernel-mode address before requesting the target service node to isolate it. Specifically, before requesting isolation, the central control node can also determine the process identifier (PID) of the target process using the target memory page based on its address. Then, the target process's status information can be obtained based on its identifier. Since each process's information is stored in a corresponding directory under the ` / proc` filesystem, various status information about the process can be obtained by reading the ` / proc / [pid] / status` file (the status information file corresponding to the PID). The ` / proc / [pid] / status` file records the status information of the process corresponding to the PID. The value of the `VmRSS` field can then be checked from this ` / proc / [pid] / status` file.
[0108] The VmRSS field represents the actual amount of physical memory currently occupied by the process, typically in kilobytes (KB). It reflects the amount of physical memory resources actually used by the process. Since kernel processes typically do not occupy physical memory, the VmRSS field value for kernel processes is 0. Therefore, if the VmRSS field value is 0, the target process is a kernel process. Correspondingly, the memory address of the target memory page is a kernel-mode address. If the VmRSS field value is not 0, the target process is a user-mode process. Correspondingly, the memory address of the target memory page is a user-mode address. Based on this, the memory address of the target memory page can be determined as either a user-mode address or a kernel-mode address based on the target process's state information.
[0109] After determining whether the memory address of the target memory page is a user-space address or a memory-space address, the corresponding isolation request method shown above can be used to request the target service node to isolate the target memory page.
[0110] In this embodiment, the central control node can directly send the page isolation request to the target service node, or, as shown in Figure 4, send the page isolation request to the operation and maintenance node of multiple service nodes, which then forwards the page isolation request to the target service node. In some service systems, service nodes do not communicate directly with other devices; instead, the operation and maintenance node is responsible for external communication. To reduce the cost of modifying existing service systems, the central control node can send the page isolation request to the operation and maintenance node of multiple service nodes, which then forwards the page isolation request to the target service node.
[0111] In some embodiments of this disclosure, before requesting the target service node to isolate the target memory page, certain filtering conditions can be set for the target memory page. These filtering conditions are used to further filter the target memory page to be isolated. If the target memory page does not meet the set filtering conditions, the operation of requesting the target service node to isolate the target memory page is executed. The implementation of the filtering conditions is illustrated below.
[0112] Filtering condition 1: If a memory page has already been isolated, performing page isolation on that page again will result in page isolation failure. Therefore, to improve the success rate of page isolation, filtering condition 1 can be implemented as follows: the target memory page has already been isolated. Accordingly, if the target memory page has already been isolated, there is no need to request the target service node to perform isolation on the target memory page again.
[0113] Filtering Condition 2: For the same memory page, if the memory page was already identified as a memory page to be isolated before it is currently identified as such, and the memory page is in use, forcibly isolating the memory page would cause the process or virtual instance using the memory page to fail. To reduce the interference of page isolation on other processes or virtual instances, the isolation operation for the memory page can be performed after it is released. Based on this, Filtering Condition 2 can be implemented as: the target memory page is marked as waiting to be released before being isolated. Specifically, if the target memory page is marked as waiting to be released before being isolated, it means that the target memory page has been marked as needing isolation and will be isolated after being released. Therefore, if the target memory page is marked as waiting to be released before being isolated, there is no need to request the target service node to perform isolation operations on the target memory page.
[0114] Filtering Condition 3: Since the memory of a service node is limited, if too many memory pages are isolated, the service node may experience server failure due to insufficient memory capacity. Furthermore, continuously reported memory errors cannot be isolated, severely impacting system stability. Therefore, flow control rules can be set to control the amount of memory that is isolated from service nodes. In some embodiments, this can be controlled by limiting the amount of memory that is isolated from a service node within a set time period. If the amount of memory isolated from a target service node within the set time period reaches the set memory capacity limit, then no further requests will be made to isolate memory pages from the target service node. Accordingly, filtering condition 3 can be implemented as: the amount of memory isolated from a target service node within the set time period reaches the set memory capacity limit.
[0115] Filtering Condition 4: Based on the analysis of Filtering Condition 3, in some other embodiments, the number of times a service node performs page isolation within a set time period can be controlled. If the target service node performs page isolation the number of times within the set time period reaches the set isolation limit, then the target service node will no longer be requested to isolate memory pages. Accordingly, Filtering Condition 4 can be implemented as: the target service node performs page isolation the number of times within the set time period reaches the set isolation limit.
[0116] Among them, filter conditions 3 and 4 can prevent service nodes from running out of memory due to excessive isolation of memory pages, which could lead to service node failure.
[0117] The aforementioned filtering conditions 1-4 are merely illustrative and do not constitute a limitation. In actual implementation, the set filtering conditions may include any one or more of the filtering conditions 1-4, or all of the filtering conditions 1-4. For embodiments where the set filtering conditions include multiple filtering conditions, if there are filtering conditions among the set filtering conditions that the target memory page satisfies, then the target service node is not requested to isolate the memory page. If the target memory page does not satisfy all filtering conditions, then the operation of requesting the target service node to isolate the memory page is performed.
[0118] When isolating a target memory page, the target service node needs to copy the data from the target memory page to a new memory page and modify the mapping between the logical and physical addresses of the memory page. During data copying, the target service device needs to lock the new memory page to prevent other processes from using it. This locking operation affects the other processing performance of the target service node. Therefore, the page isolation operation performed by the target service node has a certain impact on the performance of the service node. To reduce the impact of the page isolation operation performed by the target service node on the other processing performance of the target service node, the central control node can also determine the timing of the page isolation operation performed by the target service node.
[0119] Specifically, the idle time of the target service node can be determined based on its load information. The load information of the target service node may include CPU utilization and memory utilization, etc. In some embodiments, the load information of the target service node may be real-time load information. Accordingly, it can be determined whether the real-time load information of the target service node is less than or equal to a set load threshold; if the determination result is yes, the current time is determined as the idle time of the target service node. Since the real-time load information of the target service node can reflect its real-time performance, using the real-time load information of the target service node to determine its idle time is consistent with the actual load situation of the target service node, which helps to improve the accuracy of the determined idle time.
[0120] In other embodiments, the load on service nodes is generally considered to be periodic. For example, for a service node providing office services, office workers frequently access the service node during working hours, and access it less or not at all during rest periods. Since office workers' working hours are periodic, the load on the service node also has a certain periodicity. Therefore, the idle time of the service node can be inferred based on its historical load information.
[0121] Accordingly, the load information of the aforementioned target service node may include historical load information, which may include the historical load information of the target service node over multiple performance cycles. The performance cycle of the target service node is determined based on the actual load distribution of the target service node over time. Generally, the performance cycle is 1 day, and the performance cycle may cover 0:00 to 24:00 of a day.
[0122] Furthermore, the load distribution characteristics of the target service node over time can be determined based on historical load information. These characteristics reflect the relationship between the target service node's load and time. Based on these characteristics, the periods when the target service node's load is relatively low can be identified as its idle time. Correspondingly, the idle time of the target service node can be determined based on its load distribution characteristics. For example, the period with the lowest load can be selected from the time covered by these characteristics as the target service node's idle time. For instance, assuming the load distribution characteristics reflect the lowest load between 0:00 and 1:00 each day, this period can be identified as the target service node's idle time. In this embodiment, the idle time of the target service node is determined by analyzing its actual historical load information, rather than by assuming or estimating it. Therefore, this method can determine a more accurate idle time with higher accuracy and reliability.
[0123] After determining the idle time of the target service node, you can request the target service node to isolate the target memory pages during the corresponding idle time. Accordingly, the target service node can isolate the target memory pages during its idle time. Requesting the target service node to perform page isolation during its idle time avoids the target service node's peak load periods, which helps improve the success rate of page isolation.
[0124] In some embodiments, the central management node may provide a page isolation request carrying the memory address of the target memory page to the target service node when the idle time arrives. Accordingly, the target service node isolates the target memory page during the idle time based on its memory address. For example, the central management node may directly send the page isolation request carrying the memory address of the target memory page to the target service node when the idle time arrives. Upon receiving the page isolation request, the target service node isolates the target memory page during the idle time in response to the request, based on its memory address. Alternatively, the central management node may directly send the page isolation request carrying the memory address of the target memory page to the operations and maintenance node of multiple service nodes when the idle time arrives; the operations and maintenance node then forwards the page isolation request to the target service node. Upon receiving the page isolation request, the target service node isolates the target memory page during the idle time in response to the request, based on its memory address.
[0125] In other embodiments, the central control node can encapsulate the memory address and idle time of the target memory page into a page isolation request and send the page isolation request to the operation and maintenance nodes of multiple service nodes. When the idle time arrives, the operation and maintenance node forwards the page isolation request to the target service node. After receiving the page isolation request, the target service node responds to the request by isolating the target memory page according to its memory address during the idle time, thereby enabling the target service node to perform page isolation operations on the target memory page during its idle time.
[0126] Since the target service node's response to a page isolation request is an atomic operation—meaning it responds to the request and performs the corresponding page isolation operation as soon as it is received—sending a page isolation request carrying an idle time to the target service node, and having the target service node control the execution of page isolation operations on the target memory page when the idle time arrives, would require modification of the page isolation policy in the kernel. Therefore, the timing of the target service node's page isolation execution is controlled by the central management node or operations node, without requiring modification to the service node's kernel code, thus improving the stability of the service node side. The abnormal page prediction model determines that there can be one or more target memory pages to be isolated. "Multiple" refers to two or more (including two).
[0127] To reduce the probability of UE downtime at service nodes, target memory pages with a higher probability of causing UE downtime are isolated earlier. Based on this, the isolation priority of multiple target memory pages can be ranked. Specifically, the factors influencing the isolation urgency of multiple target memory pages can be obtained. These factors refer to those affecting the order of memory page isolation, and may include one or more of the following: real-time load information of the target service node where the target memory page resides, downtime risk information of the target service node, and the risk level of the target memory page. "Multiple factors" refers to two or more factors.
[0128] The real-time load information of the target service node reflects its resource utilization. A higher real-time load indicates greater resource consumption for other processing tasks on that node. Performing page isolation operations at this time could increase the performance burden on the target service node. Therefore, the higher the real-time load of the target service node, the lower the isolation priority of its corresponding target memory pages; that is, the higher the real-time load of a target service node, the later its corresponding target memory pages will be isolated.
[0129] The downtime risk information of the target service node where the target memory page resides refers to information reflecting the probability of the target service node experiencing downtime, which may include: the overall downtime risk of the target service node and / or the memory downtime risk of the target service node, etc. The downtime risk information of the target service node is obtained from the performance monitoring system of the target service node. How the performance monitoring system determines the downtime risk of the target service node is not the focus of this disclosure and will not be elaborated upon. Wherein, if the overall downtime risk of the target service node is high, a downtime may occur during the page isolation operation if page isolation is performed, leading to page isolation failure. Therefore, the isolation priority of target memory pages of target service nodes with low overall downtime risk is higher.
[0130] A high risk of memory failure on the target service node indicates a greater likelihood of failure due to memory errors. Therefore, early isolation of the target memory pages of this node is necessary to reduce the risk of failure caused by errors in these pages. Consequently, the higher the risk of memory failure on the target service node, the higher the priority for isolating its target memory pages.
[0131] The risk level of a target memory page reflects the likelihood that it will cause risk to the UE. The higher the risk level, the greater the likelihood of it causing risk to the UE. Therefore, the higher the risk level of a target memory page, the earlier it should be isolated; that is, the higher the risk level of a target memory page, the higher its isolation priority. For details on how to determine the risk level of a target memory page, please refer to the relevant content in the foregoing embodiments, which will not be repeated here.
[0132] Based on the factors influencing the isolation urgency of multiple target memory pages, their isolation priorities can be determined. Specifically, a pre-defined correspondence between the value range of each influencing factor and its impact on the isolation urgency of the target memory pages can be established. Then, based on the actual values of the influencing factors and this correspondence, the impact of each influencing factor on the isolation urgency of the multiple target memory pages can be determined. Next, the impact of each influencing factor on the isolation urgency of the multiple target memory pages can be weighted and summed to obtain an isolation urgency score. Finally, the multiple target memory pages can be sorted in descending order of their isolation urgency scores to determine their isolation priorities.
[0133] Furthermore, according to the isolation priority of multiple target memory pages, the target service nodes to which each target memory page belongs can be sequentially requested to isolate the corresponding target memory page. For a specific implementation method of requesting the target service node to which target memory page B belongs to isolate target memory page B from multiple target memory pages, please refer to the relevant content of the foregoing embodiments, and will not be repeated here.
[0134] In this embodiment, the isolation priority of multiple target memory pages is determined based on the factors affecting the urgency of isolation, and the target service nodes are requested to isolate the corresponding memory pages in order of priority. Memory pages with a higher risk of causing downtime can be isolated first, thereby reducing the probability of downtime due to memory page errors.
[0135] In some embodiments, before the target memory page is determined by the abnormal page prediction model to be a memory page requiring isolation, other memory pages C have already been determined as memory pages to be isolated, and these memory pages have not yet been isolated when the target memory page is determined. Therefore, it is necessary to sort the isolation priorities of the target memory page and the other memory pages C. Accordingly, the influencing factors of the isolation urgency of the target memory page and the influencing factors of the isolation urgency of the other memory pages C can be obtained. For a description of the influencing factors of the isolation urgency of memory pages, please refer to the relevant content in the foregoing embodiments.
[0136] Because memory page isolation is time-sensitive, if the time-sensitive period of a memory page to be isolated expires, that memory page will no longer be isolated. Based on this, when prioritizing the isolation of memory pages to be isolated, the target duration of other memory pages C that were previously awaiting isolation can be obtained. The target duration of memory page C's wait time for isolation can be counted from the time memory page C is determined to be an isolated memory page by the abnormal page isolation model. Then, the isolation priority between the target memory page and memory page C can be determined based on the factors influencing the isolation urgency of the target memory page and the factors influencing the isolation urgency of other memory pages C, as well as the target duration of memory page C's wait time for isolation.
[0137] Specifically, the initial isolation priority of the target memory page and memory page C can be determined based on the factors influencing the isolation urgency of the target memory page and the factors influencing the isolation urgency of other memory pages C. For a detailed implementation of determining the initial isolation priority of the target memory page and memory page C based on the factors influencing the isolation urgency of the target memory page and the factors influencing the isolation urgency of other memory pages C, please refer to the aforementioned content on determining the isolation priority of multiple target memory pages based on the factors influencing the isolation urgency of multiple target memory pages; it will not be repeated here.
[0138] Furthermore, if the difference between the target isolation time for memory page C and the set timeliness threshold is less than or equal to the set time difference threshold, then the isolation priority of memory page C is adjusted to the first position from the initial isolation priority list to obtain the isolation priority between the target memory page and memory page C. Further, according to the isolation priorities of the target memory page and memory page C, requests can be sequentially made to the target service node or the service node where memory page C resides to isolate the target memory page or memory page C. Specifically, the page isolation request for a memory page with a higher isolation priority is sent to the corresponding service node earlier. The page isolation request for the target memory page is sent to the target service node where the target memory page resides; the page isolation request for memory page C is sent to the service node where memory page C resides.
[0139] In this embodiment, when prioritizing the isolation of memory pages, not only the urgency of the isolation is considered, but also the timeliness of the isolation is taken into account. This can prevent memory page isolation from failing due to excessively long waiting times, thus helping to further improve the success rate of memory page isolation.
[0140] Because the operating environment and status of service nodes are constantly changing, the abnormal page prediction model provided in the aforementioned embodiments may become less accurate or even fail to predict UEs over time. Therefore, in some embodiments of this disclosure, the prediction accuracy of the abnormal page prediction model can be verified, and the model can be corrected or updated when its prediction accuracy is low.
[0141] Specifically, as shown in Figure 4, page isolation data for multiple service nodes can be obtained. This page isolation data includes the memory addresses of memory pages that need to be isolated within a set time period, as determined by the abnormal page prediction model. For ease of description, the memory pages that need to be isolated within a set time period, as determined by the abnormal page prediction model, are defined as the first memory page. The first memory page can be one or more. "Multiple" refers to two or more (including two). The page isolation data may also include the memory pages within the first memory page where UEs actually occurred within the set time period. The memory pages within the first memory page where UEs actually occurred within the set time period are defined as the second memory page.
[0142] To verify the prediction accuracy of the abnormal page prediction model, the total number M of the first memory pages determined by the abnormal page prediction model can be obtained; and the total number N of memory pages in the first memory pages that actually experienced UE occurrences within a set time period can be obtained. The set time period starts from the moment the abnormal page prediction model determines the first memory page, and is generally one performance cycle of the service node, such as 24 hours.
[0143] Furthermore, the prediction accuracy of the abnormal page prediction model can be determined based on the total number of first memory pages M and the total number of second memory pages N. The prediction accuracy of the abnormal page prediction model can be expressed as: N / M*100%.
[0144] If the prediction accuracy of the abnormal page prediction model is greater than or equal to the set accuracy threshold, then the prediction accuracy of the abnormal page prediction model is determined to meet the standard. If the prediction accuracy of the abnormal page prediction model is less than the set accuracy threshold, then the prediction accuracy of the abnormal page prediction model is determined to fail to meet the standard. As shown in Figure 4 "Offline Model Update", the abnormal page prediction model can be corrected based on the historical memory error information of the second memory page within a set time period.
[0145] For embodiments of an abnormal page prediction model that include multiple abnormal page prediction rules, the prediction accuracy of the abnormal page prediction model can be verified by verifying the prediction accuracy of multiple abnormal page prediction rules. Specifically, for any abnormal page prediction rule, the total number M of first memory pages that need to be isolated as determined by the abnormal page prediction rule within a set time period can be obtained; and the total number N of second memory pages in the first memory pages that have experienced UE errors within the set time period can be obtained; the prediction accuracy of the abnormal page prediction rule is determined based on the total number M of the first memory pages and the total number N of the second memory pages. The prediction accuracy of the abnormal page prediction model can be expressed as: N / M*100%. If the prediction accuracy of the abnormal page prediction rule is greater than or equal to the set accuracy threshold, then the prediction accuracy of the abnormal page prediction rule is deemed satisfactory. If the prediction accuracy of the abnormal page prediction rule is less than the set accuracy threshold, then the prediction accuracy of the abnormal page prediction rule is deemed unsatisfactory, and the abnormal page prediction rule can be corrected based on the historical memory error information corresponding to the second memory pages.
[0146] In an embodiment where the abnormal page prediction model is a neural network model, if the prediction accuracy of the abnormal page prediction model is not up to standard, the historical memory error information corresponding to the second memory page can be used to train the abnormal page prediction model, thereby correcting the abnormal page prediction model.
[0147] In this embodiment, the accuracy of the abnormal page prediction model is verified using historical data, and when the accuracy of the abnormal page prediction model is low, the abnormal page prediction model is corrected using historical data. This enables iterative updates of the abnormal page prediction model, which helps to ensure the prediction accuracy of the abnormal page prediction model.
[0148] To facilitate understanding of the memory isolation method provided in this disclosure, the memory isolation method will be described below with reference to the specific embodiment shown in FIG5. As shown in FIG5, the memory isolation method may include the following steps.
[0149] S1. Obtain memory error information from multiple service nodes.
[0150] S2. Based on the memory error information, determine the memory error characteristics of the memory page corresponding to the memory error information.
[0151] S3. Use the abnormal page prediction model to determine the target memory page to be isolated based on the characteristics of memory errors.
[0152] S4. Determine whether the memory address of the target memory page is a user-mode address. If the result is yes, proceed to step S5 and then continue to step S6; if the result is no, proceed to step S6.
[0153] S5. Determine the target page to which the target memory page belongs.
[0154] S6. Determine whether the target memory page is isolated and whether it is marked as waiting to be released after isolation. If the result of both determinations is no, proceed to step S7. If the result of both determinations is yes, proceed to step S11.
[0155] S7. Determine whether the target service node where the target memory page is located meets the set flow control rules. If the result is no, proceed to step S8; if the result is yes, proceed to step S11.
[0156] Flow control rules may include: the amount of memory isolated by the target service node within a set time period is greater than or equal to the set memory capacity limit; and / or, the number of times the target service node performs page isolation within a set time period is greater than or equal to the set isolation count limit.
[0157] S8. When the idle time of the target service node expires, or when the waiting time reaches the set duration, a page isolation request is sent to the target service node. This page isolation request is used to request the target service node to isolate the target memory page.
[0158] S9. Determine whether the page isolation request was sent successfully. If sent successfully, proceed to step S11. If sent unsuccessfully, proceed to step S10.
[0159] S10. Resend the page isolation request to the target service node. If the number of retries reaches the set maximum number of retries, proceed to step S11.
[0160] S11. End this page isolation operation.
[0161] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 301 and 302 can be device A; or the execution subject of step 301 can be device A, and the execution subject of step 302 can be device B; and so on.
[0162] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 301, 302, etc., are merely used to distinguish different operations and do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel.
[0163] Accordingly, this disclosure also provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps in the memory isolation methods provided in the foregoing embodiments.
[0164] Computer-readable storage media include volatile or non-volatile or a combination thereof, and may be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium.
[0165] This disclosure also includes a computer program product comprising a computer program that, when executed by one or more processors, causes the one or more processors to perform the steps in the memory isolation methods provided in the foregoing embodiments.
[0166] In this disclosure, the specific implementation form of the computer program product is not limited. In some embodiments, the computer program product may be implemented as an application (APP), a mini-program, a computer-side client, a program module, a plug-in, an installation package, a software development kit (SDK), an image file of an optical disc (such as an ISO file), a plug-in, or software in the form of Software as a Service (SaaS), etc., but is not limited thereto.
[0167] The computer program product should understand that each or a combination of the above-described method flow can be implemented by a computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above-described method embodiments.
[0168] Figure 6 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. As shown in Figure 6, the electronic device includes a memory 60a and a processor 60b. The memory 60a is used to store computer programs and can be configured to store various other data to support operation on a computing platform. Examples of this data include instructions for any application or method operating on the electronic device, data structures, contact data, phonebook data, messages, pictures, videos, etc.
[0169] Processor 60b is coupled to memory 60a and is used to execute computer programs to perform the steps in the memory isolation methods provided in the foregoing embodiments. Specific implementation details of each step can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0170] In some alternative embodiments, as shown in FIG6, the electronic device may further include optional components such as a communication component 60c, a power supply component 60d, a display component 60e, and an audio component 60f. FIG6 only schematically shows some components and does not mean that the electronic device must include all the components shown in FIG6, nor does it mean that the electronic device can only include the components shown in FIG6.
[0171] Furthermore, the components within the dashed boxes in Figure 6 are optional, not mandatory, and their specific requirements depend on the product form of the electronic device. The electronic device in this embodiment can be a desktop computer, laptop computer, mobile phone, or IoT device; it can also be a traditional server, cloud server, or server cluster, or other server equipment.
[0172] In embodiments of this disclosure, the memory is used to store computer programs and can be configured to store various other data to support operation on its host device. The processor can execute the computer programs stored in the memory to implement corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), electrically erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0173] In this embodiment of the disclosure, the processor can be any hardware processing device capable of executing the above-described method logic. Optionally, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), or a microcontroller unit (MCU); it can also be a programmable device such as a field-programmable gate array (FPGA), a programmable array logic (PAL), a general array logic (GAL), or a complex programmable logic device (CPLD); or it can be an advanced RISC machine (ARM) or a system on chip (SoC), etc., but is not limited thereto.
[0174] In this embodiment of the disclosure, the communication component is configured to facilitate wired or wireless communication between its host device and other devices. The device hosting the communication component can access wireless networks based on communication standards, such as 2G or 3G, 4G, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.
[0175] In embodiments of this disclosure, the display component may include a liquid crystal display (LCD) and a touch panel (TP). If the display component includes a touch panel, the display component may be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0176] In embodiments of this disclosure, a power supply component is configured to provide power to various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component resides.
[0177] In embodiments of this disclosure, the audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals. For example, in devices with voice interaction capabilities, voice interaction with a user can be achieved through the audio component.
[0178] It should be noted that the terms "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.
[0179] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the aforementioned element.
[0180] The above description is merely an embodiment of this disclosure and is not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.
Claims
1. A memory isolation method, applicable to a central control node, comprising: Obtain memory error information from the target service node; Based on the memory error information, determine the memory error characteristics of the memory page corresponding to the memory error information; Based on the memory error characteristics, the abnormal page prediction model is used to determine the target memory page to be isolated from the memory pages corresponding to the memory error information; Control the target service node to isolate the target memory page.
2. The method according to claim 1, wherein, For any memory page in the memory pages corresponding to the memory error information, determining the memory error characteristics of the memory corresponding to the memory error information based on the memory error information includes: From the memory error information of any memory page, determine the number of error bits of any memory page, and use it as the memory error feature of any memory page; And / or, From the memory error information of any memory page, obtain the burst location corresponding to the error bit; based on the burst location corresponding to the error bit, determine the burst error characteristic of any memory page, as the memory error characteristic of any memory page; and / or, From the memory error information of any memory page, obtain the identifier of the data queue line corresponding to the error bit; based on the identifier of the data queue line corresponding to the error bit, determine the error characteristics of the number queue line of any memory page, and use it as the memory error characteristics of any memory page; And / or, From the memory error information of any memory page, obtain the error time of the error bit; based on the error time of the error bit, determine the burst error characteristics of any memory page and / or the error characteristics of the number queue line of any memory page, as the memory error characteristics of any memory page; And / or, Obtain the memory location information of the error bit from the memory error information of any memory page; determine the address information of the memory unit where the error occurred in any memory page based on the memory location information of the error bit, and use it as the memory error feature.
3. The method according to claim 2, wherein, Determining the burst error characteristics of any memory page based on the burst position corresponding to the error bit includes: Based on the burst position corresponding to the error bit, determine the number of bursts in which any memory page has an error, and use this as the burst error characteristic of any memory page; And / or, Based on the burst position corresponding to the error bit, the number of error bits in the same burst is determined, which serves as the burst error characteristic of any memory page.
4. The method according to claim 2, wherein, The step of determining the error characteristics of the number queue line of any memory page based on the identifier of the data queue line corresponding to the error bit includes: Based on the identifier of the data queue line corresponding to the error bit, determine the number of data queue lines where any memory page has an error, and use this as the error characteristic of the data queue line. And / or, Based on the identifier of the data queue line corresponding to the error bit, the number of error bits in the same data queue line is determined as the error characteristic of the data queue line.
5. The method according to claim 2, wherein, The step of determining the burst error characteristics of any memory page and / or the error characteristics of the number queue line of any memory page based on the error time of the error bit includes: Based on the error time of the error bit, determine the burst error time of any memory page that causes an error, and use it as the burst error characteristic of any memory page; And / or, Based on the error time of the error bit, the error time of the data queue line where the error occurred in any memory page is determined, and this time is used as the error characteristic of the data queue line.
6. The method according to claim 1, wherein, The abnormal page prediction model includes multiple abnormal page prediction rules; the step of using the abnormal page prediction model to determine the target memory page to be isolated from the memory pages corresponding to the memory error information based on the memory error characteristics includes: For any memory page in the memory pages corresponding to the memory error information, if the memory error characteristics of any memory page satisfy some or all of the abnormal page prediction rules in the multiple abnormal page prediction rules, then the any memory page is determined to be the target memory page.
7. The method according to claim 6, wherein, If the memory error characteristics of any memory page satisfy some or all of the abnormal page prediction rules in the multiple abnormal page prediction rules, then determining any memory page as the target memory page includes: The memory error characteristics of any memory page include the number of error bits in any memory page. If the number of error bits in any memory page is greater than or equal to a set first number threshold, then any memory page is determined to be the target memory page. And / or, The error characteristics of any memory page include the number of fault bursts in the memory page. If the number of fault bursts in the memory page is greater than or equal to a set second threshold, then the memory page is determined to be the target memory page. Alternatively, the error characteristics of any memory page include the number of fault bursts in the memory page and the number of error bits in the same fault burst. If the number of fault bursts in the memory page is less than the set second threshold, and the number of error bits in the same fault burst is greater than or equal to the first threshold, then the memory page is determined to be the target memory page. And / or, The error characteristics of any memory page include the number of queue lines where errors occurred in any memory page. If the number of queue lines where errors occurred in any memory page is greater than or equal to a set third quantity threshold, then the memory page is determined to be the target memory page. Alternatively, the error characteristics of any memory page include the number of queue lines where errors occurred in any memory page and the number of error bits in the same queue line where errors occurred in any memory page. If the number of queue lines where errors occurred in any memory page is less than the set third quantity threshold, and the number of error bits in the same queue line is greater than or equal to the first quantity threshold, then the memory page is determined to be the target memory page. And / or, The error characteristics of any memory page include the error time of the burst in which the error occurs in the memory page, the same error time of the data queue line in which the error occurs in the memory page, the number of error bits in the burst in which the error occurs in the memory page, and the number of data queue lines in which the error occurs in the memory page. If the error time of the burst in which the error occurs in the memory page is the same as the error time of the data queue line in which the error occurs in the memory page, and the number of error bits in the burst in which the error occurs in the memory page is greater than or equal to the first quantity threshold, and the number of data queue lines in which the error occurs in the memory page is greater than or equal to the first quantity threshold, then the memory page is determined to be the target memory page. And / or, The error characteristics of any memory page include the address information of the memory cell in which the error occurred. If the address information of the memory cell in which the error occurred in any memory page is continuous, and the number of memory cells with continuous addresses is greater than or equal to a set fourth quantity threshold, then the memory page is determined to be the target memory page.
8. The method according to claim 1, wherein, The abnormal page prediction model includes multiple abnormal page prediction rules; the method further includes: Obtain the static resource information of the target service node; different static resource information corresponds to different abnormal page prediction rules; The step of using the abnormal page prediction model to determine the target memory page to be isolated from the memory pages corresponding to the memory error information based on the memory error characteristics includes: From the various abnormal page prediction rules, determine the target abnormal page prediction rule that matches the resource static information of the target service node; Based on the memory error characteristics of the target service node, the target memory page corresponding to the target service node is determined from the memory pages corresponding to the memory error information of the target service node using the target error page prediction rule.
9. The method according to any one of claims 1-8, further comprising: Obtain the first total number of first memory pages that need to be isolated, as determined by the abnormal page prediction model within a set time period; Obtain the second total number of second memory pages in the first memory page that actually experienced uncorrectable errors within a set time period; The prediction accuracy of the abnormal page prediction model is determined based on the first total number and the second total number. If the prediction accuracy is less than the set accuracy threshold, the abnormal page prediction model is corrected based on the historical memory error information of the second memory page within the set time period.
10. The method according to any one of claims 1-8, wherein, The step of controlling the target service node to isolate the target memory page includes: If the memory address of the target memory page is a user-mode address, determine the target large page to which the target memory page belongs; A page isolation request carrying the memory address information of the target large page is provided to the target service node to control the target service node to isolate the target large page in response to the page isolation request, thereby isolating the target memory page.
11. The method according to any one of claims 1-8, wherein, The step of controlling the target service node to isolate the target memory page includes: Based on the load information of the target service node, determine the idle time of the target service node; The target service node is requested to isolate the target memory page during the idle time.
12. The method according to claim 1, wherein, There are multiple target memory pages; the step of controlling the target service node to isolate the target memory pages includes: Factors affecting the isolation urgency of acquiring multiple target memory pages; Based on the aforementioned influencing factors, the isolation priority of the plurality of target memory pages is determined; According to the isolation priority, the target service nodes corresponding to the multiple target memory pages are requested to isolate the respective target memory pages in turn.
13. The method according to claim 1, wherein, Before the target memory page was determined, other memory pages to be isolated were identified. The other memory pages were not isolated when the target memory page was identified; the step of controlling the service node to isolate the target memory page includes: Obtain the factors influencing the isolation urgency between the target memory page and the other memory pages; Determine the target duration for the other memory pages to be isolated; Based on the influencing factors and the target duration, determine the isolation priority between the target memory page and the other memory pages; According to the isolation priority between the target memory page and the other memory pages, the target service node or the service node to which the other memory pages belong is requested to isolate the target memory page or the other memory pages in turn.
14. The method according to claim 12 or 13, wherein, The influencing factors include: the real-time load information of the target service node where the target memory page is located, the downtime risk information of the target service node where the target memory page is located, and one or more of the risk levels of multiple target memory pages.
15. The method according to any one of claims 1-8, wherein, Before requesting the target service node to isolate the target memory page, the method further includes: Determine whether the target memory page meets the set filtering conditions; the set filtering conditions include: the target memory page has been isolated; and / or, the target memory page is marked as waiting to be released and then isolated; and / or, the number of times the target service node performs page isolation within a set time period reaches the set isolation number limit; and / or, the memory capacity isolated by the target service node within a set time period reaches the set memory capacity limit. If the target memory page does not meet the set filtering conditions, the operation of requesting the target service node to isolate the target memory page is executed.
16. A memory isolation system, comprising: A central control node and multiple service nodes; the central control node is used to execute the steps in the method according to any one of claims 1-15; The target service node is the service node where the target memory page is located among the plurality of service nodes.
17. An electronic device comprising: A memory and a processor; wherein the memory is used to store computer programs; The processor is coupled to the memory for executing the computer program to perform the steps of the method according to any one of claims 1-15.
18. A computer-readable storage medium storing computer instructions, wherein, When the computer instructions are executed by one or more processors, the one or more processors are caused to perform the steps of the method according to any one of claims 1-15.
19. A computer program product comprising a computer program, wherein, When the computer program is executed by one or more processors, it causes the one or more processors to perform the steps of the method according to any one of claims 1-15.