Fault prediction method, fault processing method, and fault processing system
By extracting the basic storage unit distribution characteristics of cloud server memory error events and combining the subordinate relationship to predict instance granularity faults, the problem of large impact on operation and maintenance in the existing technology is solved, and the resource utilization and stability of the cloud service system is improved.
Patent Information
- Application Number
- PCT/IB2025/050633
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-04
- Filing Date
- 2025-01-22
- Publication Date
- 2025-08-07
AI Technical Summary
In the prior art, failure prediction through node granularity leads to a large impact range of operation and maintenance, which reduces the availability of cloud service systems, and lacks resource utilization and stability.
By obtaining the memory error event of the cloud server, the first distribution feature of the basic storage unit is extracted, and the second distribution feature of the instance is determined based on the affiliation between the basic storage unit and the instance, and the failure prediction model is used to perform the instance granular fault prediction.
The failure prediction with the granularity of instances is realized, the degree of refinement of fault prediction is improved, and the resource utilization, stability and availability of cloud service systems are improved.
Smart Images

Figure IB2025050633_07082025_PF_FP_ABST
Abstract
Description
Fault prediction method, fault handling method and fault handling system technical field
[0001] The present disclosure relates to the field of cloud computing technology, and more particularly to a fault prediction method, a fault handling method, and a fault handling system.
[0002] With the development of cloud computing technology, cloud service systems have become widely used due to their advantages such as rapid deployment, elastic scalability, high availability, and load balancing. The availability of cloud service systems is fundamental to ensuring uninterrupted user access to critical applications and data. A usable cloud service system must minimize downtime and mitigate the impact of unavailability. For example, with a cloud service system availability of at least 99.995%, users should not experience more than 130 seconds of unavailability per month. Memory failures in cloud servers are one of the main causes of reduced availability. When a memory failure occurs, the cloud server is highly likely to crash, rendering the services running on it unavailable.
[0003] Currently, memory error logs are analyzed to predict nodes where errors will occur, and then node-level maintenance is performed. However, due to the lack of precision in node-level failure prediction, subsequent maintenance operations can only be performed on the entire cloud server (the entire node). This means that maintenance operations must be performed on all instances on the cloud server, resulting in a wide range of maintenance impacts and insufficient availability of the cloud service system. Therefore, a more refined failure prediction method is urgently needed.
[0004] In light of this, embodiments of the present disclosure provide a fault prediction method. One or more embodiments of the present disclosure also relate to a fault handling method, a fault prediction apparatus, a fault handling apparatus, a fault handling system, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
[0005] In one embodiment of the present disclosure, a fault prediction method is provided, comprising: obtaining a memory error event from a cloud server, wherein the memory error event has corresponding event information; extracting, based on the event information, a first distribution feature of the memory error event in a basic storage unit of the cloud server memory; determining, based on the first distribution feature and a subordinate relationship between the basic storage unit and the instance, a second distribution feature of the memory error event in the instance; and determining, based on the second distribution feature, a predicted fault instance using a fault prediction model, wherein the fault prediction model is trained based on sample distribution features of sample memory fault events in the instance.
[0006] In one embodiment of the present disclosure, a memory error event of a cloud server is obtained, wherein the memory error event has corresponding event information. Based on the event information, a first distribution feature of the memory error event in the basic storage unit of the cloud server memory is extracted. Based on the first distribution feature and the subordinate relationship between the basic storage unit and the instance, a second distribution feature of the memory error event in the instance is determined. Based on the second distribution feature, a fault prediction model is used to determine a predicted fault instance. The fault prediction model is trained based on the sample distribution features of sample memory fault events in the instance. By extracting the first distribution feature of the memory error event in the basic storage unit of the cloud server memory, a fine-grained distribution feature of the basic storage unit is obtained. Based on this, and in combination with the subordinate relationship between the basic storage unit and the instance, a second distribution feature of the memory error event in the instance is determined, thereby implementing instance-granular feature analysis and determining the instance-granular distribution feature. Furthermore, based on the instance-granular distribution feature, a fault instance is predicted, achieving instance-granular fault prediction and improving the refinement of fault prediction. Description of the Figures
[0007] FIG1 is a flowchart of a fault prediction method provided by one embodiment of the present disclosure;
[0008] FIG2 is a flow chart of a fault prediction method provided by an embodiment of the present disclosure;
[0009] FIG3 is a schematic diagram of spatial distribution characteristics in a fault prediction method provided by an embodiment of the present disclosure;
[0010] FIG4 is a schematic diagram of spatial distribution characteristics in another fault prediction method provided by an embodiment of the present disclosure;
[0011] FIG5 is a schematic diagram of spatiotemporal distribution characteristics in a fault prediction method provided by an embodiment of the present disclosure;
[0012] FIG6 is a schematic diagram of distribution feature screening in a fault prediction method provided by an embodiment of the present disclosure;
[0013] FIG7 is a flowchart of a fault handling method provided by an embodiment of the present disclosure;
[0014] FIG8 is a schematic structural diagram of a fault prediction device provided by an embodiment of the present disclosure;
[0015] FIG9 is a schematic structural diagram of a fault handling device provided by an embodiment of the present disclosure;
[0016] FIG10 is a schematic structural diagram of a fault handling system provided by an embodiment of the present disclosure;
[0017] FIG11 is a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0018] The following description sets forth numerous specific details to facilitate a thorough understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art may make similar generalizations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.
[0019] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the one or more embodiments of the present disclosure. The singular forms "a," "an," "the," and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any and all possible combinations of one or more of the associated listed items.
[0020] It should be understood that while terms such as "first," "second," and so on may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, "first" could also be referred to as "second," and similarly, "second" could also be referred to as "first," without departing from the scope of one or more embodiments of the present disclosure. Depending on the context, the term "if" as used herein could be interpreted as "when," "when," or "in response to a determination."
[0021] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0022] In one or more embodiments of the present disclosure, a large model refers to a deep learning model with large-scale model parameters, typically including hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model or a basic model. (Foundation Model), which pre-trains large models using large unlabeled corpora to produce pre-trained models with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as large-scale language models (LLMs) and multi-modal pre-training models.
[0023] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as natural language processing (NLP) and computer vision. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image captioning (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0024] First, the terms involved in one or more embodiments of the present disclosure are explained.
[0025] Dynamic Random Access Memory (DRAM): A type of semiconductor memory that primarily uses the amount of charge stored in a capacitor to represent whether a binary bit is 1 or 0. Because transistors in reality leak current, the amount of charge stored on the capacitor is insufficient to accurately determine data, leading to data corruption. Therefore, periodic charging is an unavoidable requirement for DRAM. This requirement for periodic refresh earns it the name "dynamic" memory. Precisely because of this dynamic storage characteristic, where electrons are susceptible to external influences, memory errors are very common.
[0026] Static random-access memory (SRAM) is a type of random-access memory. The term "static" refers to the fact that the data stored in this type of memory remains permanently stored as long as the power is supplied. In contrast, the data stored in DRAM requires periodic updating.
[0027] Synchronous Dynamic Random Access Memory (SDRAM) is a type of DRAM with a synchronous interface. While DRAM typically has an asynchronous interface, allowing it to respond to changes in control inputs, SDRAM has a synchronous interface and waits for a clock signal before responding to control inputs, thus synchronizing with the computer's system bus. This clock drives a finite state machine that pipelines incoming instructions. This gives SDRAM a more complex operating mode than asynchronous DRAM (ADRAM), which lacks a synchronous interface.
[0028] Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM): This is a type of SDRAM with a double data rate, transferring data at twice the system clock frequency. This increased speed improves performance compared to traditional SDRAM. Fourth-generation DDR4 (DDR4) is a high-speed, low-power computer memory technology designed to improve system performance and storage capacity.
[0029] Memory failure: refers to failures within the memory hardware system, including failures in hardware components such as memory controllers and DRAM chips.
[0030] Memory Error: If the data bit read from the memory is different from the last written bit, a memory error is reported. If a memory fault is accessed multiple times, multiple memory errors are reported.
[0031] Error Correction Code (ECC): After the memory access is completed, the read content will be The device performs verification and error correction. The errors detected are divided into correctable errors and uncorrectable errors. Uncorrectable errors usually cause the system to crash.
[0032] Correctable Error (CE): An error that can be corrected by ECC is called a CE. The impact of a small amount of CE on system stability can usually be ignored.
[0033] Uncorrectable Error (UE): A memory error that exceeds the ECC error correction capability is called a UE. A single UE can cause the system to crash.
[0034] Error Detected and Corrected log (EDAC log or EDAC): Memory error events are recorded in multiple logs. Compared with other logs, EDAC log can directly analyze ECC-related information.
[0035] Operating system kernel log: The dmesg command in Linux and other Unix-like operating systems displays messages in the kernel buffer, including hardware initialization information and error reports. When a memory module problem occurs, even if the error is not directly reported by the EDAC subsystem, it may leave traces in the kernel log.
[0036] System Event Log (SEL): In some server hardware platforms, the SEL records hardware-level events, including memory-related errors, such as failures detected by the memory controller or memory hot-plug events.
[0037] Instance hot migration: An instance operation that allows a virtual machine to be moved from one physical host to another without shutting it down for maintenance, load balancing, or failure prevention.
[0038] Instance cold migration: An instance operation that requires shutting down a virtual machine to move it from one physical host to another.
[0039] Instance isolation: An instance maintenance operation, in a multi-tenant environment or distributed system, involves temporarily isolating a problem instance from the overall service environment to prevent it from affecting other instances. This operation can be logical, such as stopping new requests or disconnecting from other instances. It can also be physical, such as removing the instance from the cluster or assigning it to a separate resource pool for troubleshooting and repair.
[0040] Instance rollback: An instance maintenance operation that resolves an instance error or performance issue by restoring the instance to a previously known good state. This typically involves restoring the instance's data contents to a previously created backup point or snapshot, including the operating system configuration, application data, and possibly the application itself. Cloud service providers often offer scheduled automated snapshots, allowing users to quickly restore instances to a past state after an outage.
[0041] Instance restart: An instance maintenance operation involves shutting down and then restarting a virtual machine or container instance. This is a common troubleshooting technique used to resolve issues such as software errors, resource conflicts, and memory leaks caused by prolonged operation. By restarting an instance, the operating system and applications can clean up their internal state and reinitialize, potentially resolving temporary or non-persistent faults. For many services, instance restart is often the first step in quickly restoring service.
[0042] Cloud servers are virtualized servers that provide computing resources over the internet, based on cloud computing technology. Users can dynamically access and configure computing power, storage space, and network resources based on their needs, eliminating the need to own physical server hardware and paying only for what they use. Cloud servers can be quickly deployed, expanded, and configured, and support features such as high availability, load balancing, and data backup and recovery.
[0043] Node: A physical server or an independent computing unit within a computing resource pool. In distributed systems and cluster architectures, each node possesses processing power and storage space and can collaborate with other nodes. In a cloud computing service provider's data center, a node can be a physical server deployed with virtualization technology, responsible for hosting multiple virtual machines or instances. In this disclosure, the terms node, physical host, and server are different descriptions of the same thing. Node refers to the network structure and cluster system, physical host refers to the hardware device, and server refers to the functionality of the cloud computing service.
[0044] Instance: In cloud services, an instance specifically refers to an operating environment or computing unit created using virtualization technology and provided by a cloud service provider. When a user purchases and activates a cloud server, they are effectively acquiring a virtual machine instance resource allocated by the cloud service provider on a physical node. Each instance has its own operating system, central processing unit (CPU), memory, disk space, and other configurations, and can independently run applications and provide services. Instance configurations can be adjusted at any time based on demand, and can be elastically scaled based on load.
[0045] Virtual Machine (VM): A virtual machine is a complete computer system emulated on a physical server using virtualization software. In cloud server scenarios, a virtual machine is an abstract computing resource with its own independent operating system, processor, memory, hard disk, and other resources. From the operating system's perspective, a virtual machine is no different from a real physical server. A virtual machine instance is the specific manifestation of this virtualized environment in the cloud. Users can create and manage multiple virtual machine instances on the cloud platform to meet different needs. In this disclosure, the terms instance and virtual machine are different descriptions of the same thing: instance refers to the logical computing perspective, while virtual machine refers to the system resource perspective.
[0046] Dual In-line Memory Module (DIMM): A memory component package widely used in personal computers, servers, and other electronic devices. The biggest difference between DIMM memory modules and single in-line memory modules (SIMM) is that the data bus width is doubled and independent data input and output are provided. Memory supports higher data rates and larger capacities through the use of dedicated input and output lines. Each DIMM typically contains multiple DRAM chips and is controlled and managed by one or more memory controllers. The DIMM interface design allows for simultaneous read and write operations in both directions, resulting in higher bandwidth and more efficient memory access. Modern DIMM memory has different performance indicators and physical specifications based on different technical standards (such as DDR2, DDR3, and DDR4). All are compatible with corresponding motherboard slots, providing flexible memory expansion options for the system.
[0047] Memory cell or cell: A cell is the basic storage unit in DRAM (dynamic random access memory). It consists of a capacitor and a transistor. The capacitor stores binary information (0 or 1), while the transistor acts as a switch to control the reading and writing of data in the capacitor. During each refresh cycle, the memory cell must be recharged to maintain its state.
[0048] A bank is a logical concept in memory, corresponding to a group of physically contiguous memory chips or an independent address space within the same memory chip. The memory controller can independently read and write data from each bank, enabling parallel operations across multiple banks, thereby improving memory bus utilization and overall system performance. A memory bank may contain millions to billions of memory cells and has its own independent address space and control logic.
[0049] Rank: In memory, a group of DRAM chips organized in a specific manner on a DIMM (Digital Integrator Unit) share the same command and address buses and transmit data simultaneously as a single unit. A DIMM can have one or more ranks, each acting as an independent data channel capable of sending or receiving data simultaneously.
[0050] Burst: This refers to an efficient memory access mode. When the CPU initiates a memory access request, the memory controller continuously reads or writes a series of adjacent data bytes instead of transferring just one byte. For example, in DDR SDRAM, once a burst is triggered, multiple packets are transmitted continuously according to the configured burst length, reducing latency and improving bandwidth efficiency.
[0051] Data I / O (DQ): In a DRAM interface, DQ lines are used for bidirectional data transmission, i.e., input and output data signals. Each DQ line represents one bit of width. Multiple DQ lines are combined to form a data bus, which is used to transmit multi-bit data at high speed between the memory chip and the memory controller.
[0052] Windowed Aggregation: A feature processing method that performs a series of analytical calculations on the entire dataset while preserving the original row structure. Window functions not only group data but also compute specific values for each row within a group, based on the current row and other rows within the window. Specifically, in a windowing operation, the user first defines a window, which can be a range of all rows or multiple contiguous intervals defined by the values of one or more columns. Aggregate functions (such as SUM, AVG, COUNT, MIN, and MAX) are then applied within the specified window. This results in each row receiving additional information calculated based on the data within the window, while the original record remains intact. This means that the data details of the original row are not lost, as with standard grouping.
[0053] XGBoost (eXtreme Gradient Boosting): A machine learning algorithm that is an optimized implementation of Gradient Boosting Decision Trees (GBDT).
[0054] Multilayer Perceptron (MLP): A type of feedforward neural network consisting of an input layer, one or more hidden layers, and an output layer. Each layer contains multiple neurons, and adjacent layers are connected via weighted connections. An activation function is used to perform nonlinear transformations. MLPs are widely used for classification and regression problems.
[0055] Convolutional Neural Networks (CNN) model: A multi-layer deep learning model with forward propagation and backpropagation, and a convolution kernel (filter) for processing feature data.
[0056] Recurrent Neural Network (RNN) model: A recursive deep learning model that performs recursion in the processing direction of vector representation and connects each intermediate layer in a chain.
[0057] Long Short Term Memory (LSTM) model: A deep learning model that has the ability to memorize long-term and short-term information and has a convolutional filter for processing feature data.
[0058] Deep Self-Attention Model (Transformer Model): A deep learning architecture based on the attention mechanism for processing sequential data such as natural language.
[0059] Bidirectional Encoder Representations from Transformers (BERT): A specialized Transformer model trained using a bidirectional Transformer encoder and large-scale unlabeled text data.
[0060] Feature pattern (or simply pattern): A specific type of mathematical transformation of a measurement value in machine learning. For example, in image processing, a feature pattern typically includes information such as grayscale, texture, and shape. Another example is remote sensing imagery, where a feature pattern often refers to a sensor channel or the reflectivity measurement of a feature image, or its mathematical transformation.
[0061] Currently, by analyzing memory error logs, we can predict in advance the node where the error will occur and migrate all instances on that node to other nodes in advance, thereby reducing the occurrence of unavailability.
[0062] However, this approach will cause all instance services on the node to be unavailable during the migration process, and other resources must be consumed to complete the unprofitable instance migration action, which reduces the resource utilization of the cloud service. Secondly, during the migration of all instances on the node, the cloud service Services may experience jitter and damage, reducing the stability of cloud servers. Finally, migrating all instances on a node adds significant time to the entire operation and maintenance cycle and reduces the availability of cloud servers. Therefore, a fault prediction method with high resource utilization, high stability, and high availability is urgently needed.
[0063] To address the above-mentioned issues, the present disclosure provides a fault prediction method. The present disclosure also relates to a fault handling method, a fault prediction apparatus, a fault handling apparatus, a fault handling system, a computing device, a computer-readable storage medium, and a computer program product, each of which is described in detail in the following embodiments.
[0064] Referring to Figure 1, Figure 1 shows a flowchart of a fault prediction method provided by an embodiment of the present disclosure, comprising the following specific steps 102 to 108.
[0065] Step 102: Obtain a memory error event of the cloud server, wherein the memory error event has corresponding event information.
[0066] The embodiments of the present disclosure can be applied to a management node in a cloud service system. If a fault occurs on the management node, the embodiments can also be applied to a data analysis system independent of the cloud service system, which is not limited here.
[0067] A cloud server is a virtualized server within a cloud service system that uses cloud computing technology to provide computing resources over the internet. A cloud service system consists of multiple cloud servers forming a cloud server cluster and a management node that manages the cloud server cluster. A cloud server is a node in the cloud server cluster. Each cloud server (node or physical host) includes multiple instances (virtual machines). Cloud servers run these instances (virtual machines) on physical hosts and provide cloud computing services to external users. Users can manage and control these instances (virtual machines) through remote desktop or application programming interfaces (APIs) to ensure the normal operation of cloud computing services, such as data backup and recovery, software installation, and system upgrades. For example, a cloud service provider provides a cloud service system with elastic computing capabilities. Users can select the corresponding specifications (operating system, CPU specifications, memory specifications, disk specifications, etc.) on the cloud service system, create an instance (virtual machine) on the cloud server, and implement cloud computing services such as website deployment, application execution, and data management.
[0068] Memory error events are correctable memory errors. A small number of memory error events have a minimal impact on the stability of the cloud server, and correction can ensure normal operation of the cloud server. These events include, but are not limited to: Single Bit Error > Parity Error > Checksum Error > Memory Scrubbing Events. A single bit error is a data bit flip in a single memory cell (a 1 is erroneously stored as a 0, or a 0 is erroneously stored as a 1). A parity error occurs in memory systems that use parity checking technology, where the parity of the data bits does not match the expected value, indicating the presence of one or more bit errors. A checksum error occurs when the checksum calculated for a specific data packet or data structure does not match the expected value, indicating possible bit errors during data transmission or storage. Memory Scrubbing Events are errors proactively detected and corrected during memory maintenance operations, such as scrubbing. Memory error events are typically recorded in the cloud server's memory error log. This log file records memory error events that occur within the cloud server. Specifically, the memory error log contains event information related to memory error events, including but not limited to error detection and correction logs, operating system kernel logs, and system event logs. For ease of understanding, the present disclosure uses the error detection and correction log as an example. The event information for a memory error event is data that records and describes the specific circumstances of the error, including but not limited to: event type (correctable or uncorrectable error), time information, and the like. (timestamp), memory controller information (memory controller number), memory information (memory number, memory channel, physical address, memory module information), and memory unit information (memory unit information for basic storage units: cell, cell row, cell column, bank; memory read / write unit information: error check code). Taking the error detection and correction log as an example, the event information recorded for a memory error event is as follows: [2023-03-24 13: 57: 12] EDAC_MC: MCI: CE row=456, channel=l, syndrome=0x0A, address=0x89ABCDEF [2023-03-24 13: 57: 13] kernel: EDAC sbridge: UE on CPU socket 0, MCA.bank=0, chip=DDR3_DIMM0, row=456, column=12, bit=9
[0069] Among them, "[2023-03-24 13:57:12]" and "[2023-03-24 13:57:13" are timestamps, "EDAC_MC" is the message identifier of the system's memory error detection and correction module, "MCI" is the memory controller number, "CE" indicates the event type is a correctable error, "MCA.bank=0" is the bank, "chip=DDR3_DIMM0" indicates the DDR3 memory is inserted in slot 0, "row=456" is the cell row, indicating that the error occurred in cell row 456, "column=12" is the cell column, indicating that the error occurred in cell column 12, "channel=1" is the memory channel, indicating that the error occurred in channel 1. "syndrome=0x0A" is the error check code, and "address=0x89ABCDEF" is the physical address, indicating that the error occurred at physical address 0x89ABCDEF.
[0070] Obtaining a memory error event from a cloud server is done by obtaining a memory error log from the cloud server's log library. The log library collects event information about memory error events in real time, parses it, and generates a memory error log for the cloud server.
[0071] For example, a cloud server system includes a cloud server cluster consisting of 10 cloud servers (nodes or physical hosts), each of which runs five instances. The cloud server's logstore collects memory error event information from the 50 instances in real time, parses it, and generates the cloud server's EDAC log. EDAC logs are obtained from the logstore. EDAC records event information for multiple memory error events within 14 days, including: the timestamp of the memory error event, the memory controller number of the memory error event, the bank where the memory error event occurred, the memory slot in the memory error event, the cell row where the memory error event occurred, the cell column where the memory error event occurred, the error check code of the memory error event, and the physical address of the memory in the memory error event. address.
[0072] Obtain the cloud server's memory error log, which records event information about memory errors. This provides the information foundation for subsequent feature extraction.
[0073] Step 104: Based on the event information, extract a first distribution feature of the basic storage unit of the memory of the cloud server.
[0074] Memory is used to temporarily store data required and generated by cloud computing services. Memory has corresponding physical addresses. During instance creation, a virtual address for the instance is generated, and a mapping relationship between the virtual address and the physical address is established. This allows users to access physical addresses through virtual addresses even when the physical address is unknown in user mode. Memory is typically DRAM. In this disclosed embodiment, fourth-generation double data rate synchronous dynamic random access memory (DDR4) is used as an example. For example, the physical host corresponding to a cloud server includes 128GB of memory. When creating eight instances, a user requests 16GB of memory for each instance. Eight instances are running on the cloud server, and EDAC logs for the 128GB of memory are recorded on the cloud server.
[0075] A basic storage unit of a memory is a data storage unit that constitutes the memory. A memory includes multiple basic storage units, which include but are not limited to: a memory cell, a bank, and a rank. In the embodiments of the present disclosure, a memory cell is used as an example for description.
[0076] The first distribution feature is the distribution feature of memory error events occurring on basic storage units in the memory error log. It includes at least one dimension, such as the time dimension, the spatial dimension, and the spatiotemporal dimension. Correspondingly, the first distribution feature includes at least one of the first time distribution feature, the first spatial distribution feature, and the first spatiotemporal distribution feature. For example, the first time distribution feature indicates at which time a basic storage unit is more likely to experience a memory error event, the first spatial distribution feature indicates which basic storage units are more likely to experience a memory error event, and the first spatiotemporal distribution feature indicates which basic storage units are more likely to experience a memory error event at which time. The first distribution feature is the distribution feature of memory error events at the basic storage unit granularity. The granularity of the feature representation varies depending on the basic storage unit selected. Similarly, by combining and aggregating the distribution features at each granularity, a larger granularity distribution feature can be obtained. By extracting the first distribution feature, feature support at the basic storage unit granularity can be provided for subsequent fault prediction.
[0077] For example, the first spatial distribution characteristics include, but are not limited to, cell errors, cell row errors, cell column errors, cell soft errors, cell soft errors, cell row soft errors, cell column soft errors, cell hard errors, cell row hard errors, cell column hard errors, repeated cell errors, repeated cell row errors, repeated cell column errors, cell correctable errors, repeated cell row correctable errors, repeated cell column correctable errors, cell multiple bit errors, number of error checking and correction sources, number of errored data lines, maximum number of errored cells by row, minimum number of errored cells by row, total number of single data bits, total number of two data bits, and total number of multiple data bits.
[0078] For example, the first time distribution feature includes, but is not limited to, cell errors within a time window, storage cell row errors within a time window, storage cell column errors within a time window, cell soft errors within a time window, cell soft errors within a time window, cell row soft errors within a time window, cell column soft errors within a time window, cell hard errors within a time window, cell row hard errors within a time window, cell column hard errors within a time window, repeated cell errors within a time window, repeated cell row errors within a time window, repeated cell column errors within a time window, cell correctable errors within a time window, repeated storage cell row correctable errors within a time window, repeated storage cell column correctable errors within a time window, cell multiple error bit errors within a time window, the number of error checking and correction sources within a time window, the number of error storage data lines within a time window, the maximum number of error cells by row within a time window, the minimum number of error cells by row within a time window, a total of single data bits within a time window, a total of two data bits within a time window, and a total of multiple data bits within a time window.
[0079] For example, the first spatiotemporal distribution feature includes, but is not limited to: the number of cell errors in the time window, the number of storage cell row errors in the time window, the number of storage cell column errors in the time window, the number of cell soft errors in the time window, the number of cell soft errors in the time window, the number of cell row soft errors in the time window, the number of cell column soft errors in the time window, the number of cell hard errors in the time window, the number of cell row hard errors in the time window, the number of cell column hard errors in the time window, the number of repeated cell errors in the time window, the number of repeated cell row errors in the time window, the number of repeated cell column errors in the time window, the number of correctable cell errors in the time window, the number of repeated storage cell row correctable errors in the time window, the number of repeated storage cell column correctable errors in the time window, the number of cell multiple error bit errors in the time window, the number of error check and correction sources in the time window, the number of error storage data lines in the time window, the maximum number of error cells by row in the time window, the minimum number of error cells by row in the time window, the total number of single data bits in the time window, The total number of two data bits in the time window and the total number of multiple data bits in the time window.
[0080] Based on the event information, first distribution features of the basic storage units of the cloud server memory for the memory error event are extracted. Specifically, the first distribution features of the basic storage units of the cloud server memory for the memory error event are extracted based on the event information in at least one dimension. Specifically, if there are multiple dimensions, the first distribution features of the basic storage units for the memory error event can be extracted based on the event information in each dimension. Optionally, at least one first distribution feature is determined from the first distribution features in the multiple dimensions. For example, the first time distribution feature is determined as the first distribution feature from the first time distribution feature and the first spatial distribution feature. Optionally, feature fusion is performed on the first distribution features in the multiple dimensions, for example, feature aggregation is performed on the first time distribution feature and the first spatial distribution feature, to obtain the first distribution feature. This is not limited herein.
[0081] For example, based on the unit information of the memory cell (cell, cell row, cell column and error check code), the first spatial distribution feature Cell_Spat_Dist of the memory cell of the memory error event is extracted, and based on the time information (timestamp), the memory error event is extracted. The first time distribution feature Cell_Temp_Dist of the error event memory cell is aggregated to obtain the first spatiotemporal distribution feature Cell_TS_Dist, and the first spatiotemporal distribution feature is determined as the first distribution feature.
[0082] Based on the event information, we extract the first distribution features of memory error events in the basic storage units of the cloud server memory. This fine-grained distribution feature of the basic storage unit lays the foundation for the subsequent determination of instance-level distribution features.
[0083] Step 106: Determine a second distribution feature of the memory error event in the instance based on the first distribution feature and the subordinate relationship between the basic storage unit and the instance.
[0084] Because basic storage units are often not individually operable, for example, data migration is not possible for a single bit of data stored in a memory cell. An instance is the smallest operable granularity. Therefore, after extracting the first distribution feature at the basic storage unit granularity, it is necessary to determine the second distribution feature at the instance granularity for fault prediction.
[0085] The dependency relationship between a basic storage unit and an instance is a memory address dependency relationship between the basic storage unit and the instance, typically expressed as a mapping relationship between the physical address of the basic storage unit and the virtual address of the instance. For example, the basic storage unit is a cell, and the physical address of the cell (cell row, cell column) is 0x89ABCDEF. The mapping relationship between the physical address 0x89ABCDEF and the virtual address of the instance (typically represented by a mapping table, also called a page table) indicates that its page frame starting address in physical memory is 0x89AB0000. The mapping table determines the virtual page number V0x1234 corresponding to the physical page frame number 0x89AB. Adding the page offset 0xDEF yields the virtual address V 0x1234DEF, which then determines the instance to which this virtual address is pre-assigned: "Instance." The dependency relationship between the basic storage unit and the instance is recorded in the kernel interface, and step 106 is implemented in kernel mode.
[0086] The second distribution feature is the distribution feature of memory error events occurring across instances in the memory error log. It includes at least one dimension, such as the time dimension, the spatial dimension, and the spatiotemporal dimension. Accordingly, the second distribution feature includes at least one of the second time distribution feature, the second spatial distribution feature, and the second spatiotemporal distribution feature. For example, the second time distribution feature indicates at which time instances are more likely to experience memory error events, the second spatial distribution feature indicates which instances are more likely to experience memory error events, and the second spatiotemporal distribution feature indicates which instances are more likely to experience memory error events at which time. The second distribution feature is the distribution feature of memory error events at the instance granularity. By extracting the second distribution feature, feature support at the instance granularity can be provided for subsequent fault prediction.
[0087] For example, the second spatial distribution feature includes, but is not limited to, a cell error in the instance, a storage cell row error in the instance, a storage cell column error in the instance, a cell soft error in the instance, a cell soft error in the instance, a cell row soft error in the instance, a cell column soft error in the instance, a cell hard error in the instance, a cell row hard error in the instance, a cell column hard error in the instance, a repeated cell error in the instance, a repeated cell row error in the instance, a repeated cell column error in the instance, a cell correctable error in the instance, a repeated storage cell row correctable error in the instance, a repeated storage cell column correctable error in the instance, a cell multiple error bit error in the instance, the number of error checking and correction sources in the instance, the number of error storage data lines in the instance, the maximum number of error cells by row in the instance, the minimum number of error cells by row in the instance, a total of single data bits in the instance, a total of two data bits in the instance, and a total of multiple data bits in the instance.
[0088] For example, the second time distribution feature includes, but is not limited to: cell errors in the time window on the instance, storage unit row errors in the time window on the instance, storage unit column errors in the time window on the instance, cell soft errors in the time window on the instance, cell soft errors in the time window on the instance, cell row soft errors in the time window on the instance, cell column soft errors in the time window on the instance, cell hard errors in the time window on the instance, cell row hard errors in the time window on the instance, cell column hard errors in the time window on the instance, repeated cell errors in the time window on the instance, repeated cell row errors in the time window on the instance, repeated cell column errors in the time window on the instance, correctable cell errors in the time window on the instance, correctable repeated storage unit row errors in the time window on the instance, correctable repeated storage unit column errors in the time window on the instance, multiple bit errors in the cell in the time window on the instance, number of error checking and correction sources in the time window on the instance, number of error storage data lines in the time window on the instance, maximum number of error cells by row in the time window on the instance, Minimum number of error cells by row in the time window on the instance, total single data bits in the time window on the instance, total two data bits in the time window on the instance, and total multiple data bits in the time window.
[0089] For example, the second spatiotemporal distribution feature includes, but is not limited to: the number of cell errors in the time window of the instance, the number of storage cell row errors in the time window of the instance, the number of storage cell column errors in the time window of the instance, the number of cell soft errors in the time window of the instance, the number of cell soft errors in the time window of the instance, the number of cell row soft errors in the time window of the instance, the number of cell column soft errors in the time window of the instance, the number of cell hard errors in the time window of the instance, the number of cell row hard errors in the time window of the instance, the number of cell column hard errors in the time window of the instance, the number of repeated cell errors in the time window of the instance, the number of repeated cell row errors in the time window of the instance, the number of repeated cell column errors in the time window of the instance, the number of correctable cell errors in the time window of the instance, the number of correctable repeated storage cell row errors in the time window of the instance, the number of correctable repeated storage cell column errors in the time window of the instance, the number of multiple bit errors of the cell in the time window of the instance, the number of error checking and correction sources in the time window of the instance, the number of error storage data lines in the time window of the instance, Maximum number of error cells per row in the time window on the instance, Minimum number of error cells per row in the time window on the instance, Total number of single data bits in the time window on the instance, Total number of two data bits in the time window on the instance, and Total number of multiple data bits in the time window on the instance.
[0090] Based on the second distribution feature and the subordinate relationship between the basic storage unit and the instance, the second distribution feature of the memory error event in the instance is determined. Specifically, in at least one dimension, based on the event information and the subordinate relationship between the basic storage unit and the instance, the second distribution feature of the memory error event in the instance of the cloud server memory is determined. Specifically, in the case of multiple dimensions, Based on the event information and the subordinate relationship between the basic storage unit and the instance, the second distribution feature of the memory error event in each dimension of the instance of the cloud server memory can be determined. Optionally, at least one second distribution feature is determined from the second distribution features in multiple dimensions. For example, the second time distribution feature is determined as the second distribution feature from the second time distribution feature and the second space distribution feature. Optionally, feature fusion is performed on the second distribution features in multiple dimensions, for example, feature aggregation is performed on the second time distribution feature and the second space distribution feature to obtain the second distribution feature. This is not limited here.
[0091] It should be noted that, in the case of multiple dimensions, in step 104 and step 106, feature fusion may be performed first in step 104 and then in step 106, or feature fusion may not be performed in step 104 and step 106 may be performed directly, and then in step 106, both of which fall within the scope of protection of the embodiments of the present disclosure.
[0092] For example, based on the first distribution feature Cell_TS_Dist and the mapping table between the physical address of the memory cell and the virtual address of the instance, the 150 second distribution features X_(i, i~[1,150]) = Instance_TS_Dist,
[0093] Based on the first distribution characteristic and the subordinate relationship between the basic storage unit and the instance, the second distribution characteristic of the memory error event in the instance is determined. This implements instance-granular feature analysis and determines the instance-granular distribution characteristics, laying the foundation for subsequent instance-granular fault prediction.
[0094] Step 108: Based on the second distribution feature, use a fault prediction model to determine a predicted fault instance, wherein the fault prediction model is trained based on sample distribution features of sample memory fault events on the instance.
[0095] In the embodiments of the present disclosure, memory error events and memory failure events are of different severity. A small number of memory error events will not affect the normal operation of the cloud server, while a small number of memory failure events may affect the normal operation of the cloud server. There is a positive correlation between the distribution of memory error events and the occurrence of memory failure events. Specifically, the more frequent and more frequent memory error events occur, the more likely a memory failure event will occur. Therefore, in the embodiments of the present disclosure, the occurrence of memory failure events is predicted by analyzing the characteristics of memory error events. This is specifically achieved through a machine learning algorithm, namely, the pre-trained failure prediction model in step 108.
[0096] The fault prediction model is a machine learning model with fault prediction capabilities. Pre-training enables the model to predict future fault instances. Examples of fault prediction models include, but are not limited to, XGBoost models, MLP models, CNN models, RNN models, LSTM models, Transformer models, BERT models, and large models. The fault prediction model is a binary classification model. Any machine learning model capable of binary classification, trained based on the sample distribution characteristics of sample memory fault events across instances, can implement the fault prediction in step 108.
[0097] A predicted fault instance is an instance in which a memory fault event is predicted to occur. For example, if it is predicted that instance Instance_45 will have a memory fault event, then instance Instance_45 is determined as a predicted fault instance. It should be noted that a predicted fault instance is predicted within a time window, for example, a fault instance in which a memory fault event is predicted to occur within the next day.
[0098] Sample memory failure events are used to train fault prediction models. Memory failure events are those in which uncorrectable memory errors occur. Memory failure events have a significant impact on cloud server stability, making correction difficult to ensure normal cloud server operation. They are likely to cause cloud server downtime. These include, but are not limited to, multiple data bit errors (MBERs), physical damage, and interface errors. A MBER is an error in which two or more data bits occur simultaneously within a data word, exceeding the range that ECC encoding can correct. Hardware failures are permanent errors caused by hardware faults in the memory module, such as capacitor rupture or chip damage. Interface errors are communication errors between the memory controller and memory, such as signal integrity issues that cause data transmission errors for seven days.
[0099] Sample distribution features are the distribution characteristics of memory failure events occurring on instances during the sampling time. They include at least one dimension, such as the time dimension, the spatial dimension, and the spatiotemporal dimension. Correspondingly, sample distribution features include at least one of the sample time distribution feature, the sample space distribution feature, and the sample spatiotemporal distribution feature. For example, the sample time distribution feature indicates at which time instances are more likely to experience memory failure events, the sample space distribution feature indicates which instances are more likely to experience memory failure events, and the sample spatiotemporal distribution feature indicates which instances are more likely to experience memory failure events at which time. Sample distribution features are the distribution characteristics of memory failure events at the instance granularity.
[0100] For example, the sample spatial distribution features include, but are not limited to: a cell error in the instance, a storage cell row error in the instance, a storage cell column error in the instance, a cell soft error in the instance, a cell soft error in the instance, a cell row soft error in the instance, a cell column soft error in the instance, a cell hard error in the instance, a cell row hard error in the instance, a cell column hard error in the instance, a repeated cell error in the instance, a repeated cell row error in the instance, a repeated cell column error in the instance, a cell correctable error in the instance, a repeated storage cell row correctable error in the instance, a repeated storage cell column correctable error in the instance, a cell multiple error bit error in the instance, the number of error checking and correction sources in the instance, the number of error storage data lines in the instance, the maximum number of error cells by row in the instance, the minimum number of error cells by row in the instance, a total of single data bits in the instance, a total of two data bits in the instance, and a total of multiple data bits in the instance.
[0101] For example, the sample time distribution features include but are not limited to: cell errors in the time window on the instance, storage unit row errors in the time window on the instance, storage unit column errors in the time window on the instance, cell soft errors in the time window on the instance, cell soft errors in the time window on the instance, cell row soft errors in the time window on the instance, cell column soft errors in the time window on the instance, cell hard errors in the time window on the instance, cell row hard errors in the time window on the instance, cell Column hard errors, Duplicate cell errors in the time window on the instance, Duplicate cell row errors in the time window on the instance, Duplicate cell column errors in the time window on the instance, Cell correctable errors in the time window on the instance, Duplicate storage cell row correctable errors in the time window on the instance, Duplicate storage cell column correctable errors in the time window on the instance, Cell multiple-bit errors in the time window on the instance, Number of error checking and correction sources in the time window on the instance, Number of errored storage data lines in the time window on the instance, Maximum number of errored cells by row in the time window on the instance, Minimum number of errored cells by row in the time window on the instance, Total single data bits in the time window on the instance, Total two data bits in the time window on the instance, and Total multiple data bits in the time window.
[0102] For example, the spatiotemporal distribution features of the samples include, but are not limited to: the number of cell errors in the time window of the instance, the number of storage cell row errors in the time window of the instance, the number of storage cell column errors in the time window of the instance, the number of cell soft errors in the time window of the instance, the number of cell soft errors in the time window of the instance, the number of cell row soft errors in the time window of the instance, the number of cell column soft errors in the time window of the instance, the number of cell hard errors in the time window of the instance, the number of cell row hard errors in the time window of the instance, the number of cell column hard errors in the time window of the instance, the number of repeated cell errors in the time window of the instance, the number of repeated cell row errors in the time window of the instance, the number of repeated cell column errors in the time window of the instance, the number of correctable cell errors in the time window of the instance, the number of correctable repeated storage cell row errors in the time window of the instance, the number of correctable repeated storage cell column errors in the time window of the instance, the number of multiple bit errors in the cell in the time window of the instance, the number of error checking and correction sources in the time window of the instance, the number of error storage data lines in the time window of the instance, Maximum number of error cells per row in the time window on the instance, Minimum number of error cells per row in the time window on the instance, Total number of single data bits in the time window on the instance, Total number of two data bits in the time window on the instance, and Total number of multiple data bits in the time window on the instance.
[0103] Sample memory failure events are also recorded in the memory error log. The sample distribution characteristics are determined by referring to the processing of steps 104 to 106 above. The event information of the sample memory failure event refers to the event information of the memory error event, as shown below: [2023-06-07 13: 45: 32] EDAC_MC: MCI: UE: CPU O, Node O, channel=l, DIMM 0: Uncorrected error [2023-06-07 13: 45: 32] kernel: EDAC sbridge: 0x00000000: syndrome OxdO, address 0x08000000f7dfe000
[0104] In this data, ^[2023-06-07 13:45:32" and "[2023-06-07 13:45:32" are timestamps, "EDAC_MC" is the message identifier for the system's memory error detection and correction module, "MCI" is the memory controller number, "UE" indicates the event type is an uncorrectable error, and the error occurred on CPU node 0, in the 0th DIMM (dual in-line memory module) and the 1st channel of the memory controller. "syndrome=0xd0" is the error check code, and "0x08000000f7dfe000" is the physical address, indicating that the error occurred at 0x08000000f7dfe000.
[0105] Based on the second distribution feature, a fault prediction model is used to determine a predicted fault instance. Specifically, the second distribution feature is input into the fault prediction model to obtain a confidence level for each instance that a memory fault event will occur within a prediction time window. Based on the confidence level, a predicted fault instance is determined from each instance.
[0106] For example, the second distribution feature X_(i, i~[1,150]) = Instance_TS_Dist> is fed into the XGBoost model to obtain the confidence P_(j, j~[1,50]) ( Y|X_(i, i~[1,150]) ) that 50 instances j, j~[1,50] will experience a memory failure event Y within the next three days. Based on the 50 confidences, the predicted failure instance is determined from the 50 instances: Instance_12 o
[0107] Optionally, after step 108, the following specific step is further included: performing operation and maintenance processing on the predicted fault instance.
[0108] Operations and maintenance processing is performed at the instance granularity, including but not limited to instance migration, instance isolation, instance rollback, and instance restart. Instance migration includes both hot and cold migration. In the disclosed embodiments, instances are the minimum granularity for operations and maintenance processing.
[0109] Exemplarily, hot migration is performed on the predicted fault instance: Instance_12.
[0110] In the disclosed embodiments, by extracting the first distribution characteristics of memory error events in the basic storage units of the cloud server memory, a fine-grained distribution characteristic of the basic storage units is obtained. Based on this, the subordinate relationship between the basic storage units and instances is combined to determine the second distribution characteristics of memory error events in instances. This implements instance-granular feature analysis and determines the distribution characteristics at the instance granularity. Furthermore, based on the distribution characteristics at the instance granularity, fault instances are predicted, achieving instance-granular fault prediction and improving the refinement of fault prediction.
[0111] In an optional embodiment of the present disclosure, the event information includes unit information and time information associated with the memory error event.
[0112] Correspondingly, step 104 includes the following specific steps: extracting a spatial distribution feature of the memory error event in the basic storage unit based on the unit information; extracting a temporal distribution feature of the memory error event in the basic storage unit based on the spatial distribution feature and the time information; and determining a first distribution feature of the memory error event in the basic storage unit based on the spatial distribution feature and the temporal distribution feature.
[0113] The unit information associated with the memory error event is the information of the memory unit related to the memory error event. The unit information is used to locate the location of the memory unit where the memory error event occurred, from which the spatial distribution characteristics of the memory error event in the basic storage unit can be extracted. The unit information includes but is not limited to: memory controller information (memory controller number), memory information (memory number, memory channel, Physical address, memory module information), memory unit information (storage unit information of basic storage units: cell, cell row, cell column, bank; read / write unit information of memory read / write units: error check code). For example, the event information of a memory error event is as follows: [2023-03-24 13:57:12] EDAC_MC: MCI: CE row=456, channel=l, syndrome=OxOA, address=0x89ABCDEF
[0114] Among them, the cell row "row=456", the memory channel "channel=1", the error check code "syndrome=OxOA" and the physical address "address=0x89ABCDEF" are all unit information associated with the memory error event.
[0115] Time information refers to the time when a memory error event occurred. This information is crucial for analyzing the frequency and regularity of memory error events, as well as their correlation with other dimensional features. Spatial distribution features can be combined to extract the temporal distribution characteristics of memory error events within basic storage units. This is illustrated using timestamps in the disclosed embodiments. In the above example, the timestamp "[2023-03-24 13:57:12]" represents time information.
[0116] The spatial distribution feature is the spatial distribution of the basic storage units where the error occurs during a memory error event. This feature describes the distribution pattern of memory error events in physical or logical space, indicating the time at which basic storage units are more likely to experience memory error events. The granularity of the spatial distribution feature varies depending on the basic storage unit selected. For a specific example, see step 104 above and will not be repeated here.
[0117] The time distribution feature is the distribution feature of memory error events in the time dimension. It reflects the trend, periodicity, and burstiness of memory error events over time in physical or logical space. It indicates at which time the basic storage unit is "more likely" to experience a memory error event. The granularity of the time distribution feature varies depending on the basic storage unit selected. For a specific example, see step 104 above and will not be repeated here.
[0118] FIG2 shows a flow chart of a fault prediction method provided by an embodiment of the present disclosure, as shown in FIG2 .
[0119] Spatial and temporal distribution features are extracted from memory error detection and correction logs. Based on the spatial and temporal distribution features, a combination of spatiotemporal distribution features is obtained, resulting in 1,500 distribution features.
[0120] Based on the cell information, the spatial distribution characteristics of memory error events in basic storage units are extracted. Specifically, the error patterns of the memory error events are determined based on the cell information, and the error patterns of the memory error events are statistically analyzed to obtain the spatial distribution characteristics of the memory error events in the basic storage units. The error patterns are the error distribution patterns of the memory error events, including but not limited to the number of errors occurring in the basic storage units and the error patterns in the memory read / write units. Correspondingly, the spatial distribution characteristics include the spatial distribution characteristics of the basic storage units and the spatial distribution characteristics of the memory read / write units.
[0121] Based on spatial distribution features and time information, we extract the temporal distribution features of memory error events in basic storage units. Specifically, we perform statistical analysis of the spatial distribution features based on the time information to obtain the temporal distribution features of memory error events in basic storage units. Since spatial distribution features include the spatial distribution features of basic storage units and the spatial distribution features of internal read / write units, correspondingly, temporal distribution features include the temporal distribution features of basic storage units and the temporal distribution features of memory read / write units.
[0122] Based on the spatial distribution features and the temporal distribution features, a first distribution feature of the memory error event in the basic storage unit is determined. This can be determined from the spatial distribution features and the temporal distribution features, or can be obtained by fusing the spatial distribution features and the temporal distribution features, which is not limited herein. When the first distribution feature is a spatiotemporal distribution feature, since the spatial distribution feature includes the spatial distribution features of the basic storage unit and the spatial distribution features of the memory read / write unit, and the temporal distribution feature includes the temporal distribution features of the basic storage unit and the temporal distribution features of the memory read / write unit, the spatiotemporal distribution feature correspondingly includes the spatiotemporal distribution features of the basic storage unit and the spatiotemporal distribution features of the memory read / write unit.
[0123] Exemplarily, based on the unit information (cell, cell row, cell column, and error check code) of the memory cell, the error mode of the memory error event is determined: the number of error occurrences on the memory cell and the error mode on the data line are counted, the number of error occurrences on the memory cell is counted, and the spatial distribution feature Cell_Count_Spat_Dist of the memory cell is obtained; the error mode on the data line is counted, and the spatial distribution feature Bit_Spat_Dist of the data line is obtained; based on the time information (timestamp), the spatial distribution feature Cell_Count_Spat_Dist and the spatial distribution feature Bit_Spat_Dist are counted, and the temporal distribution feature Cell_Count_Temp_Dist of the memory cell and the temporal distribution feature Bit_Temp_Dist of the data line are obtained; the spatial distribution feature Cell_Count_Spat_Dist and the temporal distribution feature Cell_Count_Temp_Dist are feature fused to obtain the spatiotemporal distribution feature Cell_Count_TS_Dist of the memory cell; the spatial distribution feature Bit_Spat_Dist and the temporal distribution feature Bit_Temp_Dist are fused to obtain the spatiotemporal distribution feature Cell_Count_TS_Dist of the memory cell. _Dist is used to perform feature fusion to obtain the temporal and spatial distribution features of the data Bit_ TS_Dist.
[0124] In the embodiment of the present disclosure, the spatial distribution feature and the temporal distribution feature are extracted to determine the first distribution feature. By extracting features in two dimensions, the memory fault event is described more comprehensively, thereby improving the accuracy of subsequent fault prediction.
[0125] In an optional embodiment of the present disclosure, the unit information includes storage unit information of a basic storage unit and / or read / write unit information of a memory read / write unit, wherein the basic storage unit includes multiple memory read / write units, and the basic storage unit reads and writes data through the memory read / write units.
[0126] Correspondingly, based on the unit information, the spatial distribution characteristics of the memory error event in the basic storage unit are extracted, including the following specific steps: Step: extracting spatial distribution features of memory error events in basic storage units based on storage unit information of basic storage units; and / or extracting spatial distribution features of memory error events in basic storage units based on read / write unit information of memory read / write units.
[0127] A basic storage unit is a data storage unit that makes up a memory. A memory consists of multiple basic storage units, including but not limited to memory cells, banks, and ranks. The storage unit information of a basic storage unit is information related to a memory error event. This storage unit information is used to locate the basic storage unit where the memory error event occurred and to extract the spatial distribution characteristics of the memory error event within the basic storage unit. This storage unit information includes but is not limited to cells, cell rows, cell columns, and banks.
[0128] A memory read / write unit is a memory component that implements data reading and writing functions for basic storage cells. For example, in DRAM, a memory read / write unit is typically associated with one or more basic storage cells, responsible for receiving control signals and transferring data to or from the corresponding basic storage cells. A memory read / write unit includes, but is not limited to, a data burst and a data line (DQ). The read / write unit information of a memory read / write unit is information related to a memory error event. This storage unit information is used to locate the memory read / write unit where the memory error event occurred, from which the spatial distribution characteristics of the memory error event within the memory read / write unit can be extracted. This storage unit information includes, but is not limited to, an error checking code, such as a parity check code. It should be noted that a basic storage unit includes multiple memory read / write units, and a basic storage unit reads and writes data through the memory read / write units. Therefore, the spatial distribution characteristics of the memory error event within the memory read / write unit extracted in the embodiments of the present disclosure can also be considered as the spatial distribution characteristics of the memory error event within the basic storage units.
[0129] Based on the storage unit information of the basic storage unit, the spatial distribution characteristics of the memory error events in the basic storage unit are extracted. Specifically, the number of errors occurring in the basic storage unit is determined based on the storage unit information of the basic storage unit, and the number of errors occurring in the basic storage unit is counted to obtain the spatial distribution characteristics of the memory error events in the basic storage unit.
[0130] Based on the read / write unit information of the memory read / write unit, the spatial distribution characteristics of the memory error events in the basic storage unit are extracted. The specific method is as follows: based on the read / write unit information of the memory read / write unit, the error pattern on the memory read / write unit is determined, and the error pattern on the memory read / write unit is counted to obtain the spatial distribution characteristics of the memory error events in the basic storage unit.
[0131] For example, based on the storage unit information (cell, cell row, and cell column) of a memory cell, the number of errors occurring in the memory cell is determined, and the number of errors occurring in the memory cell is counted to obtain the spatial distribution characteristic of memory error events in the memory cell, Cell_Count_Spat_Dist. Based on the parity error check code, the error patterns in the data block (burst) and data line (DQ) are determined, and the error patterns on the data line are counted to obtain the spatial distribution characteristic of memory error events in the data line, Bit_Spat_Dist.
[0132] In the disclosed embodiments, by extracting spatial distribution features on basic storage units and / or memory read / write units, it is determined that more comprehensive spatial distribution features of memory error events are extracted, thereby improving the comprehensiveness of the first distribution features, and further improving the comprehensiveness of the second distribution features, thereby more comprehensively describing memory fault events and improving the accuracy of subsequent fault predictions.
[0133] In an optional embodiment of the present disclosure, extracting spatial distribution characteristics of memory error events in basic storage units based on storage unit information of the basic storage units includes the following specific steps: determining, based on the storage unit information of the basic storage units, a target basic storage unit where the memory error event occurs and the number of memory error event occurrences in the target basic storage unit; and performing statistics on the target basic storage unit and the number of error occurrences to obtain spatial distribution characteristics of the memory error events in the basic storage units.
[0134] The target basic storage unit is the basic storage unit in which a memory error event occurred. Target basic storage units include, but are not limited to, memory cells, banks, and ranks. In the embodiments of the present disclosure, a memory cell is used as an example for illustration. The error occurrence count is the number of memory error events that occurred within the target basic storage unit.
[0135] In an embodiment of the present disclosure, taking a memory cell as a basic storage unit, statistics are collected for target basic storage units and the number of error occurrences. Spatial distribution features obtained include, but are not limited to, the following: single soft error, faulty row error, corrupt row error > single hard error > faulty column error, and corrupt column error. FIG3 shows a schematic diagram of spatial distribution features in a fault prediction method provided in an embodiment of the present disclosure, as shown below.
[0136] In the target basic storage unit, the memory cell.
[0137] A single soft error is a memory cell that has a single memory error event ("=1"), and no other memory cell in the same row and column has a memory error event. A row error is a memory cell with at least two memory cells on the same row that have a single memory error event ("=1"). A damaged row error is a memory cell with at least two memory cells on the same row that have multiple memory error events (">1"). A single hard error is a memory cell with multiple memory error events (">1"), and no other memory cell in the same row and column has a memory error event. A column error is a memory cell with at least two memory cells on the same column that have a single memory error event ("=1"). A damaged column error is a memory cell with at least two memory cells on the same column that have multiple memory error events (">1")
[0138] For example, based on the storage unit information (cell, cell row, cell column) of the memory cell, the target memory cell where the memory error event occurs and the number of errors occurring in the target memory cell are determined, and the target memory cell and the number of errors occurring are counted to obtain the spatial distribution feature Cell_Count_Spat_Dist of the memory error event in the memory cell: Single soft error single soft error, faulty row, corrupt row, single hard error > faulty column, and corrupt column.
[0139] In the disclosed embodiment, target basic storage units are determined, and statistics are collected on the target basic storage units and the number of error occurrences to obtain the spatial distribution characteristics of memory error events in the basic storage units. This enables the extraction of spatial distribution characteristics at the basic storage unit granularity, improves the accuracy of the first distribution characteristic, and further improves the accuracy of the second distribution characteristic, thereby more accurately describing memory fault events and improving the accuracy of subsequent fault predictions.
[0140] In an optional embodiment of the present disclosure, extracting spatial distribution characteristics of memory error events in basic storage units based on read-write unit information of memory read-write units includes the following specific steps: determining a target memory read-write unit where the memory error event occurs and an error pattern of the memory error event in the target memory read-write unit based on the read-write unit information of the memory read-write unit; and performing statistics on the error pattern in the target memory read-write unit to determine the spatial distribution characteristics of the memory error events in the basic storage units.
[0141] Compared to the spatial distribution features obtained by counting the target basic storage units and the number of error occurrences in the above embodiment, in the embodiment of the present disclosure, the spatial distribution features of memory error events in the basic storage units are determined by counting the error patterns on the target memory read / write units. This is a spatial distribution feature determined by the memory read / write units and is a more fine-grained feature.
[0142] A U-labeled memory read / write unit is the memory read / write unit where a memory error event occurred. Target memory read / write units include, but are not limited to, data blocks and data lines. The error pattern is the characteristic transformation value of the memory error event occurring at the target memory read / write unit. It is used to define the error type of the memory error event and reflects the nature and regularity of the error. Error patterns can provide deeper insights into the nature of memory failures. For example, frequent parity errors on consecutive data lines may indicate a hardware defect or system design issue.
[0143] In the disclosed embodiments, taking memory read / write units as data blocks and data lines as an example, statistics are collected on error patterns on target memory read / write units. The obtained spatial distribution characteristics include, but are not limited to, single bit errors (single bit ent) > multiple bit errors (multi bit ent), adjacent bit errors (adjDQ), and burst errors (burst ent). For example, a single bit error indicates that a memory error has occurred in only one of the 32 data bits, while an adjacent error indicates that a memory error has occurred in two adjacent data bits out of the 32 data bits.
[0144] Each time a memory error event occurs, the memory error log records event information related to the memory error, including the read / write unit information of the memory read / write units. The parity error check code is converted to 32-bit binary format and arranged into 8 data blocks and 4 data lines (memory data bit width is x4), or 4 data blocks and 8 data lines (memory data bit width is x8). Figure 4 shows a schematic diagram of spatial distribution characteristics in another fault prediction method provided by an embodiment of the present disclosure, where black represents an error in the data bit. The number of error bits and error patterns vary in different data blocks and data lines, and different error patterns have different correlations with the occurrence of memory fault events. For an arrangement of 8 data blocks and 4 data lines, a memory error event occurs in data block 0 and data line 1, data block 3 and data line 2, data block 5 and data line 0, and data block 6 and data line 3. For an arrangement of 4 data blocks and 8 data lines, a memory error event occurs on data block 0 and data line 2, and a memory error event occurs on data block 0 and data line 5.
[0145] For example, based on the read / write unit information (parity error check code) of data blocks and data lines, the target data blocks and target data lines where memory errors occurred, as well as the error patterns on those blocks and lines, are determined. Statistics are collected for the target data blocks and data lines, as well as the number of error occurrences, to obtain the spatial distribution characteristics of memory error events on the data blocks and data lines (Bit_Spat_Dist): single bit errors (single bit cnt) > multiple bit errors (multi bit cnt) > adjacent errors (adjDQ) and burst errors (burst ent).
[0146] In the disclosed embodiments, target memory read / write units are identified, and statistics are collected on the target memory read / write units and error patterns to obtain the spatial distribution characteristics of memory error events in basic storage units. This enables the extraction of more fine-grained spatial distribution characteristics, improves the accuracy of the first distribution characteristics, and further improves the accuracy of the second distribution characteristics, thereby more accurately describing memory fault events and improving the accuracy of subsequent fault predictions.
[0147] In an optional embodiment of the present disclosure, extracting the temporal distribution characteristics of memory error events in basic storage units based on spatial distribution characteristics and time information includes the following specific steps: dividing the spatial distribution characteristics using at least one statistical time window based on the time information to obtain the spatial distribution characteristics within the at least one statistical time window; and performing statistics on the spatial distribution characteristics within the at least one statistical time window to obtain the temporal distribution characteristics of the memory error events in the basic storage units.
[0148] The time information of memory error events is also important for analyzing their correlation with memory failure events. For example, if a memory error event occurs on a certain instance of a cloud server at 12:00 every day, there is a high probability that a memory failure event will occur on this instance around 12:00 on a certain day in the future.
[0149] The statistical time window is a pre-set time period used to segment continuous time series data. The spatial distribution characteristics are then determined within different statistical time windows to reflect the temporal distribution patterns of memory error events in physical or logical space, such as trends, periodicity, and suddenness. In the embodiments of this disclosure, statistical time windows of one day, three days, and seven days are used as examples for illustration.
[0150] The spatial distribution features within at least one statistical time window are spatial distribution features of memory error events respectively extracted within at least one preset statistical time window, and can be represented in the form of a time series.
[0151] Perform statistics on the spatial distribution characteristics within at least one statistical time window to obtain the time distribution characteristics of the memory error event in the basic storage unit. The specific method is: perform aggregate statistics on the spatial distribution characteristics within multiple time windows to obtain the memory error event in the basic storage unit. The temporal distribution characteristics of the basic storage unit. Aggregate statistics are the number of spatial distribution characteristics counted at a specific frequency. For a frequency of minutes, aggregate statistics are obtained within the statistical time window w. For a spatial distribution characteristic x, the number of minute-level events will form a time series X(x, t), which is the temporal distribution characteristics of memory error events in the basic storage unit. Aggregate statistics include but are not limited to: sum (Sum), standard deviation (Std), delta (Diff, Delta), kurtosis (Kurtosis), and skew (Skew). Specifically, they are shown in the following formulas 1 to 6: Sum ( x, t, w
[0152] Among them, Sum () is the summation process, x is the time series of spatial distribution characteristics, w is the statistical time window, i is the variable Std(x, t, w) = JVar(x, t, w) , Mean(x, t, w) = j £-=t-w+i x i, Var(x, t, w) =ZUt-w+i ( x i - Mean(x, t, w) ) Formula 2
[0153] Where Std() is the standard deviation, Var() is the variance, Mean() is the mean, x is the spatial distribution feature, t is the time information, (x, t) is the time series of the spatial distribution feature, w is the statistical time window, and i is the variable. Diff(
[0154] Among them, Diff () is the window In summation, x represents the spatial distribution feature, t represents the temporal information, (x, t) represents the time series of the spatial distribution feature, and w represents the statistical time window. Delta(x, t, w) = Sum(x, t, w) — Sum ( x, t — w, w ) Formula 4
[0155] Delta() is the increment between windows, representing the increment of the feature in the current window relative to the previous window, Sum() is the summation process, x is the spatial distribution feature, t is the time information, (x, t) is the time series of the spatial distribution feature, and w is the statistical time window. Kurtosis (x, t, w) = E [(-y-) 4 ] Formula 5
[0156] Where Kurtosis ( ) is the peak value processing for finding the probability density, x is the spatial distribution feature, t is the time information, X ( x, t ) is the time series of the spatial distribution feature, w is the statistical time window, is the feature mean, and is the feature standard deviation. Skew(x, t, w) = 3 (dou^) Formula 6
[0157] Where Skew ( ) is the skewness processing for calculating the probability density, which is used to measure the symmetry of the feature distribution. x is the spatial distribution feature, t is the time information, X ( x, t ) is the time series of the spatial distribution feature, w is the statistical time window, □ is the feature mean, and . is the feature standard deviation.
[0158] For example, based on time information (timestamps), the spatial distribution features Cell_Count_Spat_Dist and Bit_Spat_Dist are statistically divided into three statistical time windows (1 day, 3 days, and 7 days). The spatial distribution features (Cell_Count_Spat_Dist_window_1; Cell_Count_Spat_Dist_window_3; Cell_Count_Spat_Dist_window_7) and spatial distribution features (Bit_Spat_Dist_window_1; Bit_Spat_Dist_window_3; Bit_Spat_Dist_window_7) within the three statistical time windows are obtained. The number of spatial distribution features within multiple time windows is counted at a specific frequency (per minute) to obtain the temporal distribution features of memory error events in memory cells (Cell_Count_Temp_Dist) and the temporal distribution features of data lines (Bit_Temp_Dist).
[0159] In the disclosed embodiments, spatial distribution features within at least one statistical time window are divided and statistically analyzed to obtain the temporal distribution features of memory error events within basic storage units. This enables the extraction of temporal distribution features at the granularity of basic storage units in the time dimension, improving the accuracy of the first distribution feature, which in turn improves the accuracy of the second distribution feature, thereby more accurately describing memory fault events and improving the accuracy of subsequent fault prediction.
[0160] In an optional embodiment of the present disclosure, determining a first distribution feature of memory error events in a basic storage unit based on spatial distribution features and temporal distribution features includes the following specific steps: aggregating the spatial distribution features using at least one aggregation time window to obtain a spatial aggregation feature of the at least one aggregation time window; and integrating the spatial aggregation feature and the temporal distribution feature to obtain the first distribution feature of the memory error events in the basic storage unit.
[0161] In the embodiment of the present disclosure, spatial distribution features are aggregated within an aggregation time window and integrated to obtain a first distribution feature, thereby achieving feature fusion of spatial distribution features and temporal distribution features. The obtained first distribution feature is a spatiotemporal distribution feature at the granularity of a basic storage unit.
[0162] The aggregation time window is a pre-set time period used to aggregate spatial distribution features. Spatial and temporal distribution features are then integrated within different aggregation time windows to obtain spatiotemporal distribution features. These features reflect the temporal distribution trends, periodicity, and burstiness of memory error events in physical or logical space. In this embodiment, a 5-minute aggregation time window is used as an example.
[0163] The spatial aggregation feature of the at least one aggregation time window is an aggregation feature of spatial distribution features of multiple memory error events within the at least one preset aggregation time window.
[0164] It should be noted that the statistical time window is used to extract the temporal distribution characteristics of memory error events. By setting statistical time windows of different lengths (for example, 1 day, 3 days, and 7 days), the spatial distribution characteristics of memory error events are divided and statistically analyzed to reveal the temporal variation patterns of memory error events, such as trends, periodicity, and suddenness. Within each statistical time window, the spatial distribution characteristics are statistically analyzed to form time series data. The aggregation time window is used to combine spatial distribution characteristics with temporal distribution characteristics to generate spatiotemporal distribution characteristics, achieving feature fusion. The aggregation time window is a pre-set short time period (for example, 5 minutes). Within this window, the spatial distribution characteristics of multiple consecutive memory error events are aggregated and calculated to obtain spatial aggregate features. These spatial aggregate features and the previously obtained temporal distribution features are then integrated to obtain the first distribution feature of the memory error events, namely, the spatiotemporal distribution feature at the basic storage unit granularity. Therefore, the statistical time window focuses on the changes in spatial distribution characteristics within time periods of different lengths, while the aggregation time window focuses more on merging or summarizing spatial distribution characteristics within a shorter time interval, and then combining the time distribution characteristics to construct higher-dimensional spatiotemporal distribution characteristics.
[0165] The spatial distribution features are aggregated using at least one aggregation time window to obtain spatial aggregation features for the at least one aggregation time window. Specifically, the spatial distribution features are aggregated using at least one aggregation time window to obtain spatial aggregation features for the at least one aggregation time window. Figure 5 shows a schematic diagram of spatiotemporal distribution features in a fault prediction method provided by one embodiment of the present disclosure. Figure 5 is shown below.
[0166] For multiple instances (Instance 1 ... Instance M), in any instance, using N aggregation time windows (aggregation time window 1, aggregation time window 2, aggregation time window N, where the interval between adjacent two aggregation time windows is 5 minutes), the spatial distribution features within three statistical time windows (1 day / 3 days / 7 days) are aggregated at the hourly frequency. This achieves the aggregation of nine spatial distribution features and obtains the spatial aggregation features of the N aggregation time windows.
[0167] Based on the spatial aggregation feature and the time distribution feature, a first distribution feature of the memory error event in the basic storage unit is obtained by integrating the time distribution feature based on at least one spatial aggregation feature to obtain the first distribution feature of the memory error event in the basic storage unit.
[0168] For example, 10 aggregation time windows are used with a 5-minute interval to perform window aggregation on the spatial distribution feature Cell_Count_Spat_Dist to obtain the spatial aggregation features Cell_Count_Spat_Dist_Agg of the 10 aggregation time windows. Based on the 10 spatial aggregation features Cell_Count_Spat_Dist_Agg, the temporal distribution feature Cell_Count_Temp_Dist is integrated to obtain the spatial distribution feature Cell_Count_TS_Dist of the memory error event. 10 non-aggregated time windows are used with a 5-minute interval to perform window aggregation on the spatial distribution feature Bit_Count_Spat_Dist to obtain the spatial aggregation features Bit_Count_Spat_Dist_Agg of the 10 aggregation time windows. Based on the 10 spatial aggregation features Bit_Count_Spat_Dist_Agg, the temporal distribution feature Bit_Count_Temp_Dist is integrated to obtain the spatial distribution feature Cell_Count_TS_Dist of the memory error event. Temp_Dist is used to integrate the features and obtain the temporal and spatial characteristics of the memory error event on the data line, Bit_Count_TS_Dist.
[0169] In the disclosed embodiment, the combination of temporal and spatial features is achieved through aggregation and integration. By fusing features in two dimensions, memory fault events are described more comprehensively and accurately, thereby improving the accuracy of subsequent fault prediction.
[0170] In an optional embodiment of the present disclosure, before step 108, the following specific steps are further included: obtaining event information of a sample memory failure event; extracting a first sample distribution feature of the sample memory failure event in a basic storage unit granule of a cloud server memory based on the event information of the sample memory failure event; determining a second sample distribution feature of the sample memory failure event in an instance based on the first sample distribution feature and the subordinate relationship between the basic storage unit and the instance; obtaining a memory failure event prediction result using a fault prediction model based on the second sample distribution feature; and training the fault prediction model based on the memory failure event prediction result.
[0171] Sample memory failure events are memory failure events used to train the fault prediction model. The sample memory failure events are obtained by sampling historical memory failure events. The event information of the sample memory failure events is data that records and describes the specific circumstances of the failure. The first sample distribution feature is the first distribution feature of the sample memory failure event occurring in the instance within the sampling time. The second sample distribution feature is the second distribution feature of the sample memory failure event occurring in the instance within the sampling time. The specific contents of the first and second sample distribution features have been described in detail in step 108 and will not be repeated here.
[0172] Based on the memory fault event prediction results, the fault prediction model is trained by calculating a loss value based on the memory fault event prediction results and adjusting the model parameters of the fault prediction model based on the loss value. Loss values include, but are not limited to, binary classification loss (0-1 loss), cross entropy loss, exponential loss, and Kullback-Leibler Divergence loss. The above process includes forward propagation and backward propagation. Adjusting model parameters is performed through backward propagation, and can be done using either a gradient update method or an adaptive learning rate method, which is not limited here.
[0173] It should be noted that the processing logic of each step in the embodiment of the present disclosure is consistent with that of the above steps 104-108, except for the input and output being different. Therefore, the specific methods of other steps refer to the above embodiment and will not be repeated here.
[0174] Optionally, the embodiment of the present disclosure further includes the following specific steps: obtaining event information of a sample non-memory failure event; and executing each step accordingly until a memory failure time prediction result is obtained.
[0175] Accordingly, based on the memory fault event prediction result, the loss value is calculated, and based on the loss value, the model parameters of the fault prediction model are adjusted. Specifically, based on the memory fault event prediction result, the comparative loss value is calculated, and based on the comparative loss value, the fault prediction model parameters are adjusted. Model parameters of the model.
[0176] #The memory failure event is used as a positive sample, and the non-memory failure event is used as a negative sample. The comparison loss value is any of the various loss values mentioned above. In addition, if the number of negative samples is too large, the negative samples can be downsampled.
[0177] Exemplarily, event information for 1,000 sample events constructed by sampling historical events is obtained, where the 1,000 sample events include 600 sample memory failure events and 400 sample non-memory failure events. Based on unit information (cell, cell row, cell column, and error check code) of memory cells of the sample events, first sample distribution features of the basic storage units of the 1,000 sample events are extracted. Based on the first sample distribution features and the subordinate relationship between the basic storage unit and the instance, a second sample distribution feature of the sample event in the instance is determined. Based on the second sample distribution features, an XGBoost model is used to obtain a memory failure event prediction result. A contrast loss value is calculated based on the memory failure event prediction result. Based on the contrast loss value, model parameters of the XGBoost model are adjusted using a gradient update method.
[0178] In the embodiment of the present disclosure, by pre-training the fault prediction model, the fault prediction model learns the ability to perform fault prediction based on the second distribution feature, providing model support for subsequent fault prediction and ensuring the accuracy of fault prediction.
[0179] In an optional embodiment of the present disclosure, after determining a second sample distribution feature of a sample memory failure event in an instance based on the first sample distribution feature and the subordinate relationship between the basic storage unit and the instance, the following specific steps are further included: performing a correlation analysis based on multiple second sample distribution features and the sample memory failure events to obtain a correlation between each second sample distribution feature and the sample memory failure event; and screening a target second sample distribution feature from the multiple second sample distribution features based on the correlation.
[0180] Correspondingly, obtaining a memory fault event prediction result based on the second sample distribution feature using the fault prediction model includes the following specific steps: obtaining a memory fault event prediction result based on the target second sample distribution feature using the fault prediction model.
[0181] Memory error correction logs often contain tens of thousands of event information. If distribution features were extracted for each event, the number of distribution features would be too large, resulting in inefficient feature extraction and fault prediction, as well as irrelevant interference. Among these numerous distribution features, only a small number are highly correlated with memory fault events. Therefore, correlation analysis and screening of multiple second-sample distribution features are necessary. This allows the fault prediction model to perform fault prediction based solely on the screened second-sample distribution features, improving the efficiency and accuracy of subsequent fault prediction.
[0182] Correlation analysis is a statistical analysis method used to measure the strength and direction of the relationship between two or more variables. In the disclosed embodiment, the correlation analysis is performed based on the correlation between the second sample distribution characteristics and the occurrence of sample memory failure events.
[0183] The correlation coefficient is a quantitative parameter that measures the degree of correlation between the second sample distribution feature and sample memory failure events. It is a type of correlation coefficient, including but not limited to the Pearson correlation coefficient, the Spearman correlation coefficient, and the Kendall correlation coefficient. The correlation coefficient generally ranges from -1 to +1, with positive values indicating positive correlation, negative values indicating negative correlation, and 0 indicating no correlation. The closer the value is to 1 or -1, the stronger the linear correlation between the two. For example, the Pearson correlation coefficient is calculated between a second sample distribution feature (e.g., the frequency of memory error events in a particular instance over the past week) and the actual number of sample memory failure events. If the correlation coefficient is positive and high, it indicates that the second sample distribution feature has high predictive value for predicting memory failure events.
[0184] The target second sample distribution feature is a second sample distribution feature that satisfies a preset threshold or has the highest correlation, selected from multiple second sample distribution features after correlation analysis. In the embodiment of the present disclosure, the target second sample distribution feature is shown in Table 1: Table 1
[0185] Based on the multiple second sample distribution features and the sample memory fault events, a correlation analysis is performed to obtain the correlation between each second sample distribution feature and the sample memory fault event. Based on the correlation, a target second sample distribution feature is screened from the multiple second sample distribution features. For a specific process, see FIG6 , which shows a schematic diagram of distribution feature screening in a fault prediction method provided by one embodiment of the present disclosure.
[0186] The 1500 second sample distribution features include: the number of storage cell errors within a 1-day statistical time window, the number of storage cell row errors within a 1-day statistical time window, the number of storage cell column errors within a 1-day statistical time window, the number of storage cell soft errors within a 1-day statistical time window, the number of storage cell correctable errors within a 1-day statistical time window, the number of repeated storage cell row errors within a 3-day statistical time window, the number of repeated storage cell row correctable errors within a 3-day statistical time window, the number of storage cell multiple bit errors within a 3-day statistical time window, the number of error check and correction sources within a 3-day statistical time window, the number of error storage data lines within a 3-day statistical time window, the maximum number of error storage cells counted by row within a 7-day statistical time window, the minimum number of error storage cells counted by row within a 7-day statistical time window, the total number of single data bits within a 7-day statistical time window, the total number of two data bits within a 7-day statistical time window, the total number of multiple data bits within a 7-day statistical time window, etc.
[0187] Sample memory failure events include: crash events and uncorrectable error events.
[0188] Based on the 1500 second sample distribution features and sample memory failure events, a correlation analysis was performed at the virtual machine (instance) granularity to obtain the correlation between the 1500 second sample distribution features and the sample memory failure events. Based on the correlation, 150 target second sample distribution features were screened from the 1500 second sample distribution features.
[0189] In the disclosed embodiment, the second sample distribution features are screened through correlation analysis to determine target second sample distribution features with high correlation, which are then used to train the fault prediction model. This allows the fault prediction model to perform fault prediction based solely on the screened second distribution features, thereby improving the efficiency and accuracy of subsequent fault predictions.
[0190] In an optional embodiment of the present disclosure, after step 108, the following specific step is further included: migrating the predicted fault instance from the cloud server to another cloud server, where the other cloud server is any cloud server other than the cloud server.
[0191] Other cloud servers are cloud servers in the cloud server cluster other than the current cloud server. Within the same cloud server cluster, cloud servers other than the current cloud server that can be used to host instance migrations are called other cloud servers. Other cloud servers also include instances built on virtualization technology, providing cloud computing services to external users. Other cloud servers have independent operating system environments and configurable hardware resources, and can accept migrated predicted failure instances as needed. For example, after predicting a possible failure for the predicted failure instance Instance_45, the operations and maintenance system will search for and select one or more other cloud servers (other nodes) as migration targets, migrating the predicted failure instance to these other cloud servers.
[0192] For example, a cloud server may have a specially allocated reserved cloud server or an idle cloud server on another normally operating node. When Instance_45 is predicted to fail, the system automatically or after manual confirmation initiates instance migration: Instance_45 is locked and its status is set to "pending migration." Based on a preset scheduling algorithm, a suitable cloud server is selected, with the same or higher memory capacity as Instance_45 and sufficient computing power to ensure stable operation after migration. Using efficient virtualization technologies such as storage migration, Instance_45's cleared data (including operating status and cached data) is copied from the current cloud server to the new cloud server. After the data is fully copied and verified, Instance_45 is suspended on the current cloud server, the system mapping table is updated, and the virtual address of Instance_45 is redirected to the new physical memory address. Instance_45 is then resumed on the new cloud server.
[0193] In the disclosed embodiments, fault instances are predicted based on instance-granular distribution characteristics, achieving instance-granular fault prediction and improving the refinement of fault prediction. Furthermore, instance-granular migration is achieved, improving resource utilization of cloud services, improving the stability of cloud servers, and improving the availability of cloud servers.
[0194] 7 , which shows a flowchart of a fault handling method provided by an embodiment of the present disclosure, which is applied to a management node in a cloud service system and includes the following steps 702 to 710.
[0195] Step 702: Obtain a memory error log of the cloud server, wherein the memory error log records event information of a memory error event.
[0196] Step 704: Based on the event information, extract a first distribution feature of the basic storage unit of the memory of the cloud server.
[0197] Step 706: Determine a second distribution feature of the memory error event in the instance based on the first distribution feature and the subordinate relationship between the basic storage unit and the instance.
[0198] Step 708: Based on the second distribution feature, use a fault prediction model to determine a predicted fault instance, wherein the fault prediction model is trained based on sample distribution features of sample memory fault events on the instance.
[0199] Step 710: Migrate the predicted fault instance from the cloud server to another cloud server, where the other cloud server is any cloud server other than the cloud server.
[0200] The disclosed embodiments can be applied to management nodes in cloud service systems, enabling automatic fault prediction and instance migration. A management node in a cloud service system refers to an entity or service used to centrally monitor, configure, maintain, and optimize a cloud server cluster. A management node is typically a software system or a set of services with advanced permissions. It collects key operational data from each cloud server instance in real time, including but not limited to performance metrics, log information, and hardware health status, and performs intelligent analysis and decision-making based on this data. The management node is responsible for managing and maintaining all cloud servers in the cloud server cluster. It uses automated tools and technologies to implement fault detection, performance tuning, load balancing, and high availability assurance, thus achieving operational management for the cloud server cluster.
[0201] The embodiment of the present disclosure and the embodiment of the specification of FIG. 1 are based on the same inventive concept. For the specific methods of steps 702-710, refer to steps 102-108 and the embodiment of instance migration, which will not be repeated here.
[0202] In the disclosed embodiments, by extracting the first distribution characteristics of memory error events in the basic storage units of the cloud server memory, a fine-grained distribution characteristic of the basic storage units is obtained. Based on this, the subordinate relationship between the basic storage units and instances is combined to determine the second distribution characteristics of memory error events in instances. This implements instance-granular feature analysis and determines the instance-granular distribution characteristics. Furthermore, based on the instance-granular distribution characteristics, fault instances are predicted, achieving instance-granular fault prediction and improving the refinement of fault prediction. This in turn enables instance-granular migration, improving the resource utilization of cloud services, and enhancing the stability and availability of cloud servers.
[0203] Steps 702-710 in the embodiment of the present disclosure are explained through the operation and maintenance scenario of a cloud server.
[0204] Step 1: In the cloud server, the automated log collection system retrieves memory error logs from the cloud server's log repository (e.g., Logstore) through log management tools, application programming interfaces, or other methods. These logs record key information about memory error events, including but not limited to: error type, timestamp, memory controller number, physical address, memory module information, and cell information of the affected memory cells.
[0205] Step 2: Based on the unit information of the memory cell (cell, cell row, cell column, and error check code), extract the first spatial distribution feature of the memory error event in the memory cell through aggregate statistics. Based on the time information (timestamp), extract the first temporal distribution feature of the memory error event in the memory cell through aggregate statistics. Perform feature fusion on the first temporal distribution feature and the first spatial distribution feature to obtain the first spatiotemporal distribution feature of the memory error event in the memory cell, and determine the first spatiotemporal distribution feature as the first distribution feature.
[0206] Step 3: Based on the extracted first distribution feature, combined with the virtual memory to physical memory mapping table, the first distribution feature is mapped to the specific instance level, thereby determining the second distribution feature of the memory error event on each instance.
[0207] Step 4: Use a pre-trained fault prediction model (which is trained based on a dataset of distribution characteristics of historical sample memory failure events) to input the second distribution characteristic data obtained above. Output the confidence level of each instance that a memory failure event will occur in the future. Based on the confidence level, predict the predicted fault instances that may occur in the future.
[0208] Step 5: For high-risk predicted failure instances, migrate the predicted failure instances from the current cloud server (current node or current physical host) to other cloud servers (other nodes or other physical hosts).
[0209] Through the operation and maintenance processing of steps 1 to 5, the number of unnecessary instance migrations is reduced, and other instances (virtual machines) on the damaged node continue to provide cloud computing services, thereby improving the availability of cloud servers and achieving long-term cost reduction and efficiency improvement.
[0210] Corresponding to the above method embodiments, the present disclosure also provides an embodiment of a fault prediction device. FIG8 shows a schematic structural diagram of a fault prediction device provided by one embodiment of the present disclosure. As shown in FIG8 , the device includes the following modules.
[0211] The first acquisition module 802 is configured to acquire a memory error event of the cloud server, wherein the memory error event has corresponding event information.
[0212] The first extraction module 804 is configured to extract a first distribution feature of a basic storage unit of a memory of a cloud server based on the event information.
[0213] The first determining module 806 is configured to determine a second distribution feature of the memory error event in the instance according to the first distribution feature and the subordinate relationship between the basic storage unit and the instance.
[0214] The first prediction module 808 is configured to determine a predicted fault instance based on the second distribution feature using a fault prediction model, wherein the fault prediction model is trained based on sample distribution features of sample memory fault events on the instance.
[0215] Optionally, the event information includes unit information and time information associated with the memory error event.
[0216] Correspondingly, the first extraction module 804 is further configured to: extract the spatial distribution characteristics of the memory error event in the basic storage unit based on the unit information; extract the temporal distribution characteristics of the memory error event in the basic storage unit based on the spatial distribution characteristics and the time information; and determine the first distribution characteristics of the memory error event in the basic storage unit based on the spatial distribution characteristics and the temporal distribution characteristics.
[0217] Optionally, the unit information includes storage unit information of a basic storage unit and / or read / write unit information of a memory read / write unit, wherein the basic storage unit includes multiple memory read / write units, and the basic storage unit reads and writes data through the memory read / write units.
[0218] Correspondingly, the first extraction module 804 is further configured to: extract spatial distribution characteristics of memory error events in the basic storage unit based on the storage unit information of the basic storage unit; and / or extract spatial distribution characteristics of memory error events in the basic storage unit based on the read / write unit information of the memory read / write unit.
[0219] Optionally, the first extraction module 804 is further configured to: determine, based on the storage unit information of the basic storage unit, a target basic storage unit where the memory error event occurs and the number of error occurrences of the memory error event on the target basic storage unit; and perform statistics on the target basic storage unit and the number of error occurrences to obtain spatial distribution characteristics of the memory error event in the basic storage unit.
[0220] Optionally, the first extraction module 804 is further configured to: determine the target memory read-write unit where the memory error event occurs and the error pattern of the memory error event on the target memory read-write unit based on the read-write unit information of the memory read-write unit; and perform statistics on the error pattern on the target memory read-write unit to determine the spatial distribution characteristics of the memory error event in the basic storage unit.
[0221] Optionally, the first extraction module 804 is further configured to: divide the spatial distribution characteristics using at least one statistical time window based on time information to obtain the spatial distribution characteristics within the at least one statistical time window; and perform statistics on the spatial distribution characteristics within the at least one statistical time window to obtain the time distribution characteristics of the memory error event in the basic storage unit.
[0222] Optionally, the first extraction module 804 is further configured to: aggregate the spatial distribution features using at least one aggregation time window to obtain spatial aggregation features of the at least one aggregation time window; and integrate the spatial aggregation features and the time distribution features to obtain a first distribution feature of the memory error event in the basic storage unit.
[0223] Optionally, the apparatus further includes: a training module configured to obtain event information of a sample memory fault event; extract a first sample distribution feature of the sample memory fault event in a basic storage unit of a cloud server memory based on the event information of the sample memory fault event; determine a second sample distribution feature of the sample memory fault event in an instance based on the first sample distribution feature and a subordinate relationship between the basic storage unit and the instance; obtain a memory fault event prediction result using a fault prediction model based on the second sample distribution feature; and train the fault prediction model based on the memory fault event prediction result.
[0224] Optionally, the device further includes: a screening module configured to screen out a target second sample distribution feature from the plurality of second sample distribution features based on the correlation.
[0225] Correspondingly, the training module is further configured to: obtain a memory fault event prediction result based on the target second sample distribution feature using the fault prediction model.
[0226] Optionally, the device further includes: a first migration module, configured to migrate the predicted fault instance from the cloud server to another cloud server, wherein the other cloud server is any cloud server other than the cloud server.
[0227] In the disclosed embodiments, by extracting the first distribution characteristics of memory error events in the basic storage units of the cloud server memory, a fine-grained distribution characteristic of the basic storage units is obtained. Based on this, the subordinate relationship between the basic storage units and instances is combined to determine the second distribution characteristics of memory error events in instances. This implements instance-granular feature analysis and determines the distribution characteristics at the instance granularity. Furthermore, based on the distribution characteristics at the instance granularity, fault instances are predicted, achieving instance-granular fault prediction and improving the refinement of fault prediction.
[0228] The above is a schematic diagram of a fault prediction device according to this embodiment. It should be noted that the technical solution of the fault prediction device and the technical solution of the aforementioned fault prediction method share the same concept. For details not described in detail in the technical solution of the fault prediction device, refer to the description of the technical solution of the aforementioned fault prediction method.
[0229] Corresponding to the above method embodiments, the present disclosure also provides an embodiment of a fault handling device. FIG9 shows a schematic structural diagram of a fault handling device provided by one embodiment of the present disclosure. As shown in FIG9 , the device is applied to a management node in a cloud service system and includes the following modules.
[0230] The second acquisition module 902 is configured to acquire a memory error log, wherein the memory error log records event information of a memory error event.
[0231] The second extraction module 904 is configured to extract a first distribution feature of the memory error event in the basic storage unit of the cloud server memory based on the event information.
[0232] The second determining module 906 is configured to determine, based on the first distribution feature and the subordinate relationship between the basic storage unit and the instance, Determine a second distribution characteristic of memory error events in the instance.
[0233] The second prediction module 908 is configured to determine a predicted fault instance based on the second distribution feature using a fault prediction model, wherein the fault prediction model is trained based on sample distribution features of sample memory fault events on the instance.
[0234] The second migration module 910 is configured to migrate the predicted fault instance from the cloud server to another cloud server, where the other cloud server is any cloud server other than the cloud server.
[0235] In the disclosed embodiments, by extracting the first distribution characteristics of memory error events in the basic storage units of the cloud server memory, a fine-grained distribution characteristic of the basic storage units is obtained. Based on this, the subordinate relationship between the basic storage units and instances is combined to determine the second distribution characteristics of memory error events in instances. This implements instance-granular feature analysis and determines the instance-granular distribution characteristics. Furthermore, based on the instance-granular distribution characteristics, fault instances are predicted, achieving instance-granular fault prediction and improving the refinement of fault prediction. This in turn enables instance-granular migration, improving the resource utilization of cloud services, and enhancing the stability and availability of cloud servers.
[0236] The above is a schematic diagram of a fault handling device according to this embodiment. It should be noted that the technical solution of the fault handling device and the technical solution of the fault handling method described above share the same concept. For details not described in detail in the technical solution of the fault handling device, please refer to the description of the technical solution of the fault handling method described above.
[0237] Corresponding to the above-mentioned method embodiments, the present disclosure also provides a fault handling system embodiment. Figure 10 shows a schematic diagram of the structure of a fault handling system provided by one embodiment of the present disclosure. As shown in Figure 10, the fault handling system includes a log library 1002, a fault prediction subsystem 1004, and an operation and maintenance wheel rotor system 1006. The fault prediction subsystem 1004 includes a feature extraction unit 10042, a fault prediction unit 10044, and an instruction generation unit 10046.
[0238] Feature extraction unit 10042 is configured to obtain a memory error log of the cloud server from log repository 1002, where the memory error log records event information of the memory error event; extract, based on the event information, a first distribution feature of the memory error event in a basic storage unit of the cloud server memory; and determine, based on the first distribution feature and the subordinate relationship between the basic storage unit and the instance, a second distribution feature of the memory error event in the instance.
[0239] The fault prediction unit 10044 is configured to determine a predicted fault instance based on the second distribution feature using a fault prediction model, wherein the fault prediction model is trained based on sample distribution features of sample memory fault events on the instance;
[0240] The instruction generation unit 10046 is used to generate a fault handling instruction based on the predicted fault instance and send the fault handling instruction to the maintenance wheel rotor system 1006.
[0241] The operation and maintenance rotor system 1006 is configured to migrate the predicted fault instance from the cloud server to another cloud server in response to the fault handling instruction, where the other cloud server is any cloud server other than the cloud server.
[0242] Optionally, the operation and maintenance subsystem 1006 is further configured to obtain instance operation information of migration instances of other cloud servers and send the instance operation information to the fault prediction subsystem 1004; the fault prediction subsystem 1004 is further configured to evaluate migration risks and migration benefits based on the instance operation information.
[0243] The fault prediction subsystem 1004 collects and parses memory error logs from the log repository 1002 in real time, storing them in the feature extraction unit 10042. The real-time data in the feature extraction unit 10042 undergoes relevant operations and aggregation to generate a second distribution feature of the memory error event in the instance. This second distribution feature is then input into the fault prediction model. If the current instance is judged to be high-risk, it is determined to be a predicted fault instance and, via the instruction generation unit 10046 of the fault prediction subsystem 1004, is sent to the operation and maintenance rotor system 1006 for instance migration. The operation and maintenance rotor system 1006 then feeds back the instance operation information to the fault prediction subsystem 1004 for use in assessing the migration risks and benefits.
[0244] Migration risk refers to operational issues that may occur during the migration of a predicted failure instance from its current cloud server to another cloud server. These issues include, but are not limited to: Service continuity interruption: The migration process may cause temporary unavailability of application services, impacting user experience and normal service operation. Data loss or inconsistency: If not properly handled during the migration process, some data may not be fully synchronized or updated to the new server, leading to data consistency issues. Performance degradation: The target cloud server may differ from the source server in configuration, network latency, and other aspects, potentially resulting in reduced application performance after migration. Compatibility issues: The new environment may have compatibility issues with operating system versions, middleware versions, or other components, impacting normal service startup or operation. Additional cost increases: These may include higher resource costs for the target cloud server or increased indirect costs such as technical support and manpower during the migration process. For example, if a cloud server instance is predicted to be high-risk due to frequent memory errors and is being migrated to another cloud server, but if adequate data backups and migration testing are not performed during the migration, resulting in service downtime for several hours, this constitutes migration risk.
[0245] Migration benefits refer to the operational performance improvements achieved through instance migration, including but not limited to: Improved stability: By migrating high-risk instances, potential memory failures can be effectively avoided, the stability and reliability of the cloud server can be improved, and service interruptions due to hardware or software failures can be reduced. Optimized resource utilization: Reasonable allocation of resources and removal of problematic instances can free up original resources for more efficient or more important services, while also ensuring that predicted fault instances are improved in the new environment. Reduced costs: Although migration itself may incur a one-time cost, in the long run, by promptly discovering and resolving problem instances, greater losses caused by failures, such as user churn and compensation costs, can be prevented. Enhanced disaster recovery capabilities: Distributing instances across different cloud servers can enhance The system's fault tolerance and disaster recovery capabilities. For example, if the instance mentioned above, at risk of memory errors, is successfully migrated to another cloud server with optimized configuration, not only does it eliminate the memory failure risk, but it also improves overall resource utilization and provides faster service response times in the new environment. This represents the migration benefit.
[0246] In the disclosed embodiments, by extracting the first distribution characteristics of memory error events in the basic storage units of the cloud server memory, a fine-grained distribution characteristic of the basic storage units is obtained. Based on this, the subordinate relationship between the basic storage units and instances is combined to determine the second distribution characteristics of memory error events in instances. This implements instance-granular feature analysis and determines the instance-granular distribution characteristics. Furthermore, based on the instance-granular distribution characteristics, fault instances are predicted, achieving instance-granular fault prediction and improving the refinement of fault prediction. This in turn enables instance-granular migration, improving the resource utilization of cloud services, and enhancing the stability and availability of cloud servers.
[0247] The above is a schematic diagram of a fault handling system according to this embodiment. It should be noted that the technical solution of this fault handling system shares the same concept as the technical solutions of the aforementioned fault prediction method and fault handling method. For details not described in detail in the technical solution of the fault handling system, please refer to the description of the technical solutions of the aforementioned fault prediction method or fault handling method.
[0248] Figure 11 shows a block diagram of a computing device according to an embodiment of the present disclosure. Components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.
[0249] Computing device 1100 also includes an access device 1140 that enables computing device 1100 to communicate via one or more networks 1160. Examples of these networks include a combination of a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a communication network such as the Internet. Access device 1140 may include one or more of any type of wired or wireless network interface (e.g., a network interface controller (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0250] In one embodiment of the present disclosure, the aforementioned components of the computing device 1100 and other components not shown in FIG. 11 may also be connected to one another, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG. 11 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.
[0251] Computing device 1100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1100 may also be a mobile or stationary server.
[0252] The processor 1120 is configured to execute the following computer program / instruction, which implements the steps of the above-mentioned fault prediction method or fault handling method when executed by the processor.
[0253] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device shares the same concept as the technical solutions of the aforementioned fault prediction method and fault handling method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the aforementioned fault prediction method or fault handling method.
[0254] An embodiment of the present disclosure further provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of the above-mentioned fault prediction method or fault handling method when executed by a processor.
[0255] The above is an illustrative embodiment of a computer-readable storage medium. It should be noted that the technical solution of this storage medium shares the same concept as the technical solutions of the aforementioned fault prediction method and fault handling method. For details not described in detail in the technical solution of the storage medium, refer to the description of the technical solutions of the aforementioned fault prediction method or fault handling method.
[0256] An embodiment of the present disclosure further provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned fault prediction method or fault handling method when executed by a processor.
[0257] The above is an illustrative solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product shares the same concept as the technical solutions of the aforementioned fault prediction method and fault handling method. For details not described in detail in the technical solution of the computer program product, refer to the description of the technical solutions of the aforementioned fault prediction method or fault handling method.
[0258] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are possible or may be advantageous.
[0259] The computer instructions include computer program product codes, which may be in the form of source code, object The computer program product code may be in the form of a code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program product code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunications signal, and a software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunications signals.
[0260] It should be noted that, for ease of description, the aforementioned method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present disclosure are not limited by the order of the actions described, as certain steps may be performed in a different order or simultaneously, depending on the embodiments of the present disclosure. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily required for the embodiments of the present disclosure.
[0261] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0262] The preferred embodiments disclosed above are intended only to illustrate the present disclosure. The alternative embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of the embodiments disclosed. These embodiments are selected and described in detail to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.
Claims
22 Claims 1. A fault prediction method, comprising: Obtain a memory error event of a cloud server, wherein the memory error event has corresponding event information; extract a first distribution feature of the memory error event in a basic storage unit of the cloud server memory based on the event information; determine a second distribution feature of the memory error event in an instance based on the first distribution feature and a subordinate relationship between the basic storage unit and the instance; and determine a predicted fault instance based on the second distribution feature using a fault prediction model, wherein the fault prediction model is trained based on sample distribution features of sample memory fault events on the instance.
2. The method according to claim 1, wherein the event information includes unit information and time information associated with the memory error event; The extracting, based on the event information, a first distribution feature of the memory error event in the basic storage unit of the cloud server memory includes: extracting, based on the unit information, a spatial distribution feature of the memory error event in the basic storage unit; extracting, based on the spatial distribution feature and the time information, a temporal distribution feature of the memory error event in the basic storage unit; and determining, based on the spatial distribution feature and the temporal distribution feature, a first distribution feature of the memory error event in the basic storage unit.
3. The method according to claim 2, wherein the unit information comprises storage unit information of the basic storage unit and / or read / write unit information of the memory read / write unit, wherein: The basic storage unit includes a plurality of the memory read-write units, and the basic storage unit reads and writes data through the memory read-write units; The extracting, based on the unit information, the spatial distribution characteristics of the memory error event in the basic storage unit includes: extracting, based on the storage unit information of the basic storage unit, the spatial distribution characteristics of the memory error event in the basic storage unit; Based on the read / write unit information of the memory read / write unit, a spatial distribution feature of the memory error event in the basic storage unit is extracted.
4. The method according to claim 3, wherein extracting spatial distribution characteristics of the memory error event in the basic storage unit based on the storage unit information of the basic storage unit comprises: determining, based on the storage unit information of the basic storage unit, a target basic storage unit where the memory error event occurs and a number of error occurrences of the memory error event on the target basic storage unit; The target basic storage unit and the number of error occurrences are counted to obtain the spatial distribution characteristics of the memory error event in the basic storage unit.
5. The method according to claim 3, wherein extracting spatial distribution characteristics of the memory error event in the basic storage unit based on the read / write unit information of the memory read / write unit comprises: determining, based on the read / write unit information of the memory read / write unit, a target memory read / write unit where the memory error event occurs and an error mode of the memory error event on the target memory read / write unit; Statistics are collected on the error patterns of the target memory read / write unit to determine spatial distribution characteristics of the memory error events in the basic storage unit.
6. The method according to claim 2, wherein extracting a time distribution feature of the memory error event in the basic storage unit based on the spatial distribution feature and the time information comprises: Based on the time information, dividing the spatial distribution characteristics by using at least one statistical time window to obtain the spatial distribution characteristics within the at least one statistical time window; Statistics are performed on the spatial distribution characteristics within the at least one statistical time window to obtain the time distribution characteristics of the memory error event in the basic storage unit.
7. The method according to claim 2, wherein determining a first distribution feature of the memory error event in the basic storage unit based on the spatial distribution feature and the temporal distribution feature comprises: aggregating the spatial distribution features using at least one aggregation time window to obtain spatial aggregation features of the at least one aggregation time window; Based on the spatial aggregation feature and the time distribution feature, a first distribution feature of the memory error event in the basic storage unit is obtained by integration.
8. The method according to claim 1, before determining the predicted fault instance by using a fault prediction model based on the second distribution feature, further comprising: Get event information of sample memory failure events; Extracting, based on the event information of the sample memory failure event, a first sample distribution feature of the sample memory failure event in a basic storage unit of the cloud server memory; determining, based on the first sample distribution feature and a subordinate relationship between the basic storage unit and the instance, a second sample distribution feature of the sample memory failure event in the instance; Based on the second sample distribution characteristics, using a fault prediction model, obtaining a memory fault event prediction result; The fault prediction model is trained based on the memory fault event prediction result.
9. The method according to claim 8, after determining a second sample distribution feature of the sample memory failure event in an instance based on the first sample distribution feature and the subordinate relationship between the basic storage unit and the instance, further comprising: performing a correlation analysis based on the plurality of second sample distribution features and the sample memory failure events to obtain a correlation between each of the second sample distribution features and the sample memory failure events; and selecting a target second sample distribution feature from the plurality of second sample distribution features based on the correlation. The obtaining a memory fault event prediction result based on the second sample distribution feature and using a fault prediction model includes: obtaining a memory fault event prediction result based on the target second sample distribution feature and using a fault prediction model.
10. The method according to claim 1, after determining the predicted fault instance by using a fault prediction model based on the second distribution feature, further comprising: Migrating the predicted fault instance from the cloud server to another cloud server, wherein the other cloud server is any cloud server other than the cloud server.
11. A fault handling method, applied to a management node in a cloud service system, comprising: Obtain a memory error log of a cloud server, wherein the memory error log records event information of a memory error event; extract a first distribution feature of the memory error event in a basic storage unit of the cloud server memory based on the event information; determine a second distribution feature of the memory error event in an instance based on the first distribution feature and a subordinate relationship between the basic storage unit and the instance; determine a predicted fault instance based on the second distribution feature using a fault prediction model, wherein the fault prediction model is trained based on sample distribution features of sample memory fault events on the instance; and migrate the predicted fault instance from the cloud server to another cloud server, wherein the other cloud server is any cloud server other than the cloud server.
12. A fault handling system, comprising a log library, a fault prediction subsystem, and an operation and maintenance rotor system, wherein the fault prediction subsystem comprises a feature extraction unit, a fault prediction unit, and an instruction generation unit; The feature extraction unit is configured to obtain a memory error log of the cloud server from the log library, wherein the memory error log records event information of a memory error event; based on the event information, extract a first distribution feature of the memory error event in a basic storage unit of the cloud server memory; and determine a second distribution feature of the memory error event in an instance based on the first distribution feature and a subordinate relationship between the basic storage unit and the instance; the fault prediction unit is configured to determine a predicted fault instance based on the second distribution feature using a fault prediction model, wherein the fault prediction model is trained based on sample distribution features of sample memory fault events on the instance; the instruction generation unit is configured to generate a fault handling instruction based on the predicted fault instance, and send the fault handling instruction to the operation and maintenance wheel rotor system; the operation and maintenance wheel rotor system is configured to migrate the predicted fault instance from the cloud server to another cloud server in response to the fault handling instruction, wherein the other cloud server is any cloud server other than the cloud server.
13. The system according to claim 12, wherein the operation and maintenance subsystem is further configured to obtain instance operation information of the other cloud servers and send the instance operation information to the fault prediction subsystem; and the fault prediction subsystem is further configured to evaluate migration risks and migration benefits based on the instance operation information.
14. A computing device, comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.
15. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.
16. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Early warning method and device for memory fault
CN113297046A
Memory fault prediction method, device and equipment
CN114996065A
Memory fault processing method and device and storage medium
CN115756911A
Memory bank fault prediction method and device, computing equipment and storage medium
CN115840659A
Feature extraction method and device, electronic equipment and storage medium
CN116127292A