Fault processing method and device of distributed storage system, equipment and storage medium

By assessing the status of the distributed storage system and predicting its faults, automatic repair operations are performed, solving the problem of untimely fault repair in distributed storage systems and improving system stability and fault repair efficiency.

CN122132206APending Publication Date: 2026-06-02PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-02-04
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Distributed storage systems cannot be repaired in a timely manner after a failure, causing computer systems that rely on the system to malfunction and affecting business and data retrieval in fields such as fintech and healthcare.

Method used

By collecting system operation data, performing status assessment and fault prediction, automatically selecting and executing target repair operations, and achieving early fault repair.

Benefits of technology

It reduces the manpower cost of fault repair, improves fault repair efficiency, ensures the stable operation of the distributed storage system, and avoids interruptions to business and data retrieval due to faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132206A_ABST
    Figure CN122132206A_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, device, and storage medium for fault handling in a distributed storage system, belonging to the field of artificial intelligence technology. The method includes: collecting system operation data of the distributed storage system; evaluating the status of the distributed storage system based on the system operation data to obtain the system operation status; the system operation status includes: an abnormal operation status; predicting faults in the distributed storage system based on the abnormal operation status and system operation data to obtain fault categories and fault characteristics; selecting target repair operations from preset candidate repair operations based on the fault categories and fault characteristics; and executing the target repair operation on the distributed storage system based on the fault category. This application can be applied to business systems requiring large amounts of data, such as fintech and healthcare, enabling early fault prediction of distributed storage systems and automatic repair, saving manpower and improving fault repair efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and is applied to the fields of financial technology and healthcare. In particular, it relates to a fault handling method, apparatus, device, and storage medium for a distributed storage system. Background Technology

[0002] Fault handling in distributed storage systems typically requires waiting until a fault occurs before repairs can be performed. During the fault and repair period, the distributed system is unable to provide services, causing computer systems using the distributed storage system as storage media to malfunction. For example, in fintech scenarios, insurance business data is massive and relies on distributed storage systems as the data storage medium for insurance business platforms. If the distributed storage system fails during repair, it will affect the operation of the insurance business platform and impact user experience. Similarly, in healthcare, various types of medical data, such as electronic medical records, examination reports, and laboratory reports, rely on distributed storage systems for storage. If the distributed storage system fails before repairs, it will affect the retrieval and storage of medical data. Therefore, how to predict distributed storage system faults in advance and automatically complete repairs has become a pressing technical problem to be solved. Summary of the Invention

[0003] The main objective of this application is to propose a fault handling method, apparatus, device, and storage medium for a distributed storage system, aiming to achieve early fault prediction of the distributed storage system and automatic repair, thereby saving manpower for repair and improving the efficiency of fault repair.

[0004] To achieve the above objectives, a first aspect of this application proposes a fault handling method for a distributed storage system, the method comprising: Collect system operation data from the distributed storage system; The distributed storage system is evaluated based on the system operation data to obtain the system operation status; wherein, the system operation status includes: abnormal operation status; Based on the abnormal operating state and the system operating data, fault prediction is performed on the distributed storage system to obtain fault prediction data; wherein, the fault prediction data includes: fault category and fault characteristics; The target repair operation is selected from the preset candidate repair operations based on the fault category and the fault characteristics. Perform the target repair operation on the distributed storage system according to the fault category.

[0005] In some embodiments, the system operation data includes: hardware monitoring parameters, metadata information, and storage service operation parameters; The step of evaluating the status of the distributed storage system based on the system operation data to obtain the system operation status includes: The disk status of the distributed storage system is evaluated based on the hardware monitoring parameters to obtain the disk operating status. The distributed storage system is evaluated based on the metadata information to obtain the metadata status. The service status of the distributed storage system is evaluated based on the storage service operating parameters to obtain the service operating status. The system operating status is obtained by concatenating the disk operating status, the metadata status, and the service operating status.

[0006] In some embodiments, the step of performing fault prediction on the distributed storage system based on the abnormal operating state and the system operating data to obtain fault prediction data includes: Based on the abnormal operating state, the distributed storage system is classified into abnormal categories to obtain the abnormality category; Abnormal operation data is filtered out from the system operation data according to the abnormality category; Reference cases are selected from preset historical fault repair cases based on the anomaly category; Based on the reference case and the abnormal operation data, the distributed storage system is used to predict faults, and the fault prediction data is obtained.

[0007] In some embodiments, if the anomaly category is a disk anomaly category, the abnormal operation data is disk health indicator data and service level indicator data; the step of performing fault prediction on the distributed storage system based on the reference case and the abnormal operation data to obtain the fault prediction data includes: Based on the disk health index data, the lifespan decay rate is predicted to obtain the predicted disk lifespan decay rate. Based on the service level metric data, capacity prediction is performed to obtain the predicted disk capacity value; Feature extraction is performed on the reference case to obtain reference disk failure features; Based on the predicted disk lifespan decay rate, the predicted disk capacity value, and the reference disk failure characteristics, the disk failure of the distributed storage system is predicted to be faulty, and the fault prediction data is obtained.

[0008] In some embodiments, selecting the target repair operation from a preset pool of candidate repair operations based on the fault category and the fault characteristics includes: The candidate repair operations are filtered according to the fault category to obtain preliminary repair operations; Obtain the reference features and repair success rate of the preliminary repair operation; The similarity between the reference feature and the fault feature is calculated to obtain the feature similarity. The initial repair operations are filtered based on the repair success rate and the feature similarity to obtain the target repair operations.

[0009] In some embodiments, performing the target repair operation on the distributed storage system according to the fault category includes: Based on the fault category, objects to be repaired are selected from the distributed storage system; wherein, the objects to be repaired are any one of the following: disks, storage service nodes, and metadata; The object to be repaired is risk-marked according to the fault category to obtain risk marking information; The target repair operation is performed on the object to be repaired based on the risk marking information.

[0010] In some embodiments, after performing the target repair operation on the distributed storage system according to the fault category, the method further includes: Send the target repair operation to the target terminal and receive the repair feedback information from the target terminal based on the target repair operation; Based on the repair feedback information, the fault category, the fault characteristics, and the target repair operation are used to construct cases to obtain the historical fault repair cases.

[0011] To achieve the above objectives, a second aspect of this application provides a fault handling apparatus for a distributed storage system, the apparatus comprising: The data acquisition module is used to collect system operation data from the distributed storage system. The status assessment module is used to assess the status of the distributed storage system based on the system operation data to obtain the system operation status; wherein, the system operation status includes: abnormal operation status; The fault prediction module is used to predict faults in the distributed storage system based on the abnormal operating state and the system operating data, and obtain fault prediction data; wherein, the fault prediction data includes: fault category and fault characteristics; The operation filtering module is used to filter out target repair operations from preset candidate repair operations based on the fault category and the fault characteristics. The repair execution module is used to perform the target repair operation on the distributed storage system according to the fault category.

[0012] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0013] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0014] The fault handling method, apparatus, device, and storage medium for distributed storage systems proposed in this application automatically predicts faults in advance when an anomaly occurs in the distributed storage system. This predicts the fault type and characteristics, selects a target repair operation from candidate repair operations based on the fault type and characteristics, and then executes the repair operation on the distributed storage system according to the target repair operation. Therefore, automatically and proactively predicting and repairing faults in the distributed storage system saves on manpower costs for fault repair and reduces the impact of faults on the distributed storage system, making its operation more stable. Attached Figure Description

[0015] Figure 1 This is a system architecture diagram of the fault handling method for a distributed storage system provided in the embodiments of this application; Figure 2 This is a flowchart of a fault handling method for a distributed storage system provided in an embodiment of this application; Figure 3 This is a system architecture diagram of the distributed storage system provided in the embodiments of this application; Figure 4 yes Figure 2 The flowchart of step S202 in the document; Figure 5 yes Figure 2 The flowchart of step S203 in the process; Figure 6 yes Figure 5 The flowchart of step S504 in the process; Figure 7 This is a schematic diagram of fault prediction for different anomaly categories in the fault handling method of the distributed storage system provided in the embodiments of this application; Figure 8 yes Figure 2 The flowchart of step S204 in the process; Figure 9 This is a schematic diagram illustrating the fault repair of different anomaly categories in the fault handling method of the distributed storage system provided in the embodiments of this application; Figure 10 yes Figure 2 The flowchart of step S205 in the document; Figure 11 This is a flowchart of a fault handling method for a distributed storage system provided in another embodiment of this application; Figure 12 This is an overall flowchart of the fault handling method for a distributed storage system provided in the embodiments of this application; Figure 13 This is a schematic diagram of the structure of the fault handling device for the distributed storage system provided in the embodiments of this application; Figure 14 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0017] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0019] First, let's analyze some of the terms used in this application: Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0020] Hadoop Distributed File System (HDFS) distributes data across multiple independent devices. Traditional network storage systems use a centralized data server to store all data, making the data server a bottleneck for system performance and a focal point for reliability and security, failing to meet the needs of large-scale storage applications. Distributed network storage systems employ a scalable system architecture, utilizing multiple data servers to share the storage load and location data servers to locate stored information. This not only improves system reliability, availability, and access efficiency but also facilitates expansion.

[0021] DataNode: The worker node of HDFS, responsible for storing the actual data blocks and periodically reporting the list of its stored blocks and their health status to the NameNode.

[0022] NameNode: The management node of HDFS, which stores the file system's metadata (directory tree, file-block mapping, block location, etc.).

[0023] SMART protocol: a self-monitoring, analysis and reporting technology for hard drives and solid-state drives (SSDs), built into the hard drive firmware, can monitor various indicators to assess the health of the hard drive and predict possible failures.

[0024] The HDFS fsck tool: In the Hadoop Distributed File System (HDFS), the fsck command is a very important tool used to check the data integrity of the file system. Similar to the fsck command in Unix systems, it can detect and repair problems in the file system.

[0025] SMART (Smart Disk Health Data) is data output by a built-in health monitoring system on the hard drive. The function of the health monitoring system is to record the operating status of the hard drive every day, such as the number of bad sectors, power-on time, temperature, etc.

[0026] Time series prediction models are fundamental models designed specifically for time series data analysis. Pre-trained on massive amounts of time series data using techniques such as the Transformer architecture, they can understand and generate time series data across multiple domains and can be applied to time series prediction, anomaly detection, and time series imputation. Unlike traditional time series analysis techniques, large-scale time series models possess general feature extraction capabilities and serve a wide range of analytical tasks based on zero-shot analysis and fine-tuning techniques.

[0027] KAFK is a high-throughput distributed publish-subscribe messaging system.

[0028] FLINK is a framework for real-time data processing tasks. This framework helps developers perform data processing tasks without having to worry about high availability, performance, or other issues.

[0029] Kubernetes (K8s) is a core service provider for containerized applications. Its core functionalities include automated deployment, service discovery, load balancing, and cluster resource management. It enables large-scale application operation and maintenance through the collaboration of Master and Worker nodes and supports Docker and rkt runtime environments.

[0030] The fault handling mechanism of distributed storage systems involves repairing only after a failure occurs. During the repair process, the analytical storage system cannot be used, causing platforms relying on distributed storage systems to malfunction. For example, in fintech scenarios, insurance and banking platforms handle large amounts of financial data stored in distributed storage systems. Repairing a failure in this system can easily disrupt the platform's normal operation. Similarly, in healthcare, large amounts of data such as electronic medical records, electronic examination reports, and electronic laboratory reports rely on distributed storage systems. If these systems cannot respond promptly and retrieve stored medical reports, diagnostic efficiency is affected.

[0031] Taking hardware failure repair in a distributed storage system as an example, hardware failures are typically handled by a replication mechanism, which uses multiple nodes to store data in replicas to achieve hardware failure tolerance. However, uneven replica distribution means that a single point of failure can still cause data loss. Furthermore, replica reconstruction is delayed; when a node fails, the distributed storage system must wait for the replicas to be rebuilt, reducing data availability during the repair period and impacting the normal operation of platforms that rely on the distributed storage system. Therefore, how to predict failures in distributed storage systems in advance and automatically complete repairs has become an urgent technical problem to be solved.

[0032] Based on this, embodiments of this application provide a fault handling method, apparatus, device, and storage medium for a distributed storage system. The aim is to collect system operation data of the distributed storage system and assess its operational status. When the system is determined to be in an abnormal operating state, fault prediction data is obtained in advance based on the system operation data. Then, target repair operations are extracted from candidate repair operations based on the fault prediction data, and these target repair operations are executed on the distributed storage system in advance. Therefore, the monitoring of distributed storage system operation data and fault repair are completed automatically and in advance, reducing the probability of actual failures in the distributed storage system, making its operation more stable, and reducing the occurrence of platforms dependent on the distributed storage system failing to operate.

[0033] The fault handling method, apparatus, device, and storage medium of the distributed storage system provided in this application are specifically described through the following embodiments. First, the fault handling method of the distributed storage system in this application embodiment is described.

[0034] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0035] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0036] The fault handling method for a distributed storage system provided in this application relates to the field of artificial intelligence technology. This fault handling method for a distributed storage system can be applied to a terminal, a data server, or software running on either a terminal or a data server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the data server can be configured as an independent physical data server, a data server cluster or distributed system composed of multiple physical data servers, or a cloud data server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the fault handling method for the distributed storage system, but is not limited to the above forms.

[0037] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, data server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0038] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0039] Figure 1 This is a system architecture diagram of the fault handling method of the distributed storage system in this application embodiment, which includes a distributed storage system and a target server 130, etc.

[0040] The distributed storage system consists of multiple data servers 110 and metadata servers 120. The distributed storage system also includes clients, which serve as terminals for applying the distributed storage system. The metadata server 120 is responsible for managing metadata and processing requests sent by clients, and is the core component of the entire system. The data server 130 is responsible for storing file data and ensuring the availability and integrity of the data.

[0041] Target server 130 is used to handle faults in the distributed storage system. It monitors the system's real-time operational status and predicts fault types and characteristics in advance. Based on these fault types and characteristics, it proactively repairs the distributed system, achieving automatic fault repair. Therefore, proactive fault repair makes the distributed storage system more stable and ensures more stable data storage and retrieval for clients / platforms using the distributed storage system.

[0042] Figure 2 This is an optional flowchart of a fault handling method for a distributed storage system provided in an embodiment of this application. Figure 2 The method may include, but is not limited to, steps S201 to S205.

[0043] Step S201: Collect system operation data of the distributed storage system; Step S202: Evaluate the status of the distributed storage system based on the system operation data to obtain the system operation status; wherein, the system operation status includes: abnormal operation status; Step S203: Based on the abnormal operating state and system operating data, perform fault prediction on the distributed storage system to obtain fault prediction data; wherein, the fault prediction data includes: fault category and fault characteristics; Step S204: Select the target repair operation from the preset candidate repair operations according to the fault category and fault characteristics; Step S205: Perform target repair operations on the distributed storage system according to the fault category.

[0044] Steps S201 to S205 of this embodiment involve automatically collecting system operation data of the distributed storage system and evaluating its status based on the data. If an abnormal state is detected, the fault type and characteristics are predicted, and corresponding repair operations are automatically performed to quickly restore the system from an abnormal to a normal state. Therefore, by implementing automatic status monitoring and fault prediction in the distributed storage system, fault types and characteristics can be predicted in advance, allowing for targeted fault repair and improving operational stability. This avoids reliance on the distributed storage system's service terminal, and the entire repair process is automated, requiring no manual intervention and reducing labor costs associated with fault handling.

[0045] In step S201 of some embodiments, system operation data is key data for evaluating the status of the distributed storage system. This system operation data includes: hardware monitoring parameters, metadata information, and storage service operation parameters. Hardware monitoring parameters are the hardware parameters of the distributed storage system, collected in real-time by the hardware monitoring layer. Metadata information describes the data within the distributed storage system; it determines whether the data has changed and is stored on a metadata server, which can be extracted. Storage service operation parameters represent the operating status of the data servers in the distributed storage system, used to evaluate their current operating status, and determine the status of the storage service within the distributed storage system. Therefore, collecting system operation data involves parameters from three aspects: hardware, storage service, and metadata. Thus, the evaluation of the distributed storage system will also be conducted from these three aspects, improving the comprehensiveness of the distributed storage system status assessment.

[0046] In step S202 of some embodiments, the status assessment of the distributed storage system is also a real-time assessment, updating the system's operating status in real time. The system operating status includes a normal system state and an abnormal system state, with abnormal states corresponding to situations such as slave node failure, master node switching, changes in source data, and changes in storage space. Therefore, monitoring the status of the distributed storage system allows for timely fault prediction and repair when the system's operating state is abnormal, making the distributed storage system more stable.

[0047] Please refer to Figure 3 , Figure 3 A flowchart illustrating the fault handling process of a distributed storage system is shown. Figure 3 It is evident that monitoring the status of a distributed storage system, focusing on hardware status, storage service status, and metadata status, provides a more comprehensive approach to system status monitoring. Simultaneously, hardware status is determined by collecting hardware monitoring parameters, storage service status by collecting storage service operating parameters, and metadata status by collecting metadata information. Specifically, hardware status reflects changes in disk storage space and its lifespan; an abnormal hardware status indicates insufficient disk storage space or an approaching end of life. Storage service status reflects the operational status of the distributed storage system's master or worker nodes; an abnormal storage service status indicates master node failover or worker node failure. Metadata status reflects changes in metadata; an abnormal metadata status indicates changes in metadata.

[0048] Please see Figure 4 In some embodiments, step S202 may include, but is not limited to, steps S401 to S404: Step S401: Evaluate the disk status of the distributed storage system based on hardware monitoring parameters to obtain the disk operating status; Step S402: Evaluate the metadata status of the distributed storage system based on the metadata information to obtain the metadata status; Step S403: Evaluate the service status of the distributed storage system based on the storage service operation parameters to obtain the service operation status; Step S404: Combine the disk running status, metadata status, and service running status to obtain the system running status.

[0049] In step S401 of some embodiments, the hardware monitoring parameters are collected by the hardware monitoring layer, which includes parameters related to the disk's lifespan and capacity. The disk's status is then assessed based on these hardware monitoring parameters, primarily evaluating the remaining lifespan and capacity to determine the disk's operating status. It should be noted that if the disk's operating status is abnormal, indicating that the remaining lifespan is less than a preset lifespan and / or the remaining capacity is less than a preset capacity, this serves as an early warning mode for insufficient disk capacity and / or insufficient lifespan. This further confirms the disk's fault condition and cause, allowing for proactive pre-fault repair and improving disk operational stability.

[0050] Specifically, hardware monitoring parameters include disk health metrics, file system operating parameters, and service-level metrics. Disk health metrics are collected in real time via the SMART protocol, file system operating parameters are collected using the distributed storage system's own HDFS fsck tool, and service-level metrics are collected through integrated hardware. Disk health metrics include read / write error rate, number of remapped sectors, disk temperature, and disk power-on time. It's important to note that disk health metrics are continuously output during disk operation to determine abnormal disk conditions. File system operating parameters include block loss rate and replica distribution information; these determine the disk's fault recovery rate and thus its operational status. Service-level metrics include CPU load and memory usage; these determine the disk's remaining capacity. Therefore, combining multiple parameters for disk operational status assessment allows for timely and accurate determination of the disk's operational status.

[0051] In step S402 of some embodiments, metadata information can determine the changes in data, determine the metadata status of the distributed storage system, and when the metadata status is an abnormal metadata status, it indicates that the metadata has changed and data synchronization is required to ensure data consistency.

[0052] In step S403 of some embodiments, the storage service operation parameters are the operating status of each node in the distributed storage system, specifically the operating status of the data server, to determine the service operation status of the distributed storage system. It should be noted that when the service operation status is in an abnormal state, indicating a master node switchover or worker node failure, further investigation is needed to determine the cause of the fault and promptly repair it.

[0053] In some embodiments, step S404 includes disk operating status, metadata status, and service operating status, and the repair operations performed automatically in abnormal states are different. For example... Figure 3 As shown, when the disk operating status is in an abnormal state, it is necessary to further determine whether the disk has insufficient remaining capacity or insufficient remaining lifespan; if the metadata status is in an abnormal state, it is necessary to further confirm the type of metadata change; if the service operating status is in an abnormal state, it is necessary to further determine whether it is a master node switchover or a worker node failure. Therefore, the subsequent fault analysis process differs for different system operating states.

[0054] In steps S401 to S404 of this embodiment, the system operating status of the distributed storage system is evaluated, mainly focusing on the disk operating status, metadata status, and service operating status of the distributed storage system, so as to achieve a more comprehensive and accurate evaluation of the system operating status.

[0055] In step S203 of some embodiments, if the system is operating normally, then system operation data collection and system operation status monitoring continue. If the system is operating abnormally, it is necessary to further analyze the fault categories and fault characteristics of the distributed storage system, and the fault characteristics are used to assess the factors that caused the fault, so as to achieve more accurate fault repair.

[0056] Please see Figure 5 In some embodiments, step S203 may include, but is not limited to, steps S501 to S504: Step S501: Classify the distributed storage system for anomalies based on its abnormal operating status to obtain anomaly categories; Step S502: Filter out abnormal operation data from the system operation data according to the abnormality category; Step S503: Select reference cases from the preset historical fault repair cases according to the anomaly category; Step S504: Based on reference cases and abnormal operation data, perform fault prediction on the distributed storage system to obtain fault prediction data.

[0057] In step S501 of some embodiments, the exception category characterizes the type of exception that occurs in the distributed storage system, and the exception category includes disk exception category, storage service exception category, and metadata exception category. It should be noted that the exception category is determined based on the operational exception state. If the operational exception state is a disk exception state, the exception category is determined to be the disk exception category; if the operational exception state is a service exception state, the exception category is determined to be the storage service exception category; and if the operational exception state is a metadata exception state, the exception category is determined to be the metadata exception category.

[0058] In step S502 of some embodiments, if the exception category is a disk exception category, the abnormal running data is determined to be a data value in the hardware monitoring parameters that exceeds the preset monitoring parameters; if the exception category is a storage service exception category, the abnormal running data is determined to be a data value in the storage service running parameters that exceeds the preset running parameters; if the exception category is a metadata exception category, the abnormal running data is determined to be metadata information.

[0059] In step S503 of some embodiments, historical fault repair cases are used as reference cases for fault prediction. These historical fault repair cases include historical fault data and historical fault repair operations, with the historical fault data serving as reference data for determining fault categories and fault characteristics. Category labels for historical fault repair cases are obtained, and reference cases are selected from these cases based on the anomaly category and category labels.

[0060] Specifically, if the anomaly category is a disk anomaly, a disk failure repair case is determined as a reference case; if the anomaly category is a storage service anomaly, a storage service failure repair case is determined as a reference case; and if the anomaly category is a metadata anomaly, a metadata failure repair case is determined as a reference case.

[0061] Furthermore, historical fault repair cases are stored in a designated knowledge base, and these historical fault repair cases can also be fault assessment rules, stored in the knowledge base as rule features.

[0062] In step S504 of some embodiments, the distributed storage system is combined with abnormal operation data and reference cases to perform fault prediction, identify potential fault categories and fault characteristics, trigger corresponding fault repair in advance, and improve the stability of the distributed storage system.

[0063] In steps S501 to S504 of this embodiment, when it is determined that an anomaly has occurred in the distributed storage system, the fault category and fault characteristics are predicted by combining the historical fault repair cases and abnormal operation data of the distributed storage system, so as to achieve early fault prediction and improve the accuracy of prediction.

[0064] In some embodiments, if the anomaly category is a disk anomaly category, the abnormal operation data is disk health indicator data and service level indicator data. As disclosed above, the disk health indicator data is used to predict the remaining lifespan of the disk and determine whether the disk needs to be replaced prematurely.

[0065] Please see Figure 6 In some embodiments, step S504 may include, but is not limited to, steps S601 to S604: Step S601: Based on disk health index data, predict the lifespan decay rate to obtain the predicted disk lifespan decay rate. Step S602: Perform capacity prediction based on service level indicator data to obtain the predicted disk capacity value; Step S603: Extract features from the reference case to obtain reference disk failure features; Step S604: Based on the predicted disk lifespan decay rate, predicted disk capacity value, and reference disk failure characteristics, perform disk failure prediction on the distributed storage system to obtain failure prediction data.

[0066] In step S601 of some embodiments, the lifetime degradation rate prediction also predicts the lifetime degradation of the currently used disk and determines at which node the corresponding repair operation should be performed on the disk. In this embodiment, the predicted disk lifetime degradation rate of each disk is identified in advance through a time-series prediction model and disk health indicator data. When the predicted disk lifetime degradation rate is lower than a preset degradation threshold, disk repair is performed in advance.

[0067] It's important to note that disk failure doesn't happen instantaneously, but rather undergoes a gradual process from "healthy" to "sub-healthy" and finally to "failure." During this process, the disk continuously outputs a large amount of time-series disk health indicator data, the most crucial of which are read / write error rate, remapped sector count, disk temperature, and disk power-on time. The trends, fluctuations, and correlations of these parameters over time contain rich information about potential failure symptoms. However, traditional methods often only focus on whether a threshold is exceeded at a specific moment, ignoring the evolutionary patterns over time. For example, a disk's remapped sector count might jump from 0 to 5 in a short period, which is more dangerous than a slow increase to 5. This "abrupt" or "accelerated degradation" pattern is precisely the characteristic that time-series prediction models excel at capturing. Therefore, using a time-series prediction model and disk health indicator data to predict the disk's lifetime decay rate determines that the predicted disk lifetime decay rate is a changing curve that is updated periodically. It should be noted that the time-series prediction model can be an LSTM model, a GRU model, or a Transformer model, etc.; this embodiment does not impose specific limitations.

[0068] In step S602 of some embodiments, as disclosed above, the service-level metric data includes CPU load rate and memory usage rate. The predicted disk capacity value can be directly determined by the CPU load rate and memory usage rate to determine whether the disk capacity is insufficient.

[0069] In step S603 of some embodiments, as disclosed above, if the anomaly category is a disk anomaly category and the reference case is a disk failure repair case, then the features of the disk failure repair case are extracted as reference disk failure features.

[0070] In step S604 of some embodiments, a preset large model is trained with reference to disk failure characteristics and file system state parameters to obtain a failure prediction model. Then, the disk life decay rate and the predicted disk capacity value are input into the failure prediction model to perform failure prediction and obtain failure prediction data, thereby determining the disk failure category and disk failure characteristics.

[0071] In steps S601 to S604 of this embodiment, for disk fault prediction, the disk lifespan decay rate and disk capacity value are first predicted, and then the disk fault category and fault characteristics are determined based on the disk lifespan decay rate, disk capacity value and reference case. This can accurately complete fault prediction, determine in advance whether to repair the disk fault, and make the disk operation stability in the distributed storage system higher.

[0072] like Figure 7 As shown, Figure 7 The document illustrates the fault prediction process corresponding to different anomaly categories. This embodiment of the application collects relevant data from the distributed storage system in three aspects: disk, metadata, and storage services using Kafka / FLINK. Then, it determines the anomaly category. If a disk anomaly occurs, and the disk failure category is determined to be disk lifespan expiration, the fault repair engine is activated. Simultaneously, if the anomaly category is a metadata anomaly, and the metadata failure category is determined to be metadata changes, such as capacity or field changes, the fault repair engine is also triggered. Furthermore, if the failure category is a worker node anomaly, the fault repair engine is also triggered; however, different fault repair operations are required for different fault categories.

[0073] It should be noted that fault prediction for metadata and storage services is a common method, and will not be elaborated upon in this embodiment.

[0074] In step S204 of some embodiments, after determining the fault category and fault characteristics, repair operations are selectively screened based on the fault category and fault characteristics, and the target repair operation is selected from the candidate repair operations to improve the accuracy of the repair.

[0075] Please see Figure 8In some embodiments, step S204 may include, but is not limited to, steps S801 to S804: Step S801: Filter candidate repair operations according to the fault category to obtain preliminary repair operations; Step S802: Obtain reference features and repair success rate for the initial repair operation; Step S803: Calculate the similarity between the reference features and the fault features to obtain the feature similarity. Step S804: Filter the preliminary repair operations based on the repair success rate and feature similarity to obtain the target repair operations.

[0076] In step S801 of some embodiments, the fault categories include: insufficient disk remaining capacity, insufficient disk remaining lifespan, abnormal worker node, master node switching, and metadata changes, etc. Therefore, different preliminary repair operations are determined for different fault categories; among them, candidate repair operations can be determined from historical fault repair cases.

[0077] In step S802 of some embodiments, reference features and repair success rate of the preliminary repair operation are further determined, and the repair success rate is also recorded in historical fault repair cases. The reference features serve as the fault features corresponding to the preliminary repair operation, and the repair success rate characterizes the success rate of repairing the distributed storage system.

[0078] In step S803 of some embodiments, the reference features are converted into reference vectors, the fault features are converted into fault vectors, and the similarity between the reference vectors and fault vectors is calculated to obtain the feature similarity. It should be noted that the fault features are multimodal features, and the reference features are also multimodal features. The fault features are log features in the distributed storage system, including structured indicators (fault codes, timestamps) and unstructured text (fault descriptions). The fault features obtained by log cleaning and vectorization are determined to be key features for screening repair operations.

[0079] In step S804 of some embodiments, the repair success rate is converted into a first score, the feature similarity is converted into a second score, the first score and the second score are weighted and summed to obtain a target score, and the preliminary repair operation with the highest target score is selected as the target repair operation to select the target repair operation with the most suitable repair success rate and repair fit, thereby improving the repair success rate.

[0080] Specifically, if the fault type is that the disk is full or the process has crashed, the initial repair operations are determined to be disk migration, disk expansion and restart. Therefore, the target repair operations can be further filtered based on the fault characteristics and repair success rate, so that the most suitable target repair operations with a high success rate can be selected.

[0081] In steps S801 to S804 of this embodiment, the target repair operation is screened from multiple candidate repair operations by combining the fault category, fault characteristics and repair success rate, and the target repair operation with the most suitable matching degree and repair success rate is selected to improve the repair success rate.

[0082] In some embodiments, in addition to steps S801 to S804, the determination of the target repair operation for the distributed storage system can be achieved by constructing a historical fault knowledge graph based on historical fault repair cases. Specifically, reference categories (such as disk failure, network latency, etc.), reference features (disk health index data, IO timeout frequency, etc.), and candidate repair operations (such as disk replacement, replica reconstruction) are extracted from a large number of historical fault repair cases. These three are grouped into a triple relationship, and then a fault repair screening model is trained using a large model that combines the fault features of historical fault repair cases with the candidate repair operations. When fault repair operation selection begins, the current fault features are matched with the reference features in historical fault repair cases. The fault repair screening model then completes the fault operation screening to output candidate repair operations. Real-time load data (such as disk IO queue depth, node CPU utilization) is introduced as constraints. A rule engine filters the target repair operations from the candidate repair operations (such as prioritizing hot spare disk replacement rather than full data migration in high-load scenarios). Finally, the target repair operation with the highest matching degree and lowest execution risk is output, improving the repair success rate.

[0083] Please refer to Figure 9 , Figure 9 The repair process corresponding to different anomaly categories is shown, through... Figure 9 It can be seen that if the disk reaches the end of its lifespan, disk replacement and automatic data block migration can be used. If the metadata changes, a disaster recovery cluster and a cold data archiving cluster can be set up. If the worker node is abnormal, the worker node application can be migrated and the abnormal worker node POD-1 can be replaced with a healthy worker node POD-N to restore the operation of the distributed storage system and improve the stability of the operation.

[0084] In step S205 of some embodiments, as disclosed above, the distributed storage system is repaired according to the fault category in order to complete the fault repair in a targeted manner and improve the repair success rate.

[0085] Please see Figure 10 In some embodiments, step S205 includes, but is not limited to, steps S1001 to S1003: Step S1001: Select objects to be repaired from the distributed storage system according to the fault category; wherein, the objects to be repaired are any one of the following: disk, storage service node, and metadata; Step S1002: Mark the risk of the object to be repaired according to the fault category to obtain risk marking information; Step S1003: Perform the target repair operation on the object to be repaired based on the risk marking information.

[0086] In step S1001 of some embodiments, as disclosed above, the fault monitoring of the distributed storage system mainly revolves around disks, storage services, and metadata. Therefore, for different fault categories, the object to be repaired is first determined. If the fault category is that the disk is full or the disk lifespan is insufficient, the object to be repaired is the disk that has failed. If the fault category is that the worker node is abnormal, the object to be repaired is the abnormal worker node, which is also the abnormal data server. If the fault category is that the master node has switched, the object to be repaired is the master node before and after the switch. If the fault category is that the metadata has changed, the changed metadata is determined as the object to be repaired.

[0087] In step S1002 of some embodiments, this embodiment sets a hierarchical response mechanism as a fault repair node, and the hierarchical response mechanism sets an early warning stage, a preprocessing stage, and a fault repair stage. Therefore, after determining the object to be repaired and the target repair operation, the operation to be repaired is first risk-marked to obtain risk marking information, so that subsequent fault repair can be completed according to the risk marking information, and multiple risk marking information can be accumulated and then batch repaired to improve repair efficiency. It should be noted that after risk marking, the risk marking information is recorded in the log and the operation and maintenance system is notified.

[0088] In step S1003 of some embodiments, the execution of the target repair operation is a preprocessing node and a fault repair stage. Therefore, the preprocessing operation and fault repair operation of the target repair operation are determined. The preprocessing operation generally involves migrating the data blocks in the object to be repaired with risk marking information to a data server or other disk without risk marking, and then performing the fault repair operation on the object to be repaired to reduce the possibility of data loss.

[0089] For example, if the object to be repaired is a disk, the warning phase involves marking the disk as high-risk, logging the information, and notifying the operations and maintenance system. The preprocessing phase automatically migrates data blocks from the high-risk disk to a safe disk. The fault repair phase involves triggering a replica rebuild when the disk is offline, automatically triggering a hardware replacement work order to replace the disk. Therefore, completing data migration before repair reduces data loss and makes fault repair safer.

[0090] In steps S1001 to S1003 of this embodiment, when performing the target repair operation, the object to be repaired is first determined, then the object to be repaired is marked with a risk before the target repair operation is performed, so as to ensure that the target repair operation is executed accurately and reduce the repair error rate.

[0091] Please see Figure 11 In some embodiments, after step S205, the fault handling method for the distributed storage system may also include, but is not limited to, steps S1101 to S1102: Step S1101: Send the target repair operation to the target terminal and receive the repair feedback information from the target terminal based on the target repair operation; Step S1102: Based on the repair feedback information, construct cases by combining the fault category, fault characteristics and target repair operation to obtain historical fault repair cases.

[0092] In step S1101 of some embodiments, as disclosed above, the determination of the fault category and the repair operation both depend on historical fault repair cases. Therefore, after the target repair operation is completed, repair feedback information is received from the target terminal after performing the target repair operation on the distributed storage system. The repair feedback information includes repair success rate and repair suggestion information.

[0093] In step S1102 of some embodiments, historical fault repair cases are constructed based on the fault category, fault characteristics, and target repair operation according to the repair feedback information. These cases serve as a reference for subsequent fault repairs and are also stored in a knowledge base in a structured format. Furthermore, expanding the historical fault repair cases in the knowledge base allows for retraining of the fault prediction model and the fault repair screening model using new historical fault repair cases, thereby improving the model's accuracy.

[0094] In steps S1101 to S1102 of this embodiment, after each target repair operation is completed, the fault category, fault characteristics, and target repair operation are constructed into historical fault repair cases according to the repair feedback information. This expands the historical fault repair cases in the knowledge base, which can serve as reference data for subsequent fault handling. Furthermore, the model can be retrained periodically with new historical fault repair cases (e.g., updating the LSTM model monthly). The accuracy of different rules / models is verified through A / B testing.

[0095] Please refer to Figure 12This application's embodiments address fault handling in distributed storage systems. First, hardware monitoring parameters and service logs for each data server and metadata server in the distributed storage system are collected in real-time using Kafka / Flink. The disk's operating status is determined through hardware monitoring data. Log features are extracted by analyzing service logs, and the service operating status is determined based on storage service metrics data within these log features. The metadata status is determined based on metadata information within the log features. If the disk's operating status is determined to be abnormal, the disk's fault type and characteristics need further confirmation. Specifically, historical fault repair cases are extracted from the knowledge base, and the fault type and characteristics are jointly predicted by combining these historical fault repair cases with the current abnormal disk health metrics, service-level metrics, and file system status parameters. Fault characteristics can be used as factors contributing to the fault's occurrence, and the target repair operation can be selected from candidate repair operations in historical fault repair cases, in conjunction with the fault type. If the target repair operation is disk migration, Kubernetes is triggered to schedule the disk migration. Kubernetes automatically restarts the faulty disk or migrates the task to a healthy disk. Furthermore, if changes in metadata capacity or fields are detected, a disaster recovery checkpoint is triggered, automatically synchronizing data to the disaster recovery cluster.

[0096] It should be noted that different target repair operations are performed for different types of faults: 1. If the fault self-healing engine detects that the disk is about to expire, it triggers the automatic recovery module to migrate the data blocks to a healthy disk; 2. If a worker node (also known as a data server) is detected to be down, Kubernetes will relocate the worker node to another node to restore service; 3. If a change in the capacity of the core database directory path is detected, data synchronization to the disaster recovery cluster will be automatically triggered, and disaster recovery data that has not been used for more than six months will be automatically transferred to low-cost cold data storage.

[0097] In summary, technologies such as intelligent predictive maintenance, real-time monitoring and automated scheduling, and data synchronization optimization significantly improve the fault recovery capability and stability of distributed storage systems. This solves the problems of delayed fault response and resource waste in disaster recovery of traditional distributed storage systems, while significantly reducing operation and maintenance costs and business interruption risks, making it more suitable for large-scale, high-availability scenarios.

[0098] Please see Figure 13 This application also provides a fault handling apparatus for a distributed storage system, which can implement the above-described fault handling method for a distributed storage system. The apparatus includes: Data acquisition module 1301 is used to collect system operation data of the distributed storage system; The status assessment module 1302 is used to assess the status of the distributed storage system based on system operation data to obtain the system operation status; wherein, the system operation status includes: abnormal operation status; The fault prediction module 1303 is used to predict faults in the distributed storage system based on abnormal operating conditions and system operating data, and obtain fault prediction data; wherein, the fault prediction data includes: fault category and fault characteristics; The operation filtering module 1304 is used to filter the target repair operation from the preset candidate repair operations according to the fault category and fault characteristics. Repair execution module 1305 is used to perform target repair operations on the distributed storage system according to the fault category.

[0099] The specific implementation of the fault handling device for the distributed storage system is basically the same as the specific implementation of the fault handling method for the distributed storage system described above, and will not be repeated here.

[0100] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the fault handling method of the distributed storage system described above. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0101] Please see Figure 14 , Figure 14 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1401 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1402 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1402 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1402 and is called and executed by the processor 1401 using the fault handling method of the distributed storage system of the embodiments of this application. The input / output interface 1403 is used to implement information input and output; The communication interface 1404 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1405 transmits information between various components of the device (e.g., processor 1401, memory 1402, input / output interface 1403, and communication interface 1404); The processor 1401, memory 1402, input / output interface 1403 and communication interface 1404 are connected to each other within the device via bus 1405.

[0102] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the fault handling method of the distributed storage system described above.

[0103] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0104] The fault handling method, apparatus, device, and storage medium for distributed storage systems provided in this application automatically predicts faults in advance when an anomaly occurs in the distributed storage system. This determines the fault type and characteristics, selects a target repair operation from candidate repair operations based on the fault type and characteristics, and then executes the repair operation on the distributed storage system according to the target repair operation. Therefore, automatically and proactively predicting and repairing faults in the distributed storage system saves manpower costs for fault repair and reduces the failure rate of the distributed storage system, making the distributed storage system more stable.

[0105] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0106] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0107] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0108] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0109] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0110] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0111] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0112] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0113] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0114] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, a data server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0115] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A fault handling method for a distributed storage system, characterized in that, The method includes: Collect system operation data from the distributed storage system; The distributed storage system is evaluated based on the system operation data to obtain the system operation status; wherein, the system operation status includes: abnormal operation status; Based on the abnormal operating state and the system operating data, fault prediction is performed on the distributed storage system to obtain fault prediction data; wherein, the fault prediction data includes: fault category and fault characteristics; The target repair operation is selected from the preset candidate repair operations based on the fault category and the fault characteristics. Perform the target repair operation on the distributed storage system according to the fault category.

2. The method according to claim 1, characterized in that, The system operation data includes: hardware monitoring parameters, metadata information, and storage service operation parameters; The step of evaluating the status of the distributed storage system based on the system operation data to obtain the system operation status includes: The disk status of the distributed storage system is evaluated based on the hardware monitoring parameters to obtain the disk operating status. The distributed storage system is evaluated based on the metadata information to obtain the metadata status. The service status of the distributed storage system is evaluated based on the storage service operating parameters to obtain the service operating status. The system operating status is obtained by concatenating the disk operating status, the metadata status, and the service operating status.

3. The method according to claim 2, characterized in that, The step of predicting faults in the distributed storage system based on the abnormal operating state and the system operating data to obtain fault prediction data includes: Based on the abnormal operating state, the distributed storage system is classified into abnormal categories to obtain the abnormality category; Abnormal operation data is filtered out from the system operation data according to the abnormality category; Reference cases are selected from preset historical fault repair cases based on the anomaly category; Based on the reference case and the abnormal operation data, the distributed storage system is used to predict faults, and the fault prediction data is obtained.

4. The method according to claim 3, characterized in that, If the anomaly category is a disk anomaly category, the abnormal operation data is disk health indicator data and service level indicator data; the step of performing fault prediction on the distributed storage system based on the reference case and the abnormal operation data to obtain the fault prediction data includes: Based on the disk health index data, the lifespan decay rate is predicted to obtain the predicted disk lifespan decay rate. Based on the service level metric data, capacity prediction is performed to obtain the predicted disk capacity value; Feature extraction is performed on the reference case to obtain reference disk failure features; The fault prediction data is obtained by predicting the disk failure of the distributed storage system based on the predicted disk lifespan decay rate, the predicted disk capacity value, and the reference disk failure characteristics.

5. The method according to any one of claims 1 to 4, characterized in that, The step of selecting the target repair operation from the preset candidate repair operations based on the fault category and the fault characteristics includes: The candidate repair operations are filtered according to the fault category to obtain preliminary repair operations; Obtain the reference features and repair success rate of the preliminary repair operation; The similarity between the reference feature and the fault feature is calculated to obtain the feature similarity. The initial repair operations are filtered based on the repair success rate and the feature similarity to obtain the target repair operations.

6. The method according to any one of claims 1 to 4, characterized in that, The step of performing the target repair operation on the distributed storage system according to the fault category includes: Based on the fault category, objects to be repaired are selected from the distributed storage system; wherein, the objects to be repaired are any one of the following: disks, storage service nodes, and metadata; The object to be repaired is risk-marked according to the fault category to obtain risk-marking information; The target repair operation is performed on the object to be repaired based on the risk marking information.

7. The method according to claim 3, characterized in that, After performing the target repair operation on the distributed storage system according to the fault category, the method further includes: Send the target repair operation to the target terminal and receive the repair feedback information from the target terminal based on the target repair operation; Based on the repair feedback information, the fault category, the fault characteristics, and the target repair operation are used to construct cases to obtain the historical fault repair cases.

8. A fault repair device for a distributed storage system, characterized in that, The device includes: The data acquisition module is used to collect system operation data from the distributed storage system. The status assessment module is used to assess the status of the distributed storage system based on the system operation data to obtain the system operation status; wherein, the system operation status includes: abnormal operation status; The fault prediction module is used to predict faults in the distributed storage system based on the abnormal operating state and the system operating data, and obtain fault prediction data; wherein, the fault prediction data includes: fault category and fault characteristics; The operation filtering module is used to filter out target repair operations from preset candidate repair operations based on the fault category and the fault characteristics. The repair execution module is used to perform the target repair operation on the distributed storage system according to the fault category.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the fault handling method of the distributed storage system according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the fault handling method of the distributed storage system as described in any one of claims 1 to 7.