System anomaly detection method and device, storage medium and electronic equipment

Through dynamic adaptation of detection time and multi-level detection strategies, abnormal detection is carried out on the distributed file storage system, solving the problem of low abnormal detection efficiency, and achieving rapid fault location and reducing service interruption time.

CN120029805APending Publication Date: 2025-05-23CHINA CONSTRUCTION BANK

Patent Information

Application Number
CN202510079332.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The abnormality detection efficiency of distributed file storage systems is low, the traditional methods are time-consuming and lack of automated fault isolation mechanisms, making it difficult to meet the needs of real-time monitoring and rapid recovery in massive server environments.

Method used

Using dynamic adaptation detection time and targeted multi-level detection strategy, the service type of business services provided by the distributed file storage system is obtained, the detection time is determined according to the service type, and specific detection operations are performed on the access layer, the first control layer and the second control layer, and a target prompt message is generated to indicate the fault type.

Benefits of technology

It realizes rapid acquisition of health status data at each layer, rapid location of failures, reduce service interruption time, and significantly improves the operation and maintenance efficiency and stability of large-scale storage clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029805A_ABST
    Figure CN120029805A_ABST
Patent Text Reader

Abstract

The invention discloses a system anomaly detection method and device, a storage medium and electronic equipment. The method comprises the following steps: acquiring a service type of a business service provided by a distributed file storage system, and determining a detection moment according to the service type; executing a first detection operation on the access layer, executing a second detection operation on the first management and control layer and executing a third detection operation on the second management and control layer at the detection moment to obtain detection result data; under the condition that the detection result data indicates that the distributed file storage system is abnormal, a target prompt message is generated, and the target prompt message is used for indicating a fault type causing the distributed file storage system to be abnormal. According to the method and the device, the technical problem that the distributed file storage system is relatively low in exception troubleshooting efficiency is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to a method and device for detecting anomalies in a system, a storage medium, and an electronic device. Background Art

[0002] As the scale of distributed file storage systems continues to expand, the complexity of storage clusters has also increased. Traditional anomaly detection methods, such as those that rely on script execution of numerous check items, are not only time-consuming, but also lack an automated fault isolation mechanism. Once a fault is detected, such as process anomalies, service unavailability, or hardware failure, manual intervention is still required. This is obviously inefficient in a massive server environment and is difficult to meet the needs of real-time monitoring and rapid recovery. Therefore, there is a technical problem in the related technology that the efficiency of anomaly detection in distributed file storage systems is low.

[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0004] The embodiments of the present application provide a system anomaly detection method and device, a storage medium, and an electronic device to at least solve the technical problem of low anomaly troubleshooting efficiency in a distributed file storage system.

[0005] According to one aspect of an embodiment of the present application, a method for detecting anomalies in a system is provided, comprising: obtaining a service type of a business service provided by a distributed file storage system, and determining a detection time according to the service type; performing a first detection operation on an access layer, performing a second detection operation on a first control layer, and performing a third detection operation on a second control layer at the detection time to obtain detection result data, wherein the distributed storage file system comprises an access layer, a first control layer, and a second control layer, the access layer is used to receive a front-end request and forward the front-end request to the first control layer, the first control layer is used to generate a data file according to the front-end request, and coordinate the second control layer to work, and the second control layer The layer is used to perform read, write and storage operations on the data file, the first detection operation, the second detection operation and the third detection operation are different from each other, the first detection operation includes detecting the task execution, port occupancy and process survival status of the access layer, the second detection operation includes detecting the configuration data, process survival status and business log of the first management layer, and the third detection operation includes detecting the disk usage, port occupancy and process survival status of the second management layer; when the detection result data indicates that there is an abnormality in the distributed file storage system, a target prompt message is generated, wherein the target prompt message is used to indicate the type of fault that causes the abnormality of the distributed file storage system.

[0006] According to another aspect of an embodiment of the present application, a system anomaly detection device is also provided, including: an acquisition module, used to obtain the service type of a business service provided by a distributed file storage system, and determine a detection time according to the service type; an execution module, used to perform a first detection operation on the access layer at the detection time, perform a second detection operation on the first control layer, and perform a third detection operation on the second control layer to obtain detection result data, wherein the distributed storage file system includes an access layer, a first control layer, and a second control layer, the access layer is used to receive a front-end request and forward the front-end request to the first control layer, the first control layer is used to generate a data file according to the front-end request, and coordinate the second control layer to work, the The second management and control layer is used to perform read, write and storage operations on the data files. The first detection operation, the second detection operation and the third detection operation are different from each other. The first detection operation includes detecting the task execution, port occupancy and process survival status of the access layer. The second detection operation includes detecting the configuration data, process survival status and business log of the first management and control layer. The third detection operation includes detecting the disk usage, port occupancy and process survival status of the second management and control layer. A generation module is used to generate a target prompt message when the detection result data indicates that there is an abnormality in the distributed file storage system, wherein the target prompt message is used to indicate the type of fault that causes the abnormality of the distributed file storage system.

[0007] Optionally, the device is used to generate a target prompt message when the detection result data indicates that there is an abnormality in the distributed file storage system in the following manner: when the detection result data indicates that there is an abnormality in the distributed file storage system, compare the detection result data and the data threshold to determine the abnormality detection data at the system level, wherein the detection result data includes the abnormality detection data, and the system level includes at least one of the access layer, the first management layer and the second management layer; generate the target prompt message based on at least one of the load conditions, memory utilization, disk space utilization, process status, and port status at the system level; determine the target fault type according to the target prompt information, and perform a fault repair operation based on the target fault type.

[0008] Optionally, the device is used to determine the fault type according to the target prompt information in the following manner, and perform the fault repair operation based on the fault type: when the target prompt information indicates a process crash, detect the running process at the system level, and when no heartbeat signal of the running process at the system level is received within a first preset time period, and / or when the resource occupancy change rate is 0, determine the target fault type as a process crash type; when the target prompt information indicates that a system port is unavailable, send a probe data packet through the system port at the system level to obtain a response time and a data packet loss rate of the probe data packet, and when the response time of the probe data packet is greater than or equal to a probe time threshold, and / or when the data packet loss rate is greater than or equal to a loss rate threshold, determine the target fault type as a port unavailable type; when the target prompt information indicates a disk failure, obtain the usage status, read and write error rate and disk response time of the disk at the system level, and when the read and write error rate continues to increase within a second preset time period, and / or when the disk response time continues to increase, determine the target fault type as a disk failure type.

[0009] Optionally, the device is also used for at least one of the following: when the target fault type is determined to be the process crash type, executing a system restart instruction; when the target fault type is determined to be the port unavailable type, executing the system restart instruction; when the target fault type is determined to be the process crash type, starting a repair script, wherein the repair script is used to adjust the system level; when the target fault type is determined to be the port unavailable type, starting the repair script.

[0010] Optionally, the device is also used for at least one of the following: when the target failure type is determined to be a disk failure type and the faulty disk has logical bad sectors, using a disk repair tool to scan and repair the faulty disk, wherein the logical bad sectors indicate that the faulty disk has data corruption; when the target failure type is determined to be a disk failure type and the faulty disk has physical bad sectors, migrating the data files to a healthy disk and marking or replacing the faulty disk, wherein the physical bad sectors indicate that the faulty disk has physical damage.

[0011] Optionally, the device is also used to: when the detection result data indicates that there is an abnormality in the distributed file storage system, remove the faulty node from the load balancing list and start the backup node, wherein the nodes in the load balancing list are used to process the front-end request, the faulty node represents a node that is identified as unable to provide services normally, and the backup node is used to take over the work of the faulty node.

[0012] Optionally, the device is also used to: respectively detect the number of central processing unit cores, memory, disk size and partition, operating system version, clock synchronization, domain name resolution service, network card status, scheduled tasks, system logs, and third-party software library configuration data of the access layer, the first control layer, and the second control layer.

[0013] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned system anomaly detection method when running.

[0014] According to another aspect of the embodiment of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the anomaly detection method of the above system.

[0015] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-mentioned system abnormality detection method through the computer program.

[0016] In the embodiment of the present application, the service type of the business service provided by the distributed file storage system is obtained, and the detection time is determined according to the service type. This strategy can perform timely health checks based on the characteristics of different services, avoiding the waste of resources and inspection blind spots that may be caused by unified inspection time points. At the detection time, specific detection operations are performed on different layers within the system: the first detection operation is performed on the access layer, focusing on its task execution efficiency, port occupancy and the survival status of the core process; the second detection operation is performed on the first management layer, focusing on checking the correctness of the configuration data, the running status of key processes and whether there are abnormal records in the business log; the third detection operation is performed on the second management layer, in-depth detection of disk usage, port monitoring status and process operation closely related to data reading and writing. These detection operations are independent of each other and are designed according to the characteristics of different levels to ensure the comprehensiveness and pertinence of the inspection.

[0017] Through this series of tests, the health status data of each layer can be quickly obtained. Once the test result data indicates that there is an abnormality in the distributed file storage system, such as process abnormality, service unavailability or hardware failure, a target prompt message is immediately generated. The message accurately indicates the type of fault that caused the system abnormality, whether it is a communication failure in the access layer, a configuration error in the first management layer, or a disk problem in the second management layer. In general, through dynamic adaptation of detection time and targeted multi-level detection strategies, the purpose of improving the efficiency of abnormality troubleshooting is achieved, thereby achieving the technical effect of quickly locating faults and reducing service interruption time, thereby effectively solving the technical problem of low abnormality troubleshooting efficiency in distributed file storage systems, and significantly improving the operation and maintenance efficiency and stability of large-scale storage clusters. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0019] Figure 1 is a schematic diagram of an application environment of an optional system anomaly detection method according to an embodiment of the present application;

[0020] Figure 2 is a flow chart of an optional method for detecting abnormalities in a system according to an embodiment of the present application;

[0021] Figure 3 is a schematic diagram of an optional system anomaly detection method according to an embodiment of the present application;

[0022] Figure 4 is a schematic structural diagram of an optional system abnormality detection device according to an embodiment of the present application;

[0023] Figure 5 is a schematic structural diagram of an optional system anomaly detection product according to an embodiment of the present application;

[0024] Figure 6 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] In order to more clearly understand the technical solution provided by the embodiments of the present application, the key terms involved in the embodiments of the present application are first introduced here:

[0028] Distributed file storage: Distributed file storage provides a scalable shared file storage service that can be used with cloud virtual machines and other services. It provides a standard NFS file system access protocol, provides a shared data source for multiple cloud virtual machines or other computing services, supports elastic capacity and performance expansion, and can be mounted and used without modification for existing applications. It features high availability and high reliability, and is suitable for a variety of scenarios such as big data analysis, media processing, and content management.

[0029] NFS: Net File System, a network protocol that implements file sharing over the network. It allows computers in the network to share resources. In NFS applications, local NFS client applications can transparently read and write files on remote NFS servers, just like accessing local files.

[0030] Nas_Agent: Jibiao control component, the second control layer mentioned above, the production component of the file storage core. Jibiao provides core services for the distributed file system on it, and cooperates with the Jibiao pre-installed plug-in installed when the component is installed. It is responsible for managing the core service startup on Jibiao and mounting the user's distributed file system.

[0031] Process: A running program instance. The distributed file storage cluster is a software-defined cloud storage product. The core form of providing services is the software process deployed on the server. Multiple instances communicate with each other through the network.

[0032] Dial test: A test method that calls cloud computing product functions from the caller's perspective to check the service status of cloud computing products.

[0033] The present application is described below in conjunction with embodiments:

[0034] According to one aspect of an embodiment of the present application, a method for detecting anomalies of a system is provided. Optionally, in this embodiment, the method for detecting anomalies of the system can be applied to: Figure 1 In the hardware environment composed of the server 101 and the terminal device 103 shown in FIG. Figure 1 As shown, the server 101 is connected to the terminal device 103 via a network, and can be used to provide services for the terminal device or an application 107 installed on the terminal device. The application can be a video application, an instant messaging application, a browser application, an educational application, a game application, etc. A database 105 may be set up on the server or independently of the server to provide data storage services for the server 101, for example, a game data storage server. The above-mentioned network may include, but is not limited to, a wired network and a wireless network, wherein the wired network includes, a local area network, a metropolitan area network and a wide area network; the wireless network includes, Bluetooth, WIFI and other networks that implement wireless communication; the terminal device 103 may be a terminal configured with an application, and may include, but is not limited to, at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop computer, a tablet computer, a PDA, a MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a mixed reality (MR) terminal and other computer devices; the above-mentioned server may be a single server, a server cluster consisting of multiple servers, or a cloud server.

[0035] Combination Figure 1 As shown, the anomaly detection method of the above system can be executed by an electronic device, which can be a terminal device or a server. The anomaly detection method of the above system can be implemented by the terminal device or the server respectively, or by the terminal device and the server together.

[0036] The above is only an example and is not specifically limited in this embodiment.

[0037] Optionally, as an optional implementation, as Figure 2 As shown, the anomaly detection method of the above system includes:

[0038] S202, obtaining a service type of a business service provided by the distributed file storage system, and determining a detection time according to the service type;

[0039] Optionally, in an embodiment of the present application, the above detection time refers to the time point for performing a health check, which can be adjusted according to different service types. For example, for services that are used frequently and have a significant impact on system stability, such as file reading and writing, the detection time can be set more frequently; while for services that are used less frequently, such as creating or deleting file systems, the detection time can be appropriately extended. At the same time, the detection time can also take into account the peak and trough of system usage to avoid resource-intensive detection operations during high load periods.

[0040] It should be noted that the method for setting the detection time is not limited to this, and can also be adjusted according to specific application scenarios and operation and maintenance strategies, such as dynamically adjusting the detection time based on historical data and AI prediction technology, or performing instant detection when triggered by a specific event. This application does not limit this and aims to provide a flexible, efficient and adaptable health check mechanism.

[0041] It should also be noted that the specific layering and composition of each layer of the distributed file storage system may vary depending on the specific implementation. This application is not limited to the above-mentioned layered structure. It aims to improve the efficiency of abnormality detection and enhance the stability and reliability of the system through layered health checks. This application does not limit this.

[0042] S204, at the detection time, performing a first detection operation on the access layer, performing a second detection operation on the first control layer, and performing a third detection operation on the second control layer to obtain detection result data, wherein the distributed storage file system includes an access layer, a first control layer, and a second control layer, the access layer is used to receive a front-end request and forward the front-end request to the first control layer, the first control layer is used to generate a data file according to the front-end request, and coordinate the second control layer to work, the second control layer is used to perform read, write and storage operations on the data file, the first detection operation, the second detection operation and the third detection operation are different from each other, the first detection operation includes detecting the task execution, port occupancy and process survival status of the access layer, the second detection operation includes detecting the configuration data, process survival status and business log of the first control layer, and the third detection operation includes detecting the disk usage, port occupancy and process survival status of the second control layer;

[0043] Optionally, in an embodiment of the present application, the access layer refers to a stateless service layer in the system responsible for receiving front-end user requests and forwarding them to back-end processing, including but not limited to the Nas_Access component, whose main function is to achieve rapid distribution and processing of requests.

[0044] Optionally, in an embodiment of the present application, the first management and control layer refers to the part of the system that performs core logic processing and data management, including but not limited to the Nas_Master component, which is responsible for the generation of data files, coordination of business processes and management of resource allocation.

[0045] Optionally, in an embodiment of the present application, the second management and control layer refers to the storage layer in the system that directly interacts with data, including but not limited to Nas_Agent and storage nodes, whose main task is to perform reading, writing and storage operations of data files.

[0046] It should be noted that the specific content and execution method of the detection operation may vary depending on the actual deployment environment and technical architecture. For example, the first detection operation may also include the detection of network latency, the second detection operation may also check the data consistency status, and the third detection operation may also involve the evaluation of storage redundancy.

[0047] In addition, with the development of technology, more detection items may be added in the future, such as security audits and performance bottleneck analysis. These are all part of the health check. This application does not limit this and aims to provide a flexible and scalable health check framework to meet the needs of different distributed file storage systems.

[0048] S206: When the detection result data indicates that the distributed file storage system is abnormal, generate a target prompt message, wherein the target prompt message is used to indicate the type of fault that causes the distributed file storage system to be abnormal.

[0049] Optionally, in an embodiment of the present application, the detection result data refers to a set of information about the health status of the access layer, the first control layer and the second control layer collected after executing the first detection operation, the second detection operation and the third detection operation, including but not limited to CPU load, memory usage, disk space, network port status, process status, configuration data accuracy, business log abnormal records, disk health status, etc.

[0050] It should be noted that the content and form of the target prompt message may vary according to different application scenarios and system requirements.

[0051] For example, the prompt message can include the specific location information of the fault, such as server ID or process ID, and the severity level of the fault, ranging from warning to serious error; at the same time, the message may also include preliminary analysis results of the fault and recommended processing steps to help operation and maintenance personnel respond and resolve faster.

[0052] In addition, the generation frequency of the target prompt message, the sending method (such as email, SMS, system notification) and the level of detail of the message are also configurable. This application does not limit this. The purpose is to provide an efficient and flexible exception notification mechanism to adapt to the operation and maintenance needs of various distributed file storage systems.

[0053] In an exemplary embodiment, it is assumed that a distributed file storage system of an enterprise is providing large-scale data storage services, including frequent read and write operations, file creation and deletion, and permission group management and other business services. According to these service types, the detection time is first determined: for read and write operations, because they are frequent and critical, the detection time is set to once every five minutes; for file creation and deletion services, because they are relatively infrequent, the detection time is set to once an hour; and for permission group management services, because they have low requirements for system real-time performance, the detection time is set to once a day.

[0054] When the detection time arrives, the system automatically starts the detection operation for each layer. For the access layer (Nas_Access), by performing the first detection operation, the efficiency of task execution, the occupancy of key communication ports (such as port 22000), and the survival status of core processes (such as task scheduling processes) were checked. For the first management layer (Nas_Master), the second detection operation was performed, which included checking the accuracy and completeness of the configuration data, monitoring the running status of key processes (such as data processing and coordination processes), and analyzing business logs to find potential abnormal records. For the second management layer (Nas_Agent and storage node Cell), the third detection operation was performed, focusing on the usage of data disks (such as disk space utilization, read and write performance), the listening status of ports (especially ports related to data reading and writing), and the status of core processes responsible for data reading, writing and storage (such as nfsd, rpcbind, etc.).

[0055] Furthermore, after the above-mentioned first, second, and third detection operations are completed, the system collects all the detection result data and analyzes them. If the detection result data indicates that there is an abnormality in the distributed file storage system, for example, it is found that the disk space of a server in the second management layer is close to full or the process on Nas_Agent is abnormal, the system will generate a target prompt message. This message not only clearly indicates the type of fault, such as "insufficient disk space" or "core process abnormality", but can also include specific server identification, detailed information on the abnormal process, and recommended troubleshooting paths to help operation and maintenance personnel quickly locate and solve problems, thereby reducing the time for abnormality troubleshooting and improving system stability and operation and maintenance efficiency.

[0056] In the embodiment of the present application, the service type of the business service provided by the distributed file storage system is obtained, and the detection time is determined according to the service type. This strategy can perform timely health checks based on the characteristics of different services, avoiding the waste of resources and inspection blind spots that may be caused by unified inspection time points. At the detection time, specific detection operations are performed on different layers within the system: the first detection operation is performed on the access layer, focusing on its task execution efficiency, port occupancy and the survival status of the core process; the second detection operation is performed on the first management layer, focusing on checking the correctness of the configuration data, the running status of key processes and whether there are abnormal records in the business log; the third detection operation is performed on the second management layer, in-depth detection of disk usage, port monitoring status and process operation closely related to data reading and writing. These detection operations are independent of each other and are designed according to the characteristics of different levels to ensure the comprehensiveness and pertinence of the inspection.

[0057] Through this series of tests, the health status data of each layer can be quickly obtained. Once the test result data indicates that there is an abnormality in the distributed file storage system, such as process abnormality, service unavailability or hardware failure, a target prompt message is immediately generated. The message accurately indicates the type of fault that caused the system abnormality, whether it is a communication failure in the access layer, a configuration error in the first management layer, or a disk problem in the second management layer. In general, through dynamic adaptation of detection time and targeted multi-level detection strategies, the purpose of improving the efficiency of abnormality troubleshooting is achieved, thereby achieving the technical effect of quickly locating faults and reducing service interruption time, thereby effectively solving the technical problem of low abnormality troubleshooting efficiency in distributed file storage systems, and significantly improving the operation and maintenance efficiency and stability of large-scale storage clusters.

[0058] As an optional scheme, when the above-mentioned detection result data indicates that the above-mentioned distributed file storage system has an abnormality, a target prompt message is generated, including: when the above-mentioned detection result data indicates that the above-mentioned distributed file storage system has an abnormality, the above-mentioned detection result data and the data threshold are compared to determine the abnormality detection data of the system level, wherein the above-mentioned detection result data includes the above-mentioned abnormality detection data, and the above-mentioned system level includes at least one of the above-mentioned access layer, the first management layer and the second management layer; the above-mentioned target prompt message is generated based on at least one of the load situation, memory utilization, disk space utilization, process status, and port status of the above-mentioned system level; the target fault type is determined according to the above-mentioned target prompt information, and a fault repair operation is performed based on the above-mentioned target fault type.

[0059] Optionally, in an embodiment of the present application, the data threshold refers to the upper or lower limit of the normal range set for each detection item, for example, the CPU load is higher than 80%, the disk space utilization rate exceeds 90%, the memory usage exceeds 70% of the total memory, the specific port is not in the listening state or the process is not running, etc. The system-level abnormal detection data is the indicator data that exceeds the normal range determined in the detection result data by comparing with the data threshold. The target fault type refers to the specific fault type that may cause system abnormality and is classified and identified according to the specific circumstances of the abnormal detection data.

[0060] It should be noted that the system-level status information considered when generating the target prompt message can be set to be different depending on the specific implementation and monitoring requirements of the system.

[0061] For example, prompt messages at the access layer may focus on port status and process status, while prompt messages at the second management layer may focus more on disk space utilization and process status.

[0062] Similarly, the manner and specific steps of performing fault repair operations may also vary depending on the fault type and system architecture, including but not limited to automatically restarting abnormal processes, adjusting system configuration, clearing disk space, or manual intervention for more complex troubleshooting and repair, which is not limited in this application.

[0063] In an exemplary embodiment, the system found in the detection result data collected at the time of detection that the disk space utilization rate of a server in the second management layer (Nas_Agent) reached 95%, exceeding the preset data threshold of 90%. Based on this abnormal detection data, the system combines the server's load, memory usage and other information to generate a target prompt message, which describes in detail the insufficient disk space, the data services that may be affected, and the recommended treatment measures. Based on this target prompt message, the operation and maintenance personnel quickly determined that the target fault type was "insufficient disk space", and took actions to clean up the cache and release disk space, successfully avoiding the interruption of data services.

[0064] Through the embodiments of the present application, a dynamic detection time and a layered health check mechanism are adopted to achieve refined monitoring of the status of components at each layer in the distributed file storage system, so as to timely discover and locate system anomalies, thereby effectively improving operation and maintenance efficiency and system stability.

[0065] As an optional solution, the above-mentioned determining the fault type according to the above-mentioned target prompt information and performing the above-mentioned fault repair operation based on the above-mentioned fault type include: when the above-mentioned target prompt information indicates a process crash, detecting the running process of the above-mentioned system level, and determining the above-mentioned target fault type as a process crash type when the heartbeat signal of the running process of the above-mentioned system level is not received within a first preset time period, and / or when the resource occupancy change rate is 0; when the above-mentioned target prompt information indicates that the system port is unavailable, sending a detection data packet through the system port of the above-mentioned system level to obtain the response time and data packet loss rate of the above-mentioned detection data packet, and when the response time of the above-mentioned detection data packet is greater than or equal to the detection time threshold, and / or when the above-mentioned data packet loss rate is greater than or equal to the loss rate threshold, determining the above-mentioned target fault type as a port unavailable type; when the above-mentioned target prompt information indicates a disk failure, obtaining the usage status, read and write error rate and disk response time of the disk at the above-mentioned system level, and when the above-mentioned read and write error rate continues to increase within a second preset time period, and / or when the above-mentioned disk response time continues to increase, determining the above-mentioned target fault type as a disk failure type.

[0066] Optionally, in the embodiment of the present application, the heartbeat signal refers to the self-status report periodically sent by the running process in the system level to the monitoring system, which is used to indicate the survival status and basic operation status of the process. The resource occupancy change rate refers to the ratio of the CPU, memory and other resource occupancy of the running process in the system level over time, which is used to determine whether the process is in an active state.

[0067] Optionally, in an embodiment of the present application, the system port refers to a network communication port that provides services within the distributed file storage system and to the outside world, and is used for inter-process communication, user data transmission, etc.

[0068] For example, the response time and packet loss rate of the detection data packet are used to evaluate the communication performance and stability of the system port. The longer the response time and the higher the packet loss rate, the worse the port performance. The system-level disk usage status, read and write error rate, and disk response time are used to comprehensively evaluate the health of the disk. A continuous increase in the read and write error rate and a continuous increase in the disk response time may mean a disk failure.

[0069] It should be noted that the specific duration of the first preset period and the second preset period, resource occupancy change rate, detection time threshold, packet loss rate threshold and other parameters can be set to be adjusted according to different actual application scenarios and system configurations.

[0070] For example, for a high-load system environment, the first preset period may be set shorter to more quickly detect process crashes; and for data-intensive operations, the second preset period may need to be set longer to accurately determine the increasing trend of read and write error rates. In addition, the definition of the system level is not limited to the access layer, the first control layer, and the second control layer mentioned in this application, and can be appropriately adjusted according to different distributed file storage system architectures.

[0071] In an exemplary embodiment, the target prompt information indicates that a core process of the first management layer (Nas_Master) is abnormal. The system does not receive the heartbeat signal of the process within a preset period of time, and finds that its resource occupancy change rate is 0 for a period of time, and thus determines the target fault type as a process crash type.

[0072] Then, after receiving the fault type prompt, the operation and maintenance personnel immediately restarted the crashed process and restored the normal service of the first control layer.

[0073] In another example, the target prompt information indicates that a system port of the second management layer (Nas_Agent) is unavailable. By sending a detection data packet, the system detects that the response time of the port exceeds the preset detection time threshold, and the packet loss rate is higher than the preset threshold, thereby determining that the target fault type is a port unavailable type.

[0074] In response to this failure, the operation and maintenance personnel adjusted the network configuration, fixed the system port problem, and ensured smooth data reading and writing.

[0075] In another example, if the target prompt information indicates a disk failure, the system will continue to monitor the read and write error rate and response time of the disk. When these indicators continue to deteriorate within the second preset time period, the system will determine the target failure type as a disk failure type. Based on this information, the operation and maintenance personnel will take measures to replace or isolate the faulty disk to prevent data loss.

[0076] Through the embodiments of the present application, by adopting refined fault type identification and automatic or semi-automatic repair mechanism based on fault type, accurate positioning and efficient processing of distributed file storage system faults are achieved, thereby achieving the purpose of improving system stability and reducing business interruption time.

[0077] As an optional scheme, the above method also includes at least one of the following: when the above target fault type is determined to be the above process crash type, executing the system restart instruction; when the above target fault type is determined to be the above port unavailable type, executing the above system restart instruction; when the above target fault type is determined to be the process crash type, starting the repair script, wherein the above repair script is used to adjust the above system level; when the above target fault type is determined to be the above port unavailable type, starting the above repair script.

[0078] Optionally, in an embodiment of the present application, the system restart instruction refers to a restart command sent to a certain level or specific component of the distributed file storage system to restore a crashed process or unavailable port to a normal operating state.

[0079] Optionally, in an embodiment of the present application, the repair script refers to a preset series of automated operation commands used to automatically adjust the system configuration when a specific fault type is detected, such as restarting services, adjusting port settings, clearing memory, etc., to restore the health of the system or components.

[0080] It should be noted that the specific operational details of executing a system restart command or starting a repair script, such as the format of the restart command, the writing language and content of the repair script, etc., are different from the specific architecture and operation and maintenance requirements of the distributed file storage system.

[0081] For example, some systems may need to perform specific pre-processing and post-processing steps to ensure the safety and effectiveness of the restart or repair operation. Adjustments at the system level may involve modifying network configuration, adjusting resource allocation strategies, or optimizing process scheduling strategies, which are not limited in this application.

[0082] In an exemplary embodiment, when the target fault type is determined to be a process crash type, the system automatically sends a system restart instruction to the affected layer to restart the crashed process. In another event, the target fault type is a port unavailable type, and the system also executes a system restart instruction, this time restarting the service related to the faulty port and restoring normal communication of the port.

[0083] In addition, in some cases, the system chooses to start a repair script to handle the fault type. For example, when the process crash type is confirmed, the automatically run repair script adjusts the startup configuration of the process to ensure its stable operation; when the port unavailable type is determined, the repair script modifies the port settings to solve the communication problem.

[0084] Through the embodiments of the present application, an automated system restart instruction and repair script mechanism are adopted to achieve immediate response and autonomous repair of distributed file storage system failures, thereby achieving the purpose of reducing manual intervention, improving the system's self-recovery capability, and enhancing service continuity.

[0085] As an optional solution, the above method also includes at least one of the following: when the above target failure type is determined to be a disk failure type and the faulty disk has logical bad sectors, using a disk repair tool to scan and repair the above faulty disk, wherein the above logical bad sectors indicate that the above faulty disk has data corruption; when the above target failure type is determined to be a disk failure type and the above faulty disk has physical bad sectors, migrating the above data files to a healthy disk, marking or replacing the above faulty disk, wherein the above physical bad sectors indicate that the above faulty disk has physical damage.

[0086] Optionally, in an embodiment of the present application, the disk repair tool refers to a software tool used to detect and repair logical bad sectors on the disk, including but not limited to SMART tools, bad sector scanning and repair tools, etc. These tools can attempt to repair logical errors on the disk without destroying data.

[0087] Optionally, in an embodiment of the present application, the data file refers to a user file or system metadata stored in a distributed file storage system and managed by the second management and control layer (Nas_Agent).

[0088] Optionally, in an embodiment of the present application, disk marking or replacement means that when a physical bad sector is detected on the disk, the system marks the disk as unavailable to prevent further read and write operations, and automatically or manually removes it from the service. At the same time, the data files are migrated to healthy disks in the cluster to ensure data integrity and service continuity.

[0089] It should be noted that the way disk failures are handled is set to vary according to the specific disk type, system configuration, and business requirements.

[0090] For example, for systems with high availability requirements, a disk redundancy mechanism may be configured. When a disk failure is detected, the system automatically reads data from the backup disk, marks the failed disk as unavailable, and performs data recovery and disk replacement operations in the background. For some application scenarios with high business continuity requirements, the system may choose to repair or replace the failed disk during the off-peak period of business to reduce the impact on the business.

[0091] In addition, the use of disk repair tools and data migration strategies may also be flexibly adjusted according to the severity of the fault and the usage of the disk, which is not limited in this application.

[0092] In an exemplary embodiment, the system found a disk on the second management layer (Nas_Agent) with logical bad sectors during the detection period, and immediately started the disk repair tool to perform a comprehensive scan and repair on the disk, avoiding the risk of data corruption. In another incident, the system detected that the disk in the storage node (Cell) had physical bad sectors, and the system immediately performed a data migration operation, migrating the affected data files to healthy disks in the cluster, marking the faulty disk to prevent it from participating in the data storage service again, and then physically replaced it during the maintenance window, ensuring the stability of the system and the security of the data.

[0093] Through the embodiments of the present application, intelligent disk fault detection and automated repair or data migration mechanisms are adopted to achieve efficient response and processing of disk failures in distributed file storage systems, thereby ensuring data integrity and service continuity, while reducing operation and maintenance costs and improving system availability.

[0094] As an optional solution, the above method also includes: when the above detection result data indicates that there is an abnormality in the above distributed file storage system, the above faulty node is removed from the load balancing list and the backup node is started, wherein the nodes in the above load balancing list are used to process the above front-end requests, the above faulty node represents a node that is identified as unable to provide services normally, and the above backup node is used to take over the work of the above faulty node.

[0095] Optionally, in an embodiment of the present application, the load balancing list refers to a dynamic node list used by the system to schedule and distribute front-end requests to each processing node, including but not limited to access layer (Nas_Access) nodes, first management and control layer (Nas_Master) nodes, and second management and control layer (Nas_Agent) nodes, etc. These nodes jointly participate in processing front-end requests to ensure that the system can respond to user needs efficiently and evenly.

[0096] Optionally, in an embodiment of the present application, a faulty node refers to a node that is identified as being unable to provide services normally during a health check, due to a process crash, hardware failure, or other system anomalies. A spare node is a group of healthy nodes prepared in advance in the system, which are used to quickly take over the work of a faulty node when it occurs, to ensure service continuity and high availability.

[0097] It should be noted that the operational details of removing the failed node from the load balancing list and starting the standby node vary according to the specific architecture and fault recovery strategy of the distributed file storage system.

[0098] For example, some systems may use an automatic switching mechanism. When a node failure is detected, a node is immediately selected from the backup node list for hot switching to seamlessly take over the work of the failed node. For some systems that require manual intervention, the operation and maintenance personnel can manually remove the failed node from the load balancing list based on the fault report and manually start the backup node.

[0099] In addition, there may be multiple implementations of the backup node selection mechanism and the fault node diagnosis and repair process, which are not limited in this application.

[0100] In an exemplary embodiment, the system finds that a node in the first control layer cannot process front-end requests normally due to hardware failure during the detection period, and is identified as a faulty node. The system then automatically removes the faulty node from the load balancing list and quickly starts the preset backup node to take over its work. After the backup node is started, it seamlessly assumes the workload of the faulty node, ensuring that the system's ability to process front-end requests is not affected. At the same time, after receiving the fault report, the operation and maintenance personnel further diagnose and repair the faulty node, ensuring the overall stability of the system and service continuity.

[0101] Through the embodiments of the present application, a mechanism of dynamically adjusting the load balancing list and automatically starting the backup node is adopted to achieve rapid fault isolation and recovery when anomalies occur in the distributed file storage system, thereby achieving the purpose of ensuring high system availability, improving user service experience, and reducing operation and maintenance complexity.

[0102] As an optional solution, the above method also includes: respectively detecting the number of central processing unit cores, memory, disk size and partition, operating system version, clock synchronization, domain name resolution service, network card status, scheduled tasks, system logs, and third-party software library configuration data of the above access layer, the above first management and control layer, and the above second management and control layer.

[0103] Optionally, in an embodiment of the present application, the number of CPU cores refers to the number of cores available in a server processor, which is used to evaluate the computing power of a node; memory refers to the amount of random access memory (RAM) available on a node, which is used to monitor the memory usage of a node; disk size and partitions refer to the total capacity of the disk on a node and how it is divided for use by different services; the operating system version refers to the type of operating system and its version running on the node, which is used to ensure software compatibility and timely system updates; clock synchronization refers to keeping the time consistent between nodes through mechanisms such as the Network Time Protocol (NTP) to avoid data inconsistencies or service anomalies caused by time differences; domain name resolution service refers to the DNS configuration of a node, which is used to ensure correct network communication and resource positioning; the network card status refers to the working status of the node network interface, including whether it is running normally, whether it is in bonding mode, etc.; scheduled tasks refer to periodically executed tasks or daemons running on a node, which are used to monitor the normal execution of tasks; system logs refer to log files such as / var / log / messages that record system operation information, which are used to detect abnormal operations or software errors; third-party software library configuration data refers to the version and configuration information of third-party software libraries installed through software package managers such as yum and rpm, which are used to ensure the correctness of software dependencies.

[0104] It should be noted that the specific methods and tools for detecting each component may vary. For example, the number of CPU cores and memory information can be obtained by reading the / proc / cpuinfo and / proc / meminfo files; the disk size and partition information can be read through the lsblk or df-Ph commands; the operating system version can be obtained through the uname-r command; the clock synchronization status and configuration can be checked through NTP tools such as ntpq-p; the DNS configuration can be read through the cat / etc / resolv.conf command; the network card status and configuration can be obtained through commands such as ip addr and ethtool; scheduled task information can be read through the crontab-l command; abnormal detection of system logs can be performed through keyword search through tools such as grep; third-party software library configuration data can be obtained by querying the software library list and version information of managers such as yum and rpm. This application does not limit this.

[0105] In an exemplary embodiment, during the daily health check period, the system first performs a comprehensive check on the nodes of the access layer, the first control layer, and the second control layer, including checking the number of CPU cores, memory usage, disk partition status, operating system version, clock synchronization status, DNS configuration, network card status, scheduled task execution status, whether there are abnormal records in the system log, and whether the configuration data of the third-party software library is correct. After the comprehensive detection, a layered detection is performed, and at the detection time, the first detection operation is performed on the access layer, the second detection operation is performed on the first control layer, and the third detection operation is performed on the second control layer to obtain the detection result data.

[0106] Through the embodiments of the present application, a comprehensive health check mechanism for components at all levels of the distributed file storage system is adopted to achieve detailed monitoring of system resource status, service configuration and operating status, thereby achieving the purpose of timely discovering potential problems, preventing service interruptions, and improving the overall stability of the system and operation and maintenance efficiency.

[0107] In an exemplary embodiment, the current health check script takes a long time to execute due to the large number of check items, and it takes about 2 minutes to execute. At the same time, no automatic fault isolation means is provided, and manual intervention is required when a fault is found.

[0108] It should be noted that in the era of big data, Cloud File Storage (CFS), as a new type of distributed storage service, provides a scalable shared file storage service that can be used with services such as Cloud Virtual Machine (CVM) on the cloud platform. CFS provides a standard NFS file system access protocol to provide a shared data source for multiple CVM instances or other computing services. It helps users meet the needs of high-performance and shared storage, supports linear expansion of capacity and performance, and can be mounted and used without modification for existing applications. As the application of file storage becomes more and more extensive, the scale of private file storage clusters running within enterprises is also increasing. The number of physical servers used at the bottom layer can reach tens of millions. The stable operation of physical servers involved in file storage products is crucial. The large number of physical servers brings huge pressure to daily machine health checks and inspections. At the same time, the complexity of the architecture and the rapid growth of scale have also brought huge challenges to the healthy operation of file storage clusters. Failure of the core business-related component node head in the file storage system will affect the availability of several file system instances, causing service failure in a short period of time. In addition, if the fault detection mechanism within the cluster is not sound enough, nodes in sub-health status cannot be dealt with in a timely manner, and comprehensive daily inspection methods are required to detect sub-health problems.

[0109] Based on this, this application proposes a systematic anomaly detection method. Figure 3is a schematic diagram of an optional system abnormality detection method according to an embodiment of the present application, such as Figure 3 As shown, it should be noted that different components of the distributed file storage cluster rely on different key server hardware.

[0110] For the access layer production component Nas_Access, since it is a stateless service VM node and does not store persistent data, the key to the health of its service lies in the normal service of the software process. It is necessary to monitor basic monitoring such as CPU load, memory usage, and disk usage space to keep the whole machine in a normal state;

[0111] For the production component Nas_Master of the management and control layer (the first management and control layer mentioned above), it is a virtual machine node in the master and standby state. The key to the health of its service lies in whether the core software process service of the current master and standby nodes is normal, whether there is a master node, and whether the core process has abnormal log output. At the same time, it is necessary to monitor basic monitoring such as CPU load, memory usage and disk usage space to keep the whole machine in a normal state;

[0112] For the core production component head responsible for reading and writing (the second management layer mentioned above), it provides the core service for the file system on the head, cooperates with the head pre-installed plug-in installed when the component is installed, manages the nfsd, rpcbind and other core services responsible for reading and writing on the head, mounts the user's file system, receives and executes the commands of the management component nas_master, and multiple heads form a cluster, which is load-balanced. The key to the health of its services lies in whether the core processes responsible for reading and writing, such as rpcbind, nfs, iscsi and other services are normal, whether the ports related to the first and second types of services such as reading, writing and mounting can be listened normally, and whether there is abnormal output in the message log. At the same time, it is necessary to monitor basic monitoring such as CPU load, memory usage and disk usage space to keep the whole machine in a normal state.

[0113] Furthermore, for the underlying storage nodes, due to their multi-copy data storage format, the key to the health of their services lies in the normal operation of the data disk, so the focus is on hard disk and process detection. In addition, since distributed file storage is end-to-end user-oriented, a dial-up test method that detects the read, write, mount and other states of distributed file storage from the user's perspective is also very necessary. The embodiments of the present application provide standard specifications and devices for health checks of the process running status, service provision status, operating system running status and configuration, and server running status and configuration of components at different levels of distributed file storage, as well as a dial-up test tool for calling the full functional interface of file storage.

[0114] For example, since all components at each layer are deployed on x86 or arm servers, the common inspection contents for each component are to check the server operating system running status, file system status, and disk partition status. The business logs of each component are usually serialized in the service provision process and written to the data disk or disk data partition on the server, generally in the / data directory. For servers that do disk arrays, the device identification results under the operating system and whether its partition capacity is correct should be checked.

[0115] In addition, it is necessary to check the ntp service status and dns configuration of the server to ensure clock synchronization between nodes. Since the nodes communicate through the network, the service status of the network card should also be checked. In summary, for a distributed file storage cluster, general inspection contents include but are not limited to: Number of CPU cores, which is achieved by checking the / proc / cpuinfo file. Memory size, which is achieved by checking the / proc / meminfo file. Disk size, partition, raid level, slot matching, etc., which are achieved by using the lsblk command or the df–Ph command. Operating system version. This is achieved by using the uname–r command. Ntp clock synchronization service status and configuration. DNS domain name resolution service status and configuration. Network card service status and configuration, whether it is in bonding mode. Scheduled task daemon crond running status. Whether the system log / var / log / messages contains abnormal content. Third-party dependent software library configuration (yum, rpm, etc.).

[0116] For example, for the access layer production component Nas_Access, the access layer of production control is responsible for receiving task requests issued by the front end (such as file system additions, deletions, modifications, etc.), providing queries upward, and forwarding task requests downward to the control layer production component Nas_Master. Since it is a stateless service virtual machine node and does not store persistent data, the key to the health of its service lies in the normal service of the software process. It is necessary to monitor basic general monitoring such as CPU load, memory usage and disk usage space, ntp offset, dns service, etc. to keep the whole machine in normal state. Therefore, for the access layer forwarding component, the main contents of the health check (abnormal detection) include but are not limited to:

[0117] Check the scheduled task items. This can be done by using the operating system crontab-l command and the pipe character command grep. Check the connectivity of port 22000. This can be done by using the netstat command. Check the survival status of the core processing process. This can be done by using the operating system ps command.

[0118] Exemplarily, for the production component Nas_Master of the management and control layer, this layer component is the core logic implementation component of distributed file storage. Therefore, for the components of the logic processing layer, the main contents of the health check include but are not limited to: the core configuration of each component, such as the upper limit traffic of a single process, the data deletion speed, etc. The survival status of the core process, including the receiving process, the intermediate processing process, and the sending process. View the implementation through the operating system ps command. The survival status of a specific service port, such as the service port of the object index management process of the signaling processing module. View the port status through the netstat command and access the corresponding port implementation with the curl command. Specific business log check, such as whether the data deletion is successful.

[0119] Exemplarily, for the head control component Nas_Agent, this layer of components stores all metadata of the file storage system, including metadata of user data, metadata of data management within the cluster, etc. The general architecture is access forwarding component + bypass management component + data storage component. Because this layer of components involves data storage, it is necessary to check the health status of the disk. Sub-health or damage to the disk will cause sub-health of the service. Although there are three copies of redundancy to ensure data reliability, failure to deal with sub-health or damage to the disk in a timely manner will also have immeasurable consequences. At present, the industry checks the health status of the disk by viewing relevant data indicators such as power-on time, total read and write data, logical bad sectors, and physical bad sectors through smartctl. Therefore, for the metadata storage layer, the main contents of the health check include but are not limited to:

[0120] The survival status of key processes, the receiving, processing and sending processes of access components, the bypass management component processes and the data storage processes. This is achieved through the operating system ps command. The status of the key process listening port is checked through the netstat command and the curl command accesses the corresponding port. The health status of the disk is judged by pulling the disk operation data through the smartctl tool suite. Whether the operation and maintenance system detects other anomalies, such as damage to the kv small table, and the removal of the disk due to damage. This is achieved by pulling the content of the operation and maintenance system.

[0121] Furthermore, the distributed file storage system dial-up test and inspection needs to consider the functions provided by the distributed file storage system, and divide the services into Class I and Class II. Class I services involve core functions such as reading and writing data, and Class II services include functions such as creation, deletion, configuration modification, query, and modification of permission groups. Different types of functions should be treated differently, with Class I inspection cycles being frequent and Class II inspections being on demand. The classification of this functional interface can be understood as: unmounting file system function test, querying permission group rules, querying permission group lists, querying file systems, querying regional availability, querying CFS service status, checking file system mounted clients, creating new permission groups, mounting file system function test, batch modifying permission group rules, deleting permission groups, deleting file systems, creating permission group rules, creating file systems, modifying file system names, and modifying capacity limits.

[0122] In an exemplary embodiment, the present application also provides a detailed technical solution for automatic fault isolation, including but not limited to:

[0123] S1, Monitoring system design: Build a highly centralized and intelligent monitoring system to achieve real-time and accurate monitoring of the status of each component of the distributed file storage cluster. In terms of technology selection, open source tools Prometheus and Grafana can be fully utilized for construction. By carefully designing customized indicator collection solutions and alarm rules, ensure all-round and multi-angle monitoring of each component. For each component, clarify the key indicator system, including but not limited to CPU load, memory usage, disk space utilization, process status (such as whether the process is running, running time, resource usage, etc.), port status (whether the port is open, number of connections, data transmission rate, etc.). The monitoring system regularly collects these indicator data at set time intervals and stores them in a high-performance time series database for subsequent in-depth historical analysis and accurate trend prediction.

[0124] For example, by analyzing historical data, we can discover the abnormal fluctuation patterns of certain indicators, thereby providing early warning of potential failures.

[0125] S2, fault detection mechanism, scientifically and reasonably sets thresholds and alarm rules based on the rich data collected by the monitoring system. When a certain indicator exceeds the preset threshold range, the alarm mechanism is immediately triggered to quickly notify the operation and maintenance personnel for timely processing. For different types of faults, highly targeted detection methods are adopted. For example: for process crash faults, in addition to regularly checking the process list, more accurate detection can be performed by monitoring the heartbeat signal of the process, resource occupancy changes, etc. If the heartbeat signal of the process is not received within a certain period of time, or the resource occupancy of the process suddenly becomes zero, it can be judged as a process crash. In the case of unavailable ports, not only can you try to connect to the port, but you can also judge the availability of the port by sending specific detection packets and analyzing indicators such as response time and packet loss rate. For disk failures, in addition to using SMART technology to monitor the disk status, you can also combine the disk's read and write error rate, response time and other indicators for comprehensive judgment. If the disk's read and write error rate continues to increase, or the response time is significantly prolonged, it may indicate that the disk has failed.

[0126] S3, Fault isolation strategy: Once a fault is detected, strong isolation measures must be taken quickly to resolutely prevent the spread of the fault. Different isolation strategies are formulated and implemented for different types of components. For the access layer production component Nas_Access: When a faulty node is detected, it is immediately removed from the load balancing list to effectively prevent new requests from being distributed to the faulty node. At the same time, the standby node is quickly started so that it can seamlessly take over the work of the faulty node to ensure business continuity. During the implementation process, the dynamic adjustment function of the load balancer can be used to automatically direct traffic to healthy nodes. For the management layer production component Nas_Master: When a fault occurs, the master and standby nodes are switched to the standby nodes in a timely manner to ensure the normal operation of the core logic. At the same time, the faulty nodes are deeply diagnosed and repaired. The status of the master node can be monitored in real time through the heartbeat detection mechanism between the master and standby nodes. Once the master node fails, the standby node can quickly take over and realize imperceptible switching. For the head management component Nas_Agent and the underlying storage nodes: Flexible measures are taken according to the specific circumstances of the fault. For example, when a faulty disk is detected, the read and write operations of the disk are immediately stopped to prevent data damage. At the same time, data can be quickly migrated to other healthy disks to ensure data security and availability. You can use data migration tools, such as Rsync or the data migration function of the distributed storage system, to achieve efficient data migration.

[0127] S4, the automated recovery mechanism, actively attempts to automatically recover the faulty component after successfully isolating the fault to improve the reliability and stability of the system. For some simple faults, such as process crashes, you can try a variety of automatic restart methods. On the one hand, you can use the automatic restart function of the monitoring system to automatically trigger the restart operation when a process crash is detected. On the other hand, you can also write an automated script to automatically execute the restart command when the process crashes by monitoring the process status. For disk failures, you can try to repair the disk or recover the data. For example, for logical bad sectors, you can use disk repair tools to repair them; for physical bad sectors, you can migrate data to other healthy disks and then mark or replace the faulty disk. If the automatic recovery fails, immediately notify the operation and maintenance personnel to handle it manually. At the same time, record the fault information and processing process in detail to provide valuable data support for subsequent analysis and improvement.

[0128] For example, in the embodiments of the present application, Docker and Kubernetes can be selected to encapsulate the components of the distributed file storage cluster in containers. Containerization technology has many advantages. It can not only achieve rapid deployment and elastic expansion, but also provide more powerful resource isolation and fault isolation capabilities. For example, through Docker's container isolation mechanism, it can be ensured that each container has independent resource allocation and operating environment, avoiding the failure of one component from affecting other components.

[0129] At the same time, the service mesh technology Istio is used to achieve efficient management and accurate monitoring of inter-service communication. The service mesh can provide a variety of functions, including flow control, fault injection, and fuse. Through flow control, the flow distribution between services can be dynamically adjusted according to demand to ensure that the flow is directed to healthy services when a failure occurs. The fault injection function can be used to simulate various failure scenarios to help test and verify the system's fault recovery capabilities. The fuse mechanism can automatically cut off the call to the faulty service when a service fails to prevent the spread of the fault.

[0130] By writing automated execution scripts, rapid fault detection, effective isolation, and reliable recovery operations can be achieved. Shell scripts or Python scripts can be used flexibly, combined with the API of the monitoring system to give full play to the flexibility and efficiency of the script. For example, using Python's requests library, you can easily interact with the API of the monitoring system, obtain fault information, and perform corresponding operations according to preset rules. Reasonable use of automation tools Ansible or Puppet can achieve efficient configuration management and rapid deployment of distributed file storage clusters. These tools can help quickly deploy new nodes, update configuration files, etc., greatly improving operation and maintenance efficiency. Through Ansible's playbook, batch configuration and deployment of multiple nodes can be achieved, reducing errors and time costs of manual operations.

[0131] Exemplarily, continuous integration and continuous deployment (CI / CD) can also be performed to establish a complete CI / CD process to ensure that each code change is strictly tested and verified. When deploying a new version, first conduct a comprehensive verification in an independent test environment to ensure that no new faults are introduced. Automated testing tools such as Jenkins and Selenium can be used to perform unit testing, integration testing, and functional testing to ensure the quality and stability of the code. Advanced technologies such as blue-green deployment and canary release can be used to achieve non-interruption deployment and updates. These technologies can gradually promote new versions to production environments without affecting users. For example, in blue-green deployment, two identical production environments are maintained at the same time, one for the operation of the current version and the other for the deployment of the new version. When switching, a small amount of traffic can be directed to the new version for testing first, and all traffic can be switched to the new version after confirmation to ensure the stability and reliability of the system.

[0132] Furthermore, after discovering an abnormal process, the following actions may be taken:

[0133] S1, alarm mechanism, when an abnormal process is detected, an efficient alarm mechanism is triggered immediately. Alarms can be sent through multiple channels, such as email, SMS, instant messaging tools, etc., to ensure that operation and maintenance personnel can receive alarm information in a timely manner. The alarm content should include key information such as the type of fault, time of occurrence, and scope of impact in detail, so that operation and maintenance personnel can quickly understand the fault situation and make accurate judgments and decisions. Email templates and SMS templates can be used to present fault information to operation and maintenance personnel in a clear and concise manner. At the same time, the alarm level can be set, which is divided into different levels such as emergency, important, and general according to the severity of the fault, so that operation and maintenance personnel can give priority to emergency faults.

[0134] S2, automatic restart. For some simple abnormal processes, such as process crashes, you can try a variety of automatic restart methods. You can use the automatic restart function of the monitoring system to automatically trigger the restart operation when a process crash is detected. You can also write an automation script to automatically execute the restart command when the process crashes by monitoring the process status. You can use Python's psutil library to monitor the status of the process. When a process crash is found, use the subprocess library to execute the restart command. At the same time, you can set the number and time interval of automatic restarts to avoid unlimited automatic restarts that cause waste of system resources.

[0135] S3, failover, if the abnormal process cannot be restarted automatically, you can consider failover. For components in active-standby mode, you can switch the active-standby to the standby node to ensure business continuity. For components in load balancing mode, you can distribute requests to other healthy nodes to avoid single point failures. During the active-standby switching process, you can use the heartbeat detection mechanism to monitor the status of the primary node in real time. Once the primary node fails, the standby node can take over quickly to achieve imperceptible switching. In load balancing mode, you can use the health check function of the load balancer to automatically distribute requests to healthy nodes.

[0136] S4, manual processing. If the automatic processing fails, immediately notify the operation and maintenance personnel to manually process it. The operation and maintenance personnel can take appropriate measures to repair it according to the specific situation of the fault. At the same time, the fault information and processing process are recorded in detail to provide valuable data support for subsequent analysis and improvement. The operation and maintenance personnel can log in to the faulty node, view system logs, process status and other information, analyze the cause of the fault, and take appropriate repair measures. During the processing, remote login tools such as SSH can be used to remotely manage and repair the faulty node. At the same time, the processing process is recorded in the document, including the fault phenomenon, analysis process, measures taken and results, etc., for subsequent review and summary.

[0137] Through the embodiments of the present application, a set of health inspection standards and methods for specific distributed file storage systems are implemented. A relevant inspection platform or operation and maintenance platform can be used to set up a full inspection at a specific time every day and issue a corresponding inspection report, so as to discover local failures or sub-health conditions of the distributed file storage system and timely manual access and processing, which is of great significance to maintaining the stable operation of the system, such as preventing potential problems, ensuring stable performance, ensuring data integrity, enhancing security, and improving business continuity.

[0138] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0139] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0140] According to another aspect of the embodiments of the present application, there is also provided a system abnormality detection device for implementing the above-mentioned system abnormality detection method. Figure 4 As shown, the device comprises:

[0141] An acquisition module 402 is used to acquire a service type of a business service provided by the distributed file storage system, and determine a detection time according to the service type;

[0142] The execution module 404 is used to perform a first detection operation on the access layer at the detection time, perform a second detection operation on the first control layer, and perform a third detection operation on the second control layer to obtain detection result data, wherein the distributed storage file system includes an access layer, a first control layer, and a second control layer, the access layer is used to receive a front-end request and forward the front-end request to the first control layer, the first control layer is used to generate a data file according to the front-end request, and coordinate the second control layer to work, the second control layer is used to perform read, write and storage operations on the data file, the first detection operation, the second detection operation and the third detection operation are different from each other, the first detection operation includes detecting the task execution, port occupancy and process survival status of the access layer, the second detection operation includes detecting the configuration data, process survival status and business log of the first control layer, and the third detection operation includes detecting the disk usage, port occupancy and process survival status of the second control layer;

[0143] The generating module 406 is used to generate a target prompt message when the detection result data indicates that the distributed file storage system is abnormal, wherein the target prompt message is used to indicate the type of fault that causes the distributed file storage system to be abnormal.

[0144] As an optional scheme, the above-mentioned device is used to generate a target prompt message in the following manner when the detection result data indicates that there is an abnormality in the distributed file storage system: when the detection result data indicates that there is an abnormality in the distributed file storage system, compare the detection result data and the data threshold to determine the abnormality detection data at the system level, wherein the detection result data includes the abnormality detection data, and the system level includes at least one of the access layer, the first management layer and the second management layer; generate a target prompt message based on at least one of the load conditions, memory utilization, disk space utilization, process status, and port status at the system level; determine the target fault type according to the target prompt information, and perform a fault repair operation based on the target fault type.

[0145] As an optional solution, the above-mentioned device is used to determine the fault type according to the target prompt information in the following manner, and perform a fault repair operation based on the fault type: when the target prompt information indicates a process crash, detect the running process at the system level, and when no heartbeat signal of the running process at the system level is received within a first preset time period, and / or when the resource occupancy change rate is 0, determine the target fault type as a process crash type; when the target prompt information indicates that the system port is unavailable, send a detection data packet through the system port at the system level to obtain the response time and data packet loss rate of the detection data packet, and when the response time of the detection data packet is greater than or equal to the detection time threshold, and / or when the data packet loss rate is greater than or equal to the loss rate threshold, determine the target fault type as a port unavailable type; when the target prompt information indicates a disk failure, obtain the usage status, read and write error rate and disk response time of the disk at the system level, and when the read and write error rate continues to increase within a second preset time period, and / or when the disk response time continues to increase, determine the target fault type as a disk failure type.

[0146] As an optional solution, the above-mentioned device is also used for at least one of the following: when the target fault type is determined to be a process crash type, executing a system restart instruction; when the target fault type is determined to be a port unavailable type, executing a system restart instruction; when the target fault type is determined to be a process crash type, starting a repair script, wherein the repair script is used to adjust the system level; when the target fault type is determined to be a port unavailable type, starting a repair script.

[0147] As an optional solution, the above-mentioned device is also used for at least one of the following: when the target failure type is determined to be a disk failure type and the faulty disk has logical bad sectors, using a disk repair tool to scan and repair the faulty disk, wherein the logical bad sectors indicate that the faulty disk has data corruption; when the target failure type is determined to be a disk failure type and the faulty disk has physical bad sectors, migrating data files to a healthy disk and marking or replacing the faulty disk, wherein the physical bad sectors indicate that the faulty disk has physical damage.

[0148] As an optional solution, the above-mentioned device is also used to: when the detection result data indicates that there is an abnormality in the distributed file storage system, remove the faulty node from the load balancing list and start the backup node, wherein the nodes in the load balancing list are used to process front-end requests, the faulty node represents a node that is identified as unable to provide services normally, and the backup node is used to take over the work of the faulty node.

[0149] As an optional solution, the above-mentioned device is also used to: respectively detect the number of CPU cores, memory, disk size and partition, operating system version, clock synchronization, domain name resolution service, network card status, scheduled tasks, system logs, and third-party software library configuration data of the access layer, the first management layer, and the second management layer.

[0150] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0151] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0152] According to one aspect of the present application, a computer program product is provided. The computer program product includes a computer program.

[0153] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0154] Figure 5 The structure block diagram of a computer system for implementing an electronic device according to an embodiment of the present application is schematically shown.

[0155] It should be noted that Figure 5The computer system 500 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0156] like Figure 5 As shown, the computer system 500 includes a central processing unit 501 (CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 502 (ROM) or the program loaded from the storage part 508 to the random access memory 503 (RAM). Various programs and data required for system operation are also stored in the random access memory 503. The central processing unit 501, the read-only memory 502 and the random access memory 503 are connected to each other through a bus 504. The input / output interface 505 (Input / Output interface, i.e., I / O interface) is also connected to the bus 504.

[0157] The following components are connected to the input / output interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a local area network card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read therefrom is installed into the storage section 508 as needed.

[0158] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processor 501, various functions defined in the system of the present application are executed.

[0159] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processor 501, various functions provided by the embodiment of the present application are performed.

[0160] According to another aspect of the embodiment of the present application, an electronic device for implementing the abnormality detection method of the above system is also provided. The electronic device may be Figure 1 The terminal device or server shown in the figure. This embodiment is described by taking the electronic device as a terminal device as an example. Figure 6 As shown, the electronic device includes a memory 602 and a processor 604. The memory 602 stores a computer program, and the processor 604 is configured to execute the steps in any of the above method embodiments through the computer program.

[0161] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0162] Optionally, in this embodiment, the above-mentioned processor can be configured to execute the methods in each embodiment of the present application through a computer program.

[0163] Alternatively, a person skilled in the art may understand that: Figure 6 The structure shown is for illustration only. Figure 6 The structure of the electronic device is not limited. Figure 6 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 6 Different configurations are shown.

[0164] Among them, the memory 602 can be used to store software programs and modules, such as the program instructions / modules corresponding to the system anomaly detection method and device in the embodiment of the present application. The processor 604 executes various functional applications and data processing by running the software programs and modules stored in the memory 602, that is, realizing the above-mentioned system anomaly detection method. The memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 602 may further include a memory remotely located relative to the processor 604, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 602 can be specifically, but not limited to, used to store files and other information in a distributed file storage system. As an example, such as Figure 6As shown, the memory 602 may include, but is not limited to, the acquisition module 402, execution module 404, and generation module 406 in the abnormality detection device of the system. In addition, it may also include, but is not limited to, other module units in the abnormality detection device of the system, which will not be repeated in this example.

[0165] Optionally, the transmission device 606 is used to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one example, the transmission device 606 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 606 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0166] In addition, the electronic device further includes: a display 608 for displaying files in the distributed file storage system; and a connection bus 610 for connecting various module components in the electronic device.

[0167] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. The nodes may form a peer-to-peer network, and any form of computing device, such as a server, terminal or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.

[0168] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the system anomaly detection method provided in various optional implementations of the anomaly detection aspect of the above-mentioned system.

[0169] Optionally, in this embodiment, the above-mentioned computer-readable storage medium can be configured to store data for executing the methods in various embodiments of the present application.

[0170] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.

[0171] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0172] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for one or more electronic devices to execute all or part of the steps of the methods described in each embodiment of the present application.

[0173] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0174] In the several embodiments provided in the present application, it should be understood that the disclosed application can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0175] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0176] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0177] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for detecting anomalies of a system, characterized in that: include: Obtaining a service type of a business service provided by the distributed file storage system, and determining a detection time according to the service type; At the detection time, a first detection operation is performed on the access layer, a second detection operation is performed on the first control layer, and a third detection operation is performed on the second control layer to obtain detection result data, wherein the distributed storage file system includes an access layer, a first control layer, and a second control layer, the access layer is used to receive a front-end request and forward the front-end request to the first control layer, the first control layer is used to generate a data file according to the front-end request, and coordinate the second control layer to work, the second control layer is used to perform read, write and storage operations on the data file, the first detection operation, the second detection operation and the third detection operation are different from each other, the first detection operation includes detecting the task execution, port occupancy and process survival status of the access layer, the second detection operation includes detecting the configuration data, process survival status and business log of the first control layer, and the third detection operation includes detecting the disk usage, port occupancy and process survival status of the second control layer; When the detection result data indicates that an abnormality exists in the distributed file storage system, a target prompt message is generated, wherein the target prompt message is used to indicate the type of fault that causes the abnormality of the distributed file storage system.

2. The method according to claim 1, characterized in that When the detection result data indicates that the distributed file storage system is abnormal, generating a target prompt message includes: When the detection result data indicates that the distributed file storage system is abnormal, comparing the detection result data with a data threshold to determine abnormal detection data at the system level, wherein the detection result data includes the abnormal detection data, and the system level includes at least one of the access layer, the first management and control layer, and the second management and control layer; Generate the target prompt message based on at least one of the system-level load, memory usage, disk space utilization, process status, and port status; A target fault type is determined according to the target prompt information, and a fault repair operation is performed based on the target fault type.

3. The method according to claim 2, characterized in that The determining the fault type according to the target prompt information, and performing the fault repair operation based on the fault type, includes: In the case where the target prompt information indicates a process crash, detecting the running process at the system level, and if no heartbeat signal of the running process at the system level is received within a first preset time period, and / or if the resource occupancy change rate is 0, determining the target fault type as a process crash type; When the target prompt information indicates that the system port is unavailable, sending a detection data packet through the system port of the system level to obtain a response time and a data packet loss rate of the detection data packet, and when the response time of the detection data packet is greater than or equal to a detection time threshold, and / or the data packet loss rate is greater than or equal to a loss rate threshold, determining the target fault type as a port unavailable type; When the target prompt information indicates a disk failure, the usage status, read / write error rate and disk response time of the disk at the system level are obtained. If the read / write error rate continues to increase and / or the disk response time continues to increase within a second preset time period, the target failure type is determined as a disk failure type.

4. The method according to claim 3, characterized in that The method further comprises at least one of the following: When the target fault type is determined to be the process crash type, executing a system restart instruction; When the target fault type is determined to be the port unavailable type, executing the system restart instruction; When the target fault type is determined to be a process crash type, starting a repair script, wherein the repair script is used to adjust the system level; When the target fault type is determined to be the port unavailable type, the repair script is started.

5. The method according to claim 3, characterized in that: The method further comprises at least one of the following: When the target failure type is determined to be a disk failure type and the failed disk has logical bad sectors, a disk repair tool is used to scan and repair the failed disk, wherein the logical bad sectors indicate that data on the failed disk is damaged; When the target failure type is determined to be a disk failure type and the failed disk has physical bad sectors, the data file is migrated to a healthy disk and the failed disk is marked or replaced, wherein the physical bad sectors indicate that the failed disk has physical damage.

6. The method according to claim 1, characterized in that The method further comprises: When the detection result data indicates that there is an abnormality in the distributed file storage system, the faulty node is removed from the load balancing list and the backup node is started, wherein the nodes in the load balancing list are used to process the front-end requests, the faulty node represents a node that is identified as being unable to provide services normally, and the backup node is used to take over the work of the faulty node.

7. The method according to claim 1, characterized in that The method further comprises: The number of CPU cores, memory, disk size and partition, operating system version, clock synchronization, domain name resolution service, network card status, scheduled tasks, system logs, and third-party software library configuration data of the access layer, the first control layer, and the second control layer are detected respectively.

8. A system abnormality detection device, characterized in that: include: An acquisition module, used to acquire a service type of a business service provided by the distributed file storage system, and determine a detection time according to the service type; an execution module, configured to perform a first detection operation on the access layer, a second detection operation on the first control layer, and a third detection operation on the second control layer at the detection time, to obtain detection result data, wherein the distributed storage file system includes an access layer, a first control layer, and a second control layer, the access layer is configured to receive a front-end request and forward the front-end request to the first control layer, the first control layer is configured to generate a data file according to the front-end request and coordinate the second control layer to work, the second control layer is configured to perform read, write and storage operations on the data file, the first detection operation, the second detection operation and the third detection operation are all different from each other, the first detection operation includes detecting the task execution, port occupancy and process survival status of the access layer, the second detection operation includes detecting the configuration data, process survival status and business log of the first control layer, and the third detection operation includes detecting the disk usage, port occupancy and process survival status of the second control layer; A generation module is used to generate a target prompt message when the detection result data indicates that the distributed file storage system is abnormal, wherein the target prompt message is used to indicate the type of fault that causes the distributed file storage system to be abnormal.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein the computer program can be executed by an electronic device to perform the method described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

11. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.

Citation Information

Patent Citations

  • Monitoring method, system and device of distributed storage system and medium

    CN111858240A

  • Monitoring method and system of information system, storage medium and electronic equipment

    CN116795630A

  • System health degree detection method and device, electronic equipment and storage medium

    CN118467281A

Cited By

  • Fault diagnosis method and device of storage system, electronic equipment and storage medium

    CN120743716A

  • Abnormality detection method and storage device

    CN121116697A

  • Diagnosis result-based solution recommendation method and system

    CN121262066A

  • Hardware resource release method, device and equipment facing transaction system and medium

    CN122547554A

  • Hardware resource release method, device and equipment facing transaction system and medium

    CN122547554B