Fault processing method and device for distributed file storage system
By analyzing and fusion of data information of distributed file storage systems, automatically locate and dealing with faults, the problem of inefficient manual judgment in the existing technology is solved, and fast and accurate fault recovery is achieved.
Patent Information
- Application Number
- CN202510497533.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-22
AI Technical Summary
The fault handling methods of existing distributed file storage systems rely on manual judgment, resulting in inaccurate and inefficient judgments, affecting the system recovery efficiency and response speed.
By collecting data information from the storage system, the data information of the client, network file system server and storage pool are respectively analyzed, and analysis results are generated, and fusion analysis is carried out to locate the faults and design an automatic recovery strategy for processing.
It realizes accurate, fast positioning and automatic recovery of distributed file storage system failures, reduces uncertainty and delay caused by human operations, and improves the system's processing efficiency.
Smart Images

Figure CN120353772A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and in particular, to a method and device for fault handling in a distributed file storage system. Background Art
[0002] A distributed shared file storage system provides persistent storage of data by a storage cluster composed of multiple servers, and uses the NFS protocol through an NFS server to share the data to one or more clients for mounting. Such a system provides shared file services for users, but faces many complex and variable technical challenges in terms of fault detection and automatic recovery.
[0003] In existing fault management methods, the fault judgment and fault handling of a distributed shared file storage system are often carried out manually. However, through the manual judgment method, problems such as inaccurate fault discrimination are likely to occur, and the efficiency of manually handling faults is relatively low, which is likely to affect the recovery efficiency of the system, and also reduces the overall response speed and service quality of the system. Summary of the Invention
[0004] Based on the above deficiencies of the prior art, the present application provides a method and device for fault handling in a distributed file storage system to solve the problem of relatively low efficiency in handling faults brought by the prior art.
[0005] To achieve the above object, the present application provides the following technical solutions:
[0006] The first aspect of the present application provides a method for fault handling in a distributed file storage system, including:
[0007] Collecting data information of the storage system; wherein, the data information at least includes first data information of a client, second data information of a network file system server, and third data information of a storage pool;
[0008] Respectively performing fault analysis on the first data information, the second data information, and the third data information to obtain a first analysis result corresponding to the first data information, a second analysis result corresponding to the second data information, and a third analysis result corresponding to the third data information;
[0009] Performing fusion analysis on the first analysis result, the second analysis result, and the third analysis result to obtain a fault location result;
[0010] Performing fault handling on the storage system according to the fault location result.
[0011] Optionally, in the above-mentioned fault handling method of the distributed file storage system, the step of respectively performing fault analysis on the first data information, the second data information, and the third data information to obtain a first analysis result corresponding to the first data information, a second analysis result corresponding to the second data information, and a third analysis result corresponding to the third data information includes:
[0012] Divide the first data information to obtain numerical data and log data, and divide the third data information to obtain target numerical data and target log data;
[0013] Based on a preset threshold, perform fault analysis on the numerical data and the target numerical data, and perform keyword scanning on the log data and the target log data;
[0014] When the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data pass the keyword scanning, or when the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data do not pass the keyword scanning, generate a first analysis result and a third analysis result according to the first data information and the third data information;
[0015] When the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data pass the keyword scanning, generate a first analysis result and a third analysis result based on the numerical data and the target numerical data;
[0016] When the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data do not pass the keyword scanning, generate a first analysis result and a third analysis result according to the log data and the target log data;
[0017] Judge whether there is data exceeding the preset threshold in the second data information;
[0018] If there is data exceeding the preset threshold in the second data information, generate a second analysis result based on the data exceeding the preset threshold;
[0019] If there is no data exceeding the preset threshold in the second data information, generate a second analysis result based on the second data information.
[0020] Optionally, in the above-mentioned fault handling method of the distributed file storage system, the step of performing fusion analysis on the first analysis result, the second analysis result, and the third analysis result to obtain a fault location result includes:
[0021] Fuse the first analysis result, the second analysis result, and the third analysis result to obtain a fused analysis result;
[0022] Determine whether the fault analysis in the fused analysis result is concentrated on the client;
[0023] If the fault analysis in the fused analysis result is concentrated on the client, then determine the fault location result as a single client fault;
[0024] If the fault analysis in the fused analysis result is not concentrated on the client, then detect whether the fault analysis in the fused analysis result is in different clients of the same file system;
[0025] If the fault analysis in the fused analysis result is in different clients of the same file system, then determine the fault location result as a single file system fault;
[0026] If the fault analysis in the fused analysis result is not in different clients of the same file system, then determine whether the fault analysis in the fused analysis result is in different file systems of the same network file system server;
[0027] If the fault analysis in the fused analysis result is in different file systems of the same network file system server, then determine the fault location result as a network file system server fault;
[0028] If the fault analysis in the fused analysis result is not in different file systems of the same network file system server, then detect whether the fault analysis in the fused analysis result is in different file systems of the same storage;
[0029] If the fault analysis in the fused analysis result is in different file systems of the same storage, then determine the fault location result as a storage pool fault.
[0030] Optionally, in the above-mentioned fault handling method for a distributed file storage system, when the fault location result is a single client fault, the fault handling of the storage system according to the fault location result includes:
[0031] Obtain a pre-set fault recovery program;
[0032] Use the fault recovery program to automatically recover the fault of the client and feedback the fault information of the client to the front end.
[0033] Optionally, in the above method for handling faults in a distributed file storage system, when the fault location result is a single file system fault, the handling of the storage system according to the fault location result includes:
[0034] Trigger the migration operation of the network file system server to migrate the faulty file system to other available network file system servers;
[0035] When the migration operation is completed, feedback the information of the single file system fault to the front end.
[0036] Optionally, in the above method for handling faults in a distributed file storage system, when the fault location result is a network file system server fault, the handling of the storage system according to the fault location result includes:
[0037] Trigger the migration operation of the network file system server to migrate the data on the network file system server to a standby server;
[0038] When the migration operation is completed, feedback the alarm information of the network file system server to the front end.
[0039] Optionally, in the above method for handling faults in a distributed file storage system, the handling of the storage system according to the fault location result includes:
[0040] When the fault location result is a storage pool fault, call the alarm module to feedback the alarm information of the storage pool to the front end based on the fault location result.
[0041] The second aspect of this application provides a device for handling faults in a distributed file storage system, including:
[0042] An information collection unit for collecting data information of the storage system; wherein, the data information at least includes first data information of a client, second data information of a network file system server, and third data information of a storage pool;
[0043] A fault analysis unit for respectively performing fault analysis on the first data information, the second data information, and the third data information to obtain a first analysis result corresponding to the first data information, a second analysis result corresponding to the second data information, and a third analysis result corresponding to the third data information;
[0044] A fusion analysis unit for performing fusion analysis on the first analysis result, the second analysis result, and the third analysis result to obtain a fault location result;
[0045] A fault handling unit for performing fault handling on the storage system according to the fault location result.
[0046] Optionally, in the above-mentioned fault handling device of the distributed file storage system, the fault analysis unit includes:
[0047] A partitioning unit for partitioning the first data information to obtain numerical data and log data, and partitioning the third data information to obtain target numerical data and target log data;
[0048] A processing unit for performing fault analysis on the numerical data and the target numerical data based on a preset threshold, and performing keyword scanning on the log data and the target log data;
[0049] A first generating unit for generating a first analysis result and a third analysis result according to the first data information and the third data information when the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data pass the keyword scanning, or when the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data do not pass the keyword scanning;
[0050] A second generating unit for generating a first analysis result and a third analysis result based on the numerical data and the target numerical data when the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data pass the keyword scanning;
[0051] A third generating unit for generating a first analysis result and a third analysis result according to the log data and the target log data when the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data do not pass the keyword scanning;
[0052] A first judging unit for judging whether there is data exceeding a preset threshold in the second data information;
[0053] A fourth generating unit for generating a second analysis result based on the data exceeding the preset threshold if there is data exceeding the preset threshold in the second data information;
[0054] A fifth generating unit for generating a second analysis result based on the second data information if there is no data exceeding the preset threshold in the second data information.
[0055] Optionally, in the above-mentioned fault handling device of the distributed file storage system, the fusion analysis unit includes:
[0056] A fusion unit, configured to fuse the first analysis result, the second analysis result, and the third analysis result to obtain a fused analysis result;
[0057] A second determination unit, configured to determine whether the fault analysis in the fused analysis result is concentrated on the client;
[0058] A first determination unit, configured to, if the fault analysis in the fused analysis result is concentrated on the client, determine that the fault location result is a single client fault;
[0059] A first detection unit, configured to, if the fault analysis in the fused analysis result is not concentrated on the client, detect whether the fault analysis in the fused analysis result is in different clients of the same file system;
[0060] A second determination unit, configured to, if the fault analysis in the fused analysis result is in different clients of the same file system, determine that the fault location result is a single file system fault;
[0061] A third determination unit, configured to, if the fault analysis in the fused analysis result is not in different clients of the same file system, determine whether the fault analysis in the fused analysis result is in different file systems of the same network file system server;
[0062] A third determination unit, configured to, if the fault analysis in the fused analysis result is in different file systems of the same network file system server, determine that the fault location result is a network file system server fault;
[0063] A second detection unit, configured to, if the fault analysis in the fused analysis result is not in different file systems of the same network file system server, detect whether the fault analysis in the fused analysis result is in different file systems of the same storage;
[0064] A fourth determination unit, configured to, if the fault analysis in the fused analysis result is in different file systems of the same storage, determine that the fault location result is a storage pool fault.
[0065] Optionally, in the above-mentioned fault handling device of the distributed file storage system, when the fault location result is a single client fault, the fault handling unit includes:
[0066] A program acquisition unit, configured to acquire a preset fault recovery program;
[0067] A fault recovery unit, configured to use the fault recovery program to automatically recover the fault of the client and feed back the fault information of the client to the front end.
[0068] Optionally, in the fault handling device of the above-mentioned distributed file storage system, when the fault location result is a single file system fault, the fault handling unit includes:
[0069] A first migration unit, configured to trigger a migration operation of the network file system server to migrate the faulty file system to other available network file system servers;
[0070] A first feedback unit, configured to, when the migration operation is completed, feedback information about the single file system fault to the front end.
[0071] Optionally, in the fault handling device of the above-mentioned distributed file storage system, when the fault location result is a network file system server fault, the fault handling unit includes:
[0072] A second migration unit, configured to trigger a migration operation of the network file system server to migrate the data on the network file system server to a standby server;
[0073] A second feedback unit, configured to, when the migration operation is completed, feedback alarm information about the network file system server to the front end.
[0074] Optionally, in the fault handling device of the above-mentioned distributed file storage system, the fault handling unit includes:
[0075] A third feedback unit, configured to, when the fault location result is a storage pool fault, call an alarm module to feedback alarm information about the storage pool to the front end based on the fault location result.
[0076] A fault handling method for a distributed file storage system provided by the present application collects data information of the storage system, where the data information at least includes first data information of a client, second data information of a network file system server, and third data information of a storage pool. Secondly, fault analysis is respectively performed on the first data information, the second data information, and the third data information to obtain a first analysis result corresponding to the first data information, a second analysis result corresponding to the second data information, and a third analysis result corresponding to the third data information. Then, fusion analysis is performed on the first analysis result, the second analysis result, and the third analysis result to obtain a fault location result. Finally, according to the fault location result, fault handling is performed on the storage system. Thus, by collecting operation status data from the client, the server, and the storage pool, the fault type is analyzed in real time for the collected data, and then corresponding automatic recovery strategies are designed according to different fault types to reduce the uncertainty and delay caused by manual operations, effectively solving the problem of relatively low efficiency in handling faults. Description of the Drawings
[0077] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.
[0078] Figure 1 It is a schematic flow chart of a fault handling method for a distributed file storage system provided by an embodiment of the present application;
[0079] Figure 2 It is a schematic flow chart of a fault fusion analysis method for a storage system provided by an embodiment of the present application;
[0080] Figure 3 It is a schematic flow chart of a fault handling method for a single client provided by another embodiment of the present application;
[0081] Figure 4 It is a schematic flow chart of a fault handling method for a single file system provided by another embodiment of the present application;
[0082] Figure 5 It is a schematic flow chart of a fault handling method for a network file system server provided by another embodiment of the present application;
[0083] Figure 6 It is a schematic structural diagram of a fault handling device for a distributed file storage system provided by another embodiment of the present application. Detailed implementation manners
[0084] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0085] In this application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0086] The present application embodiment provides a method for troubleshooting a distributed file storage system. Figure 1 As shown, the specific steps include:
[0087] S101. Collect data information of a storage system.
[0088] Specifically, in the embodiment of the present application, by designing a special collection program, data information of the storage system is regularly collected at a preset time interval to perform regular abnormality detection on the storage system. In addition, the storage system includes a client, a network file system server (NFS server) and a storage pool. Therefore, the data information of the storage system includes at least the first data information of the client, the second data information of the network file system server and the third data information of the storage pool.
[0089] It should be noted that in the embodiment of the present application, the information of all mounted clients of each file system will be collected regularly. The collected information covers the client's business detection situation, namely (1) understanding the client's detection results of various services in detail to determine whether the business is running normally. (2) Read and write response delay, accurately record the response time delay when the client performs read and write operations, so as to analyze the performance bottleneck of the system. (3) Logs containing error keywords, carefully collect the log information generated by the client, and filter out the log content containing specific error keywords. These keywords can often indicate potential problems.
[0090] For each NFS server where the file system is located, the same method of regular collection is also adopted. Regularly collect the second data information of the NFS server, including (1) in terms of network connectivity, check the network connection status between the server and other devices through network detection tools to ensure network stability. (2) CPU load, monitor the CPU usage of the server in real time to understand whether the CPU workload is too high. (3) Memory usage rate, pay attention to the occupancy ratio of the server's memory to judge whether the memory resources are sufficient. (4) Logs containing error keywords, scan the logs generated by the server, and extract the parts containing error keywords. These logs may reflect the abnormalities that occur during the operation of the server.
[0091] Finally, for each storage pool where the file system is located, a mechanism for regularly collecting information is also formulated. The content collected includes (1) network connectivity, ensure the normal network connection between the storage pool and other related devices to ensure smooth data transmission. (2) Availability of underlying persistent storage, check whether the underlying storage devices of the storage pool are working properly and whether the storage resources can be normally accessed and used. (3) Logs containing error keywords, analyze the logs generated by the storage pool to find the records containing error keywords. These logs may imply faults or abnormalities inside the storage pool.
[0092] It should also be noted that after collecting the first data information of the client, the second data information of the network file system server, and the third data information of the storage pool, the first data information will be sent to the client information analysis module to provide a data basis for subsequent analysis, and the second data information will be sent to the NFS server information analysis module to deeply analyze the running status of the server, and the third data information will also be sent to the storage pool information analysis module to provide a basis for evaluating the status of the storage pool.
[0093] S102. Conduct fault analysis on the first data information, the second data information, and the third data information respectively to obtain the first analysis result corresponding to the first data information, the second analysis result corresponding to the second data information, and the third analysis result corresponding to the third data information.
[0094] It should be noted that the client information analysis module will regularly receive the information reported by each mounted client, that is, the first data information, and then use specialized data analysis algorithms and tools to deeply analyze the first data information to analyze whether a specific business fails or the overall client operation has problems, so as to obtain the first analysis result.
[0095] After the NFS server information analysis module periodically receives the information reported by each NFS server, that is, the second data information, it will then conduct a detailed analysis of the information of the NFS servers where these file systems are located, analyze the fault conditions of the NFS servers, and determine whether there are network interruptions, performance bottlenecks or other fault problems in the servers, so as to obtain the second analysis result.
[0096] After the storage pool information analysis module periodically receives the information reported by the storage pool, that is, the third data information, it will use professional analysis methods to process the information of the storage pools where these file systems are located. By analyzing data such as network connectivity, underlying persistent storage availability, and error logs, analyze the fault conditions of the storage pool, and determine whether there are network faults, storage device faults, etc. in the storage pool. Thus, the third analysis result is obtained.
[0097] Optionally, in another embodiment of the present application, another specific implementation manner of step S102 includes the following steps:
[0098] Divide the first data information to obtain numerical data and log data, and divide the third data information to obtain target numerical data and target log data.
[0099] It should be noted that for the anomaly detection of the client and the storage pool, both need to be detected and judged through numerical values and logs. Therefore, the first data information and the third data information will be pre-classified respectively to obtain the numerical data and log data corresponding to the first data information, and the target numerical data and target log data corresponding to the third data information.
[0100] Based on a preset threshold, conduct fault analysis on the numerical data and the target numerical data, and perform keyword scanning on the log data and the target log data.
[0101] It can be understood that for numerical data, a threshold will be preset in advance to conduct fault analysis on the numerical data and the target numerical data, that is, to judge whether the numerical values corresponding to the numerical data and the numerical values corresponding to the target numerical data exceed the preset threshold, so as to determine whether the client and the storage pool have faults.
[0102] For log data, it will be judged whether there are incorrect keywords in the log data and the target log data, so as to determine whether the client and the storage pool have faults.
[0103] When the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data pass the keyword scan, or when the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data do not pass the keyword scan, the first analysis result and the third analysis result are generated according to the first data information and the third data information.
[0104] Specifically, if the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data pass the keyword scan, it indicates that there is no fault in the client and the storage pool. Then, the first analysis result can be generated according to the first data information, and the third analysis result can be generated according to the third data information. Among them, the first analysis result refers to the result that the client has no fault, and the third analysis result indicates the result that the storage pool has no fault.
[0105] Or if the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data do not pass the keyword scan, it indicates that there is a fault in the client and the storage pool. Therefore, the first analysis result can be generated according to the first data information and the third analysis result can be generated according to the third data information. Among them, the first analysis result refers to the result that the client has a fault, and the third analysis result is the result that the storage pool has a fault.
[0106] When the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data pass the keyword scan, the first analysis result and the third analysis result are generated based on the numerical data and the target numerical data.
[0107] It can be understood that if the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data pass the keyword scan, it means that there is a fault in the numerical data in the client and the storage pool. Therefore, at this time, the first analysis result needs to be generated based on the numerical data, and the third analysis result needs to be generated based on the target numerical data.
[0108] When the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data do not pass the keyword scan, the first analysis result and the third analysis result are generated according to the log data and the target log data.
[0109] Specifically, if the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data do not pass the keyword scan, it indicates that there is a fault in the log data in the client and the storage pool. Therefore, at this time, the first analysis result needs to be generated based on the log data, and the third analysis result needs to be generated based on the target log data.
[0110] Judge whether there is data in the second data information that exceeds the preset threshold.
[0111] Specifically, after detecting the client and the storage pool, it is also necessary to perform anomaly detection on the NFS server. It should be noted that the second data information of the NFS server is all numerical data. Therefore, it is only necessary to determine whether there is data exceeding the preset threshold in the second data information. If there is data exceeding the preset threshold in the second data information, it indicates that the NFS server has a fault. Therefore, at this time, it is necessary to generate a second analysis result based on the data exceeding the preset threshold. If there is no data exceeding the preset threshold in the second data information, it indicates that the NFS server has no fault, and a second analysis result should also be generated based on the second data information. Among them, the second analysis result at this time will indicate the result that the NFS server has no fault.
[0112] Optionally, the execution order of the fault analysis for the client, NFS server, and storage pool in the embodiments of the present application is only one optional way. Since in the embodiments of the present application, the client, NFS server, and storage pool are completely independent of each other, the execution order of the fault analysis for the client, NFS server, and storage pool can also adopt other execution orders, which can be specifically set according to requirements.
[0113] S103. Perform a fusion analysis on the first analysis result, the second analysis result, and the third analysis result to obtain a fault location result.
[0114] Specifically, after obtaining the first analysis result, the second analysis result, and the third analysis result, the client information analysis module, the NFS server information analysis module, and the storage pool information analysis module will send their respective analysis results to the comprehensive analysis module according to a preset time period. Then, the comprehensive analysis module will receive the analysis results sent by the client information analysis module, the NFS server information analysis module, and the storage pool information analysis module according to the established time period, and use complex comprehensive analysis models and algorithms to deeply fuse and analyze the first analysis result, the second analysis result, and the third analysis result. Thus, by correlating and comparing each analysis result, comprehensively sort out the possible fault clues in the system, and finally generate an accurate fault location result.
[0115] Optionally, in another embodiment of the present application, another specific implementation manner of step S103 is as Figure 2 shown, including the following steps:
[0116] S201. Fuse the first analysis result, the second analysis result, and the third analysis result to obtain a fusion analysis result.
[0117] Specifically, in order to accurately locate whether the failure occurs in the client, the NFS server, or the storage pool, the first analysis result, the second analysis result, and the third analysis result can be classified according to the cause of the failure in the analysis result to obtain multiple types of analysis results, and finally each type of analysis result is fused together to obtain the fused analysis result.
[0118] S202. Determine whether the failure analysis in the fused analysis result is concentrated on the client.
[0119] Specifically, if the failure analysis in the fused analysis result is concentrated on the client, it indicates that the failures all occur in the same client. Therefore, step S203 is executed at this time. If the failure analysis in the fused analysis result is not concentrated on the client, it indicates that the failures may not be concentrated in the same client or may occur in other clients. Therefore, further judgment is required, and step S204 is executed at this time.
[0120] S203. Determine that the failure location result is a single-client failure.
[0121] S204. Detect whether the failure analysis in the fused analysis result is in different clients of the same file system.
[0122] It can be understood that when the failure analysis in the fused analysis result is not concentrated on the client, it is necessary to know whether the failure occurs in different clients. Therefore, it is detected whether the failure analysis in the fused analysis result is in different clients of the same file system. If the failure analysis in the fused analysis result is in different clients of the same file system, it indicates that the same file system in different clients has failed. Therefore, step S205 is executed at this time. If the failure analysis in the fused analysis result is not in different clients of the same file system, it indicates that the failure does not occur on the client and may occur on the NFS server or storage. Therefore, further determination is required, and step S206 is executed at this time.
[0123] S205. Determine that the failure location result is a single-file system failure.
[0124] S206. Determine whether the failure analysis in the fused analysis result is in different file systems of the same network file system server.
[0125] It should be noted that when the fault analysis in the fusion analysis result is not in different clients of the same file system, it can be determined at this time that the client has no fault. Then, it is necessary to detect whether the NFS server has a fault, that is, to determine whether the fault analysis in the fusion analysis result is in different file systems of the same network file system server. If the fault analysis in the fusion analysis result is in different file systems of the same network file system server, it indicates that the NFS server has a fault. Therefore, step S207 is executed at this time. If the fault analysis in the fusion analysis result is not in different file systems of the same network file system server, it indicates that the NFS server has no fault. At this time, it is necessary to continue to detect whether the storage pool has a fault. Therefore, step S208 needs to be executed at this time.
[0126] S207. Determine that the fault location result is a network file system server fault.
[0127] S208. Detect whether the fault analysis in the fusion analysis result is in different file systems of the same storage.
[0128] Specifically, when the fault analysis in the fusion analysis result is not in different file systems of the same network file system server, it indicates that neither the client nor the NFS server has a fault. Then, at this time, it is only necessary to determine whether the storage pool has a fault, that is, to detect whether the fault analysis in the fusion analysis result is in different file systems of the same storage. If the fault analysis in the fusion analysis result is in different file systems of the same storage, it can be explained that the storage pool has a fault. Therefore, step S209 is executed.
[0129] Optionally, if the fault analysis in the fusion analysis result is not in different file systems of the same storage, it indicates that neither the client, the NFS server nor the storage pool has a fault. Then, step S101 is executed at a set time interval, so as to realize the real-time monitoring of the storage system, discover problems in time, and then solve problems in time.
[0130] S209. Determine that the fault location result is a storage pool fault.
[0131] S104. According to the fault location result, perform fault handling on the storage system.
[0132] It should be noted that in the embodiments of the present application, different automatic recovery strategies are adopted for different fault types, so that the storage system can be restored to the normal working state in time.
[0133] Optionally, in another embodiment of the present application, when the fault location result is a single client fault, another specific implementation manner of step S104 is as Figure 3 shown, and it includes the following steps:
[0134] S301. Obtain a pre-set fault recovery program.
[0135] Specifically, the automatic handling module will receive the fault location result sent by the comprehensive analysis module at a pre-set time interval, and immediately judge the fault situation. When the fault location result is a single client fault, the automatic handling module will call the pre-set fault recovery program from the database to facilitate subsequent fault recovery.
[0136] S302. Use the fault recovery program to automatically recover the fault of the client and feedback the fault information of the client to the front end.
[0137] Specifically, use the fault recovery program to attempt to automatically recover the fault of the client, such as restarting relevant services or performing software repair operations, etc., and will also call the alarm notification module to send an alarm notification of the client to relevant operation and maintenance personnel to inform that a fault has occurred in the client, so that the operation and maintenance personnel can pay attention and handle it in time.
[0138] Optionally, in another embodiment of the present application, when the fault location result is a single file system fault, another specific implementation manner of step S104 is as Figure 4 shown, including the following steps:
[0139] S401. Trigger the migration operation of the network file system server to migrate the faulty file system to other available network file system servers.
[0140] It can be understood that when the fault location result is a single file system fault, the automatic handling module will initiate the corresponding automatic fault recovery mechanism, that is, trigger the migration operation of the file system across NFS servers to migrate the faulty file system to other available NFS servers. In the migration process, an efficient data transfer protocol or tool can be used to minimize the downtime and impact. And during the migration process, it is also necessary to synchronize the metadata of the file system to ensure that the migrated file system can continue to be accessed and be consistent with the original file system.
[0141] Optionally, mount the migrated file system on the new NFS server to ensure that the application and users can access the newly mounted file system and ensure the data synchronization and integrity of the file system to ensure that there is no data loss or damage.
[0142] S402. When the migration operation is completed, feedback the information of the single file system fault to the front end.
[0143] Specifically, when the migration operation is completed, the alarm notification module is called to send an alarm notification to the system administrator and operation and maintenance technicians to elaborate on the file system failure and migration situation, thus facilitating subsequent inspections and handling in a timely manner.
[0144] Optionally, in another embodiment of the present application, when the fault location result is a network file system server failure, another specific implementation manner of step S104 is as Figure 5 shown and includes the following steps:
[0145] S501. Trigger the migration operation of the network file system server to migrate the data on the network file system server to the standby server.
[0146] Specifically, when the fault location result is a network file system server failure, the automatic disposal module will immediately execute the automatic fault recovery measure of the NFS server, that is, trigger the whole machine fault migration operation of the NFS server to migrate the services and data on the NFS server to the standby NFS server, thereby ensuring the continuity of the services.
[0147] S502. When the migration operation is completed, feedback the alarm information of the network file system server to the front end.
[0148] It can be understood that after the migration operation of the faulty NFS server is completed, the alarm notification module can be used to send an alarm notification to the operation and maintenance personnel to report the fault and migration situation of the NFS server, so that technicians can repair and troubleshoot the faulty NFS server in a timely manner.
[0149] Optionally, in another embodiment of the present application, when the fault location result is a storage pool failure, another specific implementation manner of step S104 includes the following steps:
[0150] When the fault location result is a storage pool failure, call the alarm module to feedback the alarm information of the storage pool to the front end based on the fault location result.
[0151] Specifically, when the fault location result is a storage pool failure, the automatic disposal module will execute corresponding operations according to the preset strategy. During the execution process, the automatic disposal module may not be able to immediately perform automatic recovery on the storage pool failure, but will promptly call the alarm notification module to send an alarm notification to relevant personnel to remind the operation and maintenance personnel that the storage pool has failed and measures need to be taken as soon as possible for repair and handling to avoid data loss or system operation being affected.
[0152] A fault handling method for a distributed file storage system provided by this application. By collecting data information of the storage system, where the data information at least includes the first data information of the client, the second data information of the network file system server, and the third data information of the storage pool. Secondly, perform fault analysis on the first data information, the second data information, and the third data information respectively to obtain the first analysis result corresponding to the first data information, the second analysis result corresponding to the second data information, and the third analysis result corresponding to the third data information. Then, perform fusion analysis on the first analysis result, the second analysis result, and the third analysis result to obtain a fault location result. Finally, according to the fault location result, perform fault handling on the storage system. Thus, by collecting the operation status data from the client, the server, and the storage pool, analyze the fault type in real time for the collected data, and then design corresponding automatic recovery strategies according to different fault types to reduce the uncertainty and delay brought by manual operations, effectively solving the problem of relatively low efficiency in handling faults.
[0153] Another embodiment of this application provides a fault handling device for a distributed file storage system, as Figure 6 shown, specifically including the following units:
[0154] An information collection unit 601, configured to collect data information of the storage system. Wherein, the data information at least includes the first data information of the client, the second data information of the network file system server, and the third data information of the storage pool.
[0155] A fault analysis unit 602, configured to perform fault analysis on the first data information, the second data information, and the third data information respectively to obtain the first analysis result corresponding to the first data information, the second analysis result corresponding to the second data information, and the third analysis result corresponding to the third data information.
[0156] A fusion analysis unit 603, configured to perform fusion analysis on the first analysis result, the second analysis result, and the third analysis result to obtain a fault location result.
[0157] A fault handling unit 604, configured to perform fault handling on the storage system according to the fault location result.
[0158] It should be noted that the specific working processes of the above modules in the embodiments of this application can refer to the steps S101 to S104 in the above method embodiment correspondingly, and will not be elaborated here.
[0159] Optionally, in a fault handling device for a distributed file storage system provided by another embodiment of this application, the fault analysis unit 602 includes:
[0160] A partitioning unit, configured to partition the first data information to obtain numerical data and log data, and partition the third data information to obtain target numerical data and target log data.
[0161] A processing unit, configured to perform fault analysis on the numerical data and the target numerical data based on a preset threshold, and perform keyword scanning on the log data and the target log data.
[0162] A first generating unit, configured to generate a first analysis result and a third analysis result according to the first data information and the third data information when the numerical data and the target numerical data pass the fault analysis and the log data and the target log data pass the keyword scanning, or when the numerical data and the target numerical data do not pass the fault analysis and the log data and the target log data do not pass the keyword scanning.
[0163] A second generating unit, configured to generate a first analysis result and a third analysis result based on the numerical data and the target numerical data when the numerical data and the target numerical data do not pass the fault analysis and the log data and the target log data pass the keyword scanning.
[0164] A third generating unit, configured to generate a first analysis result and a third analysis result according to the log data and the target log data when the numerical data and the target numerical data pass the fault analysis and the log data and the target log data do not pass the keyword scanning.
[0165] A first judging unit, configured to judge whether there is data exceeding the preset threshold in the second data information.
[0166] A fourth generating unit, configured to generate a second analysis result based on the data exceeding the preset threshold if there is data exceeding the preset threshold in the second data information.
[0167] A fifth generating unit, configured to generate a second analysis result based on the second data information if there is no data exceeding the preset threshold in the second data information.
[0168] Optionally, in a fault handling device of a distributed file storage system provided in another embodiment of the present application, the fusion analysis unit 603 includes:
[0169] A fusion unit, configured to fuse the first analysis result, the second analysis result, and the third analysis result to obtain a fusion analysis result.
[0170] A second judging unit, configured to judge whether the fault analysis in the fusion analysis result is concentrated on the client.
[0171] The first determination unit is configured to determine that the fault location result is a single client fault if the fault analysis in the fusion analysis result focuses on the client.
[0172] The first detection unit is configured to detect whether the fault analysis in the fusion analysis result is in different clients of the same file system if the fault analysis in the fusion analysis result does not focus on the client.
[0173] The second determination unit is configured to determine that the fault location result is a single file system fault if the fault analysis in the fusion analysis result is in different clients of the same file system.
[0174] The third judgment unit is configured to judge whether the fault analysis in the fusion analysis result is in different file systems of the same network file system server if the fault analysis in the fusion analysis result is not in different clients of the same file system.
[0175] The third determination unit is configured to determine that the fault location result is a network file system server fault if the fault analysis in the fusion analysis result is in different file systems of the same network file system server.
[0176] The second detection unit is configured to detect whether the fault analysis in the fusion analysis result is in different file systems of the same storage if the fault analysis in the fusion analysis result is not in different file systems of the same network file system server.
[0177] The fourth determination unit is configured to determine that the fault location result is a storage pool fault if the fault analysis in the fusion analysis result is in different file systems of the same storage.
[0178] Optionally, in a fault handling device of a distributed file storage system provided in another embodiment of the present application, when the fault location result is a single client fault, the fault handling unit 604 includes:
[0179] The program acquisition unit is configured to acquire a preset fault recovery program.
[0180] The fault recovery unit is configured to automatically recover the fault of the client by using the fault recovery program and feedback the fault information of the client to the front end.
[0181] Optionally, in a fault handling device of a distributed file storage system provided in another embodiment of the present application, when the fault location result is a single file system fault, the fault handling unit 604 includes:
[0182] The first migration unit is configured to trigger a migration operation of the network file system server to migrate the faulty file system to other available network file system servers.
[0183] The first feedback unit is used to feedback information about the single - file system failure to the front - end when the migration operation is completed.
[0184] Optionally, in a fault handling device of a distributed file storage system provided by another embodiment of the present application, when the fault location result is a network file system server failure, the fault handling unit 604 includes:
[0185] The second migration unit is used to trigger the migration operation of the network file system server to migrate the data on the network file system server to the standby server.
[0186] The second feedback unit is used to feedback the alarm information of the network file system server to the front - end when the migration operation is completed.
[0187] Optionally, in a fault handling device of a distributed file storage system provided by another embodiment of the present application, the fault handling unit 604 includes:
[0188] The third feedback unit is used to call the alarm module to feedback the alarm information of the storage pool to the front - end based on the fault location result when the fault location result is a storage pool failure.
[0189] It should be noted that the specific working processes of the respective units provided in the above - mentioned embodiments of the present application can be correspondingly referred to the corresponding steps in the above - mentioned method embodiments, and will not be elaborated here.
[0190] It should also be noted that the fault handling device of a distributed file storage system provided by the embodiments of the present application has the technical effects of any one of the above - mentioned embodiments, and will not be elaborated here.
[0191] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0192] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A fault handling method for a distributed file storage system, characterized in that Including: Collecting and storing the data information of the system; wherein, the data information at least includes the first data information of the client, the second data information of the network file system server, and the third data information of the storage pool; Respectively performing fault analysis on the first data information, the second data information, and the third data information to obtain a first analysis result corresponding to the first data information, a second analysis result corresponding to the second data information, and a third analysis result corresponding to the third data information; Performing fusion analysis on the first analysis result, the second analysis result, and the third analysis result to obtain a fault location result; Performing fault handling on the storage system according to the fault location result.
2. The method according to claim 1, characterized in that, The step of respectively performing fault analysis on the first data information, the second data information, and the third data information to obtain a first analysis result corresponding to the first data information, a second analysis result corresponding to the second data information, and a third analysis result corresponding to the third data information includes: Dividing the first data information to obtain numerical data and log data, and dividing the third data information to obtain target numerical data and target log data; Based on a preset threshold, performing fault analysis on the numerical data and the target numerical data, and performing keyword scanning on the log data and the target log data; When the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data pass the keyword scanning, or the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data do not pass the keyword scanning, generating a first analysis result and a third analysis result according to the first data information and the third data information; When the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data pass the keyword scanning, generating a first analysis result and a third analysis result based on the numerical data and the target numerical data; When the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data do not pass the keyword scanning, generating a first analysis result and a third analysis result according to the log data and the target log data; Judging whether there is data exceeding the preset threshold in the second data information; If there is data exceeding the preset threshold in the second data information, generating a second analysis result based on the data exceeding the preset threshold; If there is no data exceeding the preset threshold in the second data information, generating a second analysis result based on the second data information.
3. The method according to claim 1, characterized in that, The step of performing fusion analysis on the first analysis result, the second analysis result, and the third analysis result to obtain a fault location result includes: Fusing the first analysis result, the second analysis result, and the third analysis result to obtain a fusion analysis result; Judging whether the fault analysis in the fusion analysis result is concentrated on the client; If the fault analysis in the fusion analysis result focuses on the client, determine that the fault location result is a single client fault; If the fault analysis in the fusion analysis result does not focus on the client, detect whether the fault analysis in the fusion analysis result is in different clients of the same file system; If the fault analysis in the fusion analysis result is in different clients of the same file system, determine that the fault location result is a single file system fault; If the fault analysis in the fusion analysis result is not in different clients of the same file system, determine whether the fault analysis in the fusion analysis result is in different file systems of the same network file system server; If the fault analysis in the fusion analysis result is in different file systems of the same network file system server, determine that the fault location result is a network file system server fault; If the fault analysis in the fusion analysis result is not in different file systems of the same network file system server, detect whether the fault analysis in the fusion analysis result is in different file systems of the same storage; If the fault analysis in the fusion analysis result is in different file systems of the same storage, determine that the fault location result is a storage pool fault.
4. The method according to claim 3, wherein When the fault location result is a single client fault, the fault handling of the storage system according to the fault location result includes: Obtain a pre-set fault recovery program; Use the fault recovery program to automatically recover the fault of the client and feedback the fault information of the client to the front end.
5. The method according to claim 3, characterized in that, When the fault location result is a single file system fault, the fault handling of the storage system according to the fault location result includes: Trigger the migration operation of the network file system server to migrate the faulty file system to other available network file system servers; When the migration operation is completed, feedback the information of the single file system fault to the front end.
6. The method according to claim 3, characterized in that, When the fault location result is a network file system server fault, the fault handling of the storage system according to the fault location result includes: Trigger the migration operation of the network file system server to migrate the data on the network file system server to the standby server; When the migration operation is completed, feedback the alarm information of the network file system server to the front end.
7. The method according to claim 3, characterized in that, The fault handling of the storage system according to the fault location result includes: When the fault location result is a storage pool fault, call the alarm module to feedback the alarm information of the storage pool to the front end based on the fault location result.
8. A fault handling device for a distributed file storage system, characterized in that Include: An information collection unit for collecting data information of the storage system; wherein, the data information at least includes the first data information of the client, the second data information of the network file system server, and the third data information of the storage pool; A fault analysis unit for separately performing fault analysis on the first data information, the second data information, and the third data information to obtain a first analysis result corresponding to the first data information, a second analysis result corresponding to the second data information, and a third analysis result corresponding to the third data information; A fusion analysis unit for performing fusion analysis on the first analysis result, the second analysis result, and the third analysis result to obtain a fault location result; A fault handling unit for performing fault handling on the storage system according to the fault location result.
9. The device according to claim 8, characterized in that, The fault analysis unit includes: A partitioning unit for partitioning the first data information to obtain numerical data and log data, and partitioning the third data information to obtain target numerical data and target log data; A processing unit for performing fault analysis on the numerical data and the target numerical data based on a preset threshold, and performing keyword scanning on the log data and the target log data; A first generation unit for generating a first analysis result and a third analysis result according to the first data information and the third data information when the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data pass the keyword scanning, or when the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data do not pass the keyword scanning; A second generation unit for generating a first analysis result and a third analysis result based on the numerical data and the target numerical data when the numerical data and the target numerical data do not pass the fault analysis, and the log data and the target log data pass the keyword scanning; A third generation unit for generating a first analysis result and a third analysis result according to the log data and the target log data when the numerical data and the target numerical data pass the fault analysis, and the log data and the target log data do not pass the keyword scanning; A first judgment unit for judging whether there is data exceeding a preset threshold in the second data information; A fourth generation unit for generating a second analysis result based on the data exceeding the preset threshold if there is data exceeding the preset threshold in the second data information; A fifth generation unit for generating a second analysis result based on the second data information if there is no data exceeding the preset threshold in the second data information.
10. The device according to claim 8, characterized in that, The fusion analysis unit includes: A fusion unit for fusing the first analysis result, the second analysis result, and the third analysis result to obtain a fusion analysis result; A second judgment unit for judging whether the fault analysis in the fusion analysis result is concentrated on the client; A first determination unit for determining the fault location result as a single client fault if the fault analysis in the fusion analysis result is concentrated on the client; The first detection unit is used to detect whether the fault analysis in the fusion analysis result is in different clients of the same file system if the fault analysis in the fusion analysis result is not concentrated in the client; The second determination unit is used to determine that the fault location result is a single file system fault if the fault analysis in the fusion analysis result is in different clients of the same file system; The third judgment unit is used to judge whether the fault analysis in the fusion analysis result is in different file systems of the same network file system server if the fault analysis in the fusion analysis result is not in different clients of the same file system; The third determination unit is used to determine that the fault location result is a network file system server fault if the fault analysis in the fusion analysis result is in different file systems of the same network file system server; The second detection unit is used to detect whether the fault analysis in the fusion analysis result is in different file systems of the same storage if the fault analysis in the fusion analysis result is not in different file systems of the same network file system server; The fourth determination unit is used to determine that the fault location result is a storage pool fault if the fault analysis in the fusion analysis result is in different file systems of the same storage.