Method and device for determining slow disk in HDFS system

By comprehensively analyzing the multiple indicator data of disks in HDFS systems, slow disks are determined, which solves the problem of difficulty in discovering slow disks in a timely manner, and improves identification accuracy and cluster performance stability.

CN120029843APending Publication Date: 2025-05-23INNER MONGOLIA YILI IND GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311577819.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In HDFS systems, slow disks are difficult to detect and process in time, resulting in the stability of cluster read and write performance.

Method used

By conducting a comprehensive analysis of the fsck log files, DataNode log files, read and write performance, multiple read and write tests and delay conditions of disks in HDFS system, multiple metric data are determined, multiple analysis results are generated, and the slow disk in HDFS is finally determined.

Benefits of technology

It realizes timely discovery and accurate identification of slow disks in HDFS systems, improves the accuracy of slow disk recognition, and ensures stable read and write performance of the cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029843A_ABST
    Figure CN120029843A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for determining a slow disk in an HDFS (Hadoop Distributed File System). The method comprises the following steps: analyzing an fsck log file of a disk in the HDFS to determine fsck log index data; the method comprises the following steps: analyzing a DataNode log file of a disk in an HDFS (Hadoop Distributed File System) to determine DataNode log index data; analyzing the read-write performance of the disk in the HDFS system to determine read-write performance index data; performing a read-write test on a disk in the HDFS system to determine read-write test performance index data; analyzing the delay condition of the disk in the HDFS system to determine read-write delay index data; analyzing the possibility that the disk is a slow disk according to the index data, and generating a plurality of analysis results; according to the method, the slow disk in the HDFS can be found in time, and the slow disk recognition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and device for determining a slow disk in an HDFS system. Background Art

[0002] This section is intended to provide a background or context for embodiments of the present invention. No description herein is admitted to be prior art by virtue of its inclusion in this section.

[0003] HDFS is a distributed file system under the Hadoop ecosystem. Hadoop is mainly designed to batch process large files with large amounts of data. Most of the disks used to store data in Hadoop's DataNode are ordinary HDD disks. As time goes by, it is inevitable that the disks will age and their performance will degrade. The main manifestations are slower disk reading and writing and slower network transmission. We collectively call these disks with degraded performance slow disks, and DataNode server nodes with slow disks slow nodes. When the cluster expands to a certain scale, such as a cluster of hundreds or thousands of nodes, slow nodes and slow disks are usually not easy to be discovered. Most of the time, slow nodes are hidden among many healthy nodes. They will only be perceived when clients frequently access these problematic nodes and find that reading and writing have slowed down.

[0004] Therefore, in order to maintain stable read and write performance of the HDFS cluster, slow disks must be discovered and processed in a timely manner. Summary of the invention

[0005] An embodiment of the present invention provides a method for determining a slow disk in an HDFS system, which is used to timely discover a slow disk and improve the accuracy of slow disk identification. The method includes:

[0006] Parse the fsck log files of the disk in the HDFS system to determine the fsck log indicator data; the fsck log indicator data is the verification time of the DataNode node containing the data block;

[0007] Parse the DataNode log files on the disk in the HDFS system to determine the DataNode log indicator data; the DataNode log indicator data is the disk read and write duration;

[0008] Analyze the read and write performance of the disk in the HDFS system and determine the read and write performance index data; the read and write performance index data is the read and write performance parameters of the disk;

[0009] Perform multiple read and write tests on the disks in the HDFS system to determine the read and write test performance index data; the read and write test performance index data is the average read and write test performance data of multiple read and write tests;

[0010] Analyze the disk latency in the HDFS system and determine the read and write latency indicator data; the read and write latency indicator data includes the disk read and write operation data, read and write byte data, and read and write latency data;

[0011] According to the fsck log indicator data, DataNode log indicator data, read and write performance indicator data, read and write test performance indicator data, and read and write delay indicator data, the possibility of the disk being a slow disk is analyzed to generate multiple analysis results;

[0012] Identify slow disks in HDFS based on multiple analysis results.

[0013] The embodiment of the present invention further provides a slow disk determination device in an HDFS system, which is used to timely discover slow disks and improve the accuracy of slow disk identification. The device includes:

[0014] The first indicator data determination module is used to parse the fsck log file of the disk in the HDFS system to determine the fsck log indicator data; the fsck log indicator data is the verification time length of the DataNode node containing the data block;

[0015] The second indicator data determination module is used to parse the DataNode log file of the disk in the HDFS system to determine the DataNode log indicator data; the DataNode log indicator data is the disk read and write time;

[0016] The third indicator data determination module is used to analyze the read and write performance of the disk in the HDFS system and determine the read and write performance indicator data; the read and write performance indicator data is the read and write performance parameters of the disk;

[0017] The fourth indicator data determination module is used to perform multiple read and write tests on the disk in the HDFS system to determine the read and write test performance indicator data; the read and write test performance indicator data is the average read and write test performance data of the multiple read and write tests;

[0018] The fifth indicator data determination module is used to analyze the delay of the disk in the HDFS system and determine the read and write delay indicator data; the read and write delay indicator data is the disk's read and write operation data, read and write byte data, and read and write delay data;

[0019] An analysis module is used to analyze the possibility that the disk is a slow disk according to fsck log indicator data, DataNode log indicator data, read and write performance indicator data, read and write test performance indicator data, and read and write delay indicator data, and generate multiple analysis results;

[0020] The slow disk determination module is used to determine the slow disk in the HDFS according to multiple analysis results.

[0021] An embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for determining a slow disk in the HDFS system when executing the computer program.

[0022] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for determining a slow disk in the HDFS system is implemented.

[0023] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the method for determining a slow disk in the HDFS system is implemented.

[0024] Compared with the method of identifying slow disks by a single indicator in the prior art, the embodiment of the present invention parses the fsck log file of the disk in the HDFS system to determine the fsck log indicator data; the fsck log indicator data is the verification time of the DataNode node containing the data block; the DataNode log file of the disk in the HDFS system is parsed to determine the DataNode log indicator data; the DataNode log indicator data is the disk read and write time; the read and write performance of the disk in the HDFS system is analyzed to determine the read and write performance indicator data; the read and write performance indicator data is the read and write performance parameters of the disk; and the disk in the HDFS system is read and written multiple times. Test to determine the read-write test performance indicator data; the read-write test performance indicator data is the average read-write test performance data of multiple read-write tests; analyze the delay of the disk in the HDFS system to determine the read-write delay indicator data; the read-write delay indicator data is the disk's read-write operation data, read-write byte data, and read-write delay data; analyze the possibility of the disk being a slow disk based on the fsck log indicator data, DataNode log indicator data, read-write performance indicator data, read-write test performance indicator data, and read-write delay indicator data, and generate multiple analysis results; determine the slow disk in HDFS based on multiple analysis results, so that the slow disk can be discovered in time and the accuracy of slow disk identification can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0026] Figure 1A schematic diagram of the process of determining a slow disk in an HDFS system provided by the present invention;

[0027] Figure 2 A flowchart of a specific example of a method for determining a slow disk in an HDFS system provided by the present invention;

[0028] Figure 3 A flowchart of a specific example of a method for determining a slow disk in an HDFS system provided by the present invention;

[0029] Figure 4 A schematic diagram of the structure of a slow disk determination device in an HDFS system provided by the present invention;

[0030] Figure 5 A schematic diagram of the computer device structure provided by the present invention. DETAILED DESCRIPTION

[0031] To make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0032] First, let’s introduce the professional terms involved in this article:

[0033] Hadoop: is an open source framework for efficiently storing and processing large data sets ranging from GB to PB.

[0034] HDFS: Hadoop's distributed file system.

[0035] NameNode: NameNode is also called master node in HDFS architecture. HDFS NameNode stores metadata, such as number of data blocks, number of replicas and other details. This metadata is stored in the memory of the master node to achieve the fastest data retrieval speed. NameNode is responsible for maintaining and managing slave nodes and assigning tasks to slave nodes.

[0036] DataNode: HDFS data storage node, which interacts with NameNode and regularly reports its own storage data information.

[0037] In order to more accurately identify disks with performance bottlenecks in an HDFS system cluster, that is, slow disks, an embodiment of the present invention provides a method for determining slow disks in an HDFS system. Figure 1 A flowchart of a method for determining a slow disk in an HDFS system provided by an embodiment of the present invention is shown as follows: Figure 1 As shown, the method includes:

[0038] Step 101, parse the fsck log file of the disk in the HDFS system to determine fsck log indicator data; the fsck log indicator data is the verification time of the DataNode node containing the data block;

[0039] Step 102, parse the DataNode log file of the disk in the HDFS system to determine the DataNode log indicator data; the DataNode log indicator data is the disk read and write duration;

[0040] Step 103, analyzing the read and write performance of the disk in the HDFS system to determine read and write performance index data; the read and write performance index data is the read and write performance parameters of the disk;

[0041] Step 104, performing multiple read and write tests on the disk in the HDFS system to determine read and write test performance index data; the read and write test performance index data is the average read and write test performance data of multiple read and write tests;

[0042] Step 105, analyzing the delay of the disk in the HDFS system to determine the read and write delay index data; the read and write delay index data is the disk's read and write operation data, read and write byte data, and read and write delay data;

[0043] Step 106, analyzing the possibility that the disk is a slow disk according to the fsck log indicator data, the DataNode log indicator data, the read / write performance indicator data, the read / write test performance indicator data, and the read / write delay indicator data, and generating multiple analysis results;

[0044] Step 107: Determine the slow disk in the HDFS according to the multiple analysis results.

[0045] The embodiment of the present invention realizes accurate monitoring of slow disks in the HDFS system by comprehensively analyzing multiple indicator data, and evaluates and scores each indicator. Disks with low comprehensive scores will be judged as slow disks.

[0046] In one embodiment, Figure 2 As shown, the fsck log file of the disk in the HDFS system is parsed to determine the fsck log indicator data, which may include:

[0047] Step 201, running the fsck command on the NameNode node in the HDFS system, and outputting the running result to the fsck log file;

[0048] Step 202, determining the total time spent on verifying data blocks in the HDFS system in the fsck log file;

[0049] Step 203, determining the total verification time of the DataNode node containing the data block according to the total time spent on verifying the data block in the HDFS system;

[0050] Step 204: determine the total verification time of the DataNode node containing the data block as fsck log indicator data.

[0051] In this embodiment, a script is used to automatically parse the fsck.log log and count the check duration of each DataNode. A threshold is set, and the DataNode whose check duration exceeds the threshold can be judged to have slow disk performance.

[0052] You can run the fsck command on the NameNode node and output the results to the file: hdfs fsck / -files-blocks-locations>fsck.log; analyze the fsck.log log file and specifically check the line "Total time spent after filesystem closed". This line gives the total time spent on verifying each data block. Based on the location information, the time is allocated to each DataNode node that contains the data block. Calculate the total verification time on each DataNode. Compare the total verification time of different DataNodes. The longer the time, the slower the disk performance on the node.

[0053] In one embodiment, Figure 3 As shown, the DataNode log files on the disk in the HDFS system are parsed to determine the DataNode log indicator data, which may include:

[0054] Step 301, perform keyword search in the DataNode log file on the disk of the HDFS system to determine the slow read and write records;

[0055] Step 302, determining the read and write time of each file in the disk according to the slow read and write records;

[0056] Step 303, determining the slow file according to the read and write time of each file and the preset read and write time threshold;

[0057] Step 304, determining the disk where each slow file is located;

[0058] Step 305, determining the number of slow files on each disk and the read and write time of each disk;

[0059] Step 306: Determine the number of slow files on each disk and the read and write duration of each disk as DataNode log indicator data.

[0060] In this embodiment, a script can be used to statistically analyze DataNode logs and identify slow disks based on slow read and write disks and files. On each DataNode host, analyze the log files, which are generally located in the / var / log / hadoop / directory. Search the log keyword "slow" to detect slow IO records. Aggregate these slow IO records and count the read and write time of each file. Locate the specific disk where each slow IO file is located. You can view the disk where the file is located through the command: df-h / path / to / file. Count the number of slow IO files on each disk and the total read and write time. Compare the slow IO statistics of different disks, and the disk with a higher total read and write time can be determined as a slow disk. Set a threshold, and the disk that exceeds the threshold is confirmed as a slow disk. And give the disk a score based on the difference between the actual value and the threshold.

[0061] In one embodiment, analyzing the read and write performance of the disk in the HDFS system and determining the read and write performance indicator data may include: using the nmon tool configured in each DataNode node of the disk to determine the read and write performance parameters of each disk; the read and write performance parameters include one or any combination of read and write throughput, IOPS, and response time; and determining the read and write performance parameters as the read and write performance indicator data.

[0062] In this embodiment, disk monitoring and collection tools such as nmon can be used to count the IO performance indicators of different HDFS disks. The steps for judging slow disks are as follows: Install and configure the nmon tool on each DataNode node. Run the nmon command to count the IO read and write throughput, IOPS, response time and other parameters of each disk device. Command example: nmon-FHDFSdata-s 5-c 288. Extract the performance data of each disk from the nmon output report. Calculate the average and peak values ​​of the read and write throughput and IOPS of each disk. Compare and sort the indicator data of different disks. Determine the disk with lower performance indicator values, such as the throughput that is 50% lower than the cluster average. Combine multiple performance indicators to judge the disk with poor read and write performance. Set a reasonable threshold and determine it as a slow disk. And give the disk a score based on the difference between the actual value and the threshold.

[0063] In one embodiment, multiple read and write tests are performed on the disk in the HDFS system to determine the read and write test performance indicator data, which may include: performing multiple read and write tests on the disk using a read and write performance test tool configured on the DataNode node to generate multiple read and write test performance data; the read and write tests include random read and write tests or sequential read and write tests; determining average read and write test performance data based on multiple read and write test performance data; and determining the average read and write test performance data as the read and write test performance indicator data.

[0064] In this embodiment, tools such as dd and iouzone can be used to perform read and write tests on different disks, and the disks with poor test performance are judged as slow disks. Install disk read and write performance test tools such as dd, iouzone, and fio on the DataNode node. Use these tools to test each disk, such as testing the sequential read and write performance of the disk: dd if= / dev / sda of= / tmp / test bs=1M count=1024. Test different IO types, including random read and write, sequential read and write, etc. Repeat the test many times to calculate the average performance index. Compare the test results of different disks and calculate their relative performance differences. Disks with large performance index differences will be judged as slow disks. For example, the read and write speed is 50% lower than the average level. Set a reasonable threshold to determine the slow disk. And give the disk a score based on the difference between the actual value and the threshold.

[0065] In one embodiment, analyzing the delay of the disk in the HDFS system and determining the read and write delay index data may include: enabling statistics of the delay index, obtaining the read and write delay of the disk in the HDFS system on the DataNode index page; receiving a command to enable the delay index; analyzing the delay of the disk according to the command, recording the read and write operation data, read and write byte data, and read and write delay data of the disk; and determining the read and write operation data, read and write byte data, and read and write delay data of the disk as the read and write delay index data.

[0066] In this example, you can enable latency metric statistics and obtain the read and write latency of each disk on the DataNode Metrics page. Edit the hadoop-metrics2.properties file of the DataNode and add the following content to enable latency metrics:

[0067] dfs.datanode.metrics.logger.operation.sampling.period=1000;

[0068] dfs.datanode.metrics.logger.operation.sampling.percentage=1.0;

[0069] Restart the DataNode service and load the new configuration. Access the DataNode Web UI and open the "DataNodeMetrics" page. Select the "Operation" metric group on the Metrics page. Find the "DataNode Activity" table, which reports the number of read and write operations, the number of read and write bytes, and the read and write latency for each disk. Note that the latency metric may take some time to run before it starts reporting data. Use a script to count latency data by time period and analyze the latency of each disk. Set reasonable thresholds to identify slow disks. And give the disk a score based on the difference between the actual value and the threshold.

[0070] In one embodiment, the possibility of a disk being a slow disk is analyzed based on fsck log indicator data, DataNode log indicator data, read-write performance indicator data, read-write test performance indicator data, and read-write delay indicator data to generate multiple analysis results, which may include: determining the differences between the fsck log indicator data, DataNode log indicator data, read-write performance indicator data, read-write test performance indicator data, and read-write delay indicator data and preset corresponding thresholds to generate multiple differences; analyzing the possibility of the disk being a slow disk based on each difference to generate multiple analysis results.

[0071] In this embodiment, a reasonable threshold may be set, and the possibility of the disk being a slow disk is analyzed based on the difference between the actual value and the threshold, and the disk is given a score.

[0072] In one embodiment, determining the slow disk in HDFS according to multiple analysis results may include: assigning weights to the multiple analysis results respectively; and determining the slow disk in HDFS according to the multiple analysis results and the corresponding weights.

[0073] In this embodiment, a weight may be assigned to each of the multiple index scores, and the multiple scores may be aggregated, and a disk with a lower total score may be determined as a slow disk.

[0074] The present invention also provides a device for determining a slow disk in an HDFS system, as described in the following embodiments. Figure 4 As shown, the device comprises:

[0075] The first indicator data determination module 401 is used to parse the fsck log file of the disk in the HDFS system to determine the fsck log indicator data; the fsck log indicator data is the verification time length of the DataNode node containing the data block;

[0076] The second indicator data determination module 402 is used to parse the DataNode log file of the disk in the HDFS system to determine the DataNode log indicator data; the DataNode log indicator data is the disk read and write time;

[0077] The third indicator data determination module 403 is used to analyze the read and write performance of the disk in the HDFS system and determine the read and write performance indicator data; the read and write performance indicator data is the read and write performance parameters of the disk;

[0078] The fourth indicator data determination module 404 is used to perform multiple read and write tests on the disk in the HDFS system to determine the read and write test performance indicator data; the read and write test performance indicator data is the average read and write test performance data of the multiple read and write tests;

[0079] The fifth indicator data determination module 405 is used to analyze the delay of the disk in the HDFS system and determine the read and write delay indicator data; the read and write delay indicator data is the read and write operation data, read and write byte data and read and write delay data of the disk;

[0080] An analysis module 406 is used to analyze the possibility that the disk is a slow disk according to the fsck log indicator data, the DataNode log indicator data, the read / write performance indicator data, the read / write test performance indicator data, and the read / write delay indicator data, and generate multiple analysis results;

[0081] The slow disk determination module 407 is used to determine the slow disk in the HDFS according to the multiple analysis results.

[0082] In one embodiment, the first indicator data determination module 401 is specifically used to:

[0083] Run the fsck command on the NameNode node in the HDFS system and output the running results to the fsck log file;

[0084] Determine the total time spent verifying data blocks in the HDFS system in the fsck log file;

[0085] Determine the total verification time of the DataNode node containing the data block based on the total time spent verifying the data block in the HDFS system;

[0086] The total verification time of the DataNode node containing the data block is determined as fsck log indicator data.

[0087] In one embodiment, the second indicator data determination module 402 is specifically used to:

[0088] Perform keyword searches in the DataNode log files on the HDFS system disk to identify slow read and write records;

[0089] Determine the read and write time of each file in the disk based on the slow read and write records;

[0090] Determine the slow files based on the read and write time of each file and the preset read and write time threshold;

[0091] Determine the disk where each slow file is located;

[0092] Determine the number of slow files on each disk and the read and write time of each disk;

[0093] The number of slow files on each disk and the read and write time of each disk are determined as DataNode log indicator data.

[0094] In one embodiment, the third indicator data determination module 403 is specifically used to:

[0095] Determine the read and write performance parameters of each disk using the nmon tool configured in each DataNode node of the disk; the read and write performance parameters include one or any combination of read and write throughput, IOPS, and response time;

[0096] The read-write performance parameter is determined as read-write performance indicator data.

[0097] In one embodiment, the fourth indicator data determination module 404 is specifically used to:

[0098] Use the read / write performance test tool configured on the DataNode node to perform multiple read / write tests on the disk to generate multiple read / write test performance data; the read / write test includes random read / write test or sequential read / write test;

[0099] Determine average read-write test performance data based on multiple read-write test performance data;

[0100] The average read and write test performance data is determined as the read and write test performance indicator data.

[0101] In one embodiment, the fifth indicator data determination module 405 is specifically used to:

[0102] Enable latency metric statistics and obtain the read and write latency of the disk in the HDFS system on the DataNode metric page.

[0103] Receive a command to enable latency indicators;

[0104] Analyze the disk delay according to the command, and record the disk read and write operation data, read and write byte data, and read and write delay data;

[0105] The read and write operation data, read and write byte data and read and write delay data of the disk are determined as read and write delay indicator data.

[0106] In one embodiment, the analysis module 406 is specifically configured to:

[0107] Determine the differences between fsck log indicator data, DataNode log indicator data, read / write performance indicator data, read / write test performance indicator data, and read / write delay indicator data and preset corresponding thresholds respectively, and generate multiple differences;

[0108] The possibility of the disk being a slow disk is analyzed according to each difference value, and multiple analysis results are generated.

[0109] In one embodiment, the slow disk determination module 407 is specifically used to:

[0110] Assign weights to multiple analysis results;

[0111] Determine the slow disk in HDFS based on multiple analysis results and corresponding weights.

[0112] Since the principle of solving the problem by the device is similar to the slow disk determination method in the HDFS system, the implementation of the device can refer to the implementation of the slow disk determination method in the HDFS system, and the repeated parts will not be repeated.

[0113] Based on the above invention concept, Figure 5 As shown, the present invention also proposes a computer device 500, including a memory 510, a processor 520, and a computer program 530 stored in the memory 510 and executable on the processor 520, wherein the processor 520 implements the above-mentioned slow disk determination method in the HDFS system when executing the computer program 530.

[0114] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for determining a slow disk in the HDFS system is implemented.

[0115] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the method for determining a slow disk in the HDFS system is implemented.

[0116] Compared with the method of identifying slow disks by a single indicator in the prior art, the embodiment of the present invention parses the fsck log file of the disk in the HDFS system to determine the fsck log indicator data; the fsck log indicator data is the verification time of the DataNode node containing the data block; the DataNode log file of the disk in the HDFS system is parsed to determine the DataNode log indicator data; the DataNode log indicator data is the disk read and write time; the read and write performance of the disk in the HDFS system is analyzed to determine the read and write performance indicator data; the read and write performance indicator data is the read and write performance parameters of the disk; the disk in the HDFS system is read and written multiple times. Write test to determine the read and write test performance indicator data; the read and write test performance indicator data is the average read and write test performance data of multiple read and write tests; analyze the delay of the disk in the HDFS system to determine the read and write delay indicator data; the read and write delay indicator data is the disk's read and write operation data, read and write byte data, and read and write delay data; analyze the possibility of the disk being a slow disk based on the fsck log indicator data, DataNode log indicator data, read and write performance indicator data, read and write test performance indicator data, and read and write delay indicator data, and generate multiple analysis results; determine the slow disk in HDFS based on multiple analysis results, so that the slow disk can be discovered in time and the accuracy of slow disk identification can be improved.

[0117] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0118] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0119] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0120] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0121] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for determining slow disks in HDFS system. It is characterized in that include: Parse the fsck log files of the disk in the HDFS system to determine the fsck log indicator data; the fsck log indicator data is the verification time of the DataNode node containing the data block; Parse the DataNode log files on the disk in the HDFS system to determine the DataNode log indicator data; the DataNode log indicator data is the disk read and write duration; Analyze the read and write performance of the disk in the HDFS system and determine the read and write performance index data; the read and write performance index data is the read and write performance parameters of the disk; Perform multiple read and write tests on the disks in the HDFS system to determine the read and write test performance indicator data; The read and write test performance indicator data is the average read and write test performance data of multiple read and write tests; Analyze the disk latency in the HDFS system and determine the read and write latency indicator data; the read and write latency indicator data includes the disk read and write operation data, read and write byte data, and read and write latency data; According to the fsck log indicator data, DataNode log indicator data, read and write performance indicator data, read and write test performance indicator data, and read and write delay indicator data, the possibility of the disk being a slow disk is analyzed to generate multiple analysis results; Identify slow disks in HDFS based on multiple analysis results.

2. The method according to claim 1, It is characterized in that Parse the fsck log files of the disks in the HDFS system to determine the fsck log indicator data, including: Run the fsck command on the NameNode node in the HDFS system and output the running results to the fsck log file; Determine the total time spent verifying data blocks in the HDFS system in the fsck log file; Determine the total verification time of the DataNode node containing the data block based on the total time spent verifying the data block in the HDFS system; The total verification time of the DataNode node containing the data block is determined as fsck log indicator data.

3. The method according to claim 1, It is characterized in that Parse the DataNode log files on the disk in the HDFS system to determine the DataNode log indicator data, including: Perform keyword searches in the DataNode log files on the HDFS system disk to identify slow read and write records; Determine the read and write time of each file in the disk based on the slow read and write records; Determine the slow files based on the read and write time of each file and the preset read and write time threshold; Determine the disk where each slow file is located; Determine the number of slow files on each disk and the read and write time of each disk; The number of slow files on each disk and the read and write time of each disk are determined as DataNode log indicator data.

4. The method according to claim 1, It is characterized in that Analyze the read and write performance of the disk in the HDFS system and determine the read and write performance indicator data, including: Determine the read and write performance parameters of each disk using the nmon tool configured in each DataNode node of the disk; the read and write performance parameters include one or any combination of read and write throughput, IOPS, and response time; The read-write performance parameter is determined as read-write performance indicator data.

5. The method according to claim 1, It is characterized in that Perform multiple read and write tests on the disks in the HDFS system to determine the read and write test performance indicator data, including: Use the read / write performance test tool configured on the DataNode node to perform multiple read / write tests on the disk to generate multiple read / write test performance data; the read / write test includes random read / write test or sequential read / write test; Determine average read-write test performance data based on multiple read-write test performance data; The average read and write test performance data is determined as the read and write test performance indicator data.

6. The method according to claim 1, It is characterized in that Analyze the disk latency in the HDFS system and determine the read and write latency indicator data, including: Enable latency metric statistics and obtain the read and write latency of the disk in the HDFS system on the DataNode metric page. Receive a command to enable latency indicators; Analyze the disk delay according to the command, and record the disk read and write operation data, read and write byte data, and read and write delay data; The read and write operation data, read and write byte data and read and write delay data of the disk are determined as read and write delay indicator data.

7. The method according to claim 1, It is characterized in that Based on the fsck log indicator data, DataNode log indicator data, read and write performance indicator data, read and write test performance indicator data, and read and write delay indicator data, the possibility of the disk being a slow disk is analyzed to generate multiple analysis results, including: Determine the differences between the fsck log indicator data, the DataNode log indicator data, the read / write performance indicator data, the read / write test performance indicator data, and the read / write delay indicator data and the preset corresponding thresholds respectively, and obtain multiple differences; The possibility of the disk being a slow disk is analyzed according to each difference value, and multiple analysis results are generated.

8. The method according to claim 1, It is characterized in that Slow disks in HDFS are identified based on multiple analysis results, including: Assign weights to multiple analysis results; Determine the slow disk in HDFS based on multiple analysis results and corresponding weights.

9. A device for determining a slow disk in an HDFS system, It is characterized in that include: The first indicator data determination module is used to parse the fsck log file of the disk in the HDFS system to determine the fsck log indicator data; the fsck log indicator data is the verification time length of the DataNode node containing the data block; The second indicator data determination module is used to parse the DataNode log file of the disk in the HDFS system to determine the DataNode log indicator data; the DataNode log indicator data is the disk read and write time; The third indicator data determination module is used to analyze the read and write performance of the disk in the HDFS system and determine the read and write performance indicator data; the read and write performance indicator data is the read and write performance parameters of the disk; The fourth indicator data determination module is used to perform multiple read and write tests on the disks in the HDFS system to determine the read and write test performance indicator data; The read and write test performance indicator data is the average read and write test performance data of multiple read and write tests; The fifth indicator data determination module is used to analyze the delay of the disk in the HDFS system and determine the read and write delay indicator data; the read and write delay indicator data is the disk's read and write operation data, read and write byte data, and read and write delay data; An evaluation module is used to analyze the possibility that the disk is a slow disk according to fsck log indicator data, DataNode log indicator data, read and write performance indicator data, read and write test performance indicator data, and read and write delay indicator data, and generate multiple analysis results; The slow disk determination module is used to determine the slow disk in the HDFS according to multiple analysis results.

10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

11. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

12. A computer program product, It is characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.