Automatic performance testing method based on HDFS cluster and related equipment
Through automated performance testing methods, using built-in tools for HDFS cluster testing, the problems of limited testing coverage and difficult explanation in the existing technology are solved, and more comprehensive and efficient performance testing is achieved.
Patent Information
- Application Number
- CN202510128464.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-06-20
AI Technical Summary
When performing HDFS cluster performance testing, the test coverage is limited, and it cannot meet the testing needs of different application scenarios, and it is difficult to interpret performance indicators between tools.
Provides an automated performance testing method based on HDFS cluster. By determining test parameters and targets, matching corresponding built-in testing tools (such as TestDFSIO, NNBench), performing performance tests, collecting performance data and generating test reports.
It realizes unified testing and result analysis in different application scenarios, improves test coverage and efficiency, and simplifies result interpretation.
Smart Images

Figure CN120179519A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of electronic technologies, and particularly to an automated performance testing method, apparatus, and server based on an HDFS cluster. Background Art
[0002] In order to ensure the stable and efficient operation of an HDFS cluster in a production environment and be able to handle the growing data processing requirements, it is necessary to periodically perform performance testing on the HDFS cluster.
[0003] Currently, the tool combinations of TestDFSIO, NNBench, and TeraSort are often used to evaluate data throughput, metadata operations, and computing power respectively. However, each tool focuses on different performance metrics. For example, TestDFSIO focuses on the read and write performance of the file system, NNBench focuses on the NameNode performance, and TeraSort focuses on the sorting performance, making the result interpretation difficult. In addition, the test coverage is limited and cannot cover the tests of different application scenarios, failing to meet different test requirements. Summary of the Invention
[0004] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide an automated performance testing method, apparatus, and server based on an HDFS cluster, which can achieve unified testing and result analysis under different application scenarios based on standardized tools and configurations.
[0005] In a first aspect, an embodiment of the present application provides an automated performance testing method based on an HDFS cluster, including:
[0006] Determine the test parameters and test objectives of the HDFS cluster;
[0007] Match the built-in test tool of the HDFS cluster corresponding to the test objective as the target tool;
[0008] According to the test parameters, call the target tool to perform corresponding performance testing;
[0009] Collect the performance data during the task execution of the performance testing, determine the performance performance of the cluster under different configurations, and generate a test report.
[0010] In an embodiment, matching the built-in test tool of the HDFS cluster corresponding to the test objective includes:
[0011] If the test objective is read and write performance testing, use TestDFSIO as the target tool;
[0012] If the test objective is metadata operation testing, use NNBench as the target tool;
[0013] If the test target is high - concurrency testing, both NNBench and TestDFSIO are used as the target tools.
[0014] In one embodiment, if the test target is high - concurrency testing, then according to the test parameters, the target tools are called to perform corresponding performance tests, including:
[0015] Generate a corresponding number of MapReduce tasks according to the initial number of concurrent clients, and call TestDFSIO and NNBench to simulate concurrent clients for data reading / writing and metadata operations;
[0016] Monitor the resource consumption of NameNode and DataNode, and gradually increase the number of concurrent clients until the resources of the cluster are exhausted;
[0017] Obtain the performance data of NameNode and DataNode.
[0018] In one embodiment, it further includes:
[0019] Continuously monitor the cluster load during the performance test to obtain monitoring data;
[0020] Based on the resource optimization principle, dynamically adjust the test parameters based on the monitoring data.
[0021] In one embodiment, the dynamically adjusting the test parameters based on the monitoring data according to the resource optimization principle includes:
[0022] Obtain historical monitoring data; analyze the cluster resource utilization rates under different test parameters according to the historical monitoring data to determine the optimal test parameters; adjust the test parameters according to the optimal test parameters;
[0023] Determine the utilization rate of the cluster resources according to the monitoring data; if the utilization rate is lower than the first threshold or higher than the second threshold, adjust the test parameters.
[0024] In one embodiment, the continuously monitoring the cluster load during the performance test includes:
[0025] Continuously monitor the utilization rates of the CPU, memory, and disk I / O of the cluster during the performance test.
[0026] In one embodiment, the collecting the performance data during the execution of the tasks of the performance test includes:
[0027] Collect the execution performance metrics and resource utilization metrics of each task;
[0028] Among them, the execution performance metrics include at least one of execution time, bandwidth utilization rate, and I / O performance;
[0029] The resource utilization metrics include at least one of the consumption of CPU, memory, network, and disk I / O.
[0030] In one embodiment, it further includes:
[0031] Generate cluster configuration suggestions for optimal performance according to the performance of the cluster under different configurations.
[0032] In a second aspect, an embodiment of the present application provides an automated performance testing device based on an HDFS cluster, including:
[0033] A parameter determination unit for determining the test parameters and test objectives of the HDFS cluster;
[0034] A tool matching unit for matching the built-in test tools of the HDFS cluster corresponding to the test objectives as target tools;
[0035] A test execution unit for calling the target tool to execute corresponding performance tests according to the test parameters;
[0036] A performance evaluation unit for collecting performance data during the task execution of the performance test, determining the performance of the cluster under different configurations, and generating a test report.
[0037] In a third aspect, an embodiment of the present application provides a server, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method described in the embodiments of the present application.
[0038] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0039] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objectives, and advantages of the present application will become more apparent:
[0040] Figure 1 Shows a schematic flowchart of an automated performance testing method based on an HDFS cluster provided by an embodiment of the present application;
[0041] Figure 2 Shows a schematic overall diagram of an implementation architecture provided by an embodiment of the present application;
[0042] Figure 3Shows the implementation flowchart of intelligent parameter adjustment provided by the embodiments of the present application;
[0043] Figure 4 Shows the exemplary structural block diagram of the automated performance testing device based on the HDFS cluster provided by the embodiments of the present application;
[0044] Figure 5 Shows the schematic structural diagram of the computer system of the server suitable for implementing the embodiments of the present application. Detailed implementation manners
[0045] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that, for the sake of convenience of description, only the parts related to the invention are shown in the drawings.
[0046] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and embodiments. Although the embodiments of the present application provide the method operation instruction steps as shown in the following embodiments or drawings, more or fewer operation instruction steps may be included in the method based on routine or non-creative labor. In the steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of the present application. When the method is actually processed or the device executes, it can be executed in the method order shown in the embodiments or drawings or executed in parallel.
[0047] Please refer to Figure 1 , Figure 1 Shows the flowchart of the automated performance testing method based on the HDFS cluster provided by an embodiment of the present application. As Figure 1 shown, the method includes:
[0048] S101. Determine the test parameters and test objectives of the HDFS cluster;
[0049] Determine the HDFS (Hadoop Distributed File System) cluster to be tested, determine the number of cluster nodes (DataNode, NameNode) and the configuration of each node (CPU, memory, hard disk type, etc.) as the test object.
[0050] Test parameters refer to the configuration items when performing performance testing on the HDFS cluster, which can include file size (-fileSize), number of files (-nrFiles), and MapReduce concurrency (-maps), etc. You can set the corresponding types of test parameters according to different test needs, and there is no limitation here.
[0051] Test objectives refer to specific items or scenarios to be verified when performing performance testing on the HDFS cluster, such as testing read and write performance, metadata operations, or high concurrency scenarios. Pre-configuring various test objectives helps ensure that the testing process is more comprehensive and systematic, while improving testing efficiency.
[0052] S102, matching the HDFS cluster built-in test tool corresponding to the test target as the target tool;
[0053] HDFS cluster built-in testing tools such as TestDFSIO (this tool is mainly used to evaluate the throughput and performance of HDFS by simulating file read and write operations), NNBench (this tool is used to simulate requests to NameNode to help test its performance under different loads), etc. Use HDFS built-in tools for performance testing. Built-in testing tools are usually compatible with the HDFS environment, which can avoid environmental interference or compatibility issues caused by introducing external tools. These tools are highly integrated with the HDFS environment and can provide accurate and timely performance data, which helps to improve the stability and operation efficiency of the HDFS cluster.
[0054] Pre-matching each test target with the corresponding optimal HDFS cluster built-in test tool avoids unnecessary selection and configuration by testers during the test process, ensuring the use of the most appropriate tools for targeted testing, reducing human errors while significantly improving test efficiency and making test results more accurate. In addition, new test targets and tools can be added dynamically as requirements change. It only requires ensuring that the correspondence between new tools and targets is correct, without the need to reconfigure complex test processes, which facilitates expansion and maintenance.
[0055] Determine the HDFS cluster built-in test tool (which can be one or more) corresponding to the current target to be tested as the target tool for subsequent calling.
[0056] S103, calling the target tool to execute corresponding performance test according to the test parameters;
[0057] According to the pre-configured test parameters, the target tool is called to automatically execute the test task and perform the corresponding performance test. For example, when executing TestDFSIO, a script is used to automatically submit MapReduce (a programming model and computing framework) jobs for read and write tests.
[0058] S104. Collect the performance data during the task execution of the performance test, determine the performance of the cluster under different configurations, and generate a test report.
[0059] Collect the system performance data during the execution of each task in the performance test. By comprehensively analyzing the system performance data during the execution of each task, determine the performance of the cluster under different test configurations, and thus generate a test report on the cluster under different loads. The presentation forms of the test report can include charts, texts, and the cluster performance curves under different configurations, etc. Among them, the collection of the performance data during the task execution can be achieved by calling Prometheus (a tool responsible for data collection and storage), and the analysis process of the performance based on the performance data can be achieved by calling Grafana (a tool that visualizes and displays data in the forms of charts, dashboards, etc.).
[0060] In addition, in this embodiment, the types of the specifically collected performance data are not limited and can be set accordingly according to different test requirements. Optionally, a type of performance data includes the execution performance metrics and resource utilization metrics of each task; among them, the execution performance metrics include at least one of execution time, bandwidth usage rate, and I / O performance; the resource utilization metrics include at least one of the consumption conditions of CPU, memory, network, and disk I / O. By analyzing the execution time, the response speed of the task can be intuitively understood; the bandwidth usage rate can reveal network bottlenecks; the I / O performance can reflect the performance of the storage system. The resource utilization rate (CPU, memory, disk I / O, etc.) provides the consumption conditions of the hardware resources and can help identify the resource bottlenecks of the system. This configuration provides an all-round perspective for evaluating the system health status and performance, and thus can help comprehensively understand the performance of the system under different test scenarios. It should be noted that only the above data configuration is taken as an example for introduction in this embodiment, but it is not limited to this. The execution processes under other performance data configurations can refer to the introduction of this embodiment and will not be elaborated here.
[0061] It should be noted that although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that these operations must be performed in this specific order, or that all the shown operations must be performed to achieve the desired result.
[0062] Based on the above introduction, the method provided in this embodiment calls the corresponding built-in test tools of the HDFS cluster according to different test objectives to execute the corresponding performance test items. Different scenario simulations are realized through the configuration of different test objectives. For different cluster configurations, multi-dimensional tests can be carried out, covering aspects such as throughput, metadata operation capabilities, and high concurrency, enabling a more comprehensive test of the performance of the cluster. In addition, this method integrates the built-in test tools of HDFS and conducts unified tests based on standardized tools and configurations, facilitating performance comparison between different clusters. Through the writing of automated scripts and the assistance of intelligent algorithms, this method can maximize the test efficiency and reduce human intervention. Therefore, this method realizes a standardized, efficient, and reusable automated performance test for the HDFS cluster.
[0063] In addition, it should be noted that this method supports different hardware configurations (such as the performance difference between solid-state drives and mechanical hard drives) and can test their impact on the cluster performance.
[0064] In the above embodiment, the configured test objectives and the corresponding built-in test tools of the HDFS cluster are not limited. In order to implement automated tests for efficient customer service, a configuration of test objectives is proposed in this embodiment. Specifically, the process of step S102 for matching the built-in test tool of the HDFS cluster corresponding to the test objective can specifically include the following sub-steps:
[0065] S21: If the test objective is read-write performance test, use TestDFSIO as the target tool;
[0066] S22: If the test objective is metadata operation test, use NNBench as the target tool;
[0067] S23: If the test objective is high-concurrency test, use both NNBench and TestDFSIO as the target tools.
[0068] Such as Figure 2The figure shows an overall schematic diagram of an implementation architecture. The above method configures three test targets, namely, read-write performance test, metadata operation test, and high concurrency test, and calls two test tools, TestDFSIO and NNBench. TestDFSIO and NNBench are both test tools that come with HDFS. TestDFSIO is called to test the read-write performance of the cluster. It can generate a large number of files through MapReduce tasks and perform read-write operations to test throughput. NNBench is called to test the metadata operation capabilities of HDFS, including file creation, deletion, and renaming. TestDFSIO and NNBench are executed, and scripts are used to automatically submit MapReduce jobs for read-write or metadata operation tests. For high-concurrency test scenarios, this method calls NNBench and TestDFSIO as target tools at the same time, and uses TestDFSIO and NNBench to simulate concurrent clients for data read-write and metadata operations, avoiding the introduction of additional interference factors and improving the stability of the test.
[0069] Among them, the read and write performance test focuses on the performance of the system when processing data storage and retrieval operations. Through this type of test, the response speed, throughput and latency of the cluster can be evaluated to ensure that the system still performs well when performing large-scale data operations. Metadata is usually data about data in the system (such as file information of the file system, indexes of the database, permission settings, etc.). Through the operation test of metadata, the system's response speed to operations such as reading, modifying, and deleting metadata can be evaluated, as well as the performance under large-scale operations. The high concurrency test tests the performance under high-concurrency requests, simulating the scenario of multiple users performing operations at the same time. This test helps to evaluate the stability, response time and resource utilization of the system under high load, ensuring that the system can maintain stable operation in a high-concurrency environment to avoid performance degradation or crashes.
[0070] Through these three different types of tests, the system's performance can be comprehensively evaluated, bottlenecks can be quickly identified, and the stability, scalability, and availability of the system under different loads and operations can be ensured.
[0071] When the test target is a high-concurrency test, step S103 calls the target tool to perform a corresponding performance test according to the test parameters. A specific test step is as follows:
[0072] S31. Generate a corresponding number of MapReduce tasks according to the initial number of concurrent clients, and call TestDFSIO and NNBench to simulate concurrent clients to perform data reading and writing and metadata operations;
[0073] Set the concurrency of the initial MapReduce (e.g., 100 tasks write to a file simultaneously). Each MapReduce task simulates a client for operation. Initially set 100 concurrent clients to perform file writing and metadata operations.
[0074] S32. Monitor the resource consumption of the NameNode and DataNode, and gradually increase the number of concurrent clients until the resources of the cluster are exhausted.
[0075] In each stage, observe the CPU and memory usage of the cluster and perform real-time load adjustment. For example, if the CPU utilization rate does not exceed 60%, increase the concurrency. You can use a Python script to adjust the parameters of each MapReduce task and automatically increase it to 150 or 200 concurrent tasks. Gradually increase the number of concurrent clients until the resources of the cluster are exhausted and the NameNode starts to respond slowly or fails, thereby measuring the maximum processing capacity of the cluster.
[0076] S33. Obtain the performance data of the NameNode and DataNode.
[0077] Perform real-time performance monitoring on the NameNode and DataNode in a high-concurrency scenario, extract the response time of the NameNode and the I / O performance of the DataNode, and record key metrics such as the failure rate and response time of file creation, deletion, and writing. In this embodiment, the data types of the specific performance data obtained for the NameNode and DataNode are not limited, and corresponding data can be collected according to actual data analysis requirements.
[0078] The above high-concurrency test steps can achieve automatic concurrent load adjustment until the maximum processing capacity of the cluster is measured, and it will not affect the stable operation of the cluster. In this embodiment, only the above high-concurrency test steps are taken as an example, and the specific test steps for other test objectives are not limited in this embodiment. You can refer to the above test steps or other existing test steps, which will not be elaborated here.
[0079] In addition to the above test steps, this embodiment further proposes an adaptive parameter adjustment method, such as Figure 3 shown as a flowchart for realizing intelligent parameter adjustment in addition to performance testing. This method adjusts the test parameters in real time according to the cluster load, simulates the load fluctuations in the production environment, and realizes intelligent parameter adjustment. Specifically, the implementation steps are as follows:
[0080] S105. Continuously monitor the cluster load during performance testing to obtain monitoring data.
[0081] A method for monitoring cluster load is to continuously monitor the utilization rates of the CPU, memory, and disk I / O of the cluster during performance testing. The configuration of this type of monitoring item helps to identify and resolve performance bottlenecks, ensure the efficient utilization of resources, prevent system crashes, and provide accurate data support for system optimization. Of course, other monitoring items can also be used, which are not limited in this embodiment.
[0082] S106. Based on the resource optimization principle, dynamically adjust the test parameters based on the monitoring data.
[0083] By continuously monitoring the system operation status and dynamically adjusting the test load, ensure the efficient utilization of cluster resources, avoid overload and waste, thereby improving the system performance and stability.
[0084] Regarding the specific parameter adjustment strategy, it is not limited in this embodiment. In one embodiment, step S106 can be specifically divided into the following sub-steps:
[0085] S61. Obtain historical monitoring data; analyze the cluster resource utilization rates under different test parameters based on the historical monitoring data to determine the optimal test parameters; adjust the test parameters according to the optimal test parameters.
[0086] Use the monitoring data to analyze the cluster resource utilization rate, and use simple algorithms such as linear regression or rule-based systems to automatically generate the optimal parameters for future tests. For example, when the resource utilization rate is low, the number of files or the MapReduce concurrency can be automatically increased; when the resource utilization rate is high, the file size or concurrency can be reduced.
[0087] S62. Determine the utilization rate of the cluster resources based on the monitoring data; if the utilization rate is lower than the first threshold or higher than the second threshold, adjust the test parameters.
[0088] During the test process, continuously monitor the utilization rate of the cluster resources, such as the utilization rates of the CPU, memory, and disk I / O. Analyze these data in real time, dynamically adjust the current test load, and when the utilization rate is too low or too high, such as lower than the first threshold or higher than the second threshold, adjust the test parameters to avoid overload or waste of the cluster resources. Among them, the specific values of the first threshold and the second threshold are not limited in this embodiment.
[0089] Two methods for adjusting test parameters are provided above, namely intelligent parameter tuning based on historical data and adaptive parameter tuning based on the current state. Based on the analysis of historical monitoring data, the rules and resource utilization in past tests can be deeply explored. By analyzing the utilization rate of cluster resources under different test parameters, the optimal parameter combination can be found, thereby improving the accuracy and reliability of the test. In addition, historical data can usually provide more comprehensive trends over a long time span, reducing the impact of short-term fluctuations on test results. The adaptive parameter tuning method based on the current state can quickly identify and correct resource bottlenecks caused by changes in cluster load, and dynamically adjust according to the resource utilization of the cluster in real time, ensuring timely response to changes in cluster resources during the test and avoiding affecting the stability of test results due to excessive or insufficient resources. The combination of the two can not only ensure the scientific nature and accuracy of parameter tuning, but also enhance the system's response ability in dynamic changes, achieving the purpose of optimizing test efficiency, reducing resource waste, and improving test stability.
[0090] In one embodiment, after determining the performance of the cluster under different configurations in step S104, the cluster configuration suggestions under the optimal performance can be further generated according to the performance of the cluster under different configurations.
[0091] According to the analysis results, optimization suggestions for cluster hardware, parameter configuration, etc. are given. For example, for high concurrency bottlenecks, the network bandwidth, the configuration of the NameNode, or the addition of DataNodes can be optimized.
[0092] The optimal configuration suggestions generated according to the performance of the cluster under different configurations can not only effectively optimize resource utilization, improve overall performance, but also reduce operation and maintenance costs, and enhance the stability, scalability, and customer experience of the system. This data-driven optimization method provides a more scientific basis for decision-making, and also provides support for the efficient operation and automated management of the system.
[0093] Further referring to Figure 4 , which shows an exemplary structural block diagram of an automated performance testing device based on an HDFS cluster according to an embodiment of the present application, mainly including:
[0094] A parameter determination unit 101, configured to determine the test parameters and test objectives of the HDFS cluster;
[0095] A tool matching unit 102, configured to match the built-in test tool of the HDFS cluster corresponding to the test objective as the target tool;
[0096] A test execution unit 103, configured to call the target tool to execute the corresponding performance test according to the test parameters;
[0097] A performance evaluation unit 104 is configured to collect performance data during the execution of tasks in a performance test, determine the performance of the cluster under different configurations, and generate a test report.
[0098] It should be understood that the various units described in the above device correspond to the respective steps in the method described with reference Figure 1 Therefore, the operations and features described above for the method also apply to the device and the units contained therein, and will not be repeated here. The device can be pre-implemented in the browser or other secure applications of the server, or can be loaded into the browser or its secure applications of the server by means of downloading, etc. The corresponding units in the device can cooperate with the units in the server to implement the solutions of the embodiments of the present application.
[0099] Regarding the several units mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0100] It should be noted that for the details not disclosed in the automated performance testing device based on the HDFS cluster in the embodiments of the present application, please refer to the details disclosed in the above embodiments of the present application, and will not be repeated here.
[0101] Next, with reference to Figure 5 , Figure 5 FIG. shows a schematic structural diagram of a computer system of a server suitable for implementing the embodiments of the present application.
[0102] As Figure 5 shown, the computer system includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation instructions of the system are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0103] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 510 as needed so that a computer program read therefrom is installed into the storage section 508 as needed.
[0104] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart Figure 1 can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product which includes a computer program carried on a computer-readable medium, the computer program including program code for performing the method shown in the flowchart. In such an embodiment, the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by a central processing unit (CPU) 501, the above-described functions defined in the system of the present application are executed.
[0105] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operation instructions of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two connected blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for executing the specified functions or operation instructions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0107] The units or modules involved in the embodiments of the present application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. Among them, the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0108] The above description is only a preferred embodiment of the present application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the foregoing disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present application.
Claims
1. An automated performance testing method based on HDFS cluster, characterized in that: include: Determine the test parameters and test objectives of the HDFS cluster; Match the HDFS cluster built-in test tool corresponding to the test target as the target tool; According to the test parameters, calling the target tool to execute corresponding performance test; Collect performance data from the task execution of the performance test, determine the performance of the cluster under different configurations, and generate a test report.
2. The method according to claim 1, characterized in that The HDFS cluster built-in test tools that match the test objectives include: If the test target is a read / write performance test, TestDFSIO is used as the target tool; If the test target is a metadata operation test, NNBench is used as the target tool; If the test target is a high-concurrency test, NNBench and TestDFSIO are used as the target tools at the same time.
3. The method according to claim 2, characterized in that If the test target is a high-concurrency test, then according to the test parameters, the target tool is called to perform the corresponding performance test, including: Generates a corresponding number of MapReduce tasks according to the initial number of concurrent clients, and calls TestDFSIO and NNBench to simulate concurrent clients to perform data reading, writing, and metadata operations; Monitor the resource consumption of NameNode and DataNode, and gradually increase the number of concurrent clients until the cluster's resources are exhausted; Get the performance data of NameNode and DataNode.
4. The method according to any one of claims 1 to 3, characterized in that Also includes: Continuously monitoring the cluster load during the performance test to obtain monitoring data; According to the resource optimization principle, the test parameters are dynamically adjusted based on the monitoring data.
5. The method according to claim 4, characterized in that The dynamically adjusting the test parameters based on the monitoring data according to the resource optimization principle includes: Acquire historical monitoring data; analyze cluster resource utilization under different test parameters according to the historical monitoring data to determine optimal test parameters; adjust the test parameters according to the optimal test parameters; Determine the utilization of cluster resources according to the monitoring data; if the utilization is lower than a first threshold or higher than a second threshold, adjust the test parameter.
6. The method according to claim 4, characterized in that The cluster load is continuously monitored during the performance test, including: Continuously monitor the CPU, memory, and disk I / O utilization of the cluster during performance testing.
7. The method according to claim 1, characterized in that The collecting of performance data during the execution of the performance test tasks includes: Collect execution performance metrics and resource utilization metrics for each task; The execution performance indicator includes at least one of execution time, bandwidth usage, and I / O performance; The resource utilization indicator includes at least one of the consumption of CPU, memory, network and disk I / O.
8. The method according to claim 1, characterized in that Also includes: Based on the performance of the cluster under different configurations, generate cluster configuration recommendations for optimal performance.
9. An automated performance testing device based on HDFS cluster, characterized in that: include: A parameter determination unit, used to determine the test parameters and test targets of the HDFS cluster; A tool matching unit, used to match a built-in test tool of the HDFS cluster corresponding to the test target as a target tool; A test execution unit, used to call the target tool to execute a corresponding performance test according to the test parameters; The performance evaluation unit is used to collect performance data during the execution of the tasks of the performance test, determine the performance of the cluster under different configurations, and generate a test report.
10. A server comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Data processing method and device, nonvolatile storage medium and electronic equipment
CN120780573A
Data processing method and device, nonvolatile storage medium and electronic equipment
CN120780573B