Data processing method and device, nonvolatile storage medium and electronic equipment

By generating a test data set and encapsulating it as a task set for performance testing tools, combining it with monitoring and alarm tools to collect indicators, and using visualization tools for unified display, the problems of tool dispersion and monitoring data fragmentation in big data cluster testing are solved, achieving improved integration and testing efficiency.

CN120780573AActive Publication Date: 2025-10-14CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511247837.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-10-14
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing technologies have not yet effectively addressed the problems of tool fragmentation, high integration costs, and fragmented monitoring data in big data cluster testing.

Method used

Generate a test data set through big data benchmarking tools, encapsulate it into a set of test tasks for performance testing tools, use performance testing tools to execute and collect performance indicators, combine monitoring and alarm tools to receive resource utilization indicators, and use visualization tools to correlate and visualize them for unified display and analysis.

Benefits of technology

It improves the integration of big data testing tools, reduces testing costs, centrally processes and displays monitoring data, and improves testing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780573A_ABST
    Figure CN120780573A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, a nonvolatile storage medium and electronic equipment. The method comprises the following steps: acquiring a target test scene for a big data cluster, and generating a test data set corresponding to the target test scene; packaging an executable program used for executing the test data set into a test task set of the performance test tool, and executing test tasks in the test task set through a plurality of objects in the performance test tool; collecting a first performance index of each test task, receiving the first performance index sent by the performance test tool, and collecting a resource utilization rate index of each node in the big data cluster and a second performance index of a big data component in the node; and performing association visualization on the first performance index and the resource utilization rate index, and performing visualization on the second performance index. The technical problems of scattered tools, high integration cost and monitoring data fragmentation in big data cluster testing are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of big data testing, in particular, to a data processing method and device, a nonvolatile storage medium and an electronic device. BACKGROUND

[0002] The related big data performance testing method faces many challenges when dealing with massive, complex and diverse data, and needs professional technology beyond traditional data management solutions. The following are some examples describing the prior art: a testing solution based on a big data benchmark testing tool (such as HiBench): manually executing HiBench testing, logging in to a cluster node to run a script, collecting results through a log file, time-consuming and unable to dynamically adjust the load; using a performance testing tool (such as HMeter): tools such as JMeter need to write complex scripts to integrate test components, different big data devices need to be repeatedly adapted, and the universality is low; monitoring and alarm tools and visualization tools are independently deployed, and need to manually associate pressure testing data and resource indicators to locate bottlenecks with low efficiency.

[0003] The above method has certain effect on the purpose of big data component testing, but still has technical problems such as scattered tools, high integration cost and fragmented monitoring data.

[0004] At present, no effective solution has been proposed to solve the above problems. SUMMARY

[0005] The present application provides a data processing method and device, a nonvolatile storage medium and an electronic device to at least solve the technical problems of scattered tools, high integration cost and fragmented monitoring data in big data cluster testing.

[0006] According to one aspect of the present application, a data processing method is provided, comprising: obtaining a target test scenario for a big data cluster, generating a test data set corresponding to the target test scenario through a big data benchmark testing tool; encapsulating an executable program for executing the test data set as a test task set of a performance testing tool, and executing test tasks in the test task set through multiple objects in the performance testing tool; collecting first performance indicators of each test task through the performance testing tool, receiving the first performance indicators sent by the performance testing tool through a monitoring and alarm tool, and collecting resource utilization indicators of each node in the big data cluster and second performance indicators of big data components in the node; and associating and visualizing the first performance indicators and the resource utilization indicators through a dashboard in a visualization tool, and visualizing the second performance indicators.

[0007] Optionally, the method further comprises: constructing a target test input triggering a target performance symptom according to a preset performance symptom template and a skew heuristic rule, wherein the performance symptom template is used to determine an abnormal state meeting preset duration requirements and fluctuation intensity requirements as the target performance symptom, and the skew heuristic rule is used to guide generation of a test case with preset distribution characteristics; performing a test on the target test input, and determining an intermediate input state causing the target performance symptom if the target performance symptom is detected during the test; analyzing the intermediate input state by using a large language model to obtain a target pseudo-inverse function, wherein the target pseudo-inverse function is used to map the intermediate input state into an initial input format; re-performing the test on the initial input format, and verifying triggering efficiency of the target performance symptom and resource utilization rate indicators of each node in the big data cluster according to a test result.

[0008] Optionally, the method further comprises: generating a diagnosis report according to a space-time coupling relationship between the first performance indicator and the resource utilization rate indicator, wherein the diagnosis report includes at least one of the following: a type of performance bottleneck, an impact range indicator, and a numerical representation of the performance bottleneck; matching the diagnosis report with an optimization knowledge base to obtain an optimization measure for the diagnosis report, wherein the optimization knowledge base is a decision engine storing historical optimization cases and preset rules, used to map the diagnosis report to the optimization measure, and the optimization measure includes at least one of the following types: software parameter tuning, data governance strategy, and hardware upgrade suggestion; in a case where the optimization measure is software parameter tuning, calculating a configuration parameter value according to a cluster size and load characteristics of the big data cluster.

[0009] Optionally, the method further comprises: triggering a preset performance test task in response to a code submission or deployment event; comparing the first performance indicator, the resource utilization rate indicator, and the second performance indicator with a predefined baseline and a service level agreement threshold to obtain a comparison result after triggering the preset performance test task; and blocking a promotion process of a target pipeline based on continuous integration and continuous delivery in a case where the comparison result does not meet preset conditions, wherein the target pipeline is used to perform the preset performance test task.

[0010] Optionally, the big data benchmark test tool is deployed on the big data cluster or a first server independent of the big data cluster; the working node of the performance test tool is deployed on a working node in the big data cluster, and the control node of the performance test tool is deployed on an independent control node in the big data cluster; the working node of the performance test tool dynamically synchronizes load instructions and state snapshots with the control node of the performance test tool through a heartbeat mechanism; and the working node of the performance test tool interacts with the big data benchmark test tool through command line invocation.

[0011] Optionally, the monitoring and alarm tool is deployed on the second server, and the visualization tool is deployed on the third server, wherein the first server, the second server and the third server are different servers communicating with the big data cluster; the visualization tool connects to the second server through an application program interface, and obtains the first performance indicator sent by the performance testing tool, the resource utilization indicator of each node in the big data cluster, and the second performance indicator of the big data component in the node.

[0012] Optionally, after visualizing the association between the first performance indicator and the resource utilization indicator through the dashboard in the visualization tool, and visualizing the second performance indicator, the method also includes: aligning the first performance indicator of the target test task with the resource utilization indicator of the first node corresponding to the target test task in time series to obtain a mapping relationship between the task load and resource consumption of the target test task; based on the mapping relationship, determining the target node in the first node whose resource utilization indicator exceeds a preset threshold, and determining the target big data component in which the second performance indicator in the target node is abnormal; determining the component-level fault point that causes the abnormality of the first performance indicator of the target test task based on the performance indicator abnormal value of the target big data component and the operation log of the target big data component.

[0013] According to another aspect of the present application, a data processing device is also provided, including: an acquisition module, used to obtain a target test scenario for a big data cluster, and generate a test data set corresponding to the target test scenario through a big data benchmark testing tool; an execution module, used to encapsulate the executable program used to execute the test data set into a test task set of a performance testing tool, and execute the test tasks in the test task set respectively through multiple objects in the performance testing tool; an acquisition module, used to collect the first performance indicator of each test task through the performance testing tool, receive the first performance indicator sent by the performance testing tool through a monitoring and alarm tool, and collect the resource utilization indicator of each node in the big data cluster and the second performance indicator of the big data component in the node; a visualization module, used to associate and visualize the first performance indicator and the resource utilization indicator through a dashboard in the visualization tool, and to visualize the second performance indicator.

[0014] According to another aspect of the present application, a non-volatile storage medium is provided, which includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the above data processing method.

[0015] According to another aspect of the present application, an electronic device is provided, including: a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the above data processing method is executed when the program is run.

[0016] According to still another aspect of the present application, a computer program is also provided, wherein the computer program, when executed by a processor, implements the above data processing method.

[0017] According to still another aspect of the present application, a computer program product is also provided, which comprises a non-volatile computer readable storage medium, wherein the non-volatile computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the above data processing method.

[0018] In the present application, a target test scenario for a big data cluster is acquired, and a test data set corresponding to the target test scenario is generated by a big data benchmark test tool; an executable program for executing the test data set is encapsulated as a test task set of a performance test tool, and a test task in the test task set is executed by a plurality of objects in the performance test tool respectively; a first performance index of each test task is collected by the performance test tool, and a resource utilization index of each node in the big data cluster and a second performance index of a big data component in the node are collected by a monitoring and alarming tool receiving the first performance index sent by the performance test tool; the first performance index and the resource utilization index are associated and visualized by a dashboard in a visualization tool, and the second performance index is visualized, so as to execute the test task by using the distributed capability of the performance test tool and collect the application layer performance index in real time, monitor the resource utilization of the node and the internal index of the big data component by the monitoring and alarming tool, and display all the collected performance data by the unified visualization capability of the visualization tool, so as to improve the integration of the big data test tool, reduce the test cost, and centrally process and display the monitoring data, thereby realizing the technical effect of improving the test efficiency, and further solving the technical problems of scattered tools, high integration cost and fragmented monitoring data in the big data cluster test. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings, which are included to provide a further understanding of the present application, form a part of the present application and illustrate embodiments of the present application and together with the description serve to explain the present application. In the drawings:

[0020] Figure 1 is a flowchart of a data processing method according to an embodiment of the present application;

[0021] Figure 2 is a structural diagram of a data processing platform according to an embodiment of the present application;

[0022] Figure 3 is a structural diagram of a data processing device according to an embodiment of the present application;

[0023] Figure 4It is a hardware structure block diagram of a computer terminal according to the data processing method of the embodiment of the application. DETAILED DESCRIPTION

[0024] In order to enable personnel in the technical field to better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0025] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] According to the embodiments of the present application, a method embodiment of a data processing method is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0027] Figure 1 is a flowchart of a data processing method according to the embodiments of the present application, as shown in Figure 1 the method comprises the following steps:

[0028] Step S102, obtaining a target test scene for a big data cluster, and generating a test data set corresponding to the target test scene by a big data benchmark test tool.

[0029] The big data cluster refers to a computing architecture in which multiple computers (nodes) are connected to each other through a network and work together to process and analyze massive amounts of data. The goal of the big data cluster is to improve the efficiency, scalability and fault tolerance of data processing through distributed computing and storage technology, so that it can handle large amounts of data and complex analysis requirements that cannot be handled by traditional single machine systems.

[0030] In a big data cluster, data and computing tasks are distributed across multiple nodes, each responsible for processing a portion of the dataset, enabling parallel processing and storage. A big data cluster contains the following types of nodes: Master nodes: responsible for task scheduling, resource management, and monitoring cluster status. For example, in a Hadoop cluster, the NameNode and the ResourceManager play the role of Master nodes. Worker nodes: perform specific data processing and storage tasks. For example, DataNodes and TaskTrackers (or later NodeManagers) in a Hadoop cluster are Worker nodes. Clients: used to submit data processing tasks and query results, generally not directly involved in data storage and computation processes.

[0031] In step S102, the target test scenario for the big data cluster is first received, such as evaluating the performance of the Hadoop distributed file system, the data processing speed of Spark, or the query response time of Hive. After determining the target test scenario, the big data benchmarking tool is used to automatically generate a test dataset that matches the target test scenario.

[0032] Among them, the big data benchmarking tool is a software tool used to evaluate the performance of big data processing platforms and frameworks. By executing predefined workloads or test scenarios, it simulates real-world big data processing tasks such as large-scale data queries, data cleaning, data conversion, machine learning model training, etc., to measure key indicators such as processing speed, throughput, latency, resource usage efficiency, and scalability. Big data benchmarking tools such as HiBench provide a variety of pre-set test workloads such as WordCount, TeraSort, or DFSIO, which can be flexibly selected or customized to match specific big data components and test requirements.

[0033] Step S102 simplifies the preparation process of test data, avoiding the complexity of manually creating and managing datasets, ensuring the authenticity and representativeness of test data. Through step S102, the complexity and time cost of manually creating and managing test datasets can be significantly reduced, ensuring the authenticity and scale of test data, covering 90% of core component test scenarios.

[0034] Step S104 encapsulates the executable program for executing the test dataset as a test task set of the performance testing tool, and executes the test tasks in the test task set through multiple objects in the performance testing tool.

[0035] In step S104, executable programs of specific test data sets, such as those generated by big data benchmarking tools, are encapsulated and integrated into a test task set of a performance testing tool. The performance testing tool is a tool specially designed to evaluate the response time, processing speed, resource use efficiency, and stability of a computer system, software application, or big data cluster under a specific workload. The performance testing tool checks whether the big data cluster can maintain a good performance level under the expected load and meet the quality of service requirements by simulating user behavior or data traffic and applying pressure to the big data cluster. The performance testing tool may be, for example, Locust.

[0036] The above encapsulation process can be completed by writing a Python script containing task definition syntax specific to Locust to call various test scripts of HiBench. Specifically, the test scripts provided by HiBench (such as WordCount, TeraSort, etc.) are converted into tasks that can be understood and executed by Locust, so that each task directly corresponds to a specific big data processing scenario.

[0037] After the tasks are aggregated, Locust uses its unique distributed stress testing capability to distribute the tasks to multiple Locust objects (i.e., Locust Workers). These Workers can run in parallel on multiple machines, with each Worker executing a portion of the encapsulated test task set to simulate thousands or even tens of thousands of concurrent users or workloads, thereby generating actual pressure on the big data cluster. The above distributed execution mode not only greatly improves the concurrency and scalability of the test, ensuring the accuracy and reliability of the test results, but also effectively covers the behavior characteristics of big data applications under high load conditions.

[0038] The functions of step S104 can be implemented by the following code.

[0039] class HiBenchUser(HttpUser):

[0040] wait_time = between(1, 5) # Simulate user thinking time

[0041] host = "http: / / localhost:8089" # Locust Web UI host

[0042] @task(1)

[0043] def run_hibench_wordcount(self):

[0044] workload_name = "wordcount"

[0045] data_size = "small" # can be adjusted according to test requirements

[0046] # Simulate HiBench prepare phase

[0047] prepare_command = f". / HiBench / bin / workloads / micro / {workload_name} / prepare / prepare.sh --size {data_size}"

[0048] self.environment.runner.send_message("prepare_start", {"workload": workload_name, "size": data_size})

[0049] print(f"Running prepare: {prepare_command}")

[0050] try:

[0051] # Actual execution of the HiBench prepare script

[0052] subprocess.run(prepare_command, shell=True, capture_output=True, text=True, check=True) [54, 55]

[0053] time.sleep(2) # simulation preparation time

[0054] self.environment.runner.send_message("prepare_success", {"workload": workload_name})

[0055] except subprocess.CalledProcessError as e:

[0056] self.environment.runner.send_message("prepare_failure", {"workload": workload_name, "error": e.stderr})

[0057] print(f"HiBench prepare failed: {e.stderr}")

[0058] return # If preparation fails, subsequent operations will not be executed

[0059] # Simulate HiBench run phase

[0060] run_command = f". / HiBench / bin / workloads / micro / {workload_name} / hadoop / run.sh"

[0061] self.environment.runner.send_message("run_start", {"workload":workload_name})

[0062] print(f"Running workload: {run_command}")

[0063] start_time = time.time()

[0064] try:

[0065] # Actual execution of the HiBench run script and capture output

[0066] result = subprocess.run(run_command, shell=True, capture_output=True, text=True, check=True) [54, 55]

[0067] output = result.stdout

[0068] # In practice, we will parse HiBench's hibench.report file to get accurate throughput and latency [43, 79]

[0069] # This is simplified to simulated value

[0070] simulated_latency = (time.time() - start_time) * 1000 # milliseconds

[0071] simulated_rps = 1000 / simulated_latency * 1000 if simulated_latency > 0 else 0 # simulated RPS

[0072] print(f"HiBench {workload_name} completed in {simulated_latency:.2f}ms. Throughput: {simulated_rps:.2f} RPS")

[0073] By step S104, the originally independently running test tasks become part of the performance testing tool distributed test framework. Each object of the performance testing tool corresponds to a virtual user or data processing unit, which executes the corresponding big data test task. Step S104 realizes flexible and scalable concurrent simulation of big data workloads, overcomes the bottleneck of related tools in handling large-scale concurrency, and provides real-time application layer performance feedback.

[0074] Step S106, the performance testing tool collects the first performance indicators of each test task, and the monitoring and alarm tool receives the first performance indicators sent by the performance testing tool, and collects the resource utilization indicators of each node in the big data cluster and the second performance indicators of the big data components in the node.

[0075] In step S106, the performance testing tool is not only responsible for scheduling and executing the test tasks generated by the big data benchmark testing tool, but also collects the first performance indicators generated by the test tasks in the execution process in real time, wherein the first performance indicators include but are not limited to the number of requests per second (RPS), response time (i.e. latency) and other metrics directly related to the application layer performance of the big data component. For example, when executing a WordCount or TeraSort test task, Locust tracks and records the throughput and response time during task execution, which are important indicators for evaluating the processing capacity and efficiency of the big data component.

[0076] Further, a monitoring and alerting tool, such as Prometheus, is utilized to receive the first performance indicators sent from Locust. The monitoring and alerting tool receives the first performance indicators sent from the performance testing tool, specifically, Locust, which formats and exposes the application-level performance data (i.e., the first performance indicators) as metrics recognizable by Prometheus through its built-in Locust Exporter module. Prometheus then achieves comprehensive data monitoring by periodically scraping these exposed metrics and performance data of other big data components. For example, the metrics received by Prometheus from the Locust Exporter include "locust_requests_per_second_total" and "locust_request_failure_total", which represent the total number of requests per second and the total number of request failures, respectively.

[0077] Meanwhile, the monitoring and alerting tool actively collects resource utilization indicators of each node in the big data cluster, including but not limited to CPU, memory, disk I / O, and network I / O, etc. In addition, Prometheus also obtains internal running data of big data components such as HDFS, Spark, and Hive through specific exporters (such as JMX Exporter), i.e., the second performance indicators, such as the RPC call number of NameNode, the read / write rate of DataNode, the execution time of Spark tasks, etc.

[0078] The above multi-level performance indicator collection mechanism combines data at the application layer, system layer, and component layer, providing a more comprehensive performance analysis perspective. Through the cooperation of the monitoring and alerting tool and the performance testing tool, not only the external performance of big data components under high concurrency load is captured, but also the internal resource allocation and utilization of the big data cluster, as well as the specific running state of the components are understood. These data are crucial for performance bottleneck analysis, optimization suggestion generation, and real-time performance monitoring and early warning, ensuring the stability and high-efficiency operation of the big data platform.

[0079] Part of the functions of step S106 can be implemented by the following code.

[0080] # prometheus.yml configuration fragment

[0081] scrape_configs:

[0082] - job_name: 'hadoop_cluster'

[0083] # Static configuration of big data cluster nodes

[0084] static_configs:

[0085] - targets: ['namenode-host:9870', 'datanode1-host:9864', 'datanode2-host:9864'] # HDFS JMX Exporter port

[0086] relabel_configs:

[0087] - source_labels: [__address__]

[0088] regex: '([^:]+)(:\d+)?'

[0089] target_label: instance

[0090] replacement: '$1'

[0091] - source_labels: [__address__]

[0092] regex: '.*'

[0093] target_label: __address__

[0094] replacement: '${1}:9090' # Assuming JMX Exporter is running on port 9090

[0095] - job_name: 'os_metrics'

[0096] # Scraping OS metrics (Node Exporter)

[0097] static_configs:

[0098] - targets: ['namenode-host:9900', 'datanode1-host:9900', 'datanode2-host:9900'] # Node Exporter port

[0099] relabel_configs:

[0100] - source_labels: [__address__]

[0101] regex: '([^:]+)(:\d+)?'

[0102] target_label: instance

[0103] replacement: '$1'

[0104] - job_name: 'locust_metrics'

[0105] # Grab metrics exposed by Locust Exporter

[0106] static_configs:

[0107] - targets: ['locust-master-host:8089'] # default port for Locust Exporter

[0108] relabel_configs:

[0109] - source_labels: [__address__]

[0110] regex: '([^:]+)(:\d+)?'

[0111] target_label: instance

[0112] replacement: '$1'

[0113] Step S108, the first performance indicator and the resource utilization indicator are associated and visualized through the dashboard in the visualization tool, and the second performance indicator is visualized.

[0114] The visualization tool is, for example, Grafana. The visualization tool deeply integrates and visualizes the first performance indicator, the resource utilization indicator, and the second performance indicator collected by the performance testing tool and the monitoring and alarm tool. The visualization tool sets the monitoring and alarm tool as its data source to build a customized dashboard. The dashboard is used to visually display the performance of the big data component under different loads and the resource usage in the form of charts, time series, alarms, etc.

[0115] Grafana provides users with correlated views of system-level and application-level performance by associating these resource utilization metrics with the first performance metrics. For example, users can create a hybrid chart that displays both the throughput of HDFS (first performance metric collected by Locust) and the disk I / O usage of DataNode (resource utilization metric collected by Prometheus), allowing for quick identification of storage-level performance issues through comparative analysis.

[0116] This correlation visualization feature of Grafana enables users to analyze the performance of big data components from multiple dimensions (application layer, system layer, component layer), thereby more accurately identifying and understanding the root cause of performance bottlenecks. For example, if a significant drop in RPS (provided by Locust) is observed on the Grafana dashboard, along with an abnormal increase in CPU usage (pulled by Prometheus), users can infer that the current performance problem may be caused by a CPU bottleneck. Similarly, if the read-write rate of HDFS decreases while the disk I / O load of DataNode increases, it indicates a storage bottleneck.

[0117] Through the dashboard of Grafana, not only can these performance metrics be monitored in real time, but dynamic thresholds and alert rules can also be set. When key metrics exceed the expected range, an automatic alert is triggered to remind operations personnel or developers to handle the problem in a timely manner. This real-time, unified performance visualization greatly simplifies the performance monitoring and troubleshooting process, improves the operational efficiency and response speed of the big data platform, and ensures the stability and high performance of the platform.

[0118] Specifically, the dashboard configuration of Grafana includes the following types of visualization panels: Time series chart: shows the trend of RPS over time, as well as the usage of CPU, memory, and disk I / O resources. Statistical panel: displays key performance indicators such as average response time, error rate, and task success rate. Heatmap: visualizes the load distribution between nodes, helping to identify hotspots and load imbalance. Alert panel: displays triggered performance alerts, including RPS below the preset threshold, CPU usage too high, disk I / O busy, etc., facilitating quick response. For example, through the following PromQL query example, Grafana can display the busy time of disk I / O of HDFS DataNode and the average response time of HDFS WordCount test tasks.

[0119] The partial function of step S108 can be implemented by the following code.

[0120] # PromQL query example - Grafana panel configuration

[0121] # Query average requests per second (RPS)

[0122] rate(locust_requests_per_second_total{name="HiBench_wordcount_small"}[5m])

[0123] # Query average response time (Latency)

[0124] locust_response_time_average{name="HiBench_wordcount_small"}

[0125] # Query cluster average CPU usage

[0126] avg by (instance) (node_cpu_seconds_total{mode="idle"}[5m]) / 100 # Example, actual may be more complex

[0127] # Query HDFS DataNode disk I / O busy time

[0128] rate(node_disk_io_time_seconds_total{job="os_metrics", instance=~"datanode.*"}[5m])

[0129] The steps S102 to S108 above perform test tasks and collect application layer performance indicators in real time by utilizing the distributed capabilities of the performance testing tool, monitor the resource utilization of the nodes and the internal indicators of the big data components in combination with the monitoring and alarm tools, and display all the collected performance data in association through the unified visualization capabilities of the visualization tools, achieving the purposes of improving the integration of the big data testing tool, reducing the testing cost, and centrally processing and displaying the monitoring data, thereby realizing the technical effect of improving the testing efficiency.

[0130] The steps shown in FIG. 1 will be exemplarily described and explained below. Figure 1

[0131] ​According to some optional embodiments of the present application, the data processing method further comprises the following steps: constructing a target test input triggering a target performance symptom according to a preset performance symptom template and a skew heuristic rule, wherein the performance symptom template is used to determine an abnormal state meeting preset duration requirements and fluctuation intensity requirements as a target performance symptom, and the skew heuristic rule is used to guide the generation of a test case with preset distribution characteristics; performing a test on the target test input, and during the test, if the target performance symptom is detected, determining an intermediate input state causing the target performance symptom; analyzing the intermediate input state by using a large language model to obtain a target pseudo-inverse function, wherein the target pseudo-inverse function is used to map the intermediate input state into an initial input format; re-executing the test on the initial input format, and verifying the triggering efficiency of the target performance symptom and resource utilization rate indicators of each node in the big data cluster according to the test result.

[0132] In the present embodiment, the performance symptom template is a series of pre-defined abnormal state sets of indicators, which are used to identify and define various performance problems that may occur during testing, such as high latency, low throughput, CPU or disk I / O overload, etc. These abnormal states need to meet certain duration and fluctuation intensity requirements to ensure that the identified performance problems are indeed stable and have a wide impact, rather than random errors. The skew heuristic rule is based on the common data skew phenomenon in big data analysis, such as uneven distribution of data volume among different partitions or nodes, to guide the generation of test cases with specific distribution characteristics, with the purpose of amplifying and exposing performance bottlenecks caused by data or computing skew.

[0133] According to the preset performance symptom template and the skew heuristic rule, a target test input triggering a specific performance symptom is constructed. Specifically, existing test data sets can be used and mutated or new data sets can be generated to ensure that the skew and other phenomena that may cause performance problems are fully reflected. For example, if the performance symptom template defines high latency caused by computing skew, the skew heuristic rule guides the generation of a data set containing a large number of repetitive or similar computing requirements. Such a data set is likely to cause the load of some nodes or partitions to be much higher than that of other parts when processing, thereby exposing the computing skew problem.

[0134] After generating the target test input, a performance test is performed on the target test input. The test can be run on a big data cluster, and a large number of user concurrent operations or data processing requests can be simulated by a performance test tool while monitoring the performance indicators and resource utilization of the cluster. Assuming that the target performance symptom, such as a sudden surge in latency, is detected during the test, the intermediate input state causing this symptom is recorded, i.e., the data and operation state being processed by the cluster at a certain specific time point during test execution.

[0135] The large language model is used to analyze the intermediate input state, with the goal of generating a target pseudo-inverse function. The pseudo-inverse function is a theoretical concept that refers to an algorithm or logic that can convert the intermediate input state back to the original test input format that caused the state. In this way, the path from the original input to the performance symptom trigger point can be traced, and it can be determined which specific data properties or operation patterns caused the problem. The large language model is used to learn and predict which input patterns are most likely to cause a specific performance symptom based on a large amount of previously accumulated test data and performance symptom cases, thereby constructing the corresponding pseudo-inverse function.

[0136] The generated target pseudo-inverse function is used to convert the intermediate input state back to the initial input format, and the performance test is re-executed on the converted test case to verify the accuracy and effectiveness of the pseudo-inverse function. This cycle of testing can confirm whether the target test input can indeed trigger the predefined performance symptoms efficiently and reproducibly. At the same time, by monitoring the resource utilization indicators of each node in the big data cluster again, it can further verify whether the triggering of performance symptoms is accompanied by excessive consumption of specific resources or bottleneck phenomena.

[0137] For example, suppose that in the test, it is detected that the Hadoop MapReduce job causes a significant increase in latency due to data skew, and the intermediate state information of the job is recorded, including data distribution, the number of Map and Reduce tasks, etc. With the analysis capability of the large language model, a set of pseudo-inverse functions is derived, which reveals which data characteristics (such as key-value distribution) cause this skew. Then use this set of pseudo-inverse functions to generate a new test input that deliberately exaggerates these characteristics, and run the test again to verify whether the high latency symptom can be reproduced, and further analyze whether there are significant differences in the use of CPU and disk I / O resources of each node in the cluster in the process of reproducing the symptom, thereby helping to find the direction of resource optimization.

[0138] Through the above steps, the coverage of the test scenario and the efficiency of problem discovery are greatly improved, and test loads that are difficult to simulate by related methods and can expose deep performance problems can be generated, thereby improving the authenticity and effectiveness of the test.

[0139] Preferably, the large language model is trained by the following steps: obtaining a historical performance test dataset, which includes test inputs, corresponding performance indicators, resource usage data, and observed performance symptoms; cleaning, standardizing, and feature extracting the collected historical data; selecting a large language model framework suitable for processing sequential and structured data, or a pre-trained model suitable for specific domain tasks; training the model using historical performance test data through supervised learning or reinforcement learning mechanisms, so that it learns to identify the association patterns between specific intermediate input states and target performance symptoms; evaluating the trained model using cross-validation or indicator comparison methods to verify the accuracy and generalization ability of the model, and adjusting the model parameters or architecture based on the evaluation results to improve the precision of performance symptom prediction.

[0140] According to some optional embodiments of the present application, the data processing method further includes the following steps: generating a diagnosis report based on the spatio-temporal coupling relationship between the first performance indicator and the resource utilization indicator, wherein the diagnosis report includes at least one of the following: the type of performance bottleneck, the impact range indicator, and the numerical representation of the performance bottleneck; matching the diagnosis report with an optimization knowledge base to obtain optimization measures for the diagnosis report, wherein the optimization knowledge base is a decision engine that stores historical optimization cases and preset rules, and is used to map the diagnosis report to optimization measures, and the optimization measures include at least one of the following types: software parameter tuning, data governance strategy, and hardware upgrade suggestion; in the case of software parameter tuning, calculating the configuration parameter value based on the cluster size and load characteristics of the big data cluster.

[0141] In this embodiment, the spatio-temporal coupling relationship between the first performance indicator (such as throughput and latency collected by Locust) and the resource utilization indicator (such as CPU usage, memory occupancy, disk I / O, and network traffic pulled by Prometheus) is analyzed to generate a detailed diagnosis report. The spatio-temporal coupling refers to aligning performance indicators and resource usage data in time series, while considering their spatial distribution across different nodes in the cluster, so as to capture the trend of performance problems over time and the impact range and distribution characteristics among nodes.

[0142] The content of the diagnosis report includes key information such as the type of performance bottleneck, the impact range indicator, and the numerical representation of the performance bottleneck. For example, the report indicates that the latency of HDFS read-write operations has significantly increased in a certain time window, and is associated with the surge in disk I / O utilization of DataNode, thereby identifying a storage I / O bottleneck. The impact range indicator details which nodes or components are affected and the severity of the impact. The numerical representation of the performance bottleneck quantifies the degree of performance decline.

[0143] Next, the diagnostic report is input into the optimization knowledge base for matching to obtain targeted optimization measure suggestions. The optimization knowledge base is a decision engine integrating historical optimization cases and preset optimization rules, which can recommend optimization strategies according to the specific performance bottleneck type and characteristics in the diagnostic report, including software parameter tuning, data governance strategies, and hardware upgrade suggestions, etc. For example, facing the high latency storage I / O bottleneck, the knowledge base suggests adjusting the HDFS replication factor, data block size, or enabling data compression, etc. software parameter optimization scheme, but also can propose data deduplication, cache strategy adjustment, or use faster storage media, etc. data governance and hardware upgrade strategies. The above optimization strategies are, for example, the following. HDFS parameter adjustment: suggest adjusting HDFS block size (e.g., less than 4TB data using 256MB, greater than 4TB using 512MB or more), replication factor, namenode.handler.count and datanode.handler.count, etc. Parameters to optimize read-write efficiency and parallel RPC calls. Data compression: suggest enabling compression for data (especially intermediate MapReduce output) to reduce data transmission and storage. Cache mechanism: suggest using HDFS cache to improve the performance of frequently accessed data and reduce the load of DataNode. Hardware optimization: suggest using NVMe SSDs as the main storage for active data sets and configuring RAID to improve performance and redundancy. Data deduplication: recommend or automatically identify and eliminate duplicate data in the storage system to reduce storage requirements and improve transmission efficiency.

[0144] In particular, when the optimization measure is determined to be software parameter tuning, the optimal configuration parameter value is calculated according to the scale and load characteristics of the big data cluster. Specifically, a large amount of historical data and current cluster state analysis is required, considering factors such as the number of nodes, hardware configuration, data type and size, business load type, etc. of the cluster, based on mathematical models or machine learning algorithms, to calculate the best parameter combination that balances system performance and resource use efficiency. For example, for HDFS, parameter tuning involves adjusting the default block size from 128MB to 256MB in smaller data sets (e.g. <4TB), and increasing to 512MB or higher when dealing with larger data sets, while also considering data access patterns and redundancy strategies to ensure optimal balance between read-write performance, data redundancy, and storage efficiency.

[0145] Specifically, the configuration parameter values can be calculated according to the cluster size and load characteristics of the big data cluster by the following method: monitoring and recording the hardware resource conditions of each node in the big data cluster, the hardware resource conditions including but not limited to CPU, memory, disk type and network bandwidth, etc., and summarizing the total number of nodes and the number of processor cores of the cluster to establish the basic image of the cluster size. Analyze the type of task currently being executed by the cluster and the data processing mode, identify the load characteristics of CPU-intensive, I / O-intensive, data-intensive or communication-intensive, as the guiding basis for parameter tuning. According to the load characteristics, refer to the historical optimization case database, select the optimization strategy that best matches the current cluster load type, and preliminarily recommend the software parameter values, which include but are not limited to storage block size, executor memory, task scheduling strategy, data compression algorithm, etc. Further optimize the recommended software parameter values. Use historical performance data to train the model, with cluster size, node configuration and load characteristics as input, and optimized software parameter values as output. The model learns the performance of different configurations to predict the most suitable parameter combination for the current cluster condition; during the model training and parameter optimization process, the non-linear relationship between software parameters and cluster size is particularly considered. For example, for smaller clusters, the model tends to recommend smaller storage block sizes to reduce addressing delays, and in large-scale cluster environments, it tends to recommend larger block sizes to optimize read-write efficiency and reduce metadata management overhead; for CPU-intensive tasks, the model focuses on evaluating and recommending parameter configurations that maximize CPU utilization without excessive memory resource consumption, such as the balance point of appropriately increasing the number of concurrent executors and reasonably allocating the proportion of executor memory; for I / O-intensive tasks, the efficiency of disk I / O operations and network transmission is optimized, the model considers the access frequency and size of data, and recommends the use of appropriate caching strategies, data layout strategies and optimized storage medium selection; the calculation results of the output software parameter values include optimization suggestions for key big data components such as storage services, data processing engines and query processors, ensuring that the configuration can improve processing speed while maintaining system stability and high availability; through re-execution of performance tests, the specific improvement effect of the optimized software parameter configuration on cluster performance is verified, and the optimization strategy is continuously iterated until the expected performance indicators are reached; the calculated software parameter values are directly applied to the configuration files of the big data cluster to realize one-key operation of parameter adjustment, simplify the operation and maintenance workflow, and improve the response speed and overall efficiency of the cluster.

[0146] In some optional embodiments of the present application, the data processing method further comprises the following steps: triggering a preset performance test task in response to a code submission or deployment event; comparing the obtained first performance indicator, resource utilization indicator and second performance indicator with a predefined baseline and service level agreement threshold to obtain a comparison result after triggering the preset performance test task; and blocking the advancing process of the target pipeline based on continuous integration and continuous delivery in the case that the comparison result does not meet a preset condition, wherein the target pipeline is used to execute the preset performance test task.

[0147] Specifically, when a new code submission or deployment target action occurs, an event listener on the CI / CD server automatically identifies the target action. Based on the pre-configured test strategy, the performance test platform is automatically called, which is used to execute the method shown in steps S102 to S108. In the case that the performance test is completed, all collected performance indicators are automatically compared with the predefined baseline and SLA threshold. If all key indicators reach or exceed the set baseline and threshold, the pipeline will continue to advance, and the code change will be marked as performance qualified and can be further integrated and deployed; otherwise, if any indicator shows that the performance is lower than the baseline or violates the SLA provision, the subsequent operation of the CI / CD pipeline is blocked. This means that until the performance problem is solved, otherwise any code will not be deployed to the production environment. Further, it ensures that only the code that has passed strict performance verification is allowed to enter the production environment, avoiding business interruption and service degradation caused by performance defects.

[0148] In the present embodiment, the method shown in steps S102 to S108 can be integrated seamlessly into the existing continuous integration / continuous deployment (CI / CD) pipeline (target pipeline) as part of the automated process. That is, the target pipeline is used to execute the method shown in steps S102 to S108. After each code submission or deployment, performance testing is automatically triggered, and result evaluation is performed according to the preset performance baseline and service level agreement, and if it does not meet the standard, it is automatically blocked or alarmed. The integration into the existing continuous integration / continuous deployment (CI / CD) pipeline can be implemented through the following code.

[0149] #.gitlab-ci.yml configuration fragment

[0150] stages:

[0151] - build

[0152] - test

[0153] - deploy

[0154] - performance_test

[0155] build_job:

[0156] stage: build

[0157] script:

[0158] - echo "Building application..."

[0159] - #... build commands

[0160] unit_test_job:

[0161] stage: test

[0162] script:

[0163] - echo "Running unit tests..."

[0164] - #... unit test commands

[0165] deploy_to_staging:

[0166] stage: deploy

[0167] script:

[0168] - echo "Deploying to staging environment..."

[0169] - #... deployment commands

[0170] environment:

[0171] name: staging

[0172] url: http: / / staging.example.com

[0173] performance_test_job:

[0174] stage: performance_test

[0175] image: python:3.9-slim # Contains Python and Locust

[0176] script:

[0177] - pip install locust # Install Locust

[0178] - # Configure HiBench and Locust scripts

[0179] - cp locustfile.py.

[0180] - cp -r HiBench. # Assuming HiBench directory is present or cached via CI / CD

[0181] - # Start Locust Master and Worker

[0182] - locust --master --web-host 0.0.0.0 & # Start Master

[0183] - locust --worker --master-host localhost & # Start Worker

[0184] - sleep 10 # Wait for Locust to start

[0185] - # Trigger Locust test (e.g. via API or directly run Locust client)

[0186] - # Simplified here to directly run Locust test, in reality could be triggered via API

[0187] - locust -f locustfile.py --host http: / / your-bigdata-cluster.com --users 100 --spawn-rate 10 --run-time 5m --headless --csv=test_results

[0188] - echo "Performance test completed. Analyzing results..."

[0189] - # Can add script here to parse test_results.csv and interact with Prometheus / Grafana

[0190] - # For example, check if average response time exceeds threshold

[0191] - python -c "import pandas as pd; df = pd.read_csv('test_results_stats.csv'); assert df.iloc < 500, 'Latency too high!'"

[0192] artifacts:

[0193] paths:

[0194] - test_results_stats.csv

[0195] - test_results_failures.csv

[0196] expire_in: 1 week

[0197] allow_failure: false # If a performance test fails, the pipeline fails.

[0198] Through the above steps, performance testing is no longer an isolated, late-stage process; instead, it becomes an integral part of the software development process, beginning at the earliest possible code submission stage. This allows performance issues to be discovered and resolved early, avoiding the time-consuming backtracking and repair required to identify performance bottlenecks later in the development process. By automatically triggering performance testing after each code submission or deployment, development teams can obtain test results immediately without having to wait for dedicated testing cycles or manually deploy a test environment. This significantly shortens the cycle from code changes to performance feedback, facilitating rapid iteration and optimization.

[0199] As some optional embodiments of the present application, a big data benchmark tool is deployed on a big data cluster or a first server independent of the big data cluster; a working node of a performance testing tool is deployed on a working node in the big data cluster, and a control node of the performance testing tool is deployed on an independent control node in the big data cluster; the working node of the performance testing tool dynamically synchronizes load instructions and state snapshots with the control node of the performance testing tool through a heartbeat mechanism; the working node of the performance testing tool interacts with the big data benchmark tool through command line calls. A monitoring and alarm tool is deployed on a second server, and a visualization tool is deployed on a third server, wherein the first server, the second server, and the third server are different servers communicating with the big data cluster; the visualization tool connects to the second server through an application program interface and obtains a first performance indicator sent by the performance testing tool, a resource utilization indicator of each node in the big data cluster, and a second performance indicator of the big data component in the node.

[0200] In this embodiment, a big data benchmarking tool, such as HiBench, is deployed either inside the big data cluster or on a standalone first server. The main function of the big data benchmarking tool is to generate standardized test loads for simulating typical work scenarios in big data processing. By running on the first server inside or communicating with the cluster, the big data benchmarking tool can directly access cluster resources, generate and load test data, and ensure that the test environment is as consistent as possible with the production environment, thereby improving the accuracy and reliability of test results.

[0201] Performance testing tools, such as Locust, have their worker nodes deployed on the worker nodes in the big data cluster, while the control node is deployed in a separate control node. This distributed deployment strategy can fully utilize the computing power and network resources of each worker node, enabling large-scale concurrent stress testing and accurately simulating real user behavior and high-load scenarios. The control node maintains contact with the worker nodes through a heartbeat mechanism, dynamically issuing test instructions and collecting state snapshots to ensure that test tasks on all nodes are synchronized and can provide real-time feedback on execution status. Through command-line invocation, the worker nodes of the performance testing tool can directly interact with the big data benchmarking tool, execute pre-set test scenarios, and collect key first performance indicators such as requests per second (RPS) and latency, which reflect the performance of the application layer.

[0202] Monitoring and alerting tools, such as Prometheus, are independently deployed on a second server, responsible for collecting resource utilization indicators and performance indicators from the nodes in the big data cluster and the worker nodes of the performance testing tool. The monitoring and alerting tool's collection capabilities cover comprehensive monitoring of internal performance indicators of big data components such as HDFS and Spark, from CPU, memory, disk I / O, and network I / O usage rates. These data are stored as time series data, providing detailed information for subsequent performance analysis and bottleneck identification.

[0203] Visualization tools, such as Grafana, are deployed on a third server to present and analyze these performance data. The visualization tool connects to the monitoring and alerting tool on the second server through an application programming interface, obtaining and integrating first performance indicators from the performance testing tool, resource utilization indicators from the nodes in the big data cluster, and second performance indicators from big data components on the first server. The integrated data is displayed on a customized dashboard, providing real-time performance views that allow operations personnel to visually monitor the running status of big data components and promptly identify abnormalities or bottlenecks. The alerting function of the visualization tool can immediately notify relevant personnel when key indicators exceed pre-set thresholds, ensuring prompt response and handling of problems.

[0204] The architecture design of the above scheme ensures efficient performance testing, real-time monitoring and analysis of data, and rapid response to exceptions, providing a comprehensive, automated and scalable solution for performance guarantee and optimization of big data components. By distributing the test, monitoring, alarm and analysis functions on different servers and communicating through heartbeat mechanism, API calling and other means, the independence and flexibility of each part are guaranteed, and the reasonable allocation and utilization of resources are realized, so that the whole test and monitoring process can adapt to the dynamically changing big data processing environment, while providing stable and efficient services.

[0205] In some optional embodiments of the present application, after the first performance indicator and the resource utilization indicator are associated and visualized through the dashboard in the visualization tool, and the second performance indicator is visualized, the following steps can be performed: time series alignment of the first performance indicator of the target test task and the resource utilization indicator of the first node corresponding to the target test task is performed to obtain a mapping relationship between task load and resource consumption of the target test task; based on the mapping relationship, a target node in which the resource utilization indicator exceeds a preset threshold is determined in the first node, and a target big data component in which the second performance indicator is abnormal is determined in the target node; according to the performance indicator abnormal value of the target big data component and the operation log of the target big data component, a component-level fault point causing the first performance indicator of the target test task to be abnormal is determined.

[0206] Figure 2 is a structural diagram of a data processing platform according to an embodiment of the present application, and specifically includes the following contents.

[0207] Load generation and data preparation layer: HiBench: As a standardized big data benchmark test suite, HiBench is deployed on a big data cluster or an independent test server, and is responsible for generating preset test scenarios such as WordCount, TeraSort, etc., and large-scale data sets. The data set and the test scenario are used to simulate real-world big data processing tasks to evaluate the performance of big data components under actual workloads. PerfGen (optional extension): a framework for intelligent load generation, which automatically generates test inputs that can trigger specific performance symptoms by inputting variations of performance monitoring templates and skew heuristics, to enhance the depth and complexity of testing.

[0208] Pressure scheduling and indicator collection layer: Locust: a Python-based distributed load testing tool deployed on independent control nodes and multiple worker nodes. Through Python script calls to the load generated by HiBench, simulate thousands of concurrent users, and collect real-time application layer performance indicators such as RPS and latency. The integration of Locust and HiBench makes the execution of test scenarios and the simulation of load more flexible and efficient, and can generate realistic and representative user behavior. LocustExporter: as a plugin of Locust, used to expose the performance indicators collected by Locust to Prometheus, realizing the unified collection and storage of performance data.

[0209] Monitoring and data storage layer: Prometheus: an open-source monitoring system and time series database deployed on an independent monitoring server, responsible for pulling system-level and application-level indicators from big data components (through JMX, etc.), operating systems (through Node Exporter), and Locust Exporter, storing data as time series format, providing a solid data foundation for subsequent performance analysis. Node Exporter & JMX Exporter: deployed on each node of the big data cluster, used to collect operating system-level resource utilization indicators and internal performance indicators of big data components, respectively, and transmit these data to the Prometheus server through the pulling mechanism of Prometheus.

[0210] Visualization and analysis layer: Grafana: a visualization monitoring platform deployed on an independent server, serving as a unified interface to display performance data obtained from Prometheus. By configuring dashboards, the application layer indicators of Locust and the resource utilization indicators of Prometheus are correlated and visualized, forming a deep monitoring board to help operations and development personnel monitor the performance of big data components in real time and quickly identify bottlenecks. Alarm engine: when critical indicators (such as RPS below threshold, CPU usage too high) are triggered, the alarm rules set in Grafana can automatically send notifications, realizing timely response to performance anomalies.

[0211] Intelligent optimization and suggestion layer: optimization suggestioner: based on the performance indicators and resource monitoring data collected by the platform, it can intelligently analyze and automatically identify bottlenecks causing performance degradation, and provide targeted optimization strategies or automated tuning suggestions, especially for storage components such as HDFS parameter adjustment, data compression, and caching mechanisms.

[0212] Communication connection and control layer: API interface: The platform provides a series of API interfaces, allowing external systems (such as CI / CD tools) to trigger performance test tasks, obtain test results, adjust test parameters, etc., to realize the arrangement and monitoring of automated test processes. Command line interface: Supports invocation through the command line, executes specific test scripts or manages the running state of the platform, providing technical teams with flexible control means and script automation capabilities.

[0213] Data storage and external data source layer: Time series database: Prometheus's embedded time series database, used to store collected performance and resource metric data, supporting high-performance queries and time series analysis. External data sources: including other monitoring systems, log services or historical performance data storage, used to supplement Grafana's data display, or for more in-depth data analysis and mining.

[0214] By tightly integrating the above layers, the platform realizes an end-to-end automated process from test script writing, big data component load simulation, system-level resource monitoring, performance data visualization to intelligent optimization suggestions. This design not only reduces the complexity of big data component performance testing, improves the efficiency and accuracy of testing, but also provides the operations team with the ability to actively manage performance through real-time monitoring and intelligent analysis, ensuring that big data applications remain efficient and stable under changing workloads.

[0215] Figure 3 is a structural diagram of a data processing device according to an embodiment of the present application, as shown in Figure 3 , the device comprises:

[0216] The acquisition module 32 is configured to acquire a target test scenario for a big data cluster, and generate a test data set corresponding to the target test scenario by using a big data benchmark test tool.

[0217] The execution module 34 is configured to encapsulate an executable program for executing the test data set into a test task set of a performance test tool, and execute test tasks in the test task set through multiple objects in the performance test tool.

[0218] The acquisition module 36 is configured to acquire a first performance indicator of each test task through the performance test tool, receive the first performance indicator sent by the performance test tool through a monitoring and alarm tool, and acquire a resource utilization indicator of each node in the big data cluster and a second performance indicator of a big data component in the node.

[0219] The visualization module 38 is configured to associate and visualize the first performance indicator and the resource utilization indicator through a dashboard in a visualization tool, and visualize the second performance indicator.

[0220] Optionally, the data processing apparatus further comprises a verification module configured to perform the following steps: constructing a target test input triggering a target performance symptom according to a preset performance symptom template and a skew heuristic rule, wherein the performance symptom template is configured to determine an abnormal state satisfying preset duration requirements and fluctuation intensity requirements as the target performance symptom, and the skew heuristic rule is configured to guide generation of a test case with preset distribution characteristics; performing a test on the target test input, and determining an intermediate input state triggering the target performance symptom if the target performance symptom is detected during the test; analyzing the intermediate input state by using a large language model to obtain a target pseudo-inverse function, wherein the target pseudo-inverse function is configured to map the intermediate input state into an initial input format; re-performing the test on the initial input format, and verifying triggering efficiency of the target performance symptom and resource utilization rate indicators of each node in the big data cluster according to a test result.

[0221] Optionally, the data processing apparatus further comprises a diagnosis module configured to perform the following steps: generating a diagnosis report according to a spatiotemporal coupling relationship between the first performance indicator and the resource utilization rate indicator, wherein the diagnosis report comprises at least one of the following: a type of performance bottleneck, an impact range indicator, and a numerical representation of the performance bottleneck; matching the diagnosis report with an optimization knowledge base to obtain an optimization measure for the diagnosis report, wherein the optimization knowledge base is a decision engine storing historical optimization cases and preset rules, configured to map the diagnosis report into the optimization measure, and the optimization measure comprises at least one of the following types: software parameter tuning, data governance strategy, and hardware upgrade suggestion; in a case where the optimization measure is software parameter tuning, calculating a configuration parameter value according to a cluster size and load characteristics of the big data cluster.

[0222] Optionally, the data processing apparatus further comprises a test module configured to perform the following steps: triggering a preset performance test task in response to a code submission or deployment event; comparing the first performance indicator, the resource utilization rate indicator, and the second performance indicator with a predefined baseline and a service level agreement threshold to obtain a comparison result after triggering the preset performance test task; and blocking a promotion process of a target pipeline based on continuous integration and continuous delivery in a case where the comparison result does not satisfy a preset condition, wherein the target pipeline is configured to perform the preset performance test task.

[0223] Optionally, the big data benchmarking tool is deployed on the big data cluster or a first server independent of the big data cluster; a worker node of the performance testing tool is deployed on a worker node in the big data cluster, and a control node of the performance testing tool is deployed on an independent control node in the big data cluster; the worker node of the performance testing tool dynamically synchronizes load instructions and state snapshots with the control node of the performance testing tool through a heartbeat mechanism; and the worker node of the performance testing tool interacts with the big data benchmarking tool through command line invocation.

[0224] Optionally, the monitoring and alarming tool is deployed on a second server, and the visualization tool is deployed on a third server, wherein the first server, the second server and the third server are different servers in communication with the big data cluster; the visualization tool connects the second server through an application program interface, and acquires the first performance index, the resource utilization index of each node in the big data cluster and the second performance index of the big data component in the node sent by the performance testing tool.

[0225] Optionally, after the first performance index and the resource utilization index are associated and visualized through the dashboard in the visualization tool, and the second performance index is visualized, the data processing apparatus is further configured to perform the following steps: time series aligning the first performance index of the target test task with the resource utilization index of the first node corresponding to the target test task to obtain a mapping relationship between task load and resource consumption of the target test task; based on the mapping relationship, determining a target node in which the resource utilization index exceeds a preset threshold in the first node, and determining a target big data component in which the second performance index is abnormal; determining a component-level fault point causing the first performance index of the target test task to be abnormal according to the performance index abnormal value of the target big data component and the operation log of the target big data component.

[0226] It should be noted that the above Figure 3 Each module can be a program module (for example, a program instruction set implementing a certain specific function) or a hardware module. For the latter, it can be in the form of, but not limited to, a processor, or the functions of the above-mentioned modules are implemented by a processor.

[0227] It should be noted that the preferred embodiments of the embodiments shown in Figure 3 The related descriptions can be referred to in the related description of the embodiments shown in Figure 1 , and will not be described here.

[0228] Figure 4 A hardware structure block diagram of a computer terminal for implementing the data processing method is shown. As Figure 4 shown, the computer terminal 40 can include one or more (shown in the figure as 402a, 402b, …, 402n) processors 402 (the processor 402 can include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 404 for storing data, and a transmission module 406 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports in the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 4The illustrated configuration is merely an example and does not limit the structure of the electronic device described above. For example, the computer terminal 40 can further include more or fewer components than those shown in FIG. 4, or have a different configuration than that shown in FIG. 4. Figure 4 Figure 4

[0229] It should be noted that the one or more processors 402 and / or other data processing circuitry described above can be referred to herein generally as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part within any one of the other elements of the computer terminal 40. As referred to in embodiments of the present application, the data processing circuitry functions as a processor to control, for example, the selection of the variable resistance terminal path connected to the interface.

[0230] The memory 404 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the data processing method in embodiments of the present application. The processor 402 executes various functional applications and data processing by running the software programs and modules stored in the memory 404, i.e. implements the data processing method described above. The memory 404 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 404 can further include a memory remotely arranged with respect to the processor 402, which can be connected to the computer terminal 40 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0231] The transmission module 406 is configured to receive or send data via a network. Examples of the network include a wireless network provided by a communication provider of the computer terminal 40. In one example, the transmission module 406 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission module 406 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.

[0232] The display can be, for example, a touch screen type liquid crystal display (LCD) that enables a user to interact with the user interface of the computer terminal 40.

[0233] It should be noted that in some alternative embodiments, the above-mentioned Figure 4 ​​The illustrated computer terminal can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or a combination of both hardware and software elements. It should be noted that Figure 4 is merely one instance of a particular concrete example and is intended to show the type of components that can be present in the above computer terminal.

[0234] It should be noted that Figure 4 The illustrated computer terminal is configured to perform Figure 1 The illustrated data processing method, and therefore the related explanations in the method of executing the above commands also apply to the electronic device, which will not be repeated here.

[0235] The embodiments of the present application also provide a non-volatile storage medium, which includes a stored program, wherein the program controls a device where the storage medium is located to execute the above data processing method when running.

[0236] The non-volatile storage medium executes the program to perform the following functions: obtaining a target test scenario for a big data cluster, generating a test data set corresponding to the target test scenario through a big data benchmark test tool; encapsulating an executable program for executing the test data set into a test task set of a performance test tool, and executing test tasks in the test task set through multiple objects in the performance test tool respectively; collecting a first performance index of each test task through the performance test tool, receiving the first performance index sent by the performance test tool through a monitoring and alarming tool, and collecting a resource utilization index of each node in the big data cluster and a second performance index of a big data component in the node; and associating and visualizing the first performance index and the resource utilization index through a dashboard in a visualization tool, and visualizing the second performance index.

[0237] The embodiments of the present application also provide an electronic device, which includes a memory and a processor, and the processor is configured to run a program stored in the memory, wherein the program executes the above data processing method when running.

[0238] The processor is configured to run a program for obtaining a target test scenario for a big data cluster, generating a test data set corresponding to the target test scenario by a big data benchmark test tool, encapsulating an executable program for executing the test data set as a test task set of a performance test tool, and executing test tasks in the test task set by a plurality of objects in the performance test tool respectively, collecting a first performance index of each test task by the performance test tool, receiving the first performance index sent by the performance test tool by a monitoring and alarming tool, and collecting a resource utilization index of each node in the big data cluster and a second performance index of a big data component in the node, and performing associated visualization on the first performance index and the resource utilization index by a dashboard in a visualization tool, and performing visualization on the second performance index.

[0239] The above sequence numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0240] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0241] In the above embodiments of the present application, the collected information is information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary protection measures are taken, it does not violate public order and good customs, and corresponding operation entrances are provided for the user to choose authorization or refusal.

[0242] In the several embodiments of the present application, it should be understood that the disclosed technology can be implemented in other ways. Of course, the unit described as the division is only a description of logical function division, and there can be another division manner in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, unit or module, and can be electrical or other forms.

[0243] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.

[0244] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0245] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions, can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0246] The above is only the preferred embodiment of the present application, and it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.

Claims

1. A data processing method, characterized in that: include: Obtain a target test scenario for the big data cluster, and generate a test data set corresponding to the target test scenario using a big data benchmark testing tool; Encapsulating an executable program for executing the test data set into a test task set of a performance testing tool, and executing the test tasks in the test task set respectively through multiple objects in the performance testing tool; Collecting a first performance indicator of each of the test tasks through the performance testing tool, receiving the first performance indicator sent by the performance testing tool through the monitoring and alarm tool, and collecting a resource utilization indicator of each node in the big data cluster and a second performance indicator of the big data component in the node; The first performance indicator and the resource utilization indicator are associated and visualized through a dashboard in a visualization tool, and the second performance indicator is visualized.

2. The method according to claim 1, characterized in that The method further comprises: Constructing a target test input that triggers a target performance symptom based on a preset performance symptom template and skew heuristic rules, wherein the performance symptom template is used to determine an abnormal state that meets preset duration requirements and fluctuation intensity requirements as the target performance symptom, and the skew heuristic rules are used to guide the generation of test cases with preset distribution characteristics; performing a test on the target test input, and during the test, if the target performance symptom is detected, determining an intermediate input state that causes the target performance symptom; Analyzing the intermediate input state using a large language model to obtain a target pseudo-inverse function, wherein the target pseudo-inverse function is used to map the intermediate input state to an initial input format; The initial input format is retested, and based on the test results, the triggering efficiency of the target performance symptom and the resource utilization index of each node in the big data cluster are verified.

3. The method according to claim 1, characterized in that The method further comprises: Generate a diagnostic report based on the spatiotemporal coupling relationship between the first performance indicator and the resource utilization indicator, wherein the diagnostic report includes at least one of the following: a type of performance bottleneck, an impact range indicator, and a numerical representation of the performance bottleneck; Matching the diagnostic report with an optimization knowledge base to obtain optimization measures for the diagnostic report, wherein the optimization knowledge base is a decision engine that stores historical optimization cases and preset rules, and is used to map the diagnostic report to the optimization measures, wherein the optimization measures include at least one of the following types: software parameter tuning, data governance strategy, and hardware upgrade recommendations; In the case where the optimization measure is software parameter tuning, the configuration parameter values ​​are calculated according to the cluster scale and load characteristics of the big data cluster.

4. The method according to claim 1, wherein The method further comprises: Respond to code submission or deployment events and trigger preset performance testing tasks; After triggering the preset performance test task, comparing the obtained first performance indicator, the resource utilization indicator, and the second performance indicator with a predefined baseline and a service level agreement threshold to obtain a comparison result; When the comparison result does not meet the preset conditions, the advancement process of the target pipeline based on continuous integration and continuous delivery is blocked, wherein the target pipeline is used to execute the preset performance test task.

5. The method according to claim 1, wherein include: The big data benchmark testing tool is deployed in the big data cluster or in a first server independent of the big data cluster; The working node of the performance testing tool is deployed on a working node in the big data cluster, and the control node of the performance testing tool is deployed on an independent control node in the big data cluster; the working node of the performance testing tool dynamically synchronizes load instructions and status snapshots with the control node of the performance testing tool through a heartbeat mechanism; The working node of the performance testing tool interacts with the big data benchmark testing tool through command line calls.

6. The method according to claim 5, characterized in that include: The monitoring and alarm tool is deployed on a second server, and the visualization tool is deployed on a third server, wherein the first server, the second server, and the third server are different servers communicating with the big data cluster; The visualization tool is connected to the second server through an application program interface, and obtains the first performance indicator sent by the performance testing tool, the resource utilization indicator of each node in the big data cluster, and the second performance indicator of the big data component in the node.

7. The method according to claim 1, characterized in that After visualizing the association between the first performance indicator and the resource utilization indicator and visualizing the second performance indicator through a dashboard in a visualization tool, the method further includes: Performing time series alignment on the first performance indicator of the target test task and the resource utilization indicator of the first node corresponding to the target test task to obtain a mapping relationship between the task load and resource consumption of the target test task; Based on the mapping relationship, determining, in the first node, a target node whose resource utilization index exceeds a preset threshold, and determining, in the target node, a target big data component whose second performance index has an abnormality; According to the performance indicator abnormal value of the target big data component and the operation log of the target big data component, the component-level fault point causing the abnormality of the first performance indicator of the target test task is determined.

8. A data processing device, characterized in that: include: An acquisition module is used to acquire a target test scenario for a big data cluster and generate a test data set corresponding to the target test scenario using a big data benchmark test tool; an execution module, configured to encapsulate an executable program for executing the test data set into a test task set of a performance testing tool, and respectively execute the test tasks in the test task set through a plurality of objects in the performance testing tool; a collection module, configured to collect a first performance indicator of each of the test tasks through the performance testing tool, receive the first performance indicator sent by the performance testing tool through a monitoring and alarm tool, and collect a resource utilization indicator of each node in the big data cluster and a second performance indicator of the big data component in the node; A visualization module is used to visualize the association between the first performance indicator and the resource utilization indicator through a dashboard in a visualization tool, and to visualize the second performance indicator.

9. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored program, wherein when the program is executed, the device where the non-volatile storage medium is located is controlled to execute the data processing method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the program executes the data processing method according to any one of claims 1 to 7 when running.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Performance index monitoring method and device, equipment and storage medium

    CN115277481A

  • Computer service pressure test system and method

    CN116627799A

  • Automatic performance testing method based on HDFS cluster and related equipment

    CN120179519A

  • Performance optimization method and system of big data medium table

    CN120336176A

  • Using multiple name spaces for analyzing testing data for testing scenarios involving information technology assets

    US12147296B1