A device pressure test method, device, storage medium and program product

By collecting and integrating stress test data from the Host OS and DPU OS, and using expert systems and machine learning models for automatic diagnosis, the limitations of single-dimensional testing in traditional testing methods are solved, enabling accurate identification and report generation of performance bottlenecks across multiple resource dimensions.

CN120892271BActive Publication Date: 2025-12-12LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511367660.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-12-12
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Traditional equipment stress testing methods only focus on a single dimension and cannot fully capture and analyze the complex performance issues of multiple resource dimensions when the Host OS and DPU OS work together. Furthermore, existing tools lack the function of automatically diagnosing performance bottlenecks.

Method used

Collect stress test data from Host OS and DPU OS, extract features and fuse them, use a fusion expert system and machine learning model for correlation analysis and automatic diagnosis, and generate a diagnostic report.

Benefits of technology

It improves the comprehensiveness and accuracy of equipment stress testing diagnosis, automatically locates performance bottlenecks, reduces the inefficient workload of manual troubleshooting, and meets the needs of modern data centers for efficient server operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892271B_ABST
    Figure CN120892271B_ABST
Patent Text Reader

Abstract

The application discloses a device pressure test method, device, storage medium and program product, relates to the technical field of pressure test, and comprises the following steps: constructing a pressure test strategy of a first system and a second system in a to-be-tested device, and performing corresponding pressure test operations; collecting first test data of the first system and second test data of the second system, and extracting first feature data of the first test data and second feature data of the second test data; generating target features based on the first feature data and the second feature data; generating a diagnosis result based on the target features by using a diagnosis model constructed based on an expert system and a machine learning model, and generating a diagnosis report based on a target report template. The application can collect pressure test data of each operating system, extract features for fusion, perform correlation analysis and automatic diagnosis by using a fusion expert system and a machine learning model, realizes automatic and accurate positioning of a performance bottleneck, and improves diagnosis comprehensiveness and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of pressure testing, in particular to a device pressure testing method, device, storage medium and program product. BACKGROUND

[0002] In the face of complex performance bottleneck tests of Host OS (Host Operating System) and DPU OS (Data Processing Unit Operating System) working together, the traditional test method often only focuses on a certain aspect of the server, such as only testing CPU (Central Processing Unit) performance or memory performance, while ignoring the situation that multiple resource dimensions of the server are limited at the same time in actual operation. For example, when simultaneously performing stress testing on the Host OS and the DPU OS, complex performance problems may occur in which multiple resource dimensions such as CPU, memory, network and disk are mutually restricted, and the single-dimensional test method is difficult to fully capture and analyze these problems. Moreover, most existing test tools do not have the function of automatically diagnosing performance bottlenecks, and cannot perform intelligent analysis and judgment according to real-time collected performance data.

[0003] Therefore, there is an urgent need for a device pressure testing method that can automatically and accurately diagnose server performance bottlenecks to meet the needs of efficient operation and maintenance of modern data centers. SUMMARY

[0004] The present application provides a device pressure testing method, device, storage medium and program product, which can collect pressure test data of each operating system and extract features for fusion, and use a fusion expert system and a machine learning model for correlation analysis and automatic diagnosis, to solve the problems of single-dimensional test limitations and inefficient manual troubleshooting in the current multi-operating system collaborative scenario.

[0005] The present application provides a device pressure testing method, comprising:

[0006] Constructing a pressure test strategy corresponding to a first system and a second system in a device to be tested, and executing a pressure test operation corresponding to the first system and the second system based on the pressure test strategy; the first system is a host operating system, and the second system is an operating system of a data processing unit;

[0007] Collecting first test data corresponding to the first system and second test data corresponding to the second system, and extracting first feature data of the first test data and second feature data of the second test data; the first test data and the second test data are respectively hardware performance data generated when the pressure test operation is performed on the first system and the second system.

[0008] generate a target feature based on the first feature data and the second feature data;

[0009] generate a diagnosis result corresponding to the to-be-tested device based on the target feature by using a preset diagnosis model, and generate a corresponding diagnosis report according to the diagnosis result based on a target report template; the preset diagnosis model is a model constructed based on a preset expert system and a preset machine learning model.

[0010] The application further provides an electronic device, including a memory for storing a computer program, and a processor for executing the computer program to implement the steps of the device pressure test method.

[0011] The application further provides a computer readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of the device pressure test method are implemented.

[0012] The application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the device pressure test method are implemented.

[0013] In the application, first, a pressure test strategy corresponding to a host operating system and an operating system of a data processing unit in the to-be-tested device can be constructed, and corresponding pressure test operations are performed based on the pressure test strategy, then test data generated when the pressure test operations are performed is collected, first feature data and second feature data of the host operating system and the operating system of the data processing unit are extracted, a target feature is generated based on the first feature data and the second feature data, and a preset diagnosis model constructed based on a preset expert system and a preset machine learning model is used to generate a diagnosis result corresponding to the to-be-tested device based on the target feature, and a corresponding diagnosis report is generated according to the diagnosis result based on a target report template.

[0014] Through the application, pressure test data of a host operating system and an operating system of a data processing unit in a server can be collected, and feature fusion is performed after the features are extracted respectively, so that the internal relationship and mutual influence between different test data are mined through deep correlation analysis between the features, which helps to find complex performance problems during testing, improves diagnosis comprehensiveness and accuracy, and then helps to generate a detailed diagnosis report, improves report generation efficiency and accuracy, and automatically diagnoses by using a model of a fusion expert system and machine learning, realizes automatic and accurate positioning of a server performance bottleneck, makes up for the single dimension test limitation in the current multiple operating system collaborative scene and the low efficiency of manual troubleshooting, and improves diagnosis comprehensiveness and accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort based on these drawings.

[0016] Figure 1 A device pressure test method flow chart provided for the embodiments of the present application;

[0017] Figure 2 A pressure test flow chart provided for the embodiments of the present application;

[0018] Figure 3 A device pressure test system architecture diagram provided for the embodiments of the present application;

[0019] Figure 4 A specific device pressure test method flow chart provided for the embodiments of the present application;

[0020] Figure 5 A device pressure test device structure schematic diagram provided for the embodiments of the present application. DETAILED DESCRIPTION

[0021] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present application.

[0022] It should be noted that in the description of the present application, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, not to describe a specific order or sequence.

[0023] In the face of complex performance bottleneck tests of the cooperation of the Host OS and the DPU OS, the traditional test method often only focuses on a certain aspect of the server, such as testing only the CPU performance or the memory performance, while ignoring the case that multiple resource dimensions of the server are limited at the same time in actual operation, and the existing test tools mostly do not have the function of automatically diagnosing the performance bottleneck, and cannot intelligently analyze and judge according to the real-time collected performance data. The present application can collect the stress test data of the Host OS and the DPU OS in the server, and after extracting the features respectively, the features are fused, so as to mine the relationship between different test data through the correlation analysis between the features, and automatically diagnose by using the model of the fusion expert system and machine learning, thereby improving the comprehensive diagnosis and accuracy.

[0024] In order for those skilled in the art in this technical field to better understand the present application scheme, the present application will be further described in detail below in combination with the drawings and specific embodiments.

[0025] Next, the present embodiment will be described in detail in combination with the execution flow of the device stress test method, as shown in Figure 1 The embodiment of the present application provides a device stress test method, which comprises the following steps:

[0026] Step S11, constructing a stress test strategy corresponding to the first system and the second system in the device to be tested, and performing a stress test operation corresponding to the first system and the second system based on the stress test strategy; the first system is a host operating system, and the second system is an operating system of a data processing unit.

[0027] In the present embodiment, as shown in Figure 2 First, a stress test strategy corresponding to the first system and the second system in the device to be tested needs to be constructed, and a stress test operation corresponding to the first system and the second system is performed based on the stress test strategy; the first system is a host operating system Host OS, and the second system is an operating system of a data processing unit DPU OS. Specifically, when constructing the stress test strategy, the test requirement of the device to be tested is determined, and the test scene of the first system and the second system is determined according to the test requirement, and the stress test strategy and the stress test tool corresponding to the first system and the second system are determined, and then the stress test strategy corresponding to the first system and the second system is constructed based on the test scene, the stress test strategy and the stress test tool.

[0028] That is, in the stress test phase, first, the test scene configuration is performed, and the test scene is set according to the actual demand and test target, including but not limited to the application programs, services and expected load levels that need to be run in the Host OS and DPU OS, for example, simulating high-concurrency user access to the Web server and the database server under the big data processing scene. Then, a reasonable stress strategy is made to determine the way and strength of stress on the server and DPU, such as Figure 3 As shown, the stress mode includes but is not limited to CPU occupation realized by running multi-threaded computing tasks, memory allocation realized by cyclically applying and releasing large blocks of memory, network traffic sending parameters realized by simulating different types of network requests, and disk read-write operations realized by frequently executing file read-write tasks, and it can be understood that the stress strength can be gradually increased according to the actual demand to simulate the performance of the server under different load levels. At the same time, the stress tool selection and deployment are performed, first, the appropriate stress tool is selected, such as Stress-ng, Memtester and other tools for Host OS, and the test tool set for the network forwarding performance of DPU (such as DPDK (Data Plane Development Kit) test tool set), and the stress tool is deployed in the test environment, and connected with the server and DPU, and ready to start stress test.

[0029] Step S12, collect first test data corresponding to the first system and second test data corresponding to the second system, and extract first feature data of the first test data and second feature data of the second test data; the first test data and the second test data are hardware performance data generated when the pressure test operation is performed on the first system and the second system respectively.

[0030] In this embodiment, as shown in Figure 2 The first test data and the second test data are hardware performance data generated when the pressure test operation is performed on the first system and the second system respectively. Moreover, during the process of collecting the first test data corresponding to the first system and the second test data corresponding to the second system, the business performance data of the to-be-tested device can also be collected, and the business performance feature vector of the business performance data can be extracted.

[0031] In collecting test data, the first test data corresponding to the first system and the second test data corresponding to the second system can be collected based on a preset time interval, and the first test data and the second test data are stored in a preset memory buffer. Then, the first test data and the second test data in the preset memory buffer are preprocessed, and the preprocessed first test data and second test data are stored in a preset database. The preset database is a local database of the device under test or a database in a distributed storage system. Figure 3 As shown in the figure, a server stress testing system is disclosed in this embodiment, which can use a data collection module for real-time data collection. When the stress test is started, the data collection module collects the performance index data of the server Host OS and DPU OS in real time according to the preset time interval, including the relevant data of the CPU, memory and network of the host operating system Host OS and the corresponding special data index of the data processing unit operating system DPU OS, and stores them in the memory buffer to ensure the real-time and integrity of the data. Then, after preprocessing such as data cleaning, format conversion and other processes, the collected data is stored in the local database or the distributed storage system after removing invalid or abnormal data, and the feature extraction of the collected data is performed, so that the subsequent diagnostic analysis module can read and process it.

[0032] In this way, the data collection module can be responsible for real-time and accurate collection of multi-dimensional performance indicator data of the server Host OS and the DPU OS during the stress test process. The performance indicators include CPU usage rate accurate to each core, memory occupancy including allocation and release of physical memory and virtual memory, network bandwidth utilization distinguishing uplink and downlink bandwidth, hardware resource indicators such as disk I / O (Input / Output) operation delay and throughput, and business performance indicators such as request response time of a web server, transaction processing success rate of a database server, and query rate per second, etc. When collecting data, the data can be obtained by deep interaction with the underlying interface of the Host OS and the DPU OS, for example: for the Host OS, the performance data is collected regularly by using its built-in performance counters (such as Windows Performance Monitor API, Linux sysctl and / proc file system, etc.) and application program interface (API, Application Programming Interface); for the DPU OS, the use of internal hardware resources and network protocol processing performance data can be collected by means of the management interface and debugging tools provided by the DPU. At the same time, the embodiment can also collect business performance indicators by implanting monitoring code in the business application or using application program performance monitoring (APM, Application Performance Management) tools.

[0033] In the embodiment, when performing feature extraction of test data, first initial features of first test data and second initial features of second test data are extracted based on test requirements of the device under test, and a covariance matrix corresponding to the first initial features and the second initial features is determined, a corresponding eigenvector and eigenvalue are determined according to the covariance matrix, and then a cumulative contribution rate corresponding to each eigenvector is determined based on the eigenvalue, so as to determine first feature data of the first test data and second feature data of the second test data from the eigenvectors according to the cumulative contribution rate.

[0034] It can be understood that in the server stress test scenario, the performance bottleneck information is often hidden in the multi-dimensional and multi-source performance data. In order to more accurately mine these information, the embodiment first pre-processes the collected multi-source data, including data cleaning, normalization and feature extraction steps, to eliminate the dimensional difference and noise interference between different data sources. Then, using dimension reduction techniques such as principal component analysis (PCA), the high-dimensional performance data is mapped to a low-dimensional space, and the most representative features for performance bottleneck diagnosis are extracted. Thus, on this basis, combined with expert system rules and machine learning models, the integrated data is comprehensively analyzed to realize the automatic positioning and diagnosis of performance bottlenecks. Compared with the traditional single data source diagnosis method, the multi-source data fusion algorithm can make full use of the performance data of Host OS and DPU OS and the performance of the business layer, and avoid the one-sidedness of diagnosis caused by relying on only one type of data. For example, relying on only hardware performance indicators may miss performance problems caused by unreasonable business logic, and after integrating business performance indicators, the actual running status of the server can be more comprehensively reflected, improving the accuracy and reliability of diagnosis.

[0035] Specifically, when performing data preprocessing, first, the invalid values, missing values and abnormal values in the collected performance data are removed. For example, for CPU usage data that is obviously outside the reasonable range of 100%, it is regarded as an abnormal value and is corrected or deleted. Then the pre-processed data is normalized to convert performance indicator data of different dimensions to the same numerical range, such as the [0, 1] interval, to facilitate subsequent comprehensive analysis. Specifically, the Min-Max normalization method can be used, and the calculation formula is:

[0036] ;

[0037] Where x is the original data, x min is the minimum value of the data, x max is the maximum value of the data, and x' is the normalized value.

[0038] After that, the normalized data can be subjected to feature extraction and dimension reduction. First, according to the requirements of performance bottleneck diagnosis, key features are extracted from the preprocessed data. For example, for CPU performance bottleneck diagnosis, features such as CPU usage, process context switch frequency, and system call frequency are extracted; for network performance bottleneck diagnosis, features such as network bandwidth utilization, data packet transmission delay, and TCP retransmission rate are extracted. Then, principal component analysis is performed to reduce the dimension of the extracted high-dimensional feature vector, retain the most important feature components in the data, and reduce the computational complexity. Specifically, the covariance matrix of the feature vector can be calculated, and its eigenvalues and eigenvectors are solved, and the principal components whose cumulative contribution rate reaches a certain threshold (such as 90%) are selected as the feature representation after dimension reduction.

[0039] Step S13, generating a target feature based on the first feature data and the second feature data.

[0040] In this embodiment, the first feature data and the second feature data can be fused to generate a target feature. Based on the above steps, this embodiment can determine the weight coefficients corresponding to the first feature data, the second feature data and the service performance feature vector respectively, and generate the target feature based on the first feature data, the second feature data and the service performance feature vector based on the weight coefficients. Specifically, the reduced hardware performance feature vectors of the Host OS and the DPU OS can be fused with the service performance feature vector to construct a comprehensive performance feature vector. In this embodiment, a weighted fusion strategy can be used to assign weights according to the importance of each data source to performance bottleneck diagnosis, and the calculation formula is:

[0041] ffuse=whost⋅fhost+wdpu⋅fdpu+wbusiness⋅fbusiness;

[0042] wherein, ffuse is the fused comprehensive performance feature vector, fhost, fdpu and fbusiness are the Host OS hardware performance feature vector, the DPU OS hardware performance feature vector and the service performance feature vector respectively, whost, wdpu and wbusiness are the corresponding weight coefficients, and whost+wdpu+wbusiness=1 is satisfied. Through the above technical solution, the hardware performance indicators and service performance indicators of the server Host OS and DPU OS can be comprehensively fused, a unified data analysis framework is constructed, and all-around diagnosis of performance bottlenecks is realized. In this way, by collecting multi-dimensional performance indicators and service performance indicators of the server Host OS and DPU OS, including CPU usage, memory occupation, network bandwidth utilization, disk I / O, etc., the system performance is comprehensively reflected, and through deep correlation analysis, the internal relationship and mutual influence between different indicators are mined, complex performance problems are found, which helps to improve the accuracy of the stress test results. Meanwhile, the reduced hardware performance feature vectors and service performance feature vectors of the Host OS and DPU OS are fused to construct a comprehensive performance feature vector, the information of each data source is fully utilized to improve the comprehensiveness and accuracy of diagnosis, and a weighted fusion strategy is adopted to assign weights according to the importance of the data source to the performance bottleneck diagnosis, and the influence of key data is enhanced.

[0043] In step S14, a preset diagnosis model is used to generate a diagnosis result corresponding to the to-be-tested device based on the target feature, and a diagnosis report is generated based on the target report template according to the diagnosis result.

[0044] In this embodiment, a preset diagnosis model constructed based on a preset expert system and a preset machine learning model can be used to generate a diagnosis result corresponding to the to-be-tested device according to the target feature, and a diagnosis report is generated based on the target report template according to the diagnosis result. By inputting the fused comprehensive performance feature vector into the expert system and the machine learning model for comprehensive analysis, the expert system can preliminarily screen the comprehensive feature vector according to predefined rules to determine whether there is a potential performance bottleneck; the machine learning model further analyzes the screened feature vector in depth to determine the specific type and location of the performance bottleneck, and outputs the diagnosis result. As shown in FIG. 8, the diagnosis analysis module can receive the data sent by the data acquisition module, and analyze the data by using the expert system constructed based on the rule base and the machine learning model trained based on the training data, and then fuse the analysis results of the expert system and the machine learning model, so as to generate a report by using the report engine of the report generation module according to the diagnosis result, and present the report to the user in the form of browser display, email notification, preset interface push, etc. Figure 3 ​

[0045] In the process of generating a corresponding diagnostic report based on the target report template according to the diagnostic result, the target report template corresponding to the diagnosis can be determined from a plurality of initial report templates, and the corresponding diagnostic report is generated based on the target report template according to the diagnostic result, and the diagnostic report is converted into a target document format. Specifically, first, a suitable report template can be selected according to the type and complexity of the diagnostic analysis result. For example: for a simple performance bottleneck problem, a concise report template can be selected; for a complex multi-dimensional performance problem, a detailed report template is selected. It can be understood that the report template can be customized according to user needs in the embodiment. Then fill in the performance bottleneck position, cause analysis, affected indicators and other contents in the diagnostic analysis result according to the format of the report template to generate a complete diagnostic report. After the report is generated, it can be viewed through the report viewing tool provided by the system or exported as a common document format such as PDF (Portable Document Format), Word, etc., for easy user viewing, sharing and archiving. In this way, a detailed diagnostic report containing performance bottleneck position, cause analysis, affected indicators and optimization suggestions is automatically generated according to the diagnostic analysis result, improving the report generation efficiency and accuracy, and using a combination of charts, graphs and text for visual display, so that users can quickly understand the performance problem.

[0046] In this way, in the embodiment, a detailed and intuitive diagnostic report can be automatically generated according to the output result of the diagnostic analysis module. The report content not only includes the position of the performance bottleneck such as a specific process in the Host OS, a network forwarding module in the DPU OS, and the cause analysis such as memory leakage and unreasonable CPU task scheduling, but also covers the affected resources and business indicators, and targeted optimization suggestions such as adjusting process priority and optimizing memory allocation strategy. Moreover, a plurality of report templates are designed, and a suitable template can be selected for report generation according to different test scenarios and diagnostic results. In this way, the report uses a combination of text and pictures, and the performance problem is intuitively displayed through charts such as performance indicator trend chart and performance bottleneck distribution pie chart. The analysis result and optimization suggestion are described in detail in combination with the text description. After the report is generated, the user can obtain the report through email, browser or local file system, conveniently and quickly understand the server performance status and take optimization measures. The pressure test system and method can automatically and accurately diagnose the server performance bottleneck and provide intelligent optimization suggestions, meeting the needs of modern data centers for high performance, high reliability and efficient operation and maintenance of servers.

[0047] Based on the last embodiment, the present application can collect pressure test data of each operating system, extract features for fusion, and then automatically diagnose. Next, the process of using a fusion expert system and a machine learning model for device diagnosis will be described in detail in the embodiment. Referring toFigure 4 As shown, the embodiments of the present application disclose a specific device pressure test method, which comprises the following steps:

[0048] In step S21, a preset expert system in the preset diagnosis model is used to screen the to-be-analyzed features in the target features based on preset feature screening rules.

[0049] In this embodiment, the preset expert system in the preset diagnosis model can be used to screen the to-be-analyzed features in the target features based on preset feature screening rules, for example, Figure 3 As shown, after the diagnosis analysis module is started, the expert system rule base is first called to preliminarily screen and judge the collected performance index data. According to the predefined rules, the indexes and links that may have performance bottlenecks can be quickly identified and marked as objects to be further analyzed.

[0050] In step S22, a preset machine learning model in the preset diagnosis model is used to analyze the to-be-analyzed features to obtain the diagnosis result of the to-be-tested device.

[0051] In this embodiment, the preset machine learning model in the preset diagnosis model can be used to analyze the to-be-analyzed features to obtain the diagnosis result of the to-be-tested device. That is, the preliminarily screened data can be input into the machine learning model for deep diagnosis and analysis. The machine learning model can comprehensively consider the correlation and trend changes between multiple performance indexes to determine the specific position and cause of the performance bottleneck. For example, by analyzing the relationship between the CPU usage rate, memory occupation and business response time, it can be determined whether the abnormal increase in the CPU usage rate is caused by memory leakage or the excessive consumption of CPU resources caused by the dead loop in the business logic code. In this embodiment, in order to ensure the accuracy of the diagnosis result, the diagnosis analysis module also performs correlation analysis and verification. For example, the performance indexes of the Host OS and the DPU OS are analyzed to check whether there is a situation that the network application program in the Host OS has an excessively long response time due to the network forwarding delay of the DPU. At the same time, the influence degree of the performance bottleneck on the actual business application is verified in combination with the business performance indexes.

[0052] In this embodiment, before the preset diagnosis model is used to generate the diagnosis result corresponding to the to-be-tested device based on the target features, an initial diagnosis model needs to be constructed based on a pre-trained model, the initial diagnosis model is initialized according to preset model parameters to obtain the preset diagnosis model, and the reinforcement learning parameters corresponding to the initial diagnosis model are set. The preset model parameters are random parameters or parameters obtained based on transfer learning. Correspondingly, after the preset diagnosis model is used to generate the diagnosis result corresponding to the to-be-tested device based on the target features, the preset diagnosis model can be optimized by using the reinforcement learning parameters.

[0053] It should be noted that before optimizing the preset diagnostic model using reinforcement learning parameters, the first and second systems in the device under test need to be adjusted according to the diagnostic report. The process then proceeds to the step of collecting the first test data corresponding to the first system and the second test data corresponding to the second system to generate new target features. These new target features are then used as the features to be diagnosed, and corresponding experience tuples are constructed based on the target features and the features to be diagnosed. These experience tuples are then stored in a preset experience replay buffer. Specifically, when constructing the corresponding experience tuples based on the target features and the features to be diagnosed, the test environment for stress testing the first and second systems can be determined, and the actual reward signal generated by the test environment based on the diagnostic results can be obtained. Experience tuples are then constructed based on the target features, the features to be diagnosed, the actual reward signal, and the diagnostic results.

[0054] Specifically, this embodiment introduces reinforcement learning technology to enable the performance bottleneck diagnostic model to adapt to performance changes in servers and DPUs under different load conditions and application scenarios. A reinforcement learning-based adaptive optimization algorithm treats the diagnostic model as an agent and the server and DPU performance testing environment as an environment. Through interaction between the agent and the environment, the parameters of the diagnostic model are continuously optimized to improve diagnostic accuracy and adaptability. In this way, this embodiment combines expert systems and machine learning algorithms to automatically and accurately locate and analyze performance bottlenecks in the server's Host OS and DPU OS, reducing manual troubleshooting workload and time costs. Furthermore, it uses expert system rules to identify common performance bottleneck patterns and utilizes machine learning algorithms to learn and train on historical test data to optimize the diagnostic model, improving diagnostic accuracy and efficiency.

[0055] In this embodiment, a pre-trained machine learning model (such as a DNN (Deep Neural Network)) is first used as the initial diagnostic model. Its parameters are initialized to random values ​​or parameters transferred from relevant tasks through transfer learning. Then, relevant reinforcement learning parameters are set, such as the learning rate (α), discount factor (γ), and exploration rate (…). (e.g., learning rate controls the speed at which the agent learns new experiences, discount factor balances the agent's focus on short-term and long-term rewards, and exploration rate determines whether the agent chooses to explore new actions or utilize existing experience during the decision-making process.) Next, the diagnostic model executes diagnostic actions, outputting diagnostic actions based on the current comprehensive performance feature vector and the diagnostic model. This involves judging the performance status of the server and DPU, determining whether performance bottlenecks exist and their types, and obtaining reward signals. It can be understood that the environment can calculate reward signals by comparing the agent's diagnostic actions with the actual performance bottleneck situation. The reward signal can be defined as:

[0056] ;

[0057] wherein r is a reward signal, a is a diagnosis action of the agent, s is a current comprehensive performance feature vector, y true is an actual performance bottleneck label (obtained by manual labeling or known performance test results), y pred is a performance bottleneck label predicted by the agent. Through the above reward formula, when the diagnosis result of the agent is correct, a positive reward is given; otherwise, a negative reward is given.

[0058] And as Figure 3 shown, the preset diagnosis model can also be optimized using reinforcement learning parameters and training data. Specifically, experience tuples in the preset experience replay buffer can be randomly sampled to obtain model training data, and the preset diagnosis model can be used to generate a predicted reward signal based on the model training data. Then, based on the preset loss function, a loss value between the predicted reward signal and the actual reward signal is determined, and based on the loss value and the reinforcement learning parameters, the model parameters of the preset diagnosis model are adjusted to optimize the preset diagnosis model. That is, the embodiment can collect experience, that is, the current comprehensive performance feature vector s, the diagnosis action a, the reward signal r, and the next comprehensive performance feature vector s' (after the environment state is shifted according to the diagnosis action) are combined into an experience tuple (s, a, r, s'), and stored in the experience replay buffer (Experience Replay Buffer) for subsequent model training.

[0059] Next, model updating and optimization are realized by random sampling and batch training. First, a batch of experience tuples is randomly sampled from the experience replay buffer to form a training batch, and the loss function (such as the mean square error loss function) between the predicted value of the diagnosis model and the actual reward signal is calculated to update the parameters of the diagnosis model using the gradient descent algorithm. The loss function is defined as:

[0060] ;

[0061] wherein L is a loss function, is a TD (Temporal Difference) target, is a current predicted value of the diagnosis model, is a parameter of the diagnosis model, is a parameter of a target network (Target Network) used to stabilize the training process. After that, the parameters of the diagnosis model are copied to the target network at certain training steps to update the parameters of the target network, ensuring that the TD target is calculated based on the latest diagnosis model parameters.

[0062] It should also be noted that this embodiment employs the following during the model update process: -greedy( -Greedy) strategy balance exploration and utilization, that is, using probability Randomly select diagnostic actions for exploration, with a probability of 1− The optimal diagnostic action is selected and utilized based on the current diagnostic model. As training progresses, the optimal action is gradually reduced. The value of the algorithm is adjusted to reduce the frequency of exploration and increase the frequency of utilization, allowing the agent to gradually converge to the optimal diagnostic strategy. Through this reinforcement learning-based adaptive optimization algorithm, the diagnostic model can automatically adjust its parameters in continuous environmental interactions, adapting to performance changes in the server and DPU, improving the accuracy and adaptability of performance bottleneck diagnosis, and providing a more intelligent and efficient solution for performance optimization in server stress testing, further enhancing its practicality. In this embodiment, the diagnostic model is considered an agent, and the performance testing environment of the server and DPU is considered the environment. Through the interaction between the agent and the environment, the diagnostic model parameters are optimized to adapt to performance changes under different load conditions and application scenarios. Reinforcement learning technology is used to enable the diagnostic model to dynamically adjust its diagnostic strategy, ensuring the accuracy and timeliness of the diagnosis.

[0063] In this embodiment, within the reinforcement learning framework, the agent can take diagnostic actions (such as determining the existence and type of performance bottlenecks) based on its current performance state (i.e., the fused comprehensive performance feature vector). The environment provides corresponding reward signals based on the agent's actions (such as the accuracy of the diagnostic results and the diagnostic time). Through continuous trial-and-error learning, the agent can adjust its strategy (i.e., the parameters of the diagnostic model) to maximize long-term cumulative rewards, thereby achieving adaptive optimization of the diagnostic model. Compared to traditional static diagnostic models, this embodiment, through a reinforcement learning-based adaptive optimization algorithm, can dynamically adjust the diagnostic strategy to adapt to dynamic changes in server and DPU performance. For example, when the server's load pattern suddenly changes from low load to high load, a traditional diagnostic model may need to be retrained to adapt to the new performance characteristics. However, the reinforcement learning-based algorithm can adjust the model parameters in real time, quickly adapting to new diagnostic needs and ensuring the accuracy and timeliness of the diagnosis.

[0064] Based on the above embodiments, such as Figure 2As shown, in the pressure test in the embodiment, first, test scene configuration and pressure strategy formulation are performed, after the pressure object to be tested is selected according to the pressure strategy, the Host pressure tool and the DPU pressure tool are started to perform continuous pressure test on the Host OS and the DPU OS respectively, and in the pressure test process, the performance data is collected in real time by the data collection module and stored, so as to establish the corresponding test data set according to the collected real-time data, then the diagnostic analysis module selects the analysis trigger condition, that is, the data can be analyzed immediately during the test or after the test is completed to locate the performance bottleneck, finally, the root cause analysis of the performance bottleneck is performed, and the diagnostic report is generated by the report generation module and the report is visually presented by the preset display interface.

[0065] Through the above technical solution, the embodiment can identify common performance bottleneck patterns based on the collected performance index data, using the rules and experience models predefined in the expert system that integrates the performance expert knowledge of servers and DPUs, for example, when the CPU is frequently in full load state and the network bandwidth utilization is low, there may be a network protocol processing task with excessive CPU resource consumption. And by using machine learning algorithms to learn and train massive historical test data, a performance bottleneck diagnosis model is constructed, which dynamically adapts to different server configurations and application scenarios, and continuously optimizes the accuracy and efficiency of diagnosis. In this way, the expert system rule base is defined and maintained by performance experts according to actual experience and industry best practices, covering a variety of common performance problem scenarios; the machine learning algorithm uses a supervised learning method, taking performance indicators in historical test data as input features and known performance bottleneck types as output labels for training. In the actual diagnosis process, first, the expert system performs preliminary screening on the collected performance data to quickly lock the indicators and links that may have performance bottlenecks; then, the data preliminarily screened is input into the machine learning model for deep analysis and verification to finally determine the specific location and reason of the performance bottleneck.

[0066] Based on the above embodiment, a specific device pressure test method is disclosed in the embodiment, comprising:

[0067] First, Host OS performance indicators are collected. For example, for a Linux system, CPU usage rate ( / proc / stat file), memory occupation ( / proc / meminfo file), disk I / O operation delay and throughput ( / proc / diskstats file) and other performance indicators can be collected regularly by calling the sysctl function and reading related files in the / proc file system. For CPU usage rate, the data in the cpu line of the / proc / stat file is analyzed, including user space occupation time, kernel space occupation time, idle time, etc., to calculate the real-time CPU usage rate. For memory occupation, parameters such as MemTotal, MemFree, Buffers and Cached are obtained from the / proc / meminfo file to calculate the actual memory usage and available memory. For disk I / O operation delay and throughput, the read / write request times, read / write byte counts and read / write times of disk devices in the / proc / diskstats file are analyzed to calculate the average response time and read / write throughput per second of the disk. Then, DPU OS performance indicators are collected. According to the characteristics of the DPU, the management interface and debugging tools provided by the DPU, such as the management API of the NVIDIA BlueField DPU, are used to collect the usage of internal hardware resources, such as CPU core usage rate, memory occupation, network interface data packet transmission rate, etc. At the same time, the processing performance indicators of the network protocol stack and data processing application running on the DPU are obtained, such as network protocol processing delay, data packet forwarding rate, storage offloading task execution time, etc. Business performance indicators can also be collected. For a web server such as Apache or Nginx running on the Host OS, access log and error log recording functions are enabled in its configuration file, and third-party performance monitoring tools such as New Relic or Prometheus are used to collect business performance indicators such as web server request response time, requests per second, concurrent connection number, etc. For a database server such as MySQL or PostgreSQL, the performance status table of the database such as the tables table in the information_schema database of MySQL is queried, and the built-in performance monitoring tool of the database such as the slow query log of MySQL is used to collect business performance indicators such as database transaction processing success rate, queries per second, query response time, etc.

[0068] In the process of data collection, first, the collection frequency is set. According to the performance characteristics of the server and the DPU and the test requirements, the data collection frequency is reasonably set. For example, for relatively fast-changing indicators such as CPU usage and memory occupancy, a higher collection frequency is set, such as collecting once per second; for relatively slowly changing indicators such as disk I / O operation delay and throughput and service performance indicators, the collection frequency can be appropriately reduced, such as collecting once per minute. And in the test process, the collection frequency is dynamically adjusted according to the actual situation to balance the accuracy of data collection and the system resource occupation. Next, the collected performance indicator data is stored in the memory buffer area to ensure the real-time and integrity of the data, and in order to prevent data loss, the data in the memory buffer area is periodically written to a local database such as SQLite or MySQL or a distributed storage system such as Hadoop HDFS or Ceph for persistent storage, and at the same time, the stored data is periodically backed up and archived to meet the long-term data storage and historical data analysis requirements.

[0069] After that, the server performance experts and DPU technology experts of the embodiment organize, according to the actual performance test experience and industry best practices, define a series of expert system rules, which include but are not limited to: CPU usage is too high and memory occupancy is continuously increasing, which may be caused by memory leakage leading to excessive CPU usage; the network bandwidth utilization rate is close to saturation, but the network forwarding performance of the DPU OS is not fully utilized, which may be caused by unreasonable network protocol stack configuration of the Host OS or defects in the network driver. After that, these rules are stored in the rule base in the form of condition-conclusion, where the condition includes the threshold range of the performance indicator, the change trend, etc., and the conclusion is the type and location of the possible performance bottleneck. Then, an efficient rule matching algorithm such as Rete algorithm or its improved version is used to quickly match the collected performance indicator data, and in the matching process, the performance indicator data is compared with the conditions in the rule base one by one, and when the rule condition is met, the corresponding conclusion is triggered, that is, the type and location of the possible performance bottleneck are determined. At the same time, in order to improve the matching efficiency, the rule base can be optimized and indexed, and at the same time, according to the importance and correlation of the performance indicators, the priority of the rules is set, and the rules with high priority are matched first.

[0070] When training the machine learning algorithm, a large amount of historical test data is first collected, including performance indicator data of the server and the DPU under different load conditions, and performance bottleneck cases determined through manual investigation and analysis. The collected data is cleaned and preprocessed to remove missing values, outliers and duplicates, and normalized to the same dimension range to facilitate processing by the machine learning algorithm. At the same time, feature selection and extraction are performed on the data, selecting effective features related to performance bottlenecks such as CPU usage, memory occupancy, network bandwidth utilization, and disk I / O operation delay as input features, and performance bottleneck types such as CPU bottleneck, memory bottleneck, and network bottleneck as output labels. Then, a suitable machine learning algorithm is selected, such as random forest, support vector machine (SVM), or deep neural network (DNN), to train the preprocessed data. During the training process, cross-validation, grid search, and other methods are used to optimize the model's hyperparameters to improve the model's accuracy and generalization ability. For example, for the random forest algorithm, the number of decision trees and the depth of the tree, feature selection strategy, and other hyperparameters are adjusted; for the DNN algorithm, the network structure, activation function, learning rate, and other hyperparameters are adjusted. Through continuous iteration of training and optimization, the performance bottleneck diagnosis model is finally obtained.

[0071] In generating the report template, first, according to different test scenarios and performance bottleneck types, multiple report templates are designed. For example, for simple CPU or memory bottleneck problems, a concise report template is designed, mainly including three parts of performance bottleneck overview, cause analysis and optimization suggestion; for complex multi-dimensional performance bottleneck problems, such as mutual restriction of multiple resource dimensions such as CPU, memory, network and disk, a detailed report template is designed, including performance bottleneck overview, detailed analysis, affected resources and business indicators, optimization suggestion and conclusion and multiple parts. And allow users to customize the report template according to actual needs. Users can modify the format, content and style of the report template through a graphical interface or a configuration file. For example, users can add or delete certain parts of the report, adjust the type and display method of the chart, change the font, color and size of the text, etc. to meet different report generation needs. Then, according to the output of the diagnosis and analysis module, the position of the performance bottleneck, the cause analysis, the affected resources and business indicators and the optimization suggestions and other information are filled into the corresponding report template. In the filling process, the specific data and analysis conclusions in the diagnosis results are replaced into the placeholder positions in the report template by using text replacement and data binding, and at the same time, according to the data type and content characteristics, suitable charts such as line chart, column chart and pie chart are selected to visualize the performance indicator data, making the report content more intuitive and easy to understand. After the report is generated, multiple export formats such as PDF, Word, HTML (Hyper Text Markup Language) are provided to facilitate users to view, share and archive. Users can select the appropriate export format according to actual needs, and set the export path and file name. In the export process, ensure the integrity and accuracy of the report content, and encrypt and protect the report to prevent unauthorized access and modification.

[0072] As shown in Figure 5 The embodiments of the present application also provide a device pressure test device, comprising:

[0073] The pressure test module 11 is configured to construct a pressure test strategy corresponding to the first system and the second system in the device under test, and execute a pressure test operation corresponding to the first system and the second system based on the pressure test strategy; the first system is a host operating system, and the second system is an operating system of a data processing unit;

[0074] The feature extraction module 12 is configured to collect first test data corresponding to the first system and second test data corresponding to the second system, and extract first feature data of the first test data and second feature data of the second test data; the first test data and the second test data are hardware performance data generated when the pressure test operation is performed on the first system and the second system, respectively;

[0075] The feature generation module 13 is configured to generate target features based on the first feature data and the second feature data.

[0076] The report generation module 14 is configured to generate a diagnosis result corresponding to the device under test based on the target features by using a preset diagnosis model, and generate a corresponding diagnosis report based on the target report template and the diagnosis result. The preset diagnosis model is a model constructed based on a preset expert system and a preset machine learning model.

[0077] The features of the embodiments of the device pressure test apparatus can be referred to the related descriptions of the embodiments of the device pressure test method, which will not be repeated here.

[0078] Through the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment.

[0079] Embodiments of the present application also provide an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above device pressure test method embodiments.

[0080] Embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above device pressure test method embodiments when running.

[0081] In an example embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0082] Embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any of the above device pressure test method embodiments.

[0083] Embodiments of the present application also provide another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the above device pressure test method embodiments.

[0084] Those skilled in the art will further realize that the mere concepts, teachings, and embodiments described herein are merely meant to provide an enabling description of the claimed application. Accordingly, modifications and / or additions, other than those explicitly described herein, can be obvious to those skilled in the art in the light of this disclosure. The claimed application is intended to embrace all such modifications and / or additions.

[0085] The method, device, storage medium and program product for testing pressure of a device are described in detail above. The principles and implementation manners of the present application are described by applying specific examples in the present application. The above description of the embodiments is only applicable to help understand the method and core idea of the present application. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method of device stress testing, the method comprising: The method comprises the following steps: constructing a pressure test strategy corresponding to a first system and a second system in a to-be-tested device, and performing a pressure test operation corresponding to the first system and the second system based on the pressure test strategy; the first system is a host operating system, and the second system is an operating system of a data processing unit; collecting first test data corresponding to the first system and second test data corresponding to the second system, and extracting first feature data of the first test data and second feature data of the second test data; the first test data and the second test data are respectively hardware performance data generated when the pressure test operation is performed on the first system and the second system; generating a target feature based on the first feature data and the second feature data; generating a diagnosis result corresponding to the to-be-tested device based on the target feature by using a preset diagnosis model, and generating a corresponding diagnosis report based on the diagnosis result according to a target report template; the preset diagnosis model is a model constructed based on a preset expert system and a preset machine learning model; wherein the construction of the pressure test strategy corresponding to the first system and the second system in the to-be-tested device comprises: determining a test requirement of the to-be-tested device, and determining a test scene of the first system and the second system according to the test requirement; determining a pressurization strategy and a pressurization tool corresponding to the first system and the second system; constructing the pressure test strategy corresponding to the first system and the second system based on the test scene, the pressurization strategy and the pressurization tool; in the process of collecting the first test data corresponding to the first system and the second test data corresponding to the second system, further comprising: collecting business performance data of the to-be-tested device, and extracting a business performance feature vector of the business performance data; correspondingly, the generation of the target feature based on the first feature data and the second feature data comprises: determining weight coefficients corresponding to the first feature data, the second feature data and the business performance feature vector respectively, and generating the target feature based on the first feature data, the second feature data and the business performance feature vector based on the weight coefficients; and the generation of the diagnosis result corresponding to the to-be-tested device based on the target feature by using the preset diagnosis model comprises: filtering a to-be-analyzed feature in the target feature based on a preset feature filtering rule by using the preset expert system in the preset diagnosis model; analyzing the to-be-analyzed feature by using the preset machine learning model in the preset diagnosis model to obtain the diagnosis result of the to-be-tested device.

2. The method of pressure testing equipment according to claim 1, wherein, the collection of the first test data corresponding to the first system and the second test data corresponding to the second system comprises: collecting the first test data corresponding to the first system and the second test data corresponding to the second system based on a preset time interval; storing the first test data and the second test data into a preset memory buffer.

3. The method of pressure testing equipment of claim 2, wherein, after storing the first test data and the second test data into the preset memory buffer, further comprising: Preprocessing the first test data and the second test data in the preset memory buffer; Storing the preprocessed first test data and the second test data into a preset database; the preset database is a local database of the device under test or a database in a distributed storage system.

4. The method of claim 1, wherein, The extracting the first feature data of the first test data and the second feature data of the second test data comprises: Extracting the first initial feature of the first test data and the second initial feature of the second test data based on the test requirement of the device under test; Determining the covariance matrix corresponding to the first initial feature and the second initial feature, and determining the corresponding eigenvector and eigenvalue according to the covariance matrix; Based on the eigenvalue, determine the cumulative contribution rate corresponding to each eigenvector, and according to the cumulative contribution rate, determine the first feature data of the first test data and the second feature data of the second test data from the eigenvector.

5. The method of pressure testing an apparatus according to any one of claims 1 to 4, wherein, Before the preset diagnosis model is used to generate the diagnosis result corresponding to the device under test based on the target feature, it further comprises: Based on the pre-training model, an initial diagnosis model is constructed, and the initial diagnosis model is initialized according to the preset model parameter to obtain the preset diagnosis model; the preset model parameter is a random parameter or a parameter obtained based on transfer learning; Setting the reinforcement learning parameter corresponding to the initial diagnosis model; Correspondingly, after the preset diagnosis model is used to generate the diagnosis result corresponding to the device under test based on the target feature, it further comprises: Optimizing the preset diagnosis model by using the reinforcement learning parameter.

6. The method of pressure testing equipment of claim 5, wherein, Before the preset diagnosis model is optimized by using the reinforcement learning parameter, it further comprises: Adjusting the first system and the second system in the device under test according to the diagnosis report, and jumping to the step of collecting the first test data corresponding to the first system and the second test data corresponding to the second system to generate new target features; Taking the new target feature as a diagnosis feature, and constructing a corresponding experience tuple based on the target feature and the diagnosis feature; Store the experience tuple in a preset experience replay buffer.

7. The method of pressure testing equipment of claim 6, wherein, The experience tuple based on the target feature and the diagnosis feature comprises: Determine the test environment when the first system and the second system are subjected to stress testing; Obtain the actual reward signal generated by the test environment based on the diagnosis result, and construct the experience tuple based on the target feature, the diagnosis feature, the actual reward signal and the diagnosis result.

8. The method of pressure testing equipment of claim 7, wherein, The preset diagnosis model is optimized by using the reinforcement learning parameter, comprising: Randomly sampling the experience tuples in the preset experience replay buffer to obtain model training data; Generating a predicted reward signal according to the model training data by using the preset diagnosis model; Determine the loss value between the predicted reward signal and the actual reward signal based on a preset loss function, and adjust the model parameter of the preset diagnosis model based on the loss value and the reinforcement learning parameter to optimize the preset diagnosis model.

9. The method of pressure testing equipment of claim 1, wherein, The target report template is used to generate a corresponding diagnosis report according to the diagnosis result, including: determining the target report template corresponding to the diagnosis from a plurality of initial report templates; generating the diagnosis report according to the diagnosis result based on the target report template, and converting the diagnosis report into a target document format.

10. An electronic device, comprising: comprising: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the device stress test method according to any one of claims 1 to 9.

11. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program is executed by the processor to implement the steps of the device stress test method according to any one of claims 1 to 9.

12. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the device stress test method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • State monitoring and fault diagnosis universal platform based on CAN bus

    CN103487276A

  • Large-scale hydraulic machine remote fault diagnosis method and device based on expert system

    CN110716528A