Interface performance test method and device and electronic equipment
By preprocessing and feature extraction of interface performance test data and evaluating it in combination with machine learning models, the problem of low accuracy in performance evaluation in the existing technology is solved, and more comprehensive performance analysis and higher accuracy are achieved.
Patent Information
- Application Number
- CN202510076840.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, statistical indicators such as mean and standard deviation are used to evaluate the performance of the system, which is low in accuracy and is difficult to analyze the details during the test process.
By collecting the pressure measurement data of the interface, preprocessing is performed to determine the target data, extracting characteristic data such as response time, concurrency number and TPS, scoring based on these data, and entering the trained machine learning model to obtain performance prediction level and confidence.
It improves the accuracy of performance evaluation, can analyze interface performance more comprehensively, identify performance changes and trends under different load conditions, and reduces the workload and subjectivity of manual analysis.
Smart Images

Figure CN120011190A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of financial technology, and in particular to an interface performance testing method, device and electronic equipment. Background Art
[0002] Performance testing is a vital part of the development and maintenance of financial information systems to ensure that the system can meet business needs and regulatory requirements in actual use. Performance testing is a technology that evaluates the performance of a system under a specific workload. The main goal is to determine the response time, stability and reliability of the financial information system under its designed load capacity.
[0003] In the prior art, by analyzing statistical indicators such as average response time, average number of transactions per second (TPS), TPS fluctuation standard deviation, and then comparing the average response time with an acceptable response time threshold, the average TPS with a pre-defined performance baseline, and the TPS fluctuation standard deviation with the standard deviation threshold, the system's responsiveness and stability under specific load conditions are evaluated, and possible performance problems of the system are identified.
[0004] However, using statistical indicators such as mean and standard deviation can only reflect a rough overview of the performance test, and it is difficult to analyze the detailed changes during the test process. Therefore, using statistical indicators to evaluate system performance has low accuracy. Summary of the invention
[0005] The present application provides an interface performance testing method, device and electronic device, which are used to solve the problem of low accuracy in existing evaluation system performance.
[0006] In a first aspect, the present application provides an interface performance testing method, the method comprising:
[0007] Collect the stress test data of the interface, pre-process the stress test data, and determine the target data;
[0008] Extracting multiple feature data from the target data, and scoring based on the target data and the multiple feature data to determine at least one score data; the multiple feature data include first data for characterizing abnormal response time, second data for characterizing abnormal interface concurrency, and test data determined based on the number of transactions per second (TPS); the at least one score data includes scoring data of at least one dimension corresponding to the response time of performing stress testing at different time points;
[0009] Inputting the plurality of feature data and the at least one score data into a trained machine learning model to obtain a performance prediction level and a confidence level;
[0010] A performance test result is determined based on the performance prediction level and the confidence level.
[0011] Optionally, the stress test data is time series data; the first data includes glitch data characterizing that the response time in a first time period is greater than a first threshold and the deviation value corresponding to the response time is greater than a second threshold, third data characterizing that an oscillation interval of the glitch data occurs in a second time period, fourth data characterizing whether the glitch data in the second time period satisfies a periodic law, fifth data characterizing that an abnormal slope occurs in the response time in the second time period, and sixth data characterizing that an abnormality occurs in the response time in the second time period and the abnormal phenomenon lasts for a preset period of time.
[0012] Optionally, the feature extraction process of the second data includes:
[0013] Obtain the first concurrent number of the interface in the current gradient time and the second concurrent number in the previous gradient time, and calculate the average response time and average TPS in the current gradient time;
[0014] Based on the average response time, the average TPS, the first concurrency number, and the second concurrency number, it is determined whether the interface concurrency number is abnormal, and second data is obtained.
[0015] Optionally, scoring the target data and the plurality of feature data to determine at least one score data includes:
[0016] Performing regularization processing on a plurality of target data in a second time period to obtain a plurality of performance data, and calculating an average value and a standard deviation of the plurality of performance data;
[0017] determining a volatility for each performance data based on the mean value and the standard deviation;
[0018] determining a maximum value and a minimum value among a plurality of volatility, and determining an initial score for each performance data based on the minimum value and the maximum value;
[0019] Based on the first data and the second data, the initial score of each performance data is processed to determine at least one score data.
[0020] Optionally, the training method of the machine learning model includes:
[0021] Acquire a data set; the data set includes a training set and a validation set, the data set includes multiple samples, each sample includes multiple feature data, multiple score data and performance level corresponding to the stress test data of the interface;
[0022] Perform multiple rounds of cyclic training on the initial machine learning model using the training set, and determine the loss function value of the initial machine learning model after each round of training based on the validation set;
[0023] If the loss function values in the preset rounds are all less than the preset threshold, the trained machine learning model is determined from the initial machine learning models trained in the preset rounds.
[0024] Optionally, preprocessing the stress test data to determine target data includes:
[0025] Performing numerical anomaly processing on the stress test data to obtain seventh data;
[0026] Performing logic exception processing on the seventh data based on a preset rule to obtain eighth data;
[0027] The eighth data is subjected to data smoothing processing to obtain target data.
[0028] Optionally, the method further includes:
[0029] Determine the effective duration of the seventh data, and generate first prompt information when the effective duration is less than a duration threshold;
[0030] And / or, based on the eighth data, calculating the stress testing success rate in the first time period, and generating second prompt information when the stress testing success rate is less than a third threshold.
[0031] Optionally, determining a performance test result based on the performance prediction level and the confidence level includes:
[0032] When it is determined that the confidence level is greater than a fourth threshold, determining a performance test result based on the performance prediction level;
[0033] When it is determined that the confidence level is less than or equal to a fourth threshold, the target data and the performance prediction level are sent to a user terminal for manual review to determine a performance test result.
[0034] In a second aspect, the present application provides an interface performance testing device, the device comprising:
[0035] A preprocessing module is used to collect the stress test data of the interface, preprocess the stress test data, and determine the target data;
[0036] A feature extraction module is used to extract multiple feature data from the target data, and to score based on the target data and the multiple feature data to determine at least one score data; the multiple feature data include first data for characterizing abnormal response time, second data for characterizing abnormal interface concurrency, and test data determined based on the number of transactions per second (TPS); the at least one score data includes score data of at least one dimension corresponding to the response time of performing stress testing at different time points;
[0037] A prediction module, configured to input the plurality of feature data and the at least one score data into a trained machine learning model to obtain a performance prediction level and a confidence level;
[0038] A determination module is used to determine a performance test result based on the performance prediction level and the confidence level.
[0039] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;
[0040] The memory stores computer-executable instructions;
[0041] The processor executes the computer-executable instructions stored in the memory to implement the method as described in any one of the first aspects.
[0042] In summary, the present application provides an interface performance testing method, device and electronic device. When performing interface performance testing, a large amount of stress test data is collected, and the collected stress test data is preprocessed to clean and organize the data, remove noise and outliers, and then determine the target data for subsequent feature extraction and scoring, so as to improve the accuracy of subsequent processing data. Furthermore, multiple feature data are extracted from the target data, and these feature data include: first data for characterizing abnormal response time, second data for characterizing abnormal interface concurrency, and test data determined based on TPS. Accordingly, based on these feature data and target data, the interface performance is scored to obtain at least one score data, and the score data characterizes the abnormality of execution at different time points. Scoring data on response time during stress testing. This step can select representative feature data to better capture key indicators of interface performance. Further, the feature data and scoring data are input into a trained machine learning model, which outputs performance prediction levels and confidence levels. The performance prediction levels indicate the quality of interface performance, while the confidence levels indicate the reliability of the prediction results. Based on the performance prediction levels and confidence levels, the performance test results of the interface can be determined. In this way, the present application combines multiple performance indicators for comprehensive analysis and can perform a more comprehensive performance evaluation. This combination of multi-dimensional feature data and scoring data can help identify performance changes and trends under different load conditions, greatly improving the accuracy of performance evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0044] Figure 1 A schematic diagram of an application scenario provided for an embodiment of the present application;
[0045] Figure 2 A flowchart of an interface performance testing method provided in an embodiment of the present application;
[0046] Figure 3 A schematic diagram of a feature extraction method provided in an embodiment of the present application;
[0047] Figure 4 A schematic diagram of the structure of a machine learning model provided in an embodiment of the present application;
[0048] Figure 5 A statistical diagram of burr detection data provided by an embodiment of the present application;
[0049] Figure 6 Another statistical diagram of burr detection data provided by an embodiment of the present application;
[0050] Figure 7 A statistical schematic diagram of a detection oscillation interval provided in an embodiment of the present application;
[0051] Figure 8 A statistical schematic diagram of detecting periodic burr data provided by an embodiment of the present application;
[0052] Fig. 9 A statistical schematic diagram of a detection response time slowing sequence provided in an embodiment of the present application;
[0053] Fig.10 A statistical diagram of detecting abnormal fluctuation data provided by an embodiment of the present application;
[0054] Fig.11 A statistical diagram of VU within a detection gradient time provided in an embodiment of the present application;
[0055] Fig.12 A statistical schematic diagram of TPS within the detection gradient time provided in an embodiment of the present application;
[0056] Fig.13 An overall flow chart of an interface performance testing method provided in an embodiment of the present application;
[0057] Fig.14 A schematic diagram of the structure of an interface performance testing device provided in an embodiment of the present application;
[0058] Fig.15 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0059] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0060] In order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first", "second" and the like are used to distinguish the same items or similar items with substantially the same functions and effects. For example, the first device and the second device are only used to distinguish different devices, and their order is not limited. Those skilled in the art can understand that the words "first", "second" and the like do not limit the quantity and execution order, and the words "first", "second" and the like do not necessarily limit them to be different.
[0061] It should be noted that, in this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0062] In the present application, "at least one" means one or more, and "plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0063] In the technical solution of this application, the collection, storage, use, processing, transmission, provision and disclosure of financial data or user data and other information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals. It should be noted that in the embodiments of this application, some existing solutions in the industry such as certain software, components, models, etc. may be mentioned, and they should be considered as exemplary. Their purpose is only to illustrate the feasibility of the implementation of the technical solution of this application, but it does not mean that the applicant has or will necessarily use the solution.
[0064] The professional terms involved in this application are explained below.
[0065] Stress testing: Stress testing refers to a testing method used to evaluate the performance of a system under high load conditions, such as stress testing the performance of a financial information system, which tests the stability of the financial information system interface through high concurrent requests.
[0066] SENSE (Syncretic Empiric Numeral Stress Evaluation) refers to a stress testing digital risk perception model, which is a performance testing method for financial information systems defined in this application. It aims to discover problems in the performance data of financial information systems through statistical and deep learning methods, and automatically classify performance levels, so as to conduct digital and labeled performance stability evaluation of financial business systems.
[0067] Transaction Per Second (TPS): refers to an indicator used to measure system performance. In financial information systems, it specifically refers to the number of business transactions per second.
[0068] Virtual User (VU): refers to the concept used to simulate real user behavior in stress testing and load testing. In this application, it refers to the number of concurrent users simulated during stress testing.
[0069] Interface: refers to the way one system, device, program or component interacts with another system, device, program or component, such as the software interface for financial information systems to interact with other software systems, including the application programming interface (API) of system services such as application software, databases, and middleware.
[0070] Use case: refers to the test of an interface. Mixed scenario execution use case refers to the simultaneous testing of multiple interfaces related to a business.
[0071] Performance testing is an essential process for developing and maintaining financial information systems. Its necessity is mainly reflected in the following aspects: (1) Ensuring system stability: Financial business is highly dependent on financial information systems. Any system failure may result in significant economic losses. Therefore, performance testing is an important means to ensure that the system can operate stably under various workloads. (2) Ensuring business continuity: Financial business has very high requirements for system reliability and availability. Therefore, through performance testing, the stability and reliability of the system under various workloads can be evaluated to ensure business continuity. (3) Complying with regulatory requirements: For financial institutions, the performance and security of financial information systems are also very important regulatory requirements. Therefore, through performance testing, financial institutions can demonstrate their rigor and compliance in the management of financial information systems, thereby demonstrating to regulators their high attention to system stability and security and ensuring compliance with regulatory requirements.
[0072] In one possible implementation, analyzing the response time curve is a key step in performing performance testing. Exemplarily, performance anomalies can be discovered based on manual analysis of the response time curve. The manual analysis process includes: observing the overall trend of the response time curve over a period of time, and analyzing peaks and troughs. Further, by comparing with set thresholds, performance anomalies can be effectively analyzed and identified from the response time curve, thereby generating alarms, etc.
[0073] It should be noted that the above method requires users to comprehensively apply technical knowledge and analytical experience to ensure accurate identification and resolution of potential performance issues. Although the above method provides a basis for further system optimization, manual analysis usually requires a lot of time and effort, and may be affected by personal experience and bias. Different analysts may have different interpretations and conclusions, resulting in uncertainty in the analysis.
[0074] Moreover, the manual analysis process may be difficult to reproduce, especially without detailed records of the analysis steps and decision logic, which makes subsequent problem tracking and verification difficult. Therefore, in a dynamically changing technical environment, manual updates of analysis strategies and methods may be slow and difficult to adapt to new technical requirements and business needs in a timely manner.
[0075] In another possible implementation, performance testing is performed through statistical indicators in the performance testing process. This performance testing is a process of evaluating the responsiveness and stability of the system under specific load conditions. For example, by analyzing statistical indicators such as average response time, TPS, and TPS fluctuation standard deviation, the average response time is compared with an acceptable response time threshold, the average TPS is compared with a pre-defined performance baseline, and the TPS fluctuation standard deviation is compared with the standard deviation threshold, so as to evaluate the responsiveness and stability of the system under specific load conditions and determine possible performance problems of the system.
[0076] Among them, the average response time is lower than the acceptable response time threshold, the average TPS is lower than the performance baseline, and the TPS fluctuation standard deviation is higher than the standard deviation threshold, indicating that the system may have performance problems. In particular, a high standard deviation can indicate the instability of the system when processing loads. The response time threshold is usually set based on user expectations or business needs and is not specifically limited here.
[0077] However, it is a common practice to use statistical indicators such as average response time, average TPS, TPS fluctuation standard deviation, etc. to discover performance anomalies in performance testing. However, this practice also has some disadvantages and limitations. That is, the use of statistical indicators such as average and standard deviation can only reflect the rough overview of the performance test, and it is difficult to analyze the detailed changes during the test. For example, the slow response time caused by periodic Full GC (Full Garbage Collection) garbage collection in Java applications is difficult to quantify through statistical indicators. Therefore, the use of statistical indicators to evaluate system performance has low accuracy.
[0078] In response to the above problems, the present application provides an interface performance testing method, which aims to discover problems in interface performance data through statistics and deep learning methods. Specifically, when performing interface performance testing, a large amount of stress test data is collected, and the collected stress test data is preprocessed to clean and organize the data, remove noise and outliers, and then determine the target data for subsequent feature extraction and scoring, so as to improve the accuracy of subsequent processing data. Furthermore, multiple feature data are extracted from the target data, and these feature data include: first data for characterizing abnormal response time, second data for characterizing abnormal interface concurrency, and test data determined based on TPS. Accordingly, based on these feature data and target data, the interface performance is scored to obtain at least one score data. The score data represents the score data of the response time when performing stress testing at different time points. This step can select representative feature data to better capture the key indicators of interface performance. Furthermore, the feature data and the score data are input into a trained machine learning model, which outputs a performance prediction level and confidence. The performance prediction level indicates the quality of the interface performance, and the confidence indicates the reliability of the prediction result. Based on the performance prediction level and confidence, the performance test result of the interface can be determined. In this way, the present application combines multiple performance indicators for comprehensive analysis and can perform a more comprehensive performance evaluation. This method of combining multi-dimensional feature data and scoring data can help identify performance changes and trends under different load conditions, greatly improving the accuracy of performance evaluation.
[0079] For example, Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application, such as Figure 1 As shown, the application scenario can be applied to the performance test of the financial information system, and the application scenario includes a financial information system 101, a database 102, a data processing system 103 and a user's terminal device 104; wherein the financial information system 101 corresponds to interface 1, the database 102 corresponds to interface 2, and the data processing system 103 can collect stress test data corresponding to interface 1 and / or interface 2 to evaluate the stability of different interfaces.
[0080] Exemplarily, taking the collection of stress test data of interface 1 as an example, the financial information system 101 is performance tested. The data processing system 103 can pre-process the collected stress test data of interface 1 to obtain target data, further extract multiple feature data in the target data, and score based on the target data and the multiple feature data to determine at least one score data, and input the multiple feature data and at least one score data into a trained machine learning model to obtain a performance prediction level and confidence. Different performance prediction levels correspond to different degrees of good or bad interface performance. Further, based on the performance prediction level and confidence, the performance test result of interface 1 is determined.
[0081] Optionally, the performance prediction level and confidence level may be sent to the user's terminal device 104 for visual display, so that the user can manually review the performance prediction level and confidence level to judge the accuracy of the performance test result of the interface 1 determined by the data processing system 103 .
[0082] Optionally, when the confidence level is lower than a threshold, the target data and the performance prediction level may be sent to the user's terminal device 104 for manual review to manually determine the performance test result.
[0083] Optionally, when the confidence level is higher than the threshold, only the performance prediction level is sent to the user's terminal device 104 for visual display, so that the user can view the performance test results of interface 1.
[0084] It should be noted that the performance test process of interface 2 is similar to the performance test process of interface 1. For details, please refer to the process description of interface 1, which will not be repeated here. The embodiment of the present application does not specifically limit the number of interfaces that can evaluate interface performance.
[0085] It is understandable that the terminal device 104 can also be a display device corresponding to the data processing system 103. The embodiment of the present application does not specifically limit the device for visual display. Optionally, the terminal device can also be referred to as user equipment (UE), mobile station (MS), mobile terminal, terminal, etc. In practical applications, terminal devices are, for example, desktop computers, notebooks, personal digital assistants (PDA), smart phones, tablet computers, vehicle-mounted devices, wearable devices (such as smart watches and smart bracelets), smart home devices (such as smart display devices), etc.
[0086] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0087] Figure 2 A flowchart of an interface performance testing method provided in an embodiment of the present application is provided. The interface performance testing method can be applied to the above-mentioned data processing system, such as Figure 2 As shown, the interface performance testing method includes the following steps:
[0088] S201. Collect stress test data of an interface, pre-process the stress test data, and determine target data.
[0089] In the embodiments of the present application, stress testing data may refer to interface data collected through open source tools or self-developed tools, including response time per second, number of concurrent users, TPS, number of successful transactions, number of failed transactions, success rate, resource usage, etc. The embodiments of the present application do not specifically limit the content of the stress testing data. The stress testing data can cover different load conditions and usage scenarios, and can comprehensively evaluate the performance of the interface. For example, the stress testing data can be the response time, TPS and success rate of the financial information system interface per second.
[0090] Optionally, the stress testing data is time series data, which may be original data collected during test execution using a distributed stress testing platform deployed in a large financial institution. The original data uses timestamp as a primary key to count interface performance data every second.
[0091] Among them, time series data refers to performance indicator data collected at different time points, which can reflect the changes in interface performance over time.
[0092] Exemplarily, preprocessing of the collected stress testing data may include data cleaning and data smoothing, so as to determine the target data required for subsequent feature extraction, model use or model training. The data cleaning is used to remove duplicate or abnormal data or fill in incomplete data, and the data smoothing is used to reduce noise and fluctuations in the stress testing data, so as to better reveal the trend of the stress testing data.
[0093] Among them, the target data can be 3×T single-precision floating-point matrix data, and 3 represents the average response time, number of successful transactions, and number of failed transactions recorded per second. It can be understood that since the recording time granularity is 1 second, the successful transaction value here can also be regarded as TPS.
[0094] It should be noted that the embodiments of the present application do not specifically limit the preprocessing process, and the above is only an example.
[0095] S202. Extract multiple feature data from the target data, and score the target data and the multiple feature data to determine at least one score data; the multiple feature data include first data for characterizing an abnormal response time, second data for characterizing an abnormal interface concurrency, and test data determined based on the number of transactions per second (TPS); the at least one score data includes score data of at least one dimension corresponding to the response time of performing stress testing at different time points.
[0096] In an embodiment of the present application, the first data showing abnormal response time is used to help identify potential performance issues, and may include data of multiple dimensions that affect the stability and performance of the system. The second data showing abnormal interface concurrency is used to help identify performance bottlenecks. Test data based on TPS can be used to analyze processing power and efficiency under different loads. The embodiment of the present application does not limit the specific data corresponding to the first data, the second data, and the test data.
[0097] In this step, based on the extracted feature data and combined with the target data, the interface performance is scored and at least one score data is calculated. The scoring method can be a simple threshold judgment, a complex weighted scoring, or a predefined scoring mechanism. The embodiment of the present application does not specifically limit the scoring method.
[0098] It should be noted that the feature data affects the score of the target data. For example, a scoring mechanism may deduct ten points for each feature data hit by the target data during feature extraction. The embodiment of the present application does not specifically limit the formulation of the scoring mechanism.
[0099] It can be understood that the scoring data may include multiple dimensions, and the score of each dimension reflects the performance of the interface in that aspect. The embodiment of the present application does not specifically limit the number of dimensions corresponding to the scoring data, which can be determined based on the actual application scenario requirements.
[0100] For example, Figure 3 A schematic diagram of a feature extraction process provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, the collected stress test data is stored in the InfluxDB time series database, the stress test data to be analyzed is obtained from the InfluxDB time series database, the stress test data is preprocessed to obtain the target data, and further, multiple feature data are extracted from the target data, and the multiple feature data are stored in the feature data table, and then comprehensive scoring is performed based on the target data and the multiple feature data to obtain multiple score data, and the multiple score data are stored in the score data table, wherein the multiple feature data and the multiple score data are all single-precision floating point data, and the multiple feature data will be persisted in the feature data table of MySQL. Accordingly, the data in the table can be passed to the comprehensive scoring step for comprehensive scoring.
[0101] Optionally, after obtaining the stress testing data, multiple feature data may be directly extracted from the stress testing data, but the accuracy of the extracted multiple feature data is lower than that of the multiple feature data extracted from the target data.
[0102] S203: Input the multiple feature data and the at least one score data into a trained machine learning model to obtain a performance prediction level and confidence.
[0103] In an embodiment of the present application, the machine learning model is a model for performance prediction, which may be a regression model, a decision tree, a random forest, a support vector machine or a neural network model, etc. The embodiment of the present application does not limit the specific model corresponding to the machine learning model. Optionally, the machine learning model is a neural network model based on deep learning.
[0104] For example, Figure 4 A schematic diagram of the structure of a machine learning model provided in an embodiment of the present application is shown in FIG. Figure 4 As shown in the figure, taking the neural network model prediction performance level as an example, the neural network model has multiple one-dimensional convolution blocks (Conv1dBlock), which include: Conv1d layer (convolution layer), used to perform one-dimensional convolution operation; BatchNormalization (batch normalization), used to accelerate training and improve model stability; ReLU (Rectified LinearUnit) activation function, used to introduce nonlinearity; pooling layer, such as maximum pooling, used for dimensionality reduction and feature selection.
[0105] In the process of predicting the performance level of the neural network model, local features are extracted from the input data through Conv1dBlock, and then the dimension of the output data is reduced through convolution, batch normalization and pooling operations, thereby compressing information. For example, data with input dimension N is compressed into data with output dimension M, where N is greater than M. Nonlinearity is introduced through the ReLU activation function, so that the neural network model can learn complex nonlinear relationships. Then, by stacking multiple Conv1dBlocks, the neural network model can learn higher-level features layer by layer, thereby improving the expressiveness of the model.
[0106] Among them, before making predictions, the machine learning model has been trained using historical data sets and has learned the relationship between features and performance. Therefore, the trained machine learning model can be directly used here for performance prediction.
[0107] Optionally, the machine learning model can be updated and expanded as the amount of data increases and the features change to adapt to new performance evaluation requirements.
[0108] It should be noted that the performance prediction level output by the model is a classification result, which is used to represent different levels of interface performance, such as high risk level, medium risk level, and low risk level. Different levels can characterize the degree of quality of interface performance. For example, a high risk level indicates poor interface performance, and a low risk level indicates good interface performance. The confidence level of the model output refers to the degree of credibility or certainty of the model's prediction results. The higher the confidence level, the more accurate the model's prediction results. Optionally, low-confidence prediction results can be submitted for manual review.
[0109] S204: Determine a performance test result based on the performance prediction level and the confidence level.
[0110] In the embodiments of the present application, when determining the performance test results, not only the performance prediction level should be considered, but also the confidence level, because high-confidence prediction results are usually more valuable for reference. Optionally, a confidence threshold can be set, and prediction results exceeding the confidence threshold are considered to be reliable, thereby improving the credibility of the prediction results. Prediction results below the confidence threshold can be manually reviewed, and the embodiments of the present application do not make specific limitations on this.
[0111] Exemplarily, based on the performance prediction level and confidence, the interface performance is classified into different result categories, and the test results are organized into a report and provided to the relevant team so that corresponding measures can be taken to optimize the system performance.
[0112] Therefore, in the present application, noise and outliers are removed through preprocessing, the amount of unnecessary data is reduced, the accuracy and reliability of the data, and the efficiency of data processing are improved, which provides a solid foundation for subsequent analysis. Furthermore, by extracting multiple feature data, the interface performance can be analyzed from multiple angles, including response time, concurrency, and TPS, etc., and the present application also combines a scoring mechanism. The multi-dimensional scoring data provides a comprehensive perspective on interface performance, helping to identify performance changes and trends under different load conditions. In this way, through feature extraction and scoring, key indicators affecting interface performance can be identified. Furthermore, using a trained machine learning model, the interface performance can be automatically evaluated, reducing the workload and subjectivity of manual analysis. The machine learning model can perform multi-dimensional performance analysis by combining multiple feature data, capturing complex linear relationships, thereby improving the accuracy of performance prediction. In addition, the present application provides a quantitative basis by combining performance prediction level and confidence to determine the performance test results, thereby providing more accurate and reliable performance evaluation results and reducing the risks caused by prediction uncertainty.
[0113] Optionally, the stress test data is time series data; the first data includes glitch data characterizing that the response time in a first time period is greater than a first threshold and the deviation value corresponding to the response time is greater than a second threshold, third data characterizing that an oscillation interval of the glitch data occurs in a second time period, fourth data characterizing whether the glitch data in the second time period satisfies a periodic law, fifth data characterizing that an abnormal slope occurs in the response time in the second time period, and sixth data characterizing that an abnormality occurs in the response time in the second time period and the abnormal phenomenon lasts for a preset period of time.
[0114] Among them, the deviation value can be calculated in a variety of ways, such as by calculating the average deviation, standard deviation, relative deviation, etc. The embodiment of the present application does not specifically limit this. If the response time is greater than the first threshold within the first time period, and the deviation value corresponding to the response time is greater than the second threshold, it means that a glitch occurs.
[0115] It can be understood that the first time period can be understood as within a granularity of 1 second, or it can represent a unit of time. The embodiment of the present application does not specifically limit the definition of the first time period, and the second time period is used to indicate the time corresponding to the collection of stress testing data. For example, within the second time period, the collected stress testing data can be a stress testing response time series, that is, a sequence composed of response times collected at different time points.
[0116] Among them, the glitch in the stress test refers to the abnormal point where the response time curve of 1 second granularity has several abnormal increases, and the increase is significantly higher than the overall fluctuation of the curve. The abnormal point corresponds to the glitch data. Therefore, the glitch data can indicate a short-term performance abnormality. For example, Figure 5 A statistical diagram of burr detection data provided in an embodiment of the present application, wherein the content in the dotted box in the figure is the detected burr.
[0117] Optionally, since glitches generally appear at points where the amplitude of change far exceeds normal fluctuations, a method for detecting glitches may be to identify abnormal points using a second-order difference and an algorithm based on Chebyshev's inequality.
[0118] For example, we look for outliers on the second-order difference sequence. The outliers on the second-order difference sequence are defined as points that deviate too much from the average value of the second-order difference sequence. If the fluctuation of the entire second-order difference sequence is regarded as a normal distribution, we can apply Chebyshev's inequality: Find outliers, where μ is the mean, σ is the standard deviation, and X is the response time corresponding to different time points.
[0119] When k is a multiple of the standard deviation, σ in the above formula can be eliminated, and the probability of a point exceeding the mean value by k times the standard deviation is 1 / 2k 2 Therefore, k can be set to 3.4, that is, the probability of a point falling in this interval is about 4.33%, and it can be considered that the deviation on the second-order difference sequence does not exceed 4.33%.
[0120] It is understandable that the value of k can be set based on the actual application scenario requirements. The embodiment of the present application does not specifically limit the value of k, and the above is only an example.
[0121] It should be noted that Figure 6 Another statistical diagram of detecting burr data provided by the embodiment of the present application is as follows: Figure 6As shown in Figure 2, if only the second-order difference algorithm is used to find abnormal points, missed detection may occur. Figure 6 The corresponding average response time is about 17ms, while the maximum glitch response time is about 1000ms, which greatly increases the standard deviation of the second-order difference sequence, resulting in a decrease in the accuracy of finding abnormal points under the restriction condition of "exceeding k times the standard deviation of the average value". As shown in the figure, only the glitches indicated in the dotted box are detected, while other glitches in the figure, even if they have exceeded the average response time by several times, are still missed due to the high standard deviation of the second-order difference sequence.
[0122] Therefore, a second detection condition is introduced in the burr position detection, that is, the point with a response time higher than a times the average response time is a burr, and a is 3. The second detection condition complements the detection condition of Chebyshev's inequality, and the combination of the two is the burr position detection result.
[0123] Among them, the value of a can be set based on the actual application scenario requirements. The embodiment of the present application does not specifically limit the value of a. The above is only an example.
[0124] In the present application, the oscillation interval refers to an interval where multiple glitches are gathered. Compared with the generally distributed glitches, the appearance of an oscillation interval in the stress test indicates that the performance test results are more unstable. The logic for determining the oscillation interval is: the interval between two adjacent glitches does not exceed b seconds, and the total length of the oscillation interval d is at least c seconds, where b can be 6 and c can be 12. The embodiments of the present application do not specifically limit the values of b and c, which can be set based on the actual application scenario requirements.
[0125] For example, Figure 7 A statistical diagram of a detection oscillation interval provided in an embodiment of the present application, such as Figure 7 As shown in the figure, if the burrs are clustered in the dotted box, it is considered to be an oscillation interval, while for a single burr, it can be determined as a non-oscillation interval.
[0126] In the present application, the periodic law refers to the phenomenon that glitches appear at fixed intervals in the stress test response time series, and the logic for determining whether the glitch data satisfies the periodic law is: (1) obtaining the point positions corresponding to all glitches detected and outputted in the second time period; (2) taking the first point and the second point of the glitch from front to back in chronological order to determine the initial period, and the initial period is greater than e seconds and less than f seconds; optionally, e can be 10 and f can be 330. The embodiments of the present application do not specifically limit the values of e and f, which can be set based on the actual application scenario requirements.
[0127] (3) In chronological order, find out whether there is a reasonable next burr point. Reasonable means that the distance between the next burr point and the previous burr point is g% of the initial period, and the number of burrs between the two burr points is less than the fifth threshold. If there is a reasonable next burr point, update the initial period based on the distance between the first burr point and the reasonable next burr point found; repeat step (3) until there is no reasonable next burr point, then determine the period with periodic burrs.
[0128] Among them, g can be 10, and the fifth threshold can be 4. The embodiment of the present application does not specifically limit the values of g and the fifth threshold, which can be set based on the actual application scenario requirements.
[0129] For example, Figure 8 A statistical diagram of detecting periodic burr data provided by an embodiment of the present application, such as Figure 8 As shown in the figure, the glitch phenomenon in the dotted box has a periodic pattern. It is understandable that the periodic glitch may be caused by system scheduled tasks, JVM (Java virtual machine) periodic GC (Garbage Collection), complex requests in the test case loop, database or cache timing operations, etc. Therefore, the periodic glitch usually appears stably, so the periodic glitch is also one of the factors affecting performance that need to be considered, and in the subsequent scoring process, if the target data hits the periodic glitch, the target data will also be additionally deducted.
[0130] In the present application, the response time shows an abnormal slope in the second time period, indicating a trend of slower response. Slower response means that the response time curve with a granularity of 1 second tends to become slower and slower. Slower response is usually caused by various reasons such as increased application load pressure. The slowing phenomenon may cause the stress test time to become longer and the response performance of the application interface to become worse. Therefore, during the scoring stage, additional points will be deducted for data with abnormal slope.
[0131] For example, Fig. 9 A statistical diagram of a detection response time slowing sequence provided in an embodiment of the present application, such as Fig. 9 As shown, the response time in area A is about 1100ms, and the response time in area B is about 1300ms, showing a relatively obvious slowing trend. Therefore, areas A and B in the figure correspond to response time slowing sequences respectively, and the slope of the response time corresponding to the burr point in the response time slowing sequence is abnormal.
[0132] Optionally, the logic of detecting a slow response is: use the LinearRegression function in the open source machine learning library sklearn to fit the overall response time series to obtain the linear function coefficient coef; determine whether to perform a slow detection based on the coefficient coef. If coef is greater than a sixth threshold, it is considered that there is a certain fitting effect and the slow detection can be performed, otherwise the slow detection is not performed; optionally, the sixth threshold can be 0.31. The embodiment of the present application does not specifically limit the value of the sixth threshold, which can be set based on the actual application scenario requirements.
[0133] If slowing detection is performed, coef is multiplied by the length of the response time series. If the result is greater than 2 times the standard deviation of the sequence, it is considered to be a slowing response sequence.
[0134] It should be noted that the slowdown detection is performed only when coef is greater than the sixth threshold, which means that the slowdown detection is only performed on use cases with relatively consistent response time changes. When coef is greater than 0, the response time series has slowly increased, but the random fluctuations in the stress testing process still need to be considered. Therefore, only when the overall increase value exceeds 2 times the sequence standard deviation, it means that the speed of the response slowdown is already obvious and does not fall into the category of random fluctuations.
[0135] In the present application, if the response time is abnormal within the second time period and the abnormal phenomenon is maintained for a preset period of time, it means that an abnormal fluctuation occurs in the stress test response time series, that is, the response time significantly rises to a high level or drops to a low level and maintains the preset time. The embodiment of the present application does not specifically limit the value of the preset time, and it can be set based on the actual application scenario requirements.
[0136] For example, Fig.10 A statistical diagram of detecting abnormal fluctuation data provided by an embodiment of the present application, such as Fig.10 As shown in the figure, a test case lasts for about 10 seconds in area A and about 10 seconds in area B, and there is obvious up and down fluctuation in the response time, which is considered to be an abnormal fluctuation.
[0137] It should be noted that the causes of abnormal fluctuations may be similar to those of glitches, such as test cases with uneven complexity, load increase caused by batch logic, etc., and their impact range is often larger than that of general glitches. Therefore, abnormal fluctuations are also one of the factors affecting performance that need to be considered. In the subsequent scoring process, if the target data has abnormal fluctuations, additional points will be deducted from the target data.
[0138] Optionally, the logic for identifying abnormal fluctuations is as follows: (i) using the same principle as burr detection to detect abnormal points in the response time series, abnormal points that are more than three standard deviations larger or less than three standard deviations smaller than the average value of the second-order difference sequence will be used as fluctuation pre-selected points; (ii) calculating the mean and standard deviation of each h-second interval to the left and right of the fluctuation pre-selected point; (iii) judging the size of the mean and standard deviation of the left and right intervals, when the difference between the mean values of the left and right intervals is greater than p% of the second-order difference values of the fluctuation points, and the standard deviations of the left and right intervals are both less than q% of the second-order difference values of the fluctuation points, then the fluctuation pre-selected point is identified as an abnormal fluctuation point, thereby determining that the response time series has abnormal fluctuations.
[0139] Among them, h can be 2 to 6, p can be 90, and q can be 10. The embodiment of the present application does not specifically limit the values of h, p, and q, which can be set based on the actual application scenario requirements.
[0140] It should be noted that in the above step (2), the interval of 2 to 6 seconds is selected instead of 0 to 4 seconds in order to avoid the rising and falling stages of fluctuations and select the stable interval as much as possible. The 2-second interval can skip the 2-second offset caused by the second-order difference calculation. In the above step (3), since the rise or fall of the expected fluctuation can be maintained in the stable stage, this is reflected in the sequence curve chart, that is, the significant high and low areas, that is, Fig.10 In area A and area B.
[0141] Optionally, multiple feature data include feature extraction results and performance test TPS. The first data may include the total number of glitches, the total number of oscillation intervals, whether they are periodic glitches, the response slowdown slope and the number of abnormal fluctuations. The second data may be whether there is a gradient bottleneck. The test data may include the minimum TPS of use case execution, the minimum TPS of scenario sub-use case and the total TPS of scenario. The scoring data may include the minimum score of use case execution, the average score of use case execution, the minimum score of scenario sub-use case and the average score of scenario execution, etc. The embodiments of the present application do not specifically limit the content of the feature data and scoring data.
[0142] Exemplarily, since a stress test may include multiple interface tests or execute multiple sub-cases in mixed scenarios, the total number of glitches, the total number of oscillation intervals, and the number of abnormal fluctuations in this application are taken as the cumulative number after executing multiple tests. The feature items corresponding to whether there are periodic glitches and whether there are gradient bottlenecks are taken as the number or set of corresponding problems in each execution, and the feature item corresponding to the response slowing slope is taken as the maximum slope in each execution.
[0143] Optionally, the multiple feature data and at least one score data input into the trained machine learning model may be 13 one-dimensional features, and the 13 features are shown in the following Table 1:
[0144] Table 1
[0145]
[0146]
[0147] In the present application, through glitch detection and oscillation interval analysis, the causes of performance abnormalities can be effectively identified and located. By detecting periodic patterns and slope anomalies, the potential trends and patterns of performance problems can be identified. By identifying and analyzing the duration of anomalies, the stability and reliability of the interface can be analyzed. Therefore, by considering feature data of multiple dimensions such as glitch data, third data, fourth data, fifth data and sixth data, the performance can be accurately analyzed to help identify short-term and persistent performance problems.
[0148] Optionally, the feature extraction process of the second data includes:
[0149] Obtain the first concurrent number of the interface in the current gradient time and the second concurrent number in the previous gradient time, and calculate the average response time and average TPS in the current gradient time;
[0150] Based on the average response time, the average TPS, the first concurrency number, and the second concurrency number, it is determined whether the interface concurrency number is abnormal, and second data is obtained.
[0151] In the embodiment of the present application, the gradient time refers to a fixed time window. The embodiment of the present application does not specifically limit the size of the gradient time. The second data can be used to indicate whether there is a gradient bottleneck, that is, whether the number of concurrent interfaces is abnormal.
[0152] In this step, by comparing the first concurrency number and the second concurrency number, the changing trend of the concurrency number can be analyzed, and combined with the average response time and the average TPS, it is evaluated whether the change in the concurrency number causes performance abnormalities. For example, if the increase in the concurrency number causes a significant increase in the response time or a decrease in the TPS, there may be a performance bottleneck (concurrency bottleneck or gradient bottleneck). Based on the above analysis, it is determined whether the concurrency number is abnormal, and the result is used as the second data.
[0153] Among them, the concurrency bottleneck refers to the phenomenon that the number of concurrency increases but the TPS does not increase proportionally with the number of concurrency and the response time increases significantly.
[0154] Optionally, you can analyze the number of concurrent bottlenecks in the gradient stress test based on TPS and response time. Specifically, you can determine the concurrent bottleneck based on the following formula:
[0155] time_mean>pre_time_mean*(1+(vu / pre_vu-1)*GRAD_TIME_PATIENT);
[0156] tps_mean / vu <pre_tps_mea / pre_vu*(1-(1-pre_vu / vu)*GRAD_TPS_PATIENT);
[0157] Among them, the second concurrency number of the previous gradient time is pre_vu, the average response time is pre_time_mean, and the average TPS is pre_tps_mean, while the first concurrency number of the current gradient time is vu, the average response time is time_mean, and the average TPS is tps_mean. GRAD_TIME_PATIENT is the allowable error of the response time, and the GRAD_TIME_PATIENT can be 0.4. The embodiment of the present application does not specifically limit the value of GRAD_TIME_PATIENT, which can be set based on the actual application scenario requirements.
[0158] For example, Fig.11 A statistical diagram of VU within the detection gradient time provided by an embodiment of the present application, such as Fig.11 As shown in Figure 2, the number of concurrent connections increases with the gradient time, and the change of TPS with the gradient time is as follows: Fig.12 As shown, Fig.12 A statistical diagram of TPS within the detection gradient time provided by the embodiment of the present application, such as Fig.12 As shown in the figure, TPS encounters a bottleneck when vu increases from 400 to 500. Then, by calculating the concurrency efficiency when vu=400, we can infer that the specific bottleneck point is vu=440.
[0159] In this way, by analyzing the number of concurrency and its impact on response time and TPS, the present application can dynamically monitor the performance status, and then quickly identify performance anomalies caused by changes in the number of concurrency. By detecting anomalies in the number of concurrency, the load limit of the interface can be identified, and more reliable performance predictions can be made.
[0160] Optionally, scoring the target data and the plurality of feature data to determine at least one score data includes:
[0161] Performing regularization processing on a plurality of target data in a second time period to obtain a plurality of performance data, and calculating an average value and a standard deviation of the plurality of performance data;
[0162] determining a volatility for each performance data based on the mean value and the standard deviation;
[0163] determining a maximum value and a minimum value among a plurality of volatility, and determining an initial score for each performance data based on the minimum value and the maximum value;
[0164] Based on the first data and the second data, the initial score of each performance data is processed to determine at least one score data.
[0165] In this application, when scoring based on target data and multiple feature data, the volatility corresponding to multiple target data in the second time period is calculated. The target data can respond to the time, and the volatility P of all current response times based on the 1-second interval granularity can be calculated. all List, the P all The list includes multiple response times P, and P is regularized to obtain performance data.
[0166] Optional, based on formula P new =(PA all ) / S all Calculate each performance data P new The volatility of P all The average value is A all , standard deviation is S all .
[0167] Since the score data is between 0 and 100, we can use the normalization method to calculate the lower limit P of the volatility after all regularization processing. low and upper limit P high , and then use the scoring formula Calculate the initial score for each performance data. low is the minimum value among multiple volatility, the P high is the maximum value among multiple volatility.
[0168] Furthermore, the initial score is processed in combination with the scoring mechanism to determine at least one score data. The scoring mechanism is to determine whether the performance data hits the first data and the second data. If so, a subtraction process is performed on the initial data to determine the score data of the performance data. The subtraction process is to subtract a preset score. The embodiment of the present application does not specifically limit the size of the preset score. A hit means that the performance data has the characteristics of the first data and the second data or belongs to one of the first data and the second data.
[0169] Optionally, for different use cases, after calculating the score of each performance based on the scoring mechanism, the minimum score, average score, etc. of the use case can be determined, and then at least one score data can be determined. The embodiment of the present application does not specifically limit the content corresponding to the score data.
[0170] In this way, through regularization and volatility analysis, the stability and consistency of performance data can be evaluated more accurately, and the combination of volatility and scoring helps to identify performance anomalies, providing a basis for subsequent model prediction performance. In addition, through the first data and the second data, the score of the performance data can be continuously adjusted to improve the accuracy and reliability of the performance evaluation to adapt to new needs and scenario changes.
[0171] Optionally, the training method of the machine learning model includes:
[0172] Acquire a data set; the data set includes a training set and a validation set, the data set includes multiple samples, each sample includes multiple feature data, multiple score data and performance level corresponding to the stress test data of the interface;
[0173] Perform multiple rounds of cyclic training on the initial machine learning model using the training set, and determine the loss function value of the initial machine learning model after each round of training based on the validation set;
[0174] If the loss function values in the preset rounds are all less than the preset threshold, the trained machine learning model is determined from the initial machine learning models trained in the preset rounds.
[0175] In the present application, after obtaining the data set, the data set can be split into a training set and a validation set, which are used for model training and validation, respectively. Optionally, a test set can also be split. The embodiment of the present application does not specifically limit the splitting method, such as it can be split in random proportions.
[0176] In this step, the initial machine learning model is trained for multiple rounds using the training set. After each round of training, the model parameters are updated to minimize the loss function, which is used to measure the difference between the model prediction value and the true value. The loss function may include mean square error (MSE), cross entropy loss, etc. The embodiment of the present application does not specifically limit the type of loss function.
[0177] Among them, after each round of training, the validation set is used to evaluate the loss function value of the model to detect the generalization ability of the model. If the loss function value is less than the preset threshold within the preset rounds, the trained model is selected from these rounds as the trained machine learning model. The preset threshold is used as the standard of model performance. For example, the model corresponding to the smallest loss function value can be selected as the trained machine learning model. The embodiment of the present application does not specifically limit the method of selecting a trained machine learning model.
[0178] For example, the implementation tool of the neural network model is PyTorch, and the computing platform is a processor with six physical cores and 24GB of memory. The following experiments are conducted on a data set with a total data volume of 2,700 performance reports:
[0179] The training set, validation set, and test set were randomly divided into 2000, 300, and 400 items, respectively. There was no overlap between the three sets. The high, medium, and low risk level labels corresponding to the performance levels in the data set were derived from the expert review results. Furthermore, the Adamax optimizer provided by PyTorch was used to train the initial neural network model. The loss function was set to the cross entropy loss function, and the learning rate was 1×10 -5 , the regularization parameter betas is (0.3, 0.1), the weight decay parameter weight_decay is 0.1, the random seed is 3407, and the training process sets the upper limit of the number of epochs to 10000. After each round of training on the training set, the model will test the loss function value on the validation set and record it. If there is no better loss function value on the validation set test for 500 consecutive rounds, the training is terminated to obtain a trained neural network model.
[0180] Optionally, the model with the smallest loss function value on the validation set is tested on the test set and the results shown in Table 2 are obtained, wherein the indicators for evaluating the prediction results of the model may include accuracy, recall, F1 score, etc., which are not specifically limited in the embodiments of the present application.
[0181] Table 2
[0182]
[0183] As can be seen from Table 2, under the current data scale, the performance level prediction using the trained machine learning model is highly accurate. Therefore, most of the manual review process can be replaced by model-based use.
[0184] In this way, through multiple rounds of training and validation, the machine learning model can better fit the data and improve its prediction accuracy. The use of the validation set can help evaluate the generalization ability of the model, avoid overfitting, and improve the performance of the model on new data. The loss function value is used as the optimization criterion to automatically determine the trained machine learning model, thereby improving the efficiency and accuracy of model training.
[0185] Optionally, preprocessing the stress test data to determine target data includes:
[0186] Performing numerical anomaly processing on the stress test data to obtain seventh data;
[0187] Performing logic exception processing on the seventh data based on a preset rule to obtain eighth data;
[0188] The eighth data is subjected to data smoothing processing to obtain target data.
[0189] In the embodiment of the present application, numerical anomaly refers to the phenomenon that abnormal numerical values appear in the acquired performance test source data. Abnormal numerical values may include Nan values, the starting data and the ending data of the stress test are 0 values, etc. The embodiment of the present application does not specifically limit the type of abnormal numerical values.
[0190] For example, for abnormal values such as Nan values, if they are located at the beginning and end of the time series data, they are deleted; if they are located in the middle, the abnormal value is replaced with 0; for the 0 value at the beginning and end of the stress test data, they can be directly deleted.
[0191] Logical anomaly refers to the phenomenon that the obtained performance test source data value is normal but does not conform to the stress test logic. Logical anomaly may include a response time of 0 in the stress test, a slow response time at the beginning and end of the stress test, etc. The embodiment of the present application does not specifically limit the data corresponding to the logical anomaly type.
[0192] Optionally, the seventh data is subjected to logical exception processing using preset rules, and illogical abnormal data is removed or corrected to obtain eighth data. The embodiment of the present application does not specifically limit the preset rules.
[0193] For example, if the response time in stress testing is 0, the response time of the previous time point is used to fill the 0 value; if the response time at the beginning and end of the stress testing is slow, the sequence segment with slow response time can be deleted.
[0194] In the present application, before extracting features from the target data, the eighth data needs to be smoothed to remove random roughness on the eighth data to a certain extent. For example, the eighth data can be smoothed using moving average, exponential smoothing and other techniques to reduce random fluctuations in the data. After smoothing, the target data obtained can better reflect the actual performance trend of the interface and is suitable for subsequent feature extraction and model training.
[0195] For example, the tsa.seasonal.STL method in the machine learning algorithm statsmodels library can be used, and 3 seconds can be selected as the smoothing length to reduce the random roughness in the eighth data.
[0196] In this way, through numerical and logical exception processing, not only can outliers and errors in stress testing data be removed, improving data accuracy and reliability, but also data complexity and redundancy can be reduced, optimizing the utilization of computing resources. In addition, data smoothing reduces the impact of random fluctuations, making the data more reflective of real performance trends and improving the accuracy of analysis. Furthermore, by removing anomalies and noise, the possibility of misleading results is reduced, ensuring the effectiveness of subsequent performance predictions.
[0197] Optionally, the method further includes:
[0198] Determine the effective duration of the seventh data, and generate first prompt information when the effective duration is less than a duration threshold;
[0199] And / or, based on the eighth data, calculating the stress testing success rate in the first time period, and generating second prompt information when the stress testing success rate is less than a third threshold.
[0200] In the embodiment of the present application, the effective duration refers to the duration of data in the seventh data that is considered to be reliable and meaningful. The effective duration is compared with a preset duration threshold. If the effective duration is less than the duration threshold, it means that the data may not be sufficient for reliable analysis. For example, data with an effective duration of less than 60 seconds in stress testing may not be subjected to subsequent analysis. The effective duration refers to the length of the response time series, with a granularity of 1 second.
[0201] Optionally, when it is determined that the stress testing success rate is less than the third threshold, subsequent analysis can be performed, but a certain deduction needs to be made to the score corresponding to the eighth data. Furthermore, the eighth data corresponding to the stress testing success rate being less than the third threshold is smoothed to obtain the target data. It can be understood that performing performance testing under the above circumstances may lead to inaccurate performance evaluation, that is, the performance test results may be difficult to truly reflect the actual situation in the production environment.
[0202] Among them, the embodiment of the present application does not specifically limit the number of points deducted, which can be set based on the actual application scenario requirements.
[0203] It should be noted that the embodiment of the present application does not specifically limit the size of the duration threshold and the third threshold. If the stress test success rate is less than the third threshold, such as the stress test success rate is lower than 95%, it indicates that there may be a high failure rate in the interface test process. In this case, subsequent performance prediction may not be performed, and prompt information may be directly generated to remind the user.
[0204] Exemplarily, when the effective duration is less than the duration threshold, a first prompt message is generated to remind the user that the data may be incomplete or unstable, and it is recommended to re-collect data or adjust test parameters without the need for subsequent performance prediction.
[0205] Optionally, based on the eighth data, a stress test success rate within the first time period is calculated, where the stress test success rate refers to the proportion of requests that are successfully completed among all interface requests. The calculated success rate is compared with a preset third threshold. If the success rate is lower than the third threshold, a second prompt message is generated to remind the user that there are many failures during the test and that it may be necessary to check the system configuration, network conditions or other influencing factors.
[0206] Optionally, the first prompt information and the second prompt information can be sent to the user's terminal device, or to the display device corresponding to the data processing system. The embodiment of the present application does not specifically limit the sending form and display content of the first prompt information and the second prompt information.
[0207] It is understandable that the above-mentioned prompt information mechanism can help quickly identify possible data problems during the test process and support timely measures for adjustment and optimization.
[0208] In this way, by monitoring the effective time and success rate, the reliability and integrity of the test data can be ensured, and instant feedback on the test data and results can be provided to support more informed decision-making and resource allocation. Moreover, through timely prompts, further analysis or decision-making based on unreliable data can be avoided, thus reducing waste of resources.
[0209] Optionally, determining a performance test result based on the performance prediction level and the confidence level includes:
[0210] When it is determined that the confidence level is greater than a fourth threshold, determining a performance test result based on the performance prediction level;
[0211] When it is determined that the confidence level is less than or equal to a fourth threshold, the target data and the performance prediction level are sent to a user terminal for manual review to determine a performance test result.
[0212] In an embodiment of the present application, the fourth threshold refers to a preset confidence threshold, which is used to judge the reliability of the model prediction results. When the confidence is higher than the fourth threshold, it means that the model has a high degree of credibility in its prediction results. In this case, the performance test results can be determined directly based on the performance prediction level; wherein, high-confidence prediction results can be processed automatically to reduce manual intervention and improve efficiency.
[0213] Optionally, when the confidence level is lower than or equal to the fourth threshold, the model prediction result may not be reliable enough. In this case, the target data and the performance prediction level can be sent to the user terminal for manual review, and then the performance test result can be determined through manual review combined with the model prediction level and target data.
[0214] In this way, through confidence assessment and manual review, the reliability and accuracy of performance test results can be ensured, and the risk of misjudgment can be reduced. By automating the processing of high-confidence results, human resources can be saved, and manual review can be concentrated on complex tests with low confidence to improve resource utilization efficiency. By manually reviewing low-confidence results, the risks caused by inaccurate model predictions can be reduced, ensuring the stability and reliability of performance evaluation.
[0215] In combination with the above embodiments, Fig.13The overall flow chart of an interface performance testing method provided in an embodiment of the present application is as follows: Fig.13 As shown in the figure, the process is mainly divided into four steps: data collection, data preprocessing, feature extraction and automatic review.
[0216] Among them, data preprocessing includes data anomaly processing, logical anomaly processing and data smoothing processing; feature extraction mainly analyzes the various features in the stress testing data through expert experience rules and statistical methods; automatic review includes scoring the performance test through regularization and normalization methods, and secondly, the score data and extracted features are input into a neural network model of a three-classification task, and high, medium and low risk levels and confidence levels are predicted through deep learning technology. The high-confidence prediction results are directly used as the performance test results of the financial information system, while the low-confidence prediction results will be manually reviewed to determine the performance test results.
[0217] It should be noted that the interface performance testing method provided in this application can be integrated in SENSE to conduct a digital and labeled performance stability evaluation of the financial information system. It can be understood that under the existing organizational system, the organizational evaluation after the performance test of the financial information system is completed is decided by voting by multiple technical experts to decide whether the performance test results are passed, and the SENSE proposed in this application will replace most of the functions of expert review.
[0218] For example, based on experimental data, it can be seen that the SENSE-based scoring system can assist decision-making and promote the automation and intelligent development of the performance test review process within the bank. Among them, the manual review process has been reduced from 10 minutes per person to 2 minutes, and the fully automatic review rate is expected to reach 40%. Calculated from the labor cost, there are 477 performance testing projects supported by the method provided in this application, with a total of 1.2 million stress testing cases executed, and more than 2,700 evaluation reports generated using SENSE. The comprehensive evaluation saves 1,312 people / days in review workload, and the labor cost is greatly reduced. Among them, if an evaluation report is manually reviewed, it will require a total of 7 people, including the system manager, review administrator, senior manager, 3 deputy reviewers, and 1 main reviewer. The manual review of a report takes about 10 minutes. Therefore, about 2,700*10*7 / 6 / 24=1,312 people / days can be saved.
[0219] In terms of quality improvement and performance prediction and identification, the quantitative scoring system further clarified the stratification of performance capabilities and risk levels. The median score corresponding to the performance capabilities based on SENSE gradually increased from 63.4 points to 79.35 points. With the increasing number of subsystems and transaction interfaces using SENSE, the accuracy of actively identifying transaction execution records with various risk characteristics has been greatly improved.
[0220] In the above embodiments, the interface performance test method provided by the embodiment of the present application is introduced. In order to realize the functions in the method provided by the embodiment of the present application, the electronic device as the execution subject may include a hardware structure and / or a software module, and the above functions are realized in the form of a hardware structure, a software module, or a hardware structure plus a software module. Whether one of the above functions is executed in the form of a hardware structure, a software module, or a hardware structure plus a software module depends on the specific application and design constraints of the technical solution.
[0221] For example, Fig.14 A schematic diagram of the structure of an interface performance testing device provided in an embodiment of the present application is shown in FIG. Fig.14 As shown, the device 1400 includes: a preprocessing module 1401, which is used to collect stress test data of the interface, preprocess the stress test data, and determine target data;
[0222] The feature extraction module 1402 is used to extract multiple feature data from the target data, and score based on the target data and the multiple feature data to determine at least one score data; the multiple feature data include first data for characterizing that the response time is abnormal, second data for characterizing that the number of concurrent interfaces is abnormal, and test data determined based on the number of transactions per second TPS; the at least one score data includes score data of at least one dimension corresponding to the response time of performing stress testing at different time points;
[0223] A prediction module 1403, configured to input the plurality of feature data and the at least one score data into a trained machine learning model to obtain a performance prediction level and a confidence level;
[0224] The determination module 1404 is used to determine the performance test result based on the performance prediction level and the confidence level.
[0225] Optionally, the stress test data is time series data; the first data includes glitch data characterizing that the response time in a first time period is greater than a first threshold and the deviation value corresponding to the response time is greater than a second threshold, third data characterizing that an oscillation interval of the glitch data occurs in a second time period, fourth data characterizing whether the glitch data in the second time period satisfies a periodic law, fifth data characterizing that an abnormal slope occurs in the response time in the second time period, and sixth data characterizing that an abnormality occurs in the response time in the second time period and the abnormal phenomenon lasts for a preset period of time.
[0226] Optionally, the feature extraction module 1402 includes a feature extraction unit for the second data, the feature extraction unit being configured to:
[0227] Obtain the first concurrent number of the interface in the current gradient time and the second concurrent number in the previous gradient time, and calculate the average response time and average TPS in the current gradient time;
[0228] Based on the average response time, the average TPS, the first concurrency number, and the second concurrency number, it is determined whether the interface concurrency number is abnormal, and second data is obtained.
[0229] Optionally, the feature extraction module 1402 further includes a scoring unit, which is used to:
[0230] Performing regularization processing on a plurality of target data in a second time period to obtain a plurality of performance data, and calculating an average value and a standard deviation of the plurality of performance data;
[0231] determining a volatility for each performance data based on the mean value and the standard deviation;
[0232] determining a maximum value and a minimum value among a plurality of volatility, and determining an initial score for each performance data based on the minimum value and the maximum value;
[0233] Based on the first data and the second data, the initial score of each performance data is processed to determine at least one score data.
[0234] Optionally, the apparatus 1400 further includes a training module for a machine learning model, wherein the training module is used to:
[0235] Acquire a data set; the data set includes a training set and a validation set, the data set includes multiple samples, each sample includes multiple feature data, multiple score data and performance level corresponding to the stress test data of the interface;
[0236] Perform multiple rounds of cyclic training on the initial machine learning model using the training set, and determine the loss function value of the initial machine learning model after each round of training based on the validation set;
[0237] If the loss function values in the preset rounds are all less than the preset threshold, the trained machine learning model is determined from the initial machine learning models trained in the preset rounds.
[0238] Optionally, the preprocessing module 1401 is specifically used for:
[0239] Performing numerical anomaly processing on the stress test data to obtain seventh data;
[0240] Performing logic exception processing on the seventh data based on a preset rule to obtain eighth data;
[0241] The eighth data is subjected to data smoothing processing to obtain target data.
[0242] Optionally, the device 1400 further includes a prompt module, the prompt module being configured to:
[0243] Determine the effective duration of the seventh data, and generate first prompt information when the effective duration is less than a duration threshold;
[0244] And / or, based on the eighth data, calculating the stress testing success rate in the first time period, and generating second prompt information when the stress testing success rate is less than a third threshold.
[0245] Optionally, the determining module 1404 is specifically configured to:
[0246] When it is determined that the confidence level is greater than a fourth threshold, determining a performance test result based on the performance prediction level;
[0247] When it is determined that the confidence level is less than or equal to a fourth threshold, the target data and the performance prediction level are sent to a user terminal for manual review to determine a performance test result.
[0248] It should be noted that the specific implementation principle and effects of the above-mentioned interface performance testing device can be found in the relevant descriptions and effects corresponding to the above-mentioned embodiments, and will not be elaborated here.
[0249] The present application also provides a schematic diagram of the structure of an electronic device. Fig.15 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Fig.15 As shown, the electronic device may include: a processor 1501 and a memory 1502 that is communicatively connected to the processor 1501; the memory 1502 stores a computer program; the processor 1501 executes the computer program stored in the memory 1502, so that the processor 1501 executes the method described in any of the above embodiments.
[0250] The memory 1502 and the processor 1501 may be connected via a bus 1503 .
[0251] An embodiment of the present application further provides a computer-readable storage medium, which stores computer program execution instructions. When the computer program execution instructions are executed by a processor, they are used to implement the method described in any of the aforementioned embodiments of the present application.
[0252] An embodiment of the present application further provides a chip for executing instructions, wherein the chip is used to execute the method described in any of the aforementioned embodiments as executed by an electronic device in any of the aforementioned embodiments of the present application.
[0253] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the method described in any of the aforementioned embodiments of the present application executed by an electronic device can be implemented.
[0254] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0255] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to implement the solution of this embodiment.
[0256] In addition, each functional module in each embodiment of the present application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The above-mentioned module-composed unit can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0257] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application.
[0258] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly implemented as being executed by a hardware processor, or may be implemented by a combination of hardware and software modules in the processor.
[0259] The memory may include high-speed random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk storage, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.
[0260] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0261] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special computer.
[0262] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0263] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.
[0264] It should be further noted that, although the various steps in the flowchart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0265] In the above embodiments, the description of each embodiment has its own emphasis. For the part not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0266] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the claims.
[0267] The above is only a specific implementation of the embodiment of the present application, but the protection scope of the embodiment of the present application is not limited thereto, and any changes or replacements within the technical scope disclosed in the embodiment of the present application should be included in the protection scope of the embodiment of the present application. Therefore, the protection scope of the embodiment of the present application should be based on the protection scope of the claims.
Claims
1. An interface performance testing method, characterized in that: The method comprises: Collect the stress test data of the interface, pre-process the stress test data, and determine the target data; Extracting multiple feature data from the target data, and scoring based on the target data and the multiple feature data to determine at least one score data; the multiple feature data include first data for characterizing abnormal response time, second data for characterizing abnormal interface concurrency, and test data determined based on the number of transactions per second (TPS); the at least one score data includes scoring data of at least one dimension corresponding to the response time of performing stress testing at different time points; Inputting the plurality of feature data and the at least one score data into a trained machine learning model to obtain a performance prediction level and a confidence level; A performance test result is determined based on the performance prediction level and the confidence level.
2. The method according to claim 1, characterized in that The stress test data is time series data; the first data includes glitch data for characterizing that the response time in a first time period is greater than a first threshold and the deviation value corresponding to the response time is greater than a second threshold, third data for characterizing that an oscillation interval of the glitch data occurs in a second time period, fourth data for characterizing whether the glitch data in the second time period satisfies a periodic law, fifth data for characterizing that an abnormal slope occurs in the response time in the second time period, and sixth data for characterizing that an abnormality occurs in the response time in the second time period and the abnormal phenomenon lasts for a preset period of time.
3. The method according to claim 1, characterized in that The feature extraction process of the second data includes: Obtain the first concurrent number of the interface in the current gradient time and the second concurrent number in the previous gradient time, and calculate the average response time and average TPS in the current gradient time; Based on the average response time, the average TPS, the first concurrency number, and the second concurrency number, it is determined whether the interface concurrency number is abnormal, and second data is obtained.
4. The method according to claim 1, characterized in that: Scoring the target data and the plurality of feature data to determine at least one score data includes: Performing regularization processing on a plurality of target data in a second time period to obtain a plurality of performance data, and calculating an average value and a standard deviation of the plurality of performance data; determining a volatility for each performance data based on the mean value and the standard deviation; determining a maximum value and a minimum value among a plurality of volatility, and determining an initial score for each performance data based on the minimum value and the maximum value; Based on the first data and the second data, the initial score of each performance data is processed to determine at least one score data.
5. The method according to claim 1, characterized in that The training method of the machine learning model includes: Acquire a data set; the data set includes a training set and a validation set, the data set includes multiple samples, each sample includes multiple feature data corresponding to the stress test data of the interface, multiple score data and performance level; Perform multiple rounds of cyclic training on the initial machine learning model using the training set, and determine the loss function value of the initial machine learning model after each round of training based on the validation set; If the loss function values in the preset rounds are all less than the preset threshold, the trained machine learning model is determined from the initial machine learning models trained in the preset rounds.
6. The method according to claim 1, characterized in that Preprocess the stress test data to determine target data, including: Performing numerical anomaly processing on the stress test data to obtain seventh data; Performing logic exception processing on the seventh data based on a preset rule to obtain eighth data; The eighth data is subjected to data smoothing processing to obtain target data.
7. The method according to claim 6, characterized in that The method further comprises: Determine the effective duration of the seventh data, and generate first prompt information when the effective duration is less than a duration threshold; And / or, based on the eighth data, calculating the stress testing success rate in the first time period, and generating second prompt information when the stress testing success rate is less than a third threshold.
8. The method according to any one of claims 1 to 7, characterized in that: Determining a performance test result based on the performance prediction level and the confidence level includes: When it is determined that the confidence level is greater than a fourth threshold, determining a performance test result based on the performance prediction level; When it is determined that the confidence level is less than or equal to a fourth threshold, the target data and the performance prediction level are sent to a user terminal for manual review to determine a performance test result.
9. An interface performance testing device, characterized in that: The device comprises: A preprocessing module is used to collect the stress test data of the interface, preprocess the stress test data, and determine the target data; A feature extraction module is used to extract multiple feature data from the target data, and to score based on the target data and the multiple feature data to determine at least one score data; the multiple feature data include first data for characterizing abnormal response time, second data for characterizing abnormal interface concurrency, and test data determined based on the number of transactions per second (TPS); the at least one score data includes score data of at least one dimension corresponding to the response time of performing stress testing at different time points; A prediction module, configured to input the plurality of feature data and the at least one score data into a trained machine learning model to obtain a performance prediction level and a confidence level; A determination module is used to determine a performance test result based on the performance prediction level and the confidence level.
10. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
System interface verification method and system based on artificial intelligence
CN121070749A
Model inference server pressure measurement data-oriented filtering method and device
CN122332371A