Service performance test method and device based on rule tree and clustering
By constructing the peak proportion model and data distribution model based on rule tree and clustering, we generate real and diverse test data, solving the problems of low coverage of performance test data and single test scenarios in the existing technology, and achieving more accurate performance testing.
Patent Information
- Application Number
- CN202510376297.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the performance test data coverage is low and the test scenario is single, so the real user behavior cannot be accurately simulated, resulting in errors in the performance test results, which is difficult to reflect the performance of the system in actual business scenarios.
The business performance testing method based on rule trees and clustering is adopted. By obtaining online log data and request data, a peak proportion model and data distribution model are constructed, real and diverse test data are generated, and testing scenarios are simulated for different transaction types and load conditions.
It achieves higher test data coverage and more realistic test scenarios, can accurately simulate real user behavior, and improve the accuracy and reliability of performance testing.
Smart Images

Figure CN119988175A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of performance testing, and in particular to a business performance testing method based on rule trees and clustering, a business performance testing device based on rule trees and clustering, a computer-readable storage medium and a business system. Background Art
[0002] With the rapid development of information technology, the complexity of system architecture and deployment scale continues to increase, and business scenarios are becoming increasingly diverse, making the design and analysis of performance testing increasingly difficult. Especially in large-scale distributed systems and microservice architectures, traditional performance testing methods face many challenges. At present, most performance test design methods rely on the common field values or business expected values of the tested interface to randomly generate test data, resulting in low test data coverage and a single test scenario. Such test design often fails to fully reflect the performance of the system in the actual business environment and tends to ignore the impact of some extreme situations or complex loads.
[0003] In addition, performance testing for single or multi-transaction models usually relies on business personnel and technical personnel to estimate transaction volumes based on experience or local data. However, since these estimated data cannot fully represent real scenarios, this approach can easily lead to a large gap between performance test results and actual user behavior, and the test accuracy is insufficient. Therefore, performance testing often cannot truly simulate the operating status of the system in high-concurrency and complex transaction scenarios, affecting the effective evaluation of system performance.
[0004] Constructing rich and accurate test data scenarios and transaction ratio models is crucial for performance testing, especially in key scenarios such as system reconstruction and service migration. Performance evaluation is not only related to the stability of the system after it goes online, but also affects the scalability and robustness of the system under high load conditions.
[0005] Therefore, the existing technology has the problems of low test data coverage, single test scenario, and inability to accurately simulate real user behavior, which leads to errors in performance test results and makes it difficult to truly reflect the performance of the system in actual business scenarios. Summary of the invention
[0006] The main purpose of the present application is to provide a business performance testing method based on rule trees and clustering, a business performance testing device based on rule trees and clustering, a computer-readable storage medium and a business system, so as to at least solve the problems existing in the prior art of low test data coverage, single test scenario, and inability to accurately simulate real user behavior, which leads to errors in performance test results and makes it difficult to truly reflect the performance of the system in actual business scenarios.
[0007] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a business performance testing method based on rule tree and clustering is provided, including: obtaining online log data to obtain first target data, determining online request data corresponding to the interface to be tested based on the interface to be tested, and obtaining second target data; segmenting the first target data based on timestamp to obtain multiple third target data, clustering the third target data of different transaction types based on transaction volume to obtain clustering results, and constructing a peak share model according to the clustering results, wherein the peak share model is used to characterize the load distribution of different transaction types; performing rule matching on at least the second target data through a rule tree structure, and constructing a data distribution model based on the matching results, wherein the data distribution model is used to characterize the distribution ratio of different data features; constructing a target test scenario according to the peak share model, generating performance test data according to the data distribution model, and testing the interface to be tested according to the performance test data under the target test scenario.
[0008] Optionally, the first target data is segmented based on the timestamp to obtain multiple third target data, including: parsing the timestamp, the transaction type and the transaction volume in the first target data based on the first target data to obtain first identification data, second identification data and third identification data; taking a first preset time length as a basic unit, dividing the first target data according to the first identification data to obtain multiple third target data.
[0009] Optionally, the third target data of different transaction types are clustered based on transaction volume to obtain clustering results, and a peak share model is constructed according to the clustering results, including: sorting the third target data from large to small according to the third identification data to obtain a target sequence; constructing a peak interval window according to a second preset time length, and intercepting the target sequence based on the peak interval window to obtain fourth target data, the second preset time length is greater than the first preset time length, and the peak interval window is used to intercept the third target data within the second preset time length starting from the maximum value of the third identification data; taking the first preset time length as the minimum sample unit, and clustering based on the second identification data and the third identification data to obtain at least one target clustering cluster; calculating the mean transaction volume of samples in the cluster based on the target clustering cluster, and calculating the transaction volume proportion of different second identification data in the target clustering cluster corresponding to the maximum value of the mean transaction volume, and constructing the peak share model according to the transaction volume proportion.
[0010] Optionally, after sorting the third target data from large to small according to the third identification data to obtain a target sequence, the method further includes: determining the corresponding architecture complexity and data volume level based on the second identification data; performing weighted calculation based on the architecture complexity and the data volume level to obtain a transaction correction index, and when the transaction correction index is greater than or equal to a preset value, determining a correction weight based on the architecture complexity and the data volume level, the correction weight being proportional to the architecture complexity, and the correction weight being proportional to the data volume level; correcting the third identification data based on the correction weight, and re-sorting the third identification data from large to small according to the corrected third identification data to update the target sequence.
[0011] Optionally, rule matching is performed on at least the second target data through a rule tree structure, including: extracting business features included in the business request based on the first target data to obtain fifth target data, wherein the business features include at least one of the requested business type, request identifier, business frequency, business cycle and business rules; taking each of the fifth target data as a layer node of the rule tree to obtain a target rule tree; extracting request event data corresponding to the business request based on the second target data to obtain sixth target data; performing layer-by-layer matching based on the sixth target data and the target rule tree, and recording the matching count of the leaf nodes of the target rule tree; and determining that the current path of the target rule tree matches when the matching count is greater than or equal to the first threshold and less than or equal to the second threshold.
[0012] Optionally, after performing layer-by-layer matching based on the sixth target data and the target rule tree and recording the matching counts of the leaf nodes of the target rule tree, the method further includes: when the matching count is less than the first threshold, pruning the current path of the target rule tree; when the matching count is greater than the second threshold, splitting each layer of the target rule tree, and re-matching the sixth target data based on the split target rule tree, wherein the branches of the same layer of the split target rule tree are larger than the target rule tree before splitting.
[0013] Optionally, performance test data is generated according to the data distribution model, including: calculating the sum of the matching counts of the target rule tree corresponding to the data distribution model to obtain a total matching count; calculating the ratio of the total matching count to the total data volume of the first target data to obtain a target coverage; when the target coverage is greater than a third threshold, generating a test data set based on each path, wherein the proportion of sample data corresponding to each path in the test data set is consistent with the ratio of the matching count corresponding to the path to the total matching count.
[0014] According to another aspect of the present application, a business performance testing device based on rule trees and clustering is provided, and the device includes: a first acquisition unit, used to acquire online log data to obtain first target data, determine online request data corresponding to the interface to be tested based on the interface to be tested, and obtain second target data; a first processing unit, used to segment the first target data based on timestamps to obtain multiple third target data, cluster the third target data of different transaction types based on transaction volume to obtain clustering results, and construct a peak share model according to the clustering results, and the peak share model is used to characterize the load distribution of different transaction types; a second processing unit, used to perform rule matching on at least the second target data through a rule tree structure, and construct a data distribution model based on the matching results, and the data distribution model is used to characterize the distribution ratio of different data features; a first construction unit, used to construct a target test scenario according to the peak share model, generate performance test data according to the data distribution model, and test the interface to be tested according to the performance test data under the target test scenario.
[0015] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute any one of the methods described.
[0016] According to another aspect of the present application, a business system is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include means for executing any one of the methods described.
[0017] Applying the technical solution of the present application, the present application analyzes the online log data and the online request data respectively to construct a peak share model and a data distribution model containing data distribution characteristics, which are used to simulate the actual behavior of online users, build test scenarios based on the peak share model, and generate test data based on the data distribution model, thereby realizing the simulation of real transaction scenarios and data, and solving the problem of low test data coverage and single test scenario when conducting business system performance testing in the prior art, resulting in errors in the test results of the interface. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A hardware structure block diagram of a mobile terminal for service performance testing based on a rule tree and clustering provided in an embodiment of the present application is shown;
[0019] Figure 2A schematic diagram of a process flow of a service performance testing method based on a rule tree and clustering provided according to an embodiment of the present application is shown;
[0020] Figure 3 A schematic diagram of a process for constructing a peak share model according to an embodiment of the present application is shown;
[0021] Figure 4 A flowchart of a specific service performance testing method based on rule trees and clustering provided according to an embodiment of the present application is shown;
[0022] Figure 5 A schematic diagram of a process for constructing a data distribution model according to an embodiment of the present application is shown;
[0023] Figure 6 A rule tree logic example diagram provided according to an embodiment of the present application is shown;
[0024] Figure 7 A structural block diagram of a service performance testing device based on rule tree and clustering provided according to an embodiment of the present application is shown.
[0025] The above drawings include the following reference numerals:
[0026] 102, processor; 104, memory; 106, transmission device; 108, input and output devices. DETAILED DESCRIPTION
[0027] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0028] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0030] For the convenience of description, some nouns or terms involved in the embodiments of the present application are explained below:
[0031] Cluster analysis: Cluster analysis is a method of grouping similar objects. The goal is to automatically divide the objects in the data set into several different categories or subsets based on the similarity of their features. The objects in each subset have high internal similarity, while the objects in different subsets are quite different.
[0032] Rule tree: A rule tree is a tree structure used to represent a set of rules. It organizes rules into decision trees or condition trees according to specific logical relationships. It is usually used in scenarios such as data classification, decision making, or screening. Each node represents a condition or decision, and the tree structure helps process and analyze input data.
[0033] As introduced in the background technology, the performance testing methods in the prior art are difficult to adapt to complex architectures and diversified business needs, and the test data coverage is low and the scenarios are single, which cannot reflect the extreme situations and complex loads in actual business. Transaction volume estimation relies on experience, and the test results deviate greatly from the actual ones, which affects the accuracy and system performance evaluation, especially in reconstruction and migration scenarios, lacking comprehensive and reliable performance verification. In order to solve the problem of low test data coverage and single test scenarios in the prior art when conducting business system performance testing, resulting in errors in the test results of the interface, the embodiments of the present application provide a business performance testing method based on rule trees and clustering, a business performance testing device based on rule trees and clustering, a computer-readable storage medium, and a business system.
[0034] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0035] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 11 is a hardware structure block diagram of a mobile terminal of a service performance testing method based on a rule tree and clustering according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is for illustration only and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0036] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the display method of device information in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The transmission device 106 is used to receive or send data via a network. The above-mentioned specific examples of the network may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0037] In this embodiment, a business performance testing method based on rule trees and clustering is provided, which runs on a mobile terminal, a computer terminal or a similar computing device. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0038] Figure 2 1 is a flow chart of a service performance testing method based on rule trees and clustering according to an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:
[0039] Step S201, obtaining online log data to obtain first target data, determining online request data corresponding to the interface to be tested based on the interface to be tested, and obtaining second target data;
[0040] Specifically, by collecting log data in the online operating environment, the first target data containing user behavior and system interaction information is extracted. Subsequently, combined with the interface to be tested, the request data related to the interface is analyzed and screened to obtain the second target data. By obtaining these two types of data, it is ensured that the performance test data is highly relevant to the real business scenario and can cover the behavior patterns of actual online users.
[0041] Step S202, segmenting the first target data based on the timestamp to obtain a plurality of third target data, clustering the third target data of different transaction types based on the transaction volume to obtain a clustering result, and constructing a peak share model according to the clustering result, the peak share model being used to characterize the load distribution of different transaction types;
[0042] Specifically, by dividing the first target data into time series, the data is grouped by time period in order to analyze the system load and transaction behavior at different time nodes. On this basis, the data is clustered in combination with transaction type and transaction volume to identify the distribution pattern and load peak period of different transaction types. Based on the cluster analysis results, a peak share model is constructed to characterize the load share of each transaction type in different time periods, helping to accurately simulate the load fluctuations of different transaction types.
[0043] Step S203, performing rule matching on at least the second target data through a rule tree structure, and constructing a data distribution model based on the matching result, where the data distribution model is used to characterize the distribution ratio of different data features;
[0044] Specifically, the rule tree structure is used to match the second target data with rules. By setting a set of rules, the request data is automatically matched with the predefined rule set to extract the distribution information of different data features. Based on the matching results, a data distribution model is constructed to characterize the distribution ratio of various data features in the test data, such as the parameter type, data volume, response time and other key features of the request. By accurately matching and classifying the data through the rule tree structure, a comprehensive test data distribution model can be efficiently generated. This method can effectively characterize the distribution of different data features, providing important support for generating real and representative test data for performance testing, thereby improving the comprehensiveness and accuracy of the test.
[0045] Step S204, constructing a target test scenario according to the peak share model, generating performance test data according to the data distribution model, and testing the interface to be tested according to the performance test data in the target test scenario.
[0046] Specifically, based on the constructed peak share model, target test scenarios for different transaction types and load conditions are designed and constructed to simulate various usage scenarios that may occur in the real environment. At the same time, based on the data distribution model, performance test data that conforms to the test scenario is generated to ensure that the test data is consistent with the actual business data in terms of quantity and characteristics. Finally, these performance test data are used to conduct a comprehensive performance test on the interface to be tested in the target test scenario.
[0047] It can be seen that the embodiment of the present application provides a business performance testing method based on rule tree and clustering, which constructs a test data distribution model and a transaction model by automatically analyzing online request messages and log files, and automatically generates test data based on this. Combined with the load model and business analysis, the real scene can be simulated more accurately, so as to comprehensively evaluate the system performance. The test data distribution adopts the rule tree matching method to efficiently process large-scale data and complex rule sets, and flexibly maintain the rules. The transaction proportion analysis extracts transaction data through logs, uses cluster analysis to identify peak proportion models and trends, discovers the impact of outlier transactions on performance, and improves system stability. This application combines real log data, time series division and cluster analysis to construct an accurate peak proportion model to characterize load distribution; through rule tree matching request data, an accurate data distribution model is generated. The combination of the two generates real and diverse test scenarios, which solves the problems of low test data coverage, single test scenarios and inability to accurately simulate user behavior in the prior art.
[0048] As a possible implementation manner, dividing the first target data based on time sequence includes the following steps:
[0049] Step S301, parsing the timestamp, transaction type and transaction amount in the first target data based on the first target data to obtain first identification data, second identification data and third identification data;
[0050] Specifically, by parsing timestamps, transaction types, and transaction volumes, the time, type, and load of each piece of data can be clearly distinguished, providing an accurate basis for subsequent time series analysis and clustering.
[0051] Step S302 : taking the first preset time length as a basic unit and dividing the first target data according to the first identification data to obtain a plurality of third target data.
[0052] Specifically, based on the timestamp (first identification data) and the preset time period (first preset duration), the first target data is divided according to time. The data in each time period forms a "third target data", which can help analyze the load of the system in each time window.
[0053] As a possible implementation, Figure 3 As shown, the third target data of different transaction types are clustered based on the transaction volume to obtain clustering results, and a peak proportion model is constructed according to the clustering results, including the following steps:
[0054] Step S401, sorting the third target data from large to small according to the third identification data to obtain a target sequence;
[0055] Specifically, the third target data (i.e., the data divided by time) is sorted from large to small according to the transaction volume (third identification data). The purpose of sorting is to highlight high-load data, especially data with large transaction volumes, which usually represent the load fluctuation of the system under high concurrency and are crucial to performance testing.
[0056] It can be understood that in the previous operation, data is obtained from data source A to analyze transaction types and transaction volumes, and then grouped according to minute time points to obtain the above-mentioned third target data.
[0057] Step S402, constructing a peak interval window according to the second preset duration, and intercepting the target sequence based on the peak interval window to obtain fourth target data, the second preset duration is greater than the first preset duration, and the peak interval window is used to intercept the third target data within the second preset duration starting from the maximum value of the third identification data;
[0058] Specifically, first, a peak interval window (the daily peak transaction ratio interval) is set according to the second preset time length (greater than the first preset time length), and a longer time period is defined to capture the peak of the system load. Then, the sorted third target data is intercepted through this peak interval window, and the key data within the time window is extracted to obtain the fourth target data.
[0059] Step S403, taking the first preset time length as the minimum sample unit, and performing clustering based on the second identification data and the third identification data to obtain at least one target cluster;
[0060] Specifically, samples with similar transaction types and transaction volumes are clustered with the first preset time length as a unit to obtain the above target cluster.
[0061] Step S404, based on the target cluster, calculate the transaction volume mean of the samples in the cluster, calculate the transaction volume proportion of different second identification data in the target cluster corresponding to the maximum value of the transaction volume mean, and construct a peak proportion model according to the transaction volume proportion.
[0062] Specifically, the average transaction volume within each target cluster is calculated, and then the cluster with the largest average transaction volume is selected as the construction object of the peak share model. Then, the transaction share of each transaction type in the peak transaction model is calculated according to the total transaction volume to obtain the above-mentioned peak share model.
[0063] As a possible implementation manner, before sorting the third target data from large to small according to the third identification data corresponding to the third target data, the method further includes:
[0064] Step S501, determining the corresponding architecture complexity and data volume level based on the second identification data;
[0065] Specifically, the complexity of the system architecture and the data volume level are determined by analyzing the second identification data (i.e., the transaction type). The complexity of the architecture reflects the complexity of the system design. For example, different architecture forms such as microservice architecture and distributed architecture have different complexities. The data volume level is determined based on the scale of data processed by the system, such as low, medium, and high data volume levels. This information helps to determine the load and processing requirements of the system under different transaction types.
[0066] Step S502, performing weighted calculation based on the architecture complexity and the data volume level to obtain a transaction correction index, and when the transaction correction index is greater than or equal to a preset value, determining a correction weight based on the architecture complexity and the data volume level, wherein the correction weight is proportional to the architecture complexity and the data volume level;
[0067] Specifically, the above-mentioned architecture complexity and the above-mentioned data volume level are quantified based on pre-established rules (such as scoring according to preset standards) for calculation, and then the quantified dimensionless indicators are weighted to obtain the above-mentioned transaction correction indicators, and based on the above-mentioned transaction indicators and preset values, it is determined whether the architecture complexity and data volume level have a greater impact on the transaction volume, and when the transaction correction indicator is greater than or equal to the preset value, the correction weight is determined. The correction weight is proportional to the architecture complexity and data volume level, which means that the more complex the system architecture and the larger the data volume, the higher the correction weight assigned. This correction weight is used to adjust the subsequent third identification data (i.e., transaction volume) to reflect the actual system requirements under different architectures and load conditions. By setting appropriate correction weights for different architectures and data volumes, the transaction volume in the test data can be dynamically adjusted to ensure that in the case of complex architectures and large data volumes, the test data can better reflect the actual system load situation, thereby avoiding test deviations caused by differences in architecture complexity and data volume, and improving the accuracy and representativeness of the test data.
[0068] Step S503, modifying the third identification data based on the modified weight, and re-sorting the modified third identification data from large to small to update the target sequence.
[0069] Specifically, by applying the above-determined correction weight, the third identification data (i.e., transaction volume) is corrected. The corrected third identification data more accurately reflects the actual transaction volume of the system under different architectural complexity and data volume conditions. The corrected transaction volume data is more accurate and can better simulate the real load of the system under different architectural complexity and data volume conditions. Then, the corrected transaction volume is re-sorted to update the above target sequence.
[0070] As a possible implementation manner, performing rule matching on at least the second target data through a rule tree structure includes the following steps:
[0071] Step S601, extracting a service feature included in the service request based on the first target data to obtain fifth target data, where the service feature includes at least one of a requested service type, a request identifier, a service frequency, a service cycle, and a service rule;
[0072] Specifically, by analyzing the first target data, feature information related to the business request is extracted, which is called business features. Business features include but are not limited to: Request business type: Identify the type of business request, such as query request, submission request, etc. Request identifier: Identify different business requests in order to distinguish different requests. Business frequency: Indicates the frequency of occurrence of requests, reflecting the call intensity of requests. Business cycle: Indicates the time periodic characteristics of the request, such as every minute, every hour, etc. Business rules: are the rule information contained in the business request, such as the rules in the payment process. The extracted business features are summarized as the fifth target data as the basis for subsequent rule tree construction and matching. By extracting detailed business features, the multi-dimensional information of the business request can be fully captured, providing rich data support for the subsequent construction and matching of the rule tree. The accuracy and analysis depth of the business request data are improved, so that the rule matching can better meet the actual business scenarios.
[0073] Step S602, taking each fifth target data as a layer node of a rule tree to obtain a target rule tree;
[0074] Specifically, each extracted fifth target data is used as a node in the rule tree, and a complete target rule tree is constructed. The rule tree is a typical decision tree structure, in which each layer of nodes represents a feature of a business request, and the connection between nodes represents the logical relationship between business requests. In this way, complex business rules and business processes can be effectively represented by the rule tree. The complex logical relationships in the business requests are structured to make subsequent rule matching more intuitive and efficient. The rule tree structure helps to hierarchically process different business features, improves the flexibility and efficiency of rule matching, and is conducive to generating more accurate test data.
[0075] Step S603, extracting request event data corresponding to the business request based on the second target data to obtain sixth target data;
[0076] Specifically, by analyzing the second target data, event data related to the business request, such as the timestamp of the request, response time, request parameters, etc., is extracted to generate the sixth target data. The sixth target data is used to further analyze the processing of the request and help identify key indicators such as time and resource consumption when the system processes the business request. Extracting request event data helps to understand the processing flow of each business request more deeply and provide more comprehensive context information for subsequent rule matching. Analysis based on event data can reveal the bottleneck of the system when processing specific requests, thereby optimizing system performance.
[0077] Step S604, performing layer-by-layer matching based on the sixth target data and the target Rule Tree, and recording the matching counts of the leaf nodes of the target Rule Tree;
[0078] Specifically, based on the sixth target data and the target rule tree, each business request is checked to see if it meets the conditions of each node in the rule tree by matching layer by layer. Each time a match is made, the match count is recorded on the leaf node, which reflects the number of business request characteristics that meet the current path. Layer-by-layer matching can accurately simulate the degree of compliance between business requests and the rule tree, ensuring that the rule tree can efficiently reflect each link in the business process. Recording the match count helps quantify the matching effect of each path, making subsequent rule analysis and optimization more intuitive and operational.
[0079] Step S605: When the match count is greater than or equal to the first threshold and less than or equal to the second threshold, determine that the current path of the target rule tree matches.
[0080] Specifically, a range of matching counts is set, that is, when the matching count is greater than or equal to the first threshold and less than or equal to the second threshold, the current path is considered to have passed the match. Setting this range can avoid inaccurate matching due to some abnormal data, thereby ensuring the accuracy of rule matching. By setting a threshold for matching counts, the accuracy of rule tree matching can be controlled to avoid overmatching or mismatching. This method makes rule matching more stable and can effectively handle different types of data fluctuations that may occur in business scenarios, ensuring the effectiveness of performance testing.
[0081] As a possible implementation manner, after performing layer-by-layer matching based on the sixth target data and the target rule tree and recording the matching counts of the leaf nodes of the target rule tree, the method further includes:
[0082] Step S701, when the matching count is less than a first threshold, prune the current path of the target rule tree;
[0083] Specifically, the validity of the current path is first determined based on the match count. When the match count of a path is less than the first threshold, it means that the business request data matched by this path is small or does not meet the expected rules, so this path needs to be pruned. Pruning paths that do not meet the conditions helps reduce the amount of calculation and avoid unnecessary calculation and resource consumption. By pruning paths with low match counts, unnecessary rule matching processes can be effectively reduced and the system operation efficiency can be improved.
[0084] Step S702, when the match count is greater than the second threshold, split each layer of the target rule tree, and re-match the sixth target data according to the split target rule tree, wherein the branches of the same layer of the split target rule tree are greater than the target rule tree before the split.
[0085] Specifically, when the match count of a certain path is greater than the second threshold, it means that the business request data of this path meets more rule requirements, which means that this part of the business logic may involve more complex scenarios or more business processes, so the rule tree needs to be split. The split rule tree will divide the original path more finely, so that each layer has more branches than before the split. The split rule tree will be re-matched with the sixth target data in order to obtain a more accurate matching result.
[0086] As a possible implementation method, generating performance test data according to the data distribution model includes the following steps:
[0087] Step S801, calculating the sum of the matching counts of the target rule tree corresponding to the data distribution model to obtain the total matching count;
[0088] Specifically, the basic steps for generating performance test data based on the data distribution model. First, it is necessary to calculate the sum of the matching counts of each path in the target rule tree constructed previously. The matching count reflects the amount of data matched by each rule path, and the total matching count obtained by adding these counts represents the matching degree of all paths in the entire rule tree in the target data set. This calculation result provides a basis for subsequent test data generation.
[0089] Step S802, calculating the ratio of the total matching count to the total data volume of the first target data to obtain the target coverage;
[0090] Specifically, after calculating the total match count, it is necessary to calculate the ratio between it and the total data volume of the first target data (i.e., the actual data volume) to obtain the target coverage. The target coverage reflects the proportion of the target data matched by the rule tree to the total actual data volume. This ratio provides a quantitative indicator for the quality of test data generation, indicating whether the current rule tree covers enough data. By calculating the target coverage, the ability of the rule tree to cover data can be quantified, so that the data allocation can be adjusted according to the coverage when generating test data to ensure that the test scenario is more comprehensive and realistic. The calculation of the target coverage provides a measurable standard for the adequacy of test data, helping to evaluate the quality of the current data generation solution.
[0091] Step S803, when the target coverage is greater than the third threshold, a test data set is generated based on each path, wherein the ratio of sample data corresponding to each path in the test data set is consistent with the ratio of the matching count corresponding to the path to the total matching count.
[0092] Specifically, when the target coverage is greater than the set third threshold, the corresponding performance test data set continues to be generated based on the matching counts of each path of the target rule tree. Specifically, the number of test data samples for each path should be consistent with the ratio of the path matching count to the total matching count. That is, paths with higher matching counts correspond to more sample data to ensure that these paths are fully covered in the test. By generating test data according to the matching count ratio, it is ensured that each path in the test data set can fully reflect the matching status of the rule tree, avoiding some paths from being unable to be fully verified due to too little sample data. This method can make the test data set more balanced, improve the representativeness and effectiveness of the test data, and thus improve the reliability of performance testing.
[0093] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the service performance testing method based on rule trees and clustering of the present application will be described in detail below in combination with specific embodiments.
[0094] This embodiment relates to a specific service performance testing method based on rule trees and clustering, such as Figure 4 As shown:
[0095] The process on the left is the process of building a transaction model: first input data source A: Data source A is an online log file. By parsing the log file, transaction data is obtained and stored in the database to obtain the first target data; then the transaction type and transaction volume are analyzed: the first target data is divided based on time series, and clustered based on transaction type and transaction volume; finally, a transaction model is built: a peak share model is built based on the clustering results to characterize the load distribution of different transaction types.
[0096] The process on the right is the data distribution modeling process: first determine the interface under test and parse the interface data under test: data source B is the online request data, and the online request data corresponding to the interface under test is determined based on the interface under test to obtain the second target data, and the second target data is parsed to obtain data features and data rules; then match the rule tree and establish a data distribution model: at least the second target data is matched by the rule tree structure, and a data distribution model is constructed based on the matching results to characterize the distribution ratio of different data features; finally, generate test data and performance scripts: construct a target test scenario based on the peak share model, and generate performance test data based on the data distribution model.
[0097] After the precise test data and transaction model are designed, system performance testing and evaluation can be carried out. That is, performance testing implementation: testing the interface to be tested according to the performance test data in the target test scenario. The goal of performance testing is to simulate the user's use of the system in a realistic manner. Therefore, it is particularly important to cover as many real online scenarios as possible with the help of accurate load models and rich test data scenarios.
[0098] The peak share model constructed through cluster analysis can simulate the request pattern of real users. The steps to construct the peak share model are as follows:
[0099] Step 1: Collect transaction log data, parse the timestamp, transaction type and transaction volume fields in the transaction log, convert the timestamp to minute time points, count the total transaction volume at the minute time point, the transaction volume of each transaction type and set a weight value for each transaction type, initialize the weight value to 1, and store all parsed data in the database.
[0100] Step 2: Group the trading volume by minute time points, calculate the total trading volume at each time point, and sort in descending order based on trading volume.
[0101] Step 3: In some test scenarios, such as adding new functions to system agile engineering, it is necessary to consider the performance of transaction types with complex data access and architecture implementation and large data volume. Therefore, the present invention supports the setting of transaction type weights, adding weight factors to transaction types according to test requirements, setting weight values, and statistically analyzing group weight values. On the basis of transaction volume sorting, each time group is re-sorted in descending order according to the transaction volume corrected according to the weight value. If this test scenario does not need to be considered, step 3 can be skipped.
[0102] Step 4: Manually determine the peak interval window size based on business needs. Combined with the group sorting, select the interval with the time point in the previous sorting as the starting point and the size equal to the peak interval window as the peak interval segment. The peak interval window here can be set according to actual needs. Filter out the peak interval segments that meet the conditions and record their start time and end time points.
[0103] Step 5: The database queries the transaction type and volume at each minute time point within the peak interval, performs cluster analysis on the time points of the peak interval segment, aggregates similar transaction volumes and transaction types into clusters, obtains the average value of each cluster, represents the transaction volume and transaction type of the cluster, and takes the cluster with the largest transaction volume average as the daily peak transaction model. The transaction type, transaction volume, weight, and transaction proportion of the transaction model under the total transaction volume are counted. And the final daily transaction proportion model is determined according to the descending order of the weight and transaction volume of each transaction type.
[0104] It can be seen that the peak share model constructed through cluster analysis can be used to design and execute performance tests more accurately, making the test scenario closer to the actual application environment and improving the reliability and accuracy of the test; by analyzing the request share of time periods, the load changes in different time periods can be understood, which helps to predict the future load demand of the system and provide a reference for capacity planning and resource scheduling. In performance testing, the performance test plan can be designed based on the predicted load demand to evaluate the scalability and stability of the system under future loads.
[0105] The peak proportion model based on clustering analysis proposed in this embodiment, on the basis of the commonly used load calculation model based on transaction proportion, adds the business factors associated with the measured scenario as weights, and introduces the clustering analysis method at the same time, so as to more accurately obtain the transaction proportion within the peak range, which helps to determine a more accurate load level.
[0106] At least the second target data is matched with rules through the rule tree structure, and a data distribution model is constructed based on the matching results. The logical structure of rules and request events is as follows:
[0107] Rule tree <business type, request identifier, attribute list, rule set>;
[0108] Request event <business type, request identifier, attribute value list>;
[0109] Attribute value list {attribute1, attribute2, ..., attributeen};
[0110] attribute1{attribute name, attribute value};
[0111] The data distribution model is used to characterize the distribution ratio of different data features, such as Figure 5 As shown, the specific steps for generating the test data set are as follows:
[0112] Step 1: Preprocessing rules: Collect business type, request identifier, and attribute list according to business characteristics, and split each attribute into multiple sub-rules to form a rule set.
[0113] Step 2: Generate a rule tree: Take each feature as an internal node of the rule tree. A path of the rule tree is a complete rule, and the leaf node of each path records the number of successful matches of each rule.
[0114] Step 3: Rule tree matching: Parse the request message of data source B to obtain the request event, perform layer-by-layer search and matching for each request event according to the fixed format and rule tree, and record the matching count. Request messages that do not match the rules are counted according to the default path.
[0115] Step 4: Rule tree optimization: Set the maximum threshold and minimum threshold. When the matching count in the matching result is less than the minimum threshold, the path can be considered unimportant and can be trimmed. When the matching count is greater than the maximum threshold, the rule may not fully cover the demand or there may be over-matching. It is necessary to readjust the rules of the path and split it into multiple sub-rules to improve decision accuracy and generalization ability.
[0116] Step 5: Generate a data set: Based on the rule tree matching results, use the coverage information of the rule tree to automatically generate a data set that conforms to the data distribution characteristics.
[0117] Specifically, Figure 6 This is a rule tree logic example diagram, such as Figure 6 As shown:
[0118] Rule tree <risk business, request identifier 1, <institution, date>, {first-level institution: {recent year, recent two years...}, second-level institution: {recent year, recent two years...}}>;
[0119] Request event <Risk business, request identifier 1, <Institution A, 20230717>>;
[0120] It can be seen that by establishing a rule tree model based on business needs, matching rules on real-scene request messages through layer-by-layer tree retrieval, analyzing data distribution based on matching results and automatically generating data sets, on the one hand, a large amount of test data close to real scenarios can be quickly generated, making the data preparation steps for performance testing more efficient; on the other hand, due to the diversity of generated data, different scenarios and extreme situations in actual use can be covered, which can better help discover potential problems in system performance, thereby conducting more comprehensive performance testing.
[0121] The data distribution model proposed in this embodiment is a test data set generation method based on rule tree matching, which combines request messages and business analysis to generate a rule tree for retrieving data scenarios that match the interface under test. Different from the existing methods that often use a set of test data or individual field randomization methods, the method described in this application generates a variety of random data combinations based on the user's real business scenarios, which greatly improves the test data coverage scenarios and improves the richness and coverage of performance test scenarios.
[0122] In summary, the embodiment of the present application realizes the automated, intelligent analysis and construction of the transaction proportion model and the data distribution model through scripts. The rule tree is used to enrich the test data, so that the coverage of business scenarios is higher during the implementation of performance testing. And in the complex scenario of multiple business factors, the parameters can be configured according to the actual business scenario to realize the dynamic construction and update of the model, so that the performance test model construction is more efficient, accurate, and richer in data, which is more in line with the real business scenario of production, and the overall efficiency of performance testing is greatly improved.
[0123] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0124] The embodiment of the present application also provides a kind of service performance test device based on rule tree and clustering. It should be noted that the service performance test device based on rule tree and clustering of the embodiment of the present application can be used to perform the service performance test method based on rule tree and clustering provided by the embodiment of the present application. The device is used to implement the above-mentioned embodiment and preferred implementation mode, and the description has been made without repeating. As used below, the term "module" can implement the combination of software and / or hardware of the predetermined function. Although the device described in the following embodiments is preferably implemented with software, the implementation of hardware, or the combination of software and hardware is also possible and conceived.
[0125] The following introduces a service performance testing device based on rule tree and clustering provided in an embodiment of the present application.
[0126] Figure 7 1 is a structural block diagram of a service performance testing device based on rule tree and clustering according to an embodiment of the present application. Figure 7 As shown, the device includes: a first acquisition unit 10, a first processing unit 20, a second processing unit 30, and a first construction unit 40.
[0127] The first acquisition unit 10 is used to acquire online log data to obtain first target data, determine online request data corresponding to the interface to be tested based on the interface to be tested, and obtain second target data;
[0128] The first processing unit 20 is used to segment the first target data based on the timestamp to obtain multiple third target data, cluster the third target data of different transaction types based on the transaction volume to obtain clustering results, and build a peak share model according to the clustering results, where the peak share model is used to characterize the load distribution of different transaction types;
[0129] A second processing unit 30 is used to perform rule matching on at least the second target data through a rule tree structure, and to construct a data distribution model based on the matching result, wherein the data distribution model is used to characterize the distribution ratio of different data features;
[0130] The first construction unit 40 is used to construct a target test scenario according to the peak share model, generate performance test data according to the data distribution model, and test the interface to be tested according to the performance test data in the target test scenario.
[0131] It can be seen that the embodiment of the present application provides a business performance testing device based on rule tree and clustering. By acquiring real online log data and request data, and combining time series partitioning and clustering analysis, an accurate peak share model is constructed to characterize the load distribution of different transaction types in different time periods. At the same time, the request data is matched by rules through the rule tree structure, and a data distribution model is constructed to accurately characterize the distribution ratio of different data features. Based on these two models, more realistic and diverse test scenarios and test data can be generated, thereby overcoming the problems of low test data coverage, single test scenarios, and inability to accurately simulate user behavior in the prior art.
[0132] As a possible implementation manner, the first processing unit includes:
[0133] A parsing module, configured to parse a timestamp, a transaction type, and a transaction amount in the first target data based on the first target data to obtain first identification data, second identification data, and third identification data;
[0134] The division module is used to divide the first target data according to the first identification data with the first preset time length as a basic unit to obtain a plurality of third target data.
[0135] As a possible implementation manner, the first processing unit includes:
[0136] A sorting module, used for sorting the third target data from large to small according to the third identification data to obtain a target sequence;
[0137] An interception module is used to construct a peak interval window according to a second preset time length, and intercept the target sequence based on the peak interval window to obtain fourth target data, the second preset time length is greater than the first preset time length, and the peak interval window is used to start with the maximum value of the third identification data and intercept the third target data within the second preset time length;
[0138] A splitting module, used for taking the first preset time length as the minimum sample unit and performing clustering based on the second identification data and the third identification data to obtain at least one target cluster;
[0139] The determination module is used to calculate the transaction volume mean of the samples in the cluster based on the target cluster, calculate the transaction volume proportion of different second identification data in the target cluster corresponding to the maximum value of the transaction volume mean, and construct a peak proportion model based on the transaction volume proportion.
[0140] As a possible implementation method, the service performance testing device based on rule tree and clustering also includes:
[0141] A determination unit, configured to determine a corresponding architecture complexity and data volume level based on the second identification data;
[0142] A correction weight determination unit, configured to perform weighted calculation based on the architecture complexity and the data volume level to obtain a transaction correction index, and determine a correction weight based on the architecture complexity and the data volume level when the transaction correction index is greater than or equal to a preset value, wherein the correction weight is proportional to the architecture complexity and the correction weight is proportional to the data volume level;
[0143] The correction unit is used to correct the third identification data based on the correction weight, and re-sort the corrected third identification data from large to small to update the target sequence.
[0144] As a possible implementation manner, the second processing unit includes:
[0145] A fifth target data confirmation module, configured to extract a service feature included in the service request based on the first target data to obtain fifth target data, wherein the service feature includes at least one of a requested service type, a request identifier, a service frequency, a service cycle, and a service rule;
[0146] A target rule tree determination module, used for taking each fifth target data as a layer node of the rule tree to obtain a target rule tree;
[0147] a sixth target data confirmation module, configured to extract request event data corresponding to the business request based on the second target data to obtain sixth target data;
[0148] A matching module, used for performing layer-by-layer matching based on the sixth target data and the target rule tree, and recording matching counts of leaf nodes of the target rule tree;
[0149] The path matching determination module is used to determine that the current path matching of the target rule tree passes when the matching count is greater than or equal to a first threshold and less than or equal to a second threshold.
[0150] As a possible implementation method, the service performance testing device based on rule tree and clustering also includes:
[0151] A path pruning unit, configured to prune a current path of a target rule tree when the corresponding matching count is less than a first threshold;
[0152] The matching module is used to split each layer of the target rule tree when the matching count is greater than the second threshold, and re-match the sixth target data according to the split target rule tree, wherein the branches of the same layer of the split target rule tree are greater than the target rule tree before the split.
[0153] As a possible implementation, the first building unit includes:
[0154] A total match count determination module is used to calculate the sum of the match counts of the target rule tree corresponding to the data distribution model to obtain the total match count;
[0155] A target coverage determination module, used to calculate the ratio of the total matching count to the total data volume of the first target data to obtain the target coverage;
[0156] The test data set generation module is used to generate a test data set based on each path when the target coverage is greater than a third threshold, wherein the proportion of sample data corresponding to each path in the test data set is consistent with the ratio of the matching count corresponding to the path to the total matching count.
[0157] The above-mentioned service performance testing device based on rule tree and clustering includes a processor and a memory, and the above-mentioned first acquisition unit, first processing unit, second processing unit, first construction unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions. The above-mentioned modules are all located in the same processor; or, the above-mentioned modules are located in different processors in the form of any combination.
[0158] The processor contains a kernel, which calls the corresponding program unit from the memory. One or more kernels can be set, and the kernel parameters can be adjusted to solve the problem of low test data coverage and single test scenario in the prior art when performing business system performance testing, resulting in errors in the test results of the interface.
[0159] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0160] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute a business performance testing method based on a rule tree and clustering.
[0161] An embodiment of the present invention provides a business system, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for executing a business performance test based on a rule tree and clustering.
[0162] An embodiment of the present invention provides a processor, which is used to run a program, wherein a service performance testing method based on a rule tree and clustering is executed when the program is running.
[0163] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. The device in this article may be a server, a PC, a PAD, a mobile phone, etc. When the processor executes the program, it implements at least the steps of a business performance testing method based on a rule tree and clustering.
[0164] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program for initializing a method for testing business performance based at least on a rule tree and clustering.
[0165] Obviously, those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0166] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0167] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0168] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0170] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0171] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0172] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0173] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0174] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0175] 1) The business performance testing method based on rule tree and clustering of the present application constructs a test data distribution model and a transaction model by automatically analyzing online request messages and log files, and automatically generates test data based on this. Combined with the load model and business analysis, the real scenario can be simulated more accurately, thereby comprehensively evaluating the system performance. The test data distribution adopts the rule tree matching method to efficiently process large-scale data and complex rule sets, and flexibly maintain the rules. Transaction proportion analysis extracts transaction data through logs, uses cluster analysis to identify peak proportion models and trends, discovers the impact of outlier transactions on performance, and improves system stability. The present application combines real log data, time series division and cluster analysis to construct an accurate peak proportion model to characterize the load distribution; by matching request data with rule trees, an accurate data distribution model is generated. The combination of the two generates real and diverse test scenarios, which solves the problems of low test data coverage, single test scenarios and inability to accurately simulate user behavior in the prior art.
[0176] 2) The business performance testing device based on rule tree and clustering of the present application includes: a first acquisition unit, a first processing unit, a second processing unit, and a first construction unit. The device constructs an accurate peak share model by acquiring real online log data and request data, and combining time series division and cluster analysis, and characterizes the load distribution of different transaction types in different time periods. At the same time, the request data is matched by rules through the rule tree structure, and a data distribution model is constructed to accurately characterize the distribution ratio of different data features. Based on these two models, more realistic and diverse test scenarios and test data can be generated, thereby overcoming the problems of low test data coverage, single test scenarios, and inability to accurately simulate user behavior in the prior art.
[0177] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A service performance testing method based on rule tree and clustering, characterized in that: include: Acquire online log data to obtain first target data, determine online request data corresponding to the interface to be tested based on the interface to be tested, and obtain second target data; Segmenting the first target data based on timestamps to obtain multiple third target data, clustering the third target data of different transaction types based on transaction volumes to obtain clustering results, and constructing a peak share model based on the clustering results, wherein the peak share model is used to characterize load distribution of different transaction types; Performing rule matching on at least the second target data through a rule tree structure, and constructing a data distribution model based on the matching result, wherein the data distribution model is used to characterize the distribution ratio of different data features; A target test scenario is constructed according to the peak share model, performance test data is generated according to the data distribution model, and the interface to be tested is tested according to the performance test data under the target test scenario.
2. The method according to claim 1, characterized in that: The first target data is segmented based on the timestamp to obtain a plurality of third target data, including: parsing the timestamp, the transaction type, and the transaction amount in the first target data based on the first target data to obtain first identification data, second identification data, and third identification data; Taking the first preset time length as a basic unit, the first target data is divided according to the first identification data to obtain a plurality of the third target data.
3. The method according to claim 2, characterized in that Clustering the third target data of different transaction types based on transaction volume to obtain clustering results, and constructing a peak proportion model according to the clustering results, including: Sort the third target data from large to small according to the third identification data to obtain a target sequence; Constructing a peak interval window according to a second preset duration, and intercepting the target sequence based on the peak interval window to obtain at least one fourth target data, wherein the second preset duration is greater than the first preset duration, and the peak interval window is used to intercept the third target data within the second preset duration starting from the maximum value of the third identification data; Taking the first preset time length as the minimum sample unit, and performing clustering based on the second identification data and the third identification data, to obtain at least one target cluster; The transaction volume mean of the samples in the cluster is calculated based on the target cluster, and the transaction volume proportion of different second identification data in the target cluster corresponding to the maximum value of the transaction volume mean is calculated, and the peak proportion model is constructed according to the transaction volume proportion.
4. The method according to claim 3, characterized in that After sorting the third target data from large to small according to the third identification data to obtain a target sequence, the method further includes: Determine the corresponding architecture complexity and data volume level based on the second identification data; Performing weighted calculation based on the architecture complexity and the data volume level to obtain a transaction correction index, and determining a correction weight based on the architecture complexity and the data volume level when the transaction correction index is greater than or equal to a preset value, wherein the correction weight is proportional to the architecture complexity and the correction weight is proportional to the data volume level; The third identification data is modified based on the modified weight, and the modified third identification data is re-sorted from large to small to update the target sequence.
5. The method according to claim 1, characterized in that Performing rule matching on at least the second target data through a rule tree structure includes: Extracting a service feature included in the service request based on the first target data to obtain fifth target data, wherein the service feature includes at least one of a requested service type, a request identifier, a service frequency, a service cycle, and a service rule; Taking each of the fifth target data as a layer node of the rule tree, to obtain a target rule tree; Extracting request event data corresponding to the business request based on the second target data to obtain sixth target data; Perform layer-by-layer matching based on the sixth target data and the target Rule Tree, and record the matching counts of leaf nodes of the target Rule Tree; When the match count is greater than or equal to the first threshold and less than or equal to the second threshold, it is determined that the current path of the target rule tree matches successfully.
6. The method according to claim 5, characterized in that After performing layer-by-layer matching based on the sixth target data and the target rule tree and recording the matching counts of leaf nodes of the target rule tree, the method further includes: When the matching count is less than the first threshold, pruning the current path of the target rule tree; When the match count is greater than the second threshold, each layer of the target rule tree is split, and the split target rule tree is re-matched with the sixth target data, wherein the branches of the same layer of the split target rule tree are larger than the target rule tree before the split.
7. The method according to claim 6, characterized in that Generating performance test data according to the data distribution model includes: Calculate the sum of the matching counts of the target rule tree corresponding to the data distribution model to obtain a total matching count; Calculate the ratio of the total matching count to the total data volume of the first target data to obtain a target coverage rate; When the target coverage is greater than the third threshold, a test data set is generated based on each path, wherein the ratio of sample data corresponding to each path in the test data set is consistent with the ratio of the matching count corresponding to the path to the total matching count.
8. A service performance testing device based on rule tree and clustering, characterized in that: The device comprises: A first acquisition unit is used to acquire online log data to obtain first target data, determine online request data corresponding to the interface to be tested based on the interface to be tested, and obtain second target data; a first processing unit, configured to segment the first target data based on a timestamp to obtain a plurality of third target data, cluster the third target data of different transaction types based on transaction volumes to obtain clustering results, and construct a peak share model according to the clustering results, wherein the peak share model is used to characterize load distribution of different transaction types; A second processing unit, configured to perform rule matching on at least the second target data through a rule tree structure, and to construct a data distribution model based on the matching result, wherein the data distribution model is used to characterize the distribution ratio of different data features; The first construction unit is used to construct a target test scenario according to the peak share model, generate performance test data according to the data distribution model, and test the interface to be tested according to the performance test data under the target test scenario.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
10. A business system, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of claims 1 to 7.