Data test method and device, computer program product and data generation system

By generating test data that conforms to the distribution type of the source data using the Kolmogorov-Smirnov algorithm and the RVS generation function, the problem of mismatch between test data and source data in existing technologies is solved, enabling testing that is closer to actual business scenarios and improving the practicality of the data and the effectiveness of system testing.

CN120408114APending Publication Date: 2025-08-01中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510493575.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing technologies generate test data based on fixed data models that cannot correspond to the distribution of source data. This results in generated data that does not closely resemble the real environment, lacks flexibility and scalability, is difficult to cope with diverse business scenarios, and has poor practicality.

Method used

The Kolmogorov-Smirnov algorithm was used to identify the distribution type of the source dataset, and the RVS generation function was used to generate a standard test dataset that conforms to the distribution type. The dataset was then tested using automated testing tools to ensure that the distribution type of the data is consistent with the source data.

Benefits of technology

This improves the usability of test data, enabling more accurate simulation of real-world business scenarios, enhancing the effectiveness and accuracy of system testing, and helping financial institutions optimize their business processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408114A_ABST
    Figure CN120408114A_ABST
Patent Text Reader

Abstract

The invention provides a data testing method and device, a computer program product and a data generation system. The method comprises the steps of obtaining a source data set; according to a Kolmogorov-Smirnov algorithm, the distribution type of the source data set is determined; an rvs generation function is adopted, a standard test data set conforming to the distribution type is generated according to the source data set, and the data size of the standard test data set is larger than that of the source data set; performing automatic testing by adopting the standard test data set to obtain a test result; and sending the test result to a financial institution, so that the financial institution performs business optimization based on the test result. According to the scheme, the problems that in the prior art, test data is generally generated based on a fixed data model, distribution corresponding to source data cannot be achieved, and the practicability of the generated data is poor are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and in particular, to a method for testing data, an apparatus for testing data, a computer program product, and a data generation system. Background Art

[0002] With the rapid development of banking business, financial products based on hierarchical structures have become increasingly complex, such as fund pool products and virtual ledgers. These products usually have the characteristics of nesting and strong correlation. For example, the multi-level nested structure in fund pool products and the intricate correlation relationships of virtual ledgers based on entity accounts. At the same time, financial products with hierarchical structures integrate a large number of business rules, and each business entity or node has different-dimensional business rule manifestations. For example, fund pool products have node types, fund collection methods, fund collection directions, payment control rules, etc.

[0003] To meet the business test requirements, it is necessary to quickly construct test data that conforms to specific data distributions and business rules. Existing methods usually rely on fixed data models, cannot correspond to the source data distribution, the generated data cannot be close to the real environment, lack flexibility and scalability, and are difficult to cope with diverse business scenarios, with poor practicability. Summary of the Invention

[0004] The main purpose of the present application is to provide a method for testing data, an apparatus for testing data, a computer program product, and a data generation system, so as to at least solve the problem that in the prior art, test data is usually generated based on a fixed data model, cannot correspond to the source data distribution, and the practicability of the generated data is poor.

[0005] To achieve the above object, according to one aspect of the present application, a method for testing data is provided, including: obtaining a source data set, where the source data set is an initial data set for a financial institution to handle business, and the source data set includes at least one or more of the number of business levels, the number of first accounts for handling each business, the number of second accounts for handling the same business level for the business, the time for handling the business, and the number of business transactions. There are multiple businesses in the financial institution, and each business includes at least one sub-business, and the business level is the hierarchical relationship between the business and the sub-business; determining the distribution type of the source data set according to the Kolmogorov-Smirnov algorithm; using the rvs generation function to generate a standard test data set that conforms to the distribution type according to the source data set, where the data volume of the standard test data set is greater than the data volume of the source data set; performing an automated test using the standard test data set to obtain a test result; and sending the test result to the financial institution so that the financial institution optimizes its business based on the test result.

[0006] Optionally, according to the Kolmogorov-Smirnov algorithm, determine the distribution type of the source data set, including: constructing a distribution verification model, where the distribution verification model is trained by the Kolmogorov-Smirnov algorithm using multiple sets of training data, and each set of training data in the multiple sets of training data includes a historical source data set obtained within a historical time period and the historical distribution type corresponding to the historical source data set; input the source data set into the distribution verification model to obtain the distribution type corresponding to the source data set.

[0007] Optionally, constructing a distribution verification model includes: constructing an initial distribution verification model, where the initial distribution verification model is trained by the Kolmogorov-Smirnov algorithm using multiple sets of training data, and each set of training data in the multiple sets of training data includes the historical source data set obtained within a historical time period and the historical initial distribution type corresponding to the historical source data set; optimizing the initial distribution verification model to obtain the distribution verification model, where the optimization methods include one or more of adjusting the learning rate, Bayesian optimization, regularization, data dimensionality reduction, data augmentation, and Adam optimization.

[0008] Optionally, perform automated testing using the standard test data set to obtain a test result, including: generating service data according to the standard test data set, where the service data is data with service operation logic; using an automated testing tool to perform automated testing according to the service data to obtain the test result, where the automated testing tool includes one or more of Selenium, Appium, Katalon Studio, and Postman.

[0009] Optionally, generating service data according to the standard test data set includes: obtaining a first service feature, where the first service feature is data representing a feature predefined related to the logic or rules of the service; combining the standard test data set and the first service feature to obtain a first multi-dimensional data set; generating first service data according to the first multi-dimensional data set, where the algorithms used for data generation include one or more of decision tree, generative adversarial network, Bayesian, and K-means.

[0010] Optionally, business data is generated according to the standard test data set, including: randomly generating second business features by using a business generation algorithm, where the business generation algorithm includes one or more of an RNN algorithm, an LSTN algorithm, an Apriori algorithm, and an FP-growth algorithm, and the second business features are data representing features randomly generated related to the logic or rules of the business; combining the standard test data set and the second business features to obtain a second multi-dimensional data set; generating second business data according to the second multi-dimensional data set, where the algorithms used for data generation include one or more of a decision tree, a generative adversarial network, Bayesian, and K-means.

[0011] Optionally, a standard test data set conforming to the distribution type is generated according to the source data set by using an rvs generation function, including: generating an initial standard test data set conforming to the distribution type according to the source data set by using an rvs generation function; preprocessing the initial standard test data set to obtain the standard test data set, where the preprocessing methods include one or more of deleting duplicate values, removing null data, and data normalization.

[0012] According to another aspect of the present application, a data testing device is provided, including: an acquisition unit configured to acquire a source data set, where the source data set is an initial data set for a financial institution to handle business, and the source data set includes at least one or more of the number of business levels, the number of first accounts for handling each business, the number of second accounts for handling the business at the same business level, the time for handling the business, and the number of business transactions. There are multiple businesses in the financial institution, each business includes at least one sub-business, and the business level is the hierarchical relationship between the business and the sub-business; a determination unit configured to determine the distribution type of the source data set according to the Kolmogorov-Smirnov algorithm; a generation unit configured to generate a standard test data set conforming to the distribution type according to the source data set by using an rvs generation function, where the data volume of the standard test data set is greater than the data volume of the source data set; a testing unit configured to perform an automated test by using the standard test data set to obtain a test result; and a sending unit configured to send the test result to the financial institution so that the financial institution optimizes its business based on the test result.

[0013] According to still another aspect of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps of any one of the data testing methods are implemented.

[0014] According to another aspect of the present application, there is provided a data generation system, including: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a method for performing any one of the data testing methods.

[0015] Applying the technical solution of the present application, the Kolmogorov-Smirnov (K-S) algorithm is an algorithm used to test whether a sample set conforms to a certain data distribution. Therefore, this solution uses the Kolmogorov-Smirnov (K-S) algorithm to identify the distribution type of the source data set. The rvs function is a function used to generate random samples from a specified distribution. Therefore, after identifying the distribution type of the source data set, this solution generates a random standard test data set through the rvs function in combination with the distribution type of the source data set. In this way, the randomly generated standard test data set has the same distribution type as the source data set, solving the problem of the mismatch between the test data distribution and the source data distribution in the prior art and improving the practicality of the data. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The specification drawings forming a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0017] Figure 1 It shows a hardware structure block diagram of a mobile terminal for performing a data testing method provided in an embodiment of the present application;

[0018] Figure 2 It shows a schematic flowchart of a data testing method provided in an embodiment of the present application;

[0019] Figure 3 It shows a schematic flowchart of identifying the distribution type of the source data set;

[0020] Figure 4 It shows a schematic flowchart of generating a standard test data set;

[0021] Figure 5 It shows a schematic diagram of business logic orchestration;

[0022] Figure 6 It shows a schematic diagram of a virtual ledger structure;

[0023] Figure 7 It shows a schematic diagram of a fund pool structure;

[0024] Figure 8 It shows a schematic diagram of fund pool system parameters;

[0025] Figure 9 Shows a schematic diagram of generating a cube;

[0026] Figure 10 Shows a schematic flowchart of a test method for another type of data;

[0027] Figure 11 Shows a structural block diagram of a test device for a type of data provided according to an embodiment of the present application.

[0028] Among them, the above-mentioned drawings include the following reference numerals:

[0029] 102, a processor; 104, a memory; 106, a transmission device; 108, an input / output device. Detailed implementation manners

[0030] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0031] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances so as to describe the embodiments of the present application here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0033] For the convenience of description, some nouns or terms related to the embodiments of the present application are described below:

[0034] ZIP data distribution: The full name of the ZIP data distribution model is Zero-Inflated Poisson (ZIP) distribution, which combines the Poisson distribution and zero inflation. Data can be generated in two ways: one is to generate zero values with a certain probability, and the other is to generate non-zero values with a Poisson distribution.

[0035] ZINB data distribution: The full name of the ZINB data distribution model is Zero-Inflated Negative Binomial (ZINB) distribution, which combines the negative binomial distribution and zero inflation. The data can also be generated in two ways: one is to generate zero values with a certain probability, and the other is to generate non-zero values with the negative binomial distribution.

[0036] Kolmogorov-Smirnov test method: A non-parametric test method used to test whether a sample comes from a specific distribution. This test is carried out by comparing the difference between the empirical distribution function of the sample data and the theoretical distribution function (hypothetical distribution). This method is applicable to both continuous and discrete data and is a commonly used distribution fitting test method.

[0037] rvs (Random Variates Sampling) method: A random variable sampling method in the scipy.stats library, which is inherited by different probability distribution classes, allowing users to generate a set of random numbers according to the selected probability distribution. rvs is based on the corresponding probability density function (PDF) or probability mass function (PMF) and techniques such as inverse transform sampling and stratified sampling to extract sample points from the given probability distribution.

[0038] As introduced in the background art, in the prior art, test data is usually generated based on a fixed data model, which cannot correspond to the source data distribution, and the generated data has poor practicability. To solve the above problems, the embodiments of the present application provide a test method for data, a test device for data, a computer program product, and a data generation system.

[0039] Next, the technical solutions in the embodiments of the present invention will be described clearly and completely in conjunction with the accompanying drawings in the embodiments of the present invention.

[0040] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a test method of data in the embodiments of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in Figure 1The structure shown is only schematic and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than those shown in Figure 1 or have a different configuration from that shown in Figure 1 .

[0041] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the display method of device information in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include the wireless network provided by the communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0042] In this embodiment, a method for testing data running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that here.

[0043] Figure 2 is a flowchart diagram of a method for testing data according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:

[0044] Step S201: Obtain the source dataset. The source dataset is the initial dataset for the financial institution to handle business. The source dataset includes at least one or more of the number of business levels, the number of first accounts for handling each business, the number of second accounts for handling the business at the same business level, the time for handling the business, and the number of business transactions. There are multiple such businesses in the financial institution, and each business includes at least one sub-business. The business level is the hierarchical relationship between the business and the sub-business.

[0045] Specifically, the source dataset is collected from the operation system of the financial institution and contains historical business data and account information. For example, for a commercial bank, the source dataset may include the number of accounts for business types such as loans, deposits, and transfers, as well as information such as the distribution of these businesses at different levels (such as headquarters, branches, and sub-branches), transaction timestamps, and transaction volumes. In this way, the original dataset provides a complete picture of the business activities of the financial institution, especially the hierarchical relationship of the business and the distribution of the number of accounts.

[0046] The number of business levels: This refers to the number of levels in the internal business processing structure of the financial institution. In a financial institution, businesses are usually organized into a multi-level structure, such as headquarters, branches, departments, etc. Each level is responsible for different business types or business processing links. Counting the number of business levels can help understand the organizational structure and business process complexity of the financial institution.

[0047] The number of first accounts for handling each business: This usually refers to the initial number of accounts directly related to each business. For example, for a loan business, the first account may be the account of the lender; for a deposit business, the first account may be the account of the depositor. This parameter reflects the base number of participating accounts for different business types and helps evaluate the activity and scale of the business.

[0048] The number of second accounts for handling the business at the same business level: This parameter refers to the number of other accounts participating in handling the same business at the same business level. It emphasizes that at the same level (such as the same branch or department), in addition to the first accounts directly related to the business, how many accounts are indirectly involved or affected by this business. For example, in a fund pool product, the number of sub-accounts under the fund pool is the number of second accounts, which helps understand the horizontal influence range of the business.

[0049] The time for handling the business: This refers to the specific timestamp when the business occurs, including the initiation time, processing time, and completion time of the business. Time information is crucial for understanding the seasonality, periodicity, and real-time nature of the business and is a key parameter for evaluating system performance and optimizing resource allocation.

[0050] Number of business transactions: This refers to the number of specific business transactions that occur within a certain period of time (such as a day or a month). The number of business transactions reflects the activity and frequency of the business, and is extremely important for testing the processing capacity of the system and stress testing. For example, a high transaction volume may indicate that the system needs to be optimized to handle high concurrent requests.

[0051] Step S202, determine the distribution type of the above source data set according to the Kolmogorov-Smirnov algorithm;

[0052] Specifically, use the Kolmogorov-Smirnov test method to evaluate whether each dimension or feature in the source data set follows a certain specific probability distribution. For example, for the distribution of the number of accounts, use the K-S test to determine whether it conforms to the Poisson distribution or the negative binomial distribution. For the transaction time, test whether it follows a uniform distribution. This process ensures that the data distribution of each dimension can be accurately identified, which helps to generate test data that is more in line with the actual data characteristics in the subsequent process.

[0053] Step S203, use the rvs generation function to generate a standard test data set that conforms to the above distribution type according to the above source data set, where the data volume of the above standard test data set is greater than the data volume of the above source data set;

[0054] Specifically, according to the distribution type determined in step S202, use the.rvs() function in the scipy.stats library of Python to generate a large-scale test data set that conforms to this distribution. For example, if the number of accounts follows the negative binomial distribution, then use the rvs() function of the negative binomial distribution to generate more data points of the number of accounts to cover more possible business scenarios and ensure that the data volume is sufficient for comprehensive testing.

[0055] Step S204, conduct an automated test using the above standard test data set to obtain a test result;

[0056] Specifically, use the generated standard test data set in an automated test script or test platform to simulate the scenarios of a financial institution handling various businesses, such as large-scale loan applications, high-frequency deposit and withdrawal operations, complex fund transfers, etc. Through automated testing, the performance of the financial institution's information system in the face of large data volumes and complex business rules can be quickly evaluated, and key indicators such as the system's response time, throughput, and error rate can be obtained.

[0057] Step S205, send the above test result to the above financial institution so that the above financial institution can optimize its business based on the above test result.

[0058] Specifically, after the test, the test results (including performance metrics, discovered functional issues, and potential security vulnerabilities) are compiled into a report and sent to the relevant teams of the financial institution. Based on these feedbacks, the financial institution adjusts the system architecture, optimizes the business logic, or patches the security vulnerabilities to improve the overall performance and stability of the system.

[0059] The existing solution directly generates test data based on the source data. The existing solution simply does not consider how the data distribution of the source data is. Therefore, the distribution type of the generated test data cannot be guaranteed to be the same as that of the source data. In this solution, the Kolmogorov-Smirnov (K-S) algorithm is an algorithm used to test whether a sample set conforms to a certain data distribution. Therefore, this solution uses the Kolmogorov-Smirnov (K-S) algorithm to identify the distribution type of the source data set. The rvs function is a function used to generate random samples from a specified distribution. Therefore, after identifying the distribution type of the source data set in this solution, a random standard test data set is generated through the rvs function in combination with the distribution type of the source data set. In this way, the randomly generated standard test data set has the same distribution type as the source data set, solving the problem of the mismatch between the test data distribution and the source data distribution in the prior art and improving the practicality of the data.

[0060] In the specific implementation process, according to the Kolmogorov-Smirnov algorithm, to determine the distribution type of the above-mentioned source data set, it can be achieved through the following steps: construct a distribution verification model, where the above-mentioned distribution verification model is trained using multiple sets of training data through the Kolmogorov-Smirnov algorithm. Each set of the multiple sets of training data includes a historical source data set obtained within a historical time period and the corresponding historical distribution type of the above-mentioned historical source data set; input the above-mentioned source data set into the above-mentioned distribution verification model to obtain the above-mentioned distribution type corresponding to the above-mentioned source data set.

[0061] In this solution, through a large amount of historical training data, the data distribution laws in different business scenarios are learned, and the appropriate distribution type can be quickly and accurately matched for the new data set. By inputting the source data set into the model, its distribution type can be quickly obtained, improving the efficiency of determining the data distribution type.

[0062] In the process of building the model, multiple sets of historical source datasets and their corresponding historical distribution types are used as training data. The goodness of fit between the dataset and various predefined distribution models is evaluated through the Kolmogorov-Smirnov algorithm. The maximum distance D and the corresponding p-value between the actual data and the theoretical distribution calculated by this algorithm are used to measure the suitability of the distribution model. Through continuous optimization and adjustment, the model can automatically identify the distribution type that best matches the characteristics of the dataset. Once the model is built, when the source dataset is input into the model, the distribution type can be quickly obtained, saving the time and labor costs of the traditional one-by-one distribution test and improving the efficiency and accuracy of data construction.

[0063] During the historical time period, a large amount of business data of financial institutions is collected. These datasets cover various business types and operation instances. For each dataset, it is determined whether it conforms to the preset distribution type through the K-S test and recorded. Using the above historical source datasets and historical distribution types as training samples, a model is trained to enable it to predict the correct data distribution type based on the input source dataset. When the current source dataset is input into the trained distribution verification model, the model will automatically output the distribution type corresponding to the dataset without the need for manual K-S test, greatly improving the efficiency.

[0064] Specifically, as Figure 3 shown, identifying the distribution of the source dataset includes the following steps:

[0065] Build a distribution verification model: Establish a distribution verification model to test the distribution law of each dimension dataset and determine whether it follows a certain power law, lognormal or other special distribution. This model supports common data distributions, including uniform distribution, polynomial distribution, Poisson distribution, negative binomial distribution, geometric distribution, normal distribution, ZIP distribution, and ZINB distribution, etc. Based on the frequency analysis of actual production data, verify and optimize this model to accurately reflect the frequencies and proportions of the eigenvalue of each dimension that appear in actual business.

[0066] Identify the distribution of the source dataset: Since the business feature dataset represents the specific business feature mapping relationship, and finally whether to combine the business feature set, arrange the business logic, and generate business data will be determined according to the requirements, the dataset of the business feature dimension does not need to be identified for data distribution. Traverse the multi-dimensional datasets other than this, and identify the dataset distribution according to the dimension attributes actually represented by the dataset.

[0067] In the distribution verification model, the Kolmogorov-Smirnov test method is used to determine which preset probability distribution model the data sets of each dimension conform to. This method can accept a sample sequence and an expected distribution model, and calculate the Kolmogorov-Smirnov statistic D and the corresponding p-value. The smaller the D value, the closer the actual data is to the theoretical distribution. If the p-value is greater than the significance level (usually 0.05), it indicates that the null hypothesis cannot be rejected, that is, it is considered that the actual data is likely to come from this theoretical distribution. Traverse the data sets of each dimension, perform K-S verification on all preset distribution models, and select the distribution with the highest goodness of fit as the optimal model by comparing the statistic D and the p-value.

[0068] In some embodiments, a distribution verification model is constructed, which can be specifically implemented through the following steps: construct an initial distribution verification model, where the initial distribution verification model is trained by the Kolmogorov-Smirnov algorithm using multiple sets of training data, and each set of training data in the multiple sets of training data includes the historical source data set obtained within the historical time period and the historical initial distribution type corresponding to the historical source data set; optimize the initial distribution verification model to obtain the distribution verification model, where the optimization methods include one or more of adjusting the learning rate, Bayesian optimization, regularization, data dimensionality reduction, data augmentation, and Adam optimization.

[0069] In this solution, on the basis of constructing the initial model, the model can be tuned through various optimization methods, including adjusting the learning rate, Bayesian optimization, regularization, data dimensionality reduction, data augmentation, and using the Adam optimizer, etc. For example, adjusting the learning rate can ensure the model is more stable during training and avoid overfitting; Bayesian optimization can intelligently find the best combination of model parameters; regularization can reduce the complexity of the model and prevent overfitting; data dimensionality reduction can improve the efficiency and accuracy of the model; data augmentation can increase the generalization ability of the model; the Adam optimizer can more effectively adjust the model parameters and speed up the training speed. Through optimization, the model can better adapt to the data distribution in different business scenarios, and can maintain a high recognition accuracy even in scenarios with large data sets or complex distributions.

[0070] The purpose of optimizing the initial model is to improve the performance and accuracy of the model. Various methods can be adopted to enhance the model. For example, the learning rate can be changed to control the learning speed of the model, ensuring that it can converge quickly and avoid overfitting. The Bayesian optimization method automatically adjusts the model parameters without prior knowledge to find the optimal parameter combination. Regularization techniques reduce the complexity of the model by adding penalty terms to the loss function, thereby improving its generalization ability. Data dimensionality reduction techniques can reduce the dimensionality of the dataset, improving the training speed and prediction efficiency of the model. Data augmentation increases the diversity of the training data, enabling the model to have better adaptability when facing new data. Finally, using the Adam optimizer can more effectively adjust the parameters during model training, accelerate the convergence process, and improve the model learning efficiency.

[0071] Specifically, as Figure 4 shown, after identifying the distribution of the source data, a standard test dataset can be generated, including the following two steps:

[0072] Construct an rvs generation function: Establish a general dataset generation function that takes the distribution type, distribution parameters (mean, variance, etc.), and the required number of samples as inputs, and randomly generates a standard test dataset that meets the input conditions as required. The data distributions supported by this function are consistent with the distribution types output by the distribution test model.

[0073] Generate a standard test dataset: Use the scipy.stats library to create an object corresponding to the data distribution, and use the.rvs() method provided by the distribution object to generate a specified number of random samples according to the given data distribution and parameter values.

[0074] In the specific implementation process, the above standard test dataset is used for automated testing to obtain test results, which can be achieved through the following steps: Generate business data based on the above standard test dataset, where the above business data is data with business operation logic; use an automated testing tool to perform automated testing based on the above business data to obtain the above test results, where the above automated testing tool includes one or more of Selenium, Appium, Katalon Studio, and Postman.

[0075] In this solution, automated testing can cover a large number of test cases in a short time, while the business data generated by the standard test dataset ensures the comprehensiveness of the test scenarios and the data quality. This helps financial institutions identify and fix defects in the system in a timely manner, ensuring that new business products can run stably before going online, and improving the quality and market competitiveness of the products.

[0076] Specifically, according to the distribution type and business characteristics of the source dataset, data with actual business operation logic was generated using a standard test dataset. These data are not only statistically similar to actual business data but also logically reflect the relationships between business entities and the complexity of business processes. For example, the multi-level nested structure in the fund pool product and the virtual ledger association relationship based on entity accounts.

[0077] Next, an automated testing tool (such as Selenium, Appium, Katalon Studio, or Postman, etc.) was used to perform automated tests on the generated business data. These tools can automatically execute test cases according to predefined business scenarios and test scripts, simulate user operations and business processes, and evaluate the system's ability to handle large-scale and complex business data. The test results will include but are not limited to key performance indicators such as system response time, throughput, and data processing error rate.

[0078] The generated business data, based on the distribution and business characteristics of the standardized test dataset, ensures the statistical consistency and logical coherence of the test data. Automated testing tools such as Selenium can drive web applications and simulate user behavior; Appium can drive mobile applications for testing; Katalon Studio provides a powerful test automation solution; and Postman is used for API interface testing. The generated business data is input into these tools, the preset test scripts are executed, business operations are simulated, and data such as system response time and processing results are collected to generate a test report.

[0079] Business data refers to the data related to business logic generated by financial institutions in their daily operations, covering various aspects such as customer information, transaction records, account balances, and loan approval status. Business data is the basis for business processing and decision-making analysis in financial institutions, and its authenticity, integrity, and accuracy are crucial for business operations. The actual data generated during the daily operations of an enterprise or financial institution directly records the institution's business activities and transaction details and is the basis for analyzing, monitoring, and managing business processes.

[0080] A standard test dataset is a set of data generated under specific conditions for testing software systems or validating algorithm performance. The design purpose of the standard test dataset is to cover various boundary conditions and abnormal situations that the system may encounter to ensure the functional correctness and performance stability of the system. In the technology of this solution, the standard test dataset is created based on the statistical distribution and characteristics of business data through data augmentation or data generation algorithms and is a simulation dataset for automated testing.

[0081] Suppose a bank is developing a new loan product that involves multiple levels of approval processes and complex customer credit assessment mechanisms. The bank's business data set contains detailed information on all loan applications in the past year, including applicant basic information, application amount, approval results, etc.

[0082] Business data: This part of the data contains actual loan application and approval records. For example, on a certain day, a total of 100 loan applications were received, among which 50 applications had an amount below 100,000, 30 applications were between 100,000 and 500,000, and 20 applications were above 500,000. In terms of approval results, 60 applications were approved and 40 applications were rejected.

[0083] Standard test data set: To test the stability and performance of the new loan product, the R & D team first analyzes the above business data set and determines the distribution types of various indicators. For example, it is found that the number of loan applications follows a Poisson distribution, and the application amount follows a negative binomial distribution. Then, using the.rvs() function and these distribution models, a test data set containing 1000 loan applications is generated. The statistical characteristics of the loan applications (such as the number of applications and the amount distribution) are consistent with the original business data set, but the data volume is magnified ten times. Among these 1000 loan applications, 600 are randomly marked as "approved" according to the actual approval situation of the business data, and the rest are marked as "rejected", and the distribution of the application amount is also simulated through the negative binomial distribution model.

[0084] In some embodiments, according to the above standard test data set, business data is generated, which can be specifically achieved through the following steps: Obtain the first business feature, where the above first business feature is data representing a feature predefined related to the logic or rules of the above business; Combine the above standard test data set and the above first business feature to obtain a first multi-dimensional data set; According to the above first multi-dimensional data set, generate the first business data, where the algorithms used to generate the data include one or more of decision tree, generative adversarial network, Bayesian, and K-means.

[0085] In this solution, by combining the standard test data set and the predefined first business feature, and using advanced machine learning algorithms to generate data, first business data that is closer to reality, more representative, and more in line with the set rules can be obtained.

[0086] The above solution helps to more comprehensively test the system functions in the follow-up, especially those complex operations that rely on business logic and rules, ensuring that the system can operate stably and meet business requirements before going live.

[0087] When generating business data, first ensure that the test data set conforms to the statistical distribution of the business data. Then, introduce the first business features, which represent various aspects of predefined business rules. To generate data that is both statistically consistent and meets business logic, machine learning algorithms are adopted. For example, the decision tree algorithm can determine the data generation path based on business features and generate data that conforms to business rules; the generative adversarial network (GAN) generates data similar to the actual business data distribution through the game process between the generator and the discriminator; the Bayesian algorithm can start from the prior probability and calculate the posterior probability of the generated data in combination with business features to ensure that the data is reasonable in business logic; the K-means algorithm can be used to construct data clusters, with each cluster representing a specific business scenario, and generate data in these clusters to cover various possible business situations.

[0088] Specifically, first define a series of business features based on business rules. For example, in a fund pool product, it may include node type, fund collection method, fund collection direction, payment control rules, etc. These features represent the key attributes and logic of business data and are crucial for constructing test data that conforms to the business scenario.

[0089] Fuse the standardized test data set with the defined first business features to create a first multi-dimensional data set. This data set not only retains the statistical characteristics of the source data but also introduces additional feature dimensions according to business logic, making the test data richer and more diverse and better reflecting the complexity of the actual business scenario.

[0090] Finally, adopt machine learning algorithms such as decision tree, generative adversarial network (GAN), Bayesian algorithm, or K-means clustering algorithm to generate the first business data. These algorithms can effectively capture the internal correlations between data, ensuring that the generated data is not only statistically consistent with the business data set but also reasonable in business logic and can reflect the complex relationships between business entities.

[0091] Specifically, as Figure 5 shown, the first business data can be generated based on the standard test data set according to the choreographed business logic, and an automated call process based on the choreography logic of the first multi-dimensional data set is formulated, following the following choreography process:

[0092] Business entity feature extraction: For example, take the first feature value combination of the first multi-dimensional data set to form the feature set of the business entity, which is used as the basis for constructing the business entity.

[0093] Business logic orchestration: Define the correspondence between business characteristic values and business methods, interfaces, SQL sequences, and parameter sets. According to the input business entity characteristic set, orchestrate the specific business implementation call process, recursively construct test data layer by layer, and generate the first business data that meets the first business characteristic value.

[0094] In the specific implementation process, to generate business data based on the above standard test data set, it can be achieved through the following steps: Randomly generate the second business characteristic using a business generation algorithm, where the above business generation algorithm includes one or more of the RNN algorithm, LSTN algorithm, Apriori algorithm, and FP-growth algorithm. The above second business characteristic is data representing a randomly generated feature related to the logic or rules of the above business; Combine the above standard test data set and the above second business characteristic to obtain a second multi-dimensional data set; Generate the second business data based on the above second multi-dimensional data set, where the algorithms used to generate data include one or more of decision trees, generative adversarial networks, Bayesian, and K-means.

[0095] In this solution, by combining the standard test data set and the randomly generated second business characteristic and using advanced machine learning algorithms to generate data, randomly generated and diverse second business data can be obtained.

[0096] The randomly generated second business characteristic can randomly simulate potential data patterns that conform to the actual business logic through algorithms such as RNN, LSTM, Apriori, and FP-growth. For example, RNN and LSTM can predict the possible future fund flow sequences based on the historical transfer records of the fund pool products; Apriori and FP-growth can find the implicit association rules between different accounts based on a large amount of transaction data and be used to generate the transfer behaviors of virtual account books.

[0097] Subsequently, through the application of decision trees, GANs, Bayesian algorithms, or K-means, these feature data are combined with the standardized data set to generate the second business data. Decision trees can divide different business decision branches according to the combination of feature values, and the generated data can reflect the decision-making processes of different business entities; GAN continuously optimizes the distribution of the generated data through the iterative confrontation between the generator and the discriminator in order to be as consistent as possible with the distribution of actual business data; The Bayesian algorithm can use prior knowledge and combine new feature data to generate test cases that conform to the posterior probability distribution; K-means divides the data set into multiple subsets with similar business characteristics through clustering analysis and generates test data representing the characteristics of each subset respectively.

[0098] Specifically, business generation algorithms such as RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory Network), Apriori algorithm, or FP-growth algorithm can randomly and automatically generate feature data associated with business logic. Specifically, RNN and LSTM algorithms are suitable for the generation of sequence data, such as transfer records of fund flows or trading volumes in time series. They can capture the temporal dependence relationships between data. While Apriori and FP-growth are more adept at discovering frequent item sets and association rules and can be used to mine relationship patterns between different business entities, such as associated transfer behaviors between accounts or cross-selling trends between different product lines.

[0099] Combine the second business features obtained through the business generation algorithm with the standard test data set generated in the previous step to form a second multi-dimensional data set. This data set not only contains data with consistent statistical features but also incorporates randomly generated second business features that conform to specific business logics or rules, further enhancing the diversity and representativeness of the test data.

[0100] Generate second business data based on the second multi-dimensional data set through machine learning algorithms such as decision trees, generative adversarial networks (GANs), Bayesian algorithms, or K-means. These algorithms can help us construct test data sets that meet business rules and statistical distributions. Among them, GAN is particularly suitable for generating high-dimensional and complex-distribution data. Decision trees and Bayesian algorithms are helpful for handling classification and regression tasks, and the generated data can reflect the logical relationships between business entities. K-means is suitable for data clustering analysis and can generate data representing multiple business scenarios.

[0101] The hierarchical structure business data construction method based on multi-dimensional data sets proposed above in this solution aims to meet the test requirements of hierarchical data distribution and complex business features. This method can quickly generate data based on business entities, has high flexibility and scalability, and can adapt to the needs of different types and scales of enterprises in business product data construction. In addition, the constructed data has a wide range of applications, can improve the effectiveness and accuracy of system testing, and thus is conducive to enhancing the quality and market competitiveness of software products.

[0102] Specifically, the hierarchical relationship is as Figure 6 、 Figure 7As shown, in an achievable solution, the data of the hierarchical structure business entity at least includes the number of business entity levels (i.e., the number of business levels), the number of nodes per layer of a single business entity (i.e., the number of first accounts for handling each business), the total number of nodes per layer of all business entities (i.e., the number of second accounts for handling business at the same business level), business feature values (multiple, the first business feature or the second business feature), etc. It may also include the time for handling business, the number of business transactions, and statistical data feature values of production business entities from the above dimensions. The specific explanations are as follows:

[0103] Number of business entity levels: Query and traverse all business entities, determine the maximum hierarchical depth of each business entity, and count the distribution of the number of business entity levels to form a business entity level data set. The number of business entity levels refers to the depth of the business entity in the hierarchical structure, that is, how many levels need to be passed from the highest level to the lowest level. For example, in the bank's fund pool product, a fund pool may contain multiple sub-fund pools, and each sub-fund pool may contain more levels of fund accounts. The number of levels describes the depth of such a hierarchical relationship. The parameters of the fund pool are as Figure 8 shown.

[0104] Number of nodes per layer of a single business entity: Query and traverse all business entities, calculate the number of nodes corresponding to business instances at each layer under each business entity, and form a data set of the number of nodes per layer of a single business entity. The number of nodes per layer of a single business entity represents the number of nodes contained in this business entity at a specific layer. Taking a specific business entity (such as a fund pool) as an example, the number of nodes per layer describes the number of sub-entities contained in this fund pool at each layer. For example, there may be 3 sub-fund pools in the first layer, and each sub-fund pool in the second layer may have 2 next-level fund accounts, etc.

[0105] Total number of nodes per layer of all business entities: Count the total number of nodes per layer of all business entities to form a data set of the total number of nodes per layer of all business entities, and this data set generates a constraint on the total value of each layer of data. The total number of nodes per layer of all business entities focuses on the total number of nodes at a specific layer in the overall hierarchical structure. It is not for a single business entity, but the total number of nodes of all entities in the entire business system or product at a certain specific layer. For example, in the fund pool product, the total number of nodes of all fund pools and sub-fund pools in the first layer, and the total number of nodes in all sub-layers.

[0106] Business characteristic values (multiple): Query the business rules of business entities and express the business characteristics of the instance with specific numerical values. There are many types of business rules for an instance, and a characteristic value of one business dimension is extracted from each business rule to form a business characteristic data set. A business characteristic value refers to the quantitative representation of specific attributes or rules related to a business entity. In hierarchical business data, each entity or node may have different business characteristics, such as the node type of a fund pool product (such as the main fund pool, sub-fund pool, fund account, etc.), the fund collection method (such as daily, monthly, etc.), the fund collection direction (such as upward collection, downward collection, etc.), and the payment control rule (such as the maximum limit, minimum limit, etc.). Business characteristic values can be the numerical representation of these rules, providing more fine-grained business information for data generation and analysis.

[0107] Generate a multi-dimensional data set: As Figure 9 shown, the data structure characteristic set of business entities (business entity hierarchical data set, data set of each layer's nodes of a single business entity, data set of the total nodes of each layer of all business entities) and the business characteristic data set can be combined to construct a multi-dimensional data set MultiDimDataSet (hierarchy, number of nodes in each layer, total number of nodes in each layer, business characteristic 1, business characteristic 2,..., business characteristic n).

[0108] (I) Account management refers to the services provided by financial institutions to build a reasonable account system for customers, scientifically allocate cash resources, and facilitate their timely access to account balances or transaction information. Account management includes account sorting, account system construction, establishment of the parent-child relationship, control of account payment limits, control of account targeted income and expenditure relationships, etc.

[0109] (II) Receivables and payables management refers to various receivables and payables products or services provided by financial institutions to customers through channels such as bank counters, online banking, and direct bank-enterprise connection.

[0110] (III) Liquidity management refers to the products and services provided by financial institutions to help enterprises maintain appropriate liquidity in the daily production and operation process, improve the overall capital efficiency of enterprise customers, and reduce the capital use cost. Liquidity management is the core content of cash management, specifically including products such as single-account fund pools, entity fund pools, virtual fund pools, average fund pools, and intelligent fund pools.

[0111] (IV) Investment and financing management mainly includes investment products and financing products. Investment products include unit time deposits, unit call deposits, agreement deposits, intelligent fund pools, open-end funds, and wealth management products. Financing products include corporate account overdrafts, entrusted loans, and bill financing services.

[0112] (5) Information services refer to the various information services provided by financial institutions to customers relying on counter channels or other electronic channels. These include SMS notifications, statement and receipt printing, and various data inquiries, etc.

[0113] For hierarchical business entities, corresponding hierarchical relationships are constructed. Then, combined with specific business attributes (such as different permissions for hierarchical business entities, different account control methods, different account interest calculation methods, different disposable funds, different income and expenditure relationships, etc.), different test results are obtained under the same hierarchical relationship for different business attributes. We need to conduct corresponding tests according to the actual production business attributes.

[0114] Specifically, this solution abstracts and statistically analyzes the hierarchical structure and business rules, establishes data characteristics in different dimensions, expresses the multi-dimensional data characteristics of business entities using multi-dimensional arrays, applies a distribution verification model to perform probability fitting and standardization on the data set, and then conducts business interface orchestration based on the standardized data set to batch generate business data that meets the requirements.

[0115] In some embodiments, the rvs generation function is adopted to generate a standard test data set that conforms to the above distribution type according to the above source data set. Specifically, it can be achieved through the following steps: using the rvs generation function to generate an initial standard test data set that conforms to the above distribution type according to the above source data set; preprocessing the above initial standard test data set to obtain the above standard test data set, where the preprocessing methods include one or more of deleting duplicate values, removing null data, and data normalization.

[0116] In this solution, the generated initial standard test data set may contain non-standard elements such as duplicate values and null data, which will interfere with subsequent testing and analysis work. Therefore, preprocessing the initial data set to ensure data quality. The preprocessing process includes, but is not limited to, operations such as deleting duplicate values, removing null data, and data normalization. Deleting duplicate values can avoid redundant information in the test data, removing null data ensures the integrity and effectiveness of the data, and avoids test errors caused by data missing; data normalization scales the data to the same range, which is particularly important when using machine learning algorithms to generate business data, and can improve the convergence speed and prediction accuracy of the algorithm. Through the above solution, the effectiveness of the data can be guaranteed better.

[0117] By using the rvs generation function to generate an initial standard test data set and combining it with preprocessing steps, a high-quality standard test data set can be obtained. This not only ensures that the data is consistent with the actual business data in terms of distribution characteristics, but also optimizes the quality and format of the data through preprocessing, improving the effectiveness and accuracy of the test. The implementation of the preprocessing steps ensures the cleanliness and standardization of the test data, reduces the complexity of subsequent data analysis and model construction, makes the test process smoother, and the results more reliable.

[0118] The rvs function is a tool for generating random sample points based on a probability distribution model, which can effectively simulate the statistical characteristics similar to the actual business data. After the source data set undergoes feature statistics and distribution model identification, we use the random sample points generated by the rvs function to construct an initial standard test data set. However, due to the random generation characteristics, the initial data set may contain some samples that do not meet the test requirements, such as duplicate values and null data, which will cause unnecessary interference to the test results. Therefore, we preprocess the initial standard test data set, delete duplicate values and remove null data to ensure that each test data is independent and valid. In addition, data normalization is an important step in data preprocessing. It can convert the data into a unified scale. Especially when dealing with multi-dimensional business data sets, it can avoid certain features having a dominant effect on the model due to dimensional differences, thus ensuring better effectiveness of the model when generating business data.

[0119] In the specific implementation process, after preprocessing the initial standard test data set, the above method further includes the following steps: obtaining the data volume of the preprocessed initial standard test data set. In the case where the data volume of the preprocessed initial standard test data set is less than the preset data volume threshold, a data synthesis algorithm is used to expand the preprocessed initial standard test data set to obtain a standard test data set. The data synthesis algorithm includes one or more of SMOTE, GMM, and GAN.

[0120] In this solution, techniques such as Synthetic Minority Over-sampling Technique (SMOTE), Gaussian Mixture Models (GMM), or Generative Adversarial Networks (GAN) can be used to synthesize more sample data that conforms to business logic and statistical distribution to expand the preprocessed initial standard test data set to ensure sufficient data for subsequent automated testing.

[0121] If the amount of data in the preprocessed dataset is insufficient to meet the needs of system testing, that is, the amount of data is less than a preset data volume threshold (for example, the threshold can be set to 10,000 records, or any other feasible pre-set value), a data synthesis algorithm is used to expand the dataset. The choice of data synthesis algorithm should be based on the characteristics of the dataset and the specific requirements of the test. For example, if there is a certain minority business scenario in the business data, the SMOTE algorithm can be used to balance the number of samples in each category; if the dataset shows a multimodal distribution, GMM can be considered to generate samples that conform to the distribution; and when the data distribution pattern is complex and difficult to directly fit, GAN may be a better choice, which can generate new samples that both follow the distribution rules and have diversity. The application of the above algorithms can ensure that while expanding the amount of data, the diversity and representativeness of the data are maintained, thus constructing a standard test dataset that meets the test requirements.

[0122] In the specific implementation process, an automated testing tool is used to perform automated testing based on the business data to obtain test results, which can be specifically implemented through the following steps: When a new Nth test requirement is received, add a first label to the business data generated in the Nth time, N≥1; store the business data generated in the Nth time and the first label in the first database; use the business data generated in the Nth time in the first database to perform automated testing to obtain the Nth test result corresponding to the business data generated in the Nth time, and store it in the first database; when a test requirement for the N+1th time is received, regenerate the business data, and add a second label to the business data generated in the N+1th time; store the business data generated in the Nth time and the second label in the second database; use the business data generated in the N+1th time in the second database to perform automated testing to obtain the (N+1)th test result corresponding to the business data generated in the N+1th time, and store it in the second database.

[0123] In this solution, when a new test round, that is, the Nth test, is required, a new batch of business data will be generated, and a unique first label will be added to this batch of data. This label is used to identify that this batch of data is specific to the Nth test, which helps to quickly screen and locate this batch of data in the subsequent stage, distinguish it from the data of other test rounds, and ensure the purity of the test data. When facing the test requirement for the next round, that is, the (N+1)th test, a new batch of business data will be generated, and a second label different from the first label will be added to this batch of data to identify that this batch of data is specific to the (N+1)th test, avoiding the influence of residual historical test data on the new test, thereby improving the accuracy and reliability of the test results.

[0124] Specifically, the business data with the first tag will be stored in a pre-set first database, which is specifically used to store test data and related metadata, including tag information. When storing data, it is saved together with the tag information to facilitate quickly finding this batch of data during subsequent testing and data cleaning processes. The automated testing tool will read the business data specific to the Nth test from the first database, execute the corresponding test cases, and then record the test results, including successful test cases, failed cases, and any abnormal situations. These test results will be stored back in the first database again, associated with the corresponding data set, for later analysis and debugging. The (N + 1)th business data with the second tag is stored in a separate second database, which is dedicated to storing the data and metadata associated with the (N + 1)th test, including the test data for this time and the corresponding tags. The automated testing tool will then read the business data for the (N + 1)th test from the second database, execute a new round of tests, and store the results of this test back in the second database. The storage of test results closely associated with the data set facilitates in-depth analysis of the test results later.

[0125] Of course, in the case of not having multiple databases, it is also possible to use the business data generated in the Nth time in the first database for automated testing. After obtaining the test results for the Nth time corresponding to the business data generated in the Nth time, the business data with the first tag can be deleted, or in the case of having a test requirement for the (N + 1)th time, the business data with the first tag is not used for testing, and the business data with the second tag is used for testing.

[0126] The existing solutions mainly generate data based on a fixed data model. When facing complex hierarchical structure data, multi-dimensional business characteristics, and considering the complex distribution of business data at the same time, the data construction method of generating test data based on a fixed data model cannot flexibly handle different businesses or new scenarios. To address this challenge, based on the analysis of data, this solution introduces a design method of multi-dimensional data sets. Through the fitting of the distribution verification model and the standardized data processing, finally, business rule orchestration is carried out based on the multi-dimensional data sets to quickly generate standard test data sets.

[0127] The disadvantages of the existing technologies are that the generated business entity models are fixed, unable to cover diverse business scenarios, and when the business entity models change, the implementation methods need to be adjusted synchronously, resulting in high costs and low efficiency. The purpose of this solution is to solve these problems. By abstracting and statistically analyzing the hierarchical structure and business rules, multi-dimensional data characteristics are established, and the distribution verification model is used for fitting and standardization processing, and finally, the rapid generation of business data is realized.

[0128] The scope of this application can meet the data structure requirements of business products with different types and scales of enterprise hierarchical structures, ensuring that the generated data meets the requirements of business scenarios. By accurately simulating the actual distribution of business entities in the production environment, the effectiveness and accuracy of testing are improved, so as to discover and fix potential system performance bottlenecks and functional defects in advance.

[0129] This solution has a wide range of applicable scenarios. For example, it is applicable to the hierarchical structure business data construction scenarios involving multi-dimensional business characteristics in industries such as banking, insurance, and real estate. This method can accurately simulate the actual distribution of business entities in the production environment, improve the effectiveness and accuracy of testing, so as to discover and fix potential system performance bottlenecks and functional defects in advance.

[0130] In order to enable those skilled in the art to more clearly understand the technical solution of this application, the implementation process of the data testing method of this application will be described in detail below in combination with specific embodiments.

[0131] This embodiment relates to a specific data testing method. The hierarchical structure business data construction method based on multi-dimensional data sets is applicable to various data construction scenarios, such as Figure 10 shown, and includes the following steps:

[0132] First, determine whether there is a source data set. In the case of no source data set, directly specify the data distribution type.

[0133] In the case of having a source data set, perform data feature statistics on the source data set to determine whether it needs to be converted into a standard test data set.

[0134] In the case of needing to be converted into a standard test data set, identify the distribution type of the source data set, generate a standard test data set based on the distribution type of the source data set, and determine whether it is necessary to combine the first business feature.

[0135] In the case of needing to combine the first business feature, combine the standard test data set and the first business feature to obtain a first multi-dimensional data set, and generate first business data according to the first multi-dimensional data set.

[0136] In the case of not needing to combine the first business feature, randomly generate a second business feature, combine the standard test data set and the second business feature to obtain a second multi-dimensional data set, and generate second business data according to the second multi-dimensional data set.

[0137] In the case of not needing to be converted into a standard test data set, determine whether it is necessary to combine the first business feature.

[0138] In the case of needing to combine the first business feature, combine the source data set and the first business feature to obtain a third multi-dimensional data set, and generate third business data according to the third multi-dimensional data set.

[0139] Without the need to combine the first business feature, the second business feature is randomly generated, and the source data set and the second business feature are combined to obtain a fourth multi-dimensional data set, and fourth business data is generated based on the fourth multi-dimensional data set.

[0140] Automated testing is performed based on the business data to obtain test results.

[0141] This solution has high flexibility and scalability, can adapt to the data construction requirements of business products with different types and scales of enterprise hierarchical structures, ensures that the generated data meets the requirements of the business scenario, and the constructed data has a wide range of applications. For example, it can improve the effectiveness and accuracy of system testing, and is conducive to enhancing the quality and market competitiveness of software products.

[0142] The method for constructing hierarchical structure business data based on multi-dimensional data sets in this solution can be applied to different data generation scenarios, specifically as follows:

[0143] Based on the source data set: There is no need to perform data standardization conversion, and business data can be directly generated based on the source data set.

[0144] Based on the specified distribution type: In the absence of source data, business data is generated based on the specified data distribution type.

[0145] Combined with the business feature set: After performing distribution recognition on the source data set, data set standardization conversion is performed to generate a standard test data set. Based on the standard test data set, the business feature set is arranged to generate business data.

[0146] To further prove the effectiveness of the above solution, it can be learned from the following tests:

[0147] Performance test: Under the same hardware and software environment, compare the performance of this solution method and the traditional data construction method in terms of data generation speed. The test results show that this solution method is superior to the traditional method in terms of data generation speed, with an average speed increase of 30%.

[0148] Accuracy test: Through comparative analysis with actual business data, verify the performance of the data constructed by this solution method in terms of business rule coverage and data accuracy. The test results show that the data constructed by this solution method has a business rule coverage rate of over 95% and a data accuracy rate of 99%.

[0149] Scalability test: In different business scenarios, evaluate the flexibility and scalability of this solution method in adapting to new business requirements. The test results confirm that this solution method can quickly adapt to new business scenarios without complex adjustments.

[0150] The above test results show that the technical solution of this solution has significant advantages in practical applications and can effectively improve the efficiency and quality of data construction.

[0151] The embodiment of the present application also provides a test device for data. It should be noted that the test device for data in the embodiment of the present application can be used to execute the test method for data provided by the embodiment of the present application. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0152] The following introduces the test device for data provided by the embodiment of the present application.

[0153] Figure 11 is a structural block diagram of a test device for data according to an embodiment of the present application. As Figure 11 shown, the device includes:

[0154] An acquisition unit 10, configured to acquire a source data set, where the source data set is an initial data set for a financial institution to handle business, and the source data set includes at least one or more of the number of business levels, the number of first accounts for handling each of the above-mentioned businesses, the number of second accounts for handling the above-mentioned business at the same above-mentioned business level, the time for handling the above-mentioned business, and the number of business transactions. There are multiple above-mentioned businesses for the financial institution, and each of the above-mentioned businesses includes at least one sub-business, and the above-mentioned business level is the hierarchical relationship between the above-mentioned business and the above-mentioned sub-business;

[0155] A determination unit 20, configured to determine the distribution type of the source data set according to the Kolmogorov-Smirnov algorithm;

[0156] A generation unit 30, configured to use the rvs generation function to generate a standard test data set that conforms to the above distribution type according to the source data set, where the data volume of the standard test data set is greater than the data volume of the source data set;

[0157] A test unit 40, configured to perform an automated test using the standard test data set to obtain a test result;

[0158] A sending unit 50, configured to send the test result to the financial institution so that the financial institution can optimize its business based on the test result.

[0159] The existing solutions directly generate test data based on the source data without considering the data distribution of the source data at all. Therefore, the distribution type of the generated test data cannot be guaranteed to be the same as that of the source data. In this solution, the Kolmogorov-Smirnov (K-S) algorithm is an algorithm used to test whether a sample set conforms to a certain data distribution. Therefore, this solution uses the Kolmogorov-Smirnov (K-S) algorithm to identify the distribution type of the source data set. The rvs function is a function used to generate random samples from a specified distribution. Therefore, after identifying the distribution type of the source data set, this solution generates a random standard test data set through the rvs function in combination with the distribution type of the source data set. In this way, the randomly generated standard test data set has the same distribution type as the source data set, solving the problem of the mismatch between the test data distribution and the source data distribution in the prior art and improving the practicality of the data.

[0160] In the specific implementation process, the determination unit includes a construction module and a first processing module. The construction module is used to construct a distribution verification model. Among them, the above distribution verification model is trained by using multiple groups of training data through the Kolmogorov-Smirnov algorithm. Each group of the multiple groups of training data includes a historical source data set obtained within a historical time period and the historical distribution type corresponding to the historical source data set. The first processing module is used to input the above source data set into the above distribution verification model to obtain the above distribution type corresponding to the source data set.

[0161] In this solution, the data distribution rules in different business scenarios are learned through a large number of historical training data, and the appropriate distribution type can be quickly and accurately matched for the new data set. By inputting the source data set into the model, its distribution type can be quickly obtained, improving the efficiency of determining the data distribution type.

[0162] In some embodiments, the construction module includes a construction sub-module and an optimization sub-module. The construction sub-module is used to construct an initial distribution verification model. Among them, the above initial distribution verification model is trained by using multiple groups of training data through the Kolmogorov-Smirnov algorithm. Each group of the multiple groups of training data includes the above historical source data set obtained within a historical time period and the historical initial distribution type corresponding to the historical source data set. The optimization sub-module is used to optimize the above initial distribution verification model to obtain the above distribution verification model. Among them, the optimization methods include one or more of adjusting the learning rate, Bayesian optimization, regularization, data dimensionality reduction, data augmentation, and Adam optimization.

[0163] In this solution, based on the constructed initial model, various optimization methods can be used for model tuning, including adjusting the learning rate, Bayesian optimization, regularization, data dimensionality reduction, data augmentation, and using the Adam optimizer, etc. For example, adjusting the learning rate can ensure the model is more stable during training and avoid overfitting; Bayesian optimization can intelligently find the optimal combination of model parameters; regularization can reduce the complexity of the model and prevent overfitting; data dimensionality reduction can improve the efficiency and accuracy of the model; data augmentation can increase the generalization ability of the model; the Adam optimizer can more effectively adjust model parameters and speed up the training process. Through optimization, the model can better adapt to the data distribution in different business scenarios and maintain a high recognition accuracy even in scenarios with a large or complex dataset.

[0164] In the specific implementation process, the test unit includes a first generation module and a test module. The first generation module is used to generate business data according to the above standard test dataset, where the above business data is data with business operation logic; the test module is used to perform automated testing according to the above business data using an automated testing tool to obtain the above test results, where the above automated testing tool includes one or more of Selenium, Appium, Katalon Studio, and Postman.

[0165] In this solution, automated testing can cover a large number of test cases in a short time, while the business data generated by the standard test dataset ensures the comprehensiveness of the test scenario and data quality. This helps financial institutions identify and fix defects in the system in a timely manner, ensure the stable operation of new business products before going live, and improve the quality and market competitiveness of the products.

[0166] In some embodiments, the first generation module includes an acquisition sub-module, a first combination sub-module, and a first generation sub-module. The acquisition sub-module is used to acquire first business features, where the above first business features are data representing features predefined related to the logic or rules of the above business; the first combination sub-module is used to combine the above standard test dataset and the above first business features to obtain a first multi-dimensional dataset; the first generation sub-module is used to generate first business data according to the above first multi-dimensional dataset, where the algorithms used for generating data include one or more of decision tree, generative adversarial network, Bayesian, and K-means.

[0167] In this solution, by combining the standard test dataset and the predefined first business features and using advanced machine learning algorithms to generate data, first business data that is closer to reality, more representative, and more in line with the set rules can be obtained.

[0168] In the specific implementation process, the first generation module includes a second generation sub-module, a second combination sub-module, and a third generation sub-module. The second generation sub-module is used to randomly generate second service features by using a service generation algorithm. Among them, the above service generation algorithm includes one or more of the RNN algorithm, the LSTN algorithm, the Apriori algorithm, and the FP-growth algorithm. The above second service features are data representing features related to the logic or rules of the above service randomly generated; the second combination sub-module is used to combine the above standard test data set and the above second service features to obtain a second multi-dimensional data set; the third generation sub-module is used to generate second service data according to the above second multi-dimensional data set. Among them, the algorithms used to generate data include one or more of decision trees, generative adversarial networks, Bayesian, and K-means.

[0169] In this solution, by combining the standard test data set and the randomly generated second service features, and using advanced machine learning algorithms to generate data, random and diverse second service data can be obtained.

[0170] In some embodiments, the generation unit includes a second generation module and a second processing module. The second generation module is used to generate an initial standard test data set that conforms to the above distribution type according to the above source data set by using the rvs generation function; the second processing module is used to preprocess the above initial standard test data set to obtain the above standard test data set. Among them, the preprocessing methods include one or more of deleting duplicate values, removing null data, and data normalization.

[0171] In this solution, the generated initial standard test data set may contain non-standard elements such as duplicate values and null data, which will interfere with subsequent testing and analysis work. Therefore, the initial data set is preprocessed to ensure data quality. The preprocessing process includes, but is not limited to, operations such as deleting duplicate values, removing null data, and data normalization. Deleting duplicate values can avoid redundant information in the test data, and removing null data ensures the integrity and validity of the data, avoiding test errors caused by data missing; data normalization scales the data to the same range, which is particularly important when using machine learning algorithms to generate service data, and can improve the convergence speed and prediction accuracy of the algorithm. Through the above solution, the effectiveness of the data can be guaranteed to be better.

[0172] In the specific implementation process, the above device further includes a data volume acquisition unit and a data expansion unit. The data volume acquisition unit is used to acquire the data volume of the preprocessed initial standard test data set after preprocessing the initial standard test data set. The data expansion unit is used to, when the data volume of the preprocessed initial standard test data set is less than a preset data volume threshold, adopt a data synthesis algorithm to expand the preprocessed initial standard test data set to obtain a standard test data set. The data synthesis algorithm includes one or more of SMOTE, GMM, and GAN.

[0173] In this solution, techniques such as Synthetic Minority Over-sampling Technique (SMOTE), Gaussian Mixture Models (GMM), or Generative Adversarial Networks (GAN) can be used to synthesize more sample data that conforms to business logic and statistical distribution to expand the preprocessed initial standard test data set to ensure sufficient data for subsequent automated testing.

[0174] In the specific implementation process, the test module includes a first addition sub-module, a first storage sub-module, a first test sub-module, a second addition sub-module, a second storage sub-module, and a second test sub-module. The first addition sub-module is used to add a first label to the service data generated for the Nth time when receiving a new Nth test requirement, where N≥1. The first storage sub-module is used to store the service data generated for the Nth time and the first label in the first database. The first test sub-module is used to perform automated testing using the service data generated for the Nth time in the first database to obtain the test result for the Nth time corresponding to the service data generated for the Nth time and store it in the first database. The second addition sub-module is used to regenerate service data and add a second label to the service data generated for the (N + 1)th time when receiving the (N + 1)th test requirement. The second storage sub-module is used to store the service data generated for the Nth time and the second label in the second database. The second test sub-module is used to perform automated testing using the service data generated for the (N + 1)th time in the second database to obtain the test result for the (N + 1)th time corresponding to the service data generated for the (N + 1)th time and store it in the second database.

[0175] In this solution, when a new test round is required, that is, the Nth test, a new batch of business data is generated, and a unique first label is added to this batch of data. This label is used to identify that this batch of data is specific to the Nth test, which helps to quickly screen and locate this batch of data in subsequent stages, distinguish it from the data of other test rounds, and ensure the purity of the test data. When facing the next test requirement, that is, the (N + 1)th test, a new batch of business data is generated, and a second label different from the first label is added to this batch of data to identify that this batch of data is specific to the (N + 1)th test, avoiding the influence of residual historical test data on the new test, thereby improving the accuracy and reliability of the test results relatively high.

[0176] The test device for the above data includes a processor and a memory. The above acquisition unit, determination unit, generation unit, test unit, sending unit, etc. are all stored in the memory as program units, and the corresponding functions are implemented by the processor executing the above program units stored in the memory. The above modules are all located in the same processor; or, the above each module is located in different processors in any combination form.

[0177] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem that in the prior art, test data is usually generated based on a fixed data model, which cannot correspond to the source data distribution and the generated data has poor practicability can be solved.

[0178] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0179] The embodiment of the present invention provides a computer-readable storage medium. The above computer-readable storage medium includes a stored program. Among them, when the above program runs, it controls the device where the above computer-readable storage medium is located to execute the test method of the above data.

[0180] The embodiment of the present invention provides a processor. The above processor is used to run a program. Among them, when the above program runs, it executes the test method of the above data.

[0181] The embodiment of the present invention provides a device. The device includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the steps of the test method of the data. The device herein can be a server, a PC, a PAD, a mobile phone, etc.

[0182] A computer program product includes a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the test method for the data in each embodiment of the present application are implemented.

[0183] The present application also provides a data generation system, including one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the test methods for the data.

[0184] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described herein can be executed in a different order, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0185] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0186] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0187] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.

[0188] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to generate a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.

[0189] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0190] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0191] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0192] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0193] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for testing data, characterized in that, Including: Obtain a source data set, where the source data set is an initial data set for a financial institution to handle business. The source data set includes at least one or more of the number of business levels, the number of first accounts for handling each business, the number of second accounts for handling the business at the same business level, the time for handling the business, and the number of business transactions. There are multiple businesses in the financial institution, and each business includes at least one sub-business. The business level is the hierarchical relationship between the business and the sub-business; Determine the distribution type of the source data set according to the Kolmogorov-Smirnov algorithm; Use the rvs generation function to generate a standard test data set that conforms to the distribution type according to the source data set, where the data volume of the standard test data set is greater than the data volume of the source data set; Perform an automated test using the standard test data set to obtain a test result; Send the test result to the financial institution so that the financial institution can optimize its business based on the test result.

2. The method according to claim 1, wherein Determine the distribution type of the source data set according to the Kolmogorov-Smirnov algorithm, including: Construct a distribution verification model, where the distribution verification model is trained by the Kolmogorov-Smirnov algorithm using multiple sets of training data. Each set of training data in the multiple sets of training data includes a historical source data set obtained within a historical time period and the corresponding historical distribution type of the historical source data set; Input the source data set into the distribution verification model to obtain the corresponding distribution type of the source data set.

3. The method according to claim 2, wherein Construct a distribution verification model, including: Construct an initial distribution verification model, where the initial distribution verification model is trained by the Kolmogorov-Smirnov algorithm using multiple sets of training data. Each set of training data in the multiple sets of training data includes the historical source data set obtained within a historical time period and the corresponding historical initial distribution type of the historical source data set; Optimize the initial distribution verification model to obtain the distribution verification model, where the optimization methods include one or more of adjusting the learning rate, Bayesian optimization, regularization, data dimensionality reduction, data augmentation, and Adam optimization.

4. The method according to claim 1, wherein Perform an automated test using the standard test data set to obtain a test result, including: Generate business data according to the standard test data set, where the business data is data with business operation logic; Use an automated test tool to perform an automated test according to the business data to obtain the test result, where the automated test tool includes one or more of Selenium, Appium, Katalon Studio, and Postman.

5. The method according to claim 4, wherein Generate business data according to the standard test data set, including: Obtain a first business feature, where the first business feature is data representing a feature predefined related to the logic or rules of the business; Combine the standard test data set and the first business feature to obtain a first multi-dimensional data set; Generate first business data according to the first multi-dimensional data set, wherein the algorithms used for data generation include one or more of decision tree, generative adversarial network, Bayesian, and K-means.

6. The method according to claim 4, characterized in that, Generate business data according to the standard test data set, including: Randomly generate second business features using a business generation algorithm, wherein the business generation algorithm includes one or more of RNN algorithm, LSTN algorithm, Apriori algorithm, and FP-growth, and the second business feature is data representing a feature related to the logic or rules of the business randomly generated; Combine the standard test data set and the second business feature to obtain a second multi-dimensional data set; Generate second business data according to the second multi-dimensional data set, wherein the algorithms used for data generation include one or more of decision tree, generative adversarial network, Bayesian, and K-means.

7. The method according to any one of claims 1 to 6, characterized in that Use the rvs generation function to generate a standard test data set that conforms to the distribution type according to the source data set, including: Use the rvs generation function to generate an initial standard test data set that conforms to the distribution type according to the source data set; Preprocess the initial standard test data set to obtain the standard test data set, wherein the preprocessing methods include one or more of deleting duplicate values, removing null data, and data normalization.

8. A testing device for data, characterized in that Include: An acquisition unit for acquiring a source data set, wherein the source data set is an initial data set for a financial institution to handle business, and the source data set includes at least one or more of the number of business levels, the number of first accounts for handling each business, the number of second accounts for handling the business at the same business level, the time for handling the business, and the number of business transactions. There are multiple businesses for the financial institution, each business includes at least one sub-business, and the business level is the hierarchical relationship between the business and the sub-business; A determination unit for determining the distribution type of the source data set according to the Kolmogorov-Smirnov algorithm; A generation unit for using the rvs generation function to generate a standard test data set that conforms to the distribution type according to the source data set, wherein the data volume of the standard test data set is greater than the data volume of the source data set; A test unit for performing an automated test using the standard test data set to obtain a test result; A sending unit for sending the test result to the financial institution so that the financial institution can optimize its business based on the test result.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the data testing method described in any one of claims 1 to 7.

10. A data generation system, characterized in that, Include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a method for testing data recited in any one of claims 1 to 7.