A test data generation method and related device

By filtering strategy information and historical data, hierarchical variables and rule-based admission variables are generated, assigned values, and combined. This solves the problems of privacy risks, insufficient scenario coverage, and low efficiency in the generation of test data in existing technologies, and realizes efficient and diversified test data generation, thereby improving the stability and reliability of the risk control model.

CN120974082BActive Publication Date: 2026-01-27CHONGQING ANT CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511514864.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-27
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Existing technologies using real user data as test data pose privacy and compliance risks, have insufficient scenario coverage, are inefficient to generate manually, and the batch data generation methods do not meet the required metric values, resulting in high rework difficulty and insufficient test data richness, which cannot meet the needs of rapid iteration.

Method used

By filtering strategy information and historical data, hierarchical variables and rule admission variables are obtained, assigned values ​​and combined to determine accurate values ​​and rejection values, and test data covering normal, abnormal and extreme scenarios are generated.

Benefits of technology

It improved the efficiency and quality of test data generation, ensured data coverage of diverse scenarios, reduced the amount of data generated, met the needs of rapid iteration, and improved the stability and reliability of risk control models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974082B_ABST
    Figure CN120974082B_ABST
Patent Text Reader

Abstract

The specification discloses a test data generation method and related equipment, and relates to the field of data processing. The specification filters and analyzes policy information and a plurality of historical data to obtain a plurality of hierarchical variables and a plurality of rule access variables, performs at least one assignment for each hierarchical variable, performs full-quantity combination on the plurality of assigned hierarchical variables to obtain a plurality of initial test data, determines accurate values and rejection values that do not meet the access rules corresponding to the rule access variables, adds at least one rule access variable to each initial test data and assigns an accurate value or a rejection value, and obtains a plurality of test data. The test data is used for training a decision model that uses policy information to make decisions. The test data generation method provided in the specification can solve the problem of insufficient richness of test data used for training the decision model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of data processing, and in particular to a method for generating test data and related equipment. Background Technology

[0002] With the advancement of digitalization, the demand for high-quality test data for training risk control models is increasing. High-quality test data is crucial for the stability and reliability of the trained risk control models, enabling intelligent, compliant, and risk-controlled transaction processing. However, directly using real user data as test data presents privacy and compliance risks. Data leaks involving sensitive information (such as identity identifiers and transaction records) are easily triggered, and there are also issues with insufficient scenario coverage, as real data cannot cover extreme scenarios. Manually generating test data is inefficient, lacks richness, and suffers from delayed updates, failing to meet the requirements of rapid transaction iteration.

[0003] As for the commonly used batch data generation method, the system automatically initializes a verification table. Specifically, it initializes multiple indicators used in the strategies involved in the risk control module to obtain the verification table. Users can either construct the values ​​of each indicator on the verification table offline or automatically by the system to obtain test data. However, this method has problems such as the indicator values ​​not being of the required type or format, leading to frequent rework and high rework difficulty. Furthermore, the data generation algorithm is relatively basic, which can result in insufficient richness of test data, as it only constructs test data by filling in the indicator values ​​of each indicator on the verification table. Summary of the Invention

[0004] This specification provides a test data generation method and related equipment, which can solve the above-mentioned problems. The technical solution is as follows:

[0005] Firstly, embodiments of this specification provide a method for generating test data, the method comprising:

[0006] By filtering strategy information and multiple historical data, multiple hierarchical variables and multiple rule admission variables are obtained;

[0007] The hierarchical variables are assigned values ​​at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable. The multiple assigned hierarchical variables are combined to obtain multiple initial test data. The multiple assigned hierarchical variables included in the initial test data are obtained by assigning values ​​to different hierarchical variables.

[0008] Determine the accurate values ​​that conform to the admission rule and the rejection values ​​that do not conform to the admission rule corresponding to the rule admission variable. Add at least one rule admission variable to each of the initial test data and assign an accurate value or a rejection value to the rule admission variable to obtain multiple test data. The test data is used at least to train a decision model that makes decisions using the policy information.

[0009] Secondly, embodiments of this specification provide a test data generation apparatus, the apparatus comprising:

[0010] The data filtering module is used to filter strategy information and multiple historical data to obtain multiple hierarchical variables and multiple rule admission variables;

[0011] The first data generation module is used to assign values ​​to the hierarchical variables at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and to combine multiple assigned hierarchical variables to obtain multiple initial test data; wherein, the multiple assigned hierarchical variables included in the initial test data are obtained by assigning values ​​to different hierarchical variables;

[0012] The second data generation module is used to determine the accurate values ​​that conform to the admission rules and the rejection values ​​that do not conform to the admission rules corresponding to the rule admission variables, add at least one rule admission variable to each of the initial test data and assign an accurate value or a rejection value to the rule admission variable to obtain multiple test data; wherein, the test data is used at least to train a decision model that makes decisions using the policy information.

[0013] Thirdly, embodiments of this specification provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the above-described method steps.

[0014] Fourthly, embodiments of this specification provide a computer program product that stores multiple instructions adapted for loading by a processor and executing the above-described method steps.

[0015] Fifthly, embodiments of this specification provide an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.

[0016] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:

[0017] This embodiment of the specification improves the accuracy of the obtained stratified variables and rule-based admission variables by filtering and analyzing policy information and multiple historical data, thus providing reliable data support for the subsequent construction of test data. Furthermore, by assigning values ​​to the stratified variables at least once, at least one assigned stratified variable is obtained for each stratified variable. These assigned stratified variables are combined to obtain multiple initial test data sets. The accurate values ​​that conform to the admission rules and the rejection values ​​that do not conform to the admission rules for the rule-based admission variables are determined. At least one rule-based admission variable is added to each initial test data set and an accurate or rejection value is assigned to it, resulting in multiple test data sets sufficient for training a decision-making model that uses policy information for decision-making. By assigning values ​​to multiple stratified variables multiple times and constructing initial test data in its entirety, this embodiment of the specification enriches the final test data, ensuring that the test data covers normal, abnormal, and extreme scenarios. Furthermore, the embodiments in this specification reduce the amount of test data generated by assigning accurate and rejection values ​​only to the rule admission variables and adding them to the initial test data. This allows for the construction of test data that conforms to transaction logic, thereby improving the efficiency of test data generation and the overall quality of the test data. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the architecture of a test data generation method provided in the embodiments of this specification;

[0020] Figure 2 This is a flowchart illustrating a test data generation method provided in the embodiments of this specification;

[0021] Figure 3 This is a schematic diagram of an initial test data generation process provided in the embodiments of this specification;

[0022] Figure 4 This is a schematic diagram of a test data generation process provided in the embodiments of this specification;

[0023] Figure 5 This is a flowchart illustrating a test data generation method provided in the embodiments of this specification;

[0024] Figure 6This is a flowchart illustrating a method for obtaining a set of values ​​for a feature variable, as provided in an embodiment of this specification.

[0025] Figure 7 This is a flowchart illustrating a test data generation method provided in the embodiments of this specification;

[0026] Figure 8 This is a schematic diagram of the structure of a test data generation device provided in the embodiments of this specification;

[0027] Figure 9 This is a schematic diagram of the structure of an electronic device provided in the embodiments of this specification. Detailed Implementation

[0028] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.

[0029] In the description of this specification, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this specification based on the specific circumstances. Furthermore, in the description of this specification, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0030] The present specification will now be described in detail with reference to specific embodiments.

[0031] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the features, information, and data involved in this specification were all obtained under full authorization.

[0032] like Figure 1 As shown, Figure 1 This is a flowchart illustrating a test data generation method provided in the embodiments of this specification. Figure 1 It includes at least a server 101 that executes the test data generation method, and multiple terminals that request to upload policy information, historical data, or generate test data. These multiple terminals include at least a first terminal 1021, a second terminal 1022, and a third terminal 1023. It is understood that... Figure 1 The number of servers and terminals shown is for illustrative purposes only, and the embodiments in this specification do not impose any limitations on this.

[0033] The aforementioned server 101 can be a standalone server device, such as a rack-mount, blade, tower, or cabinet-type server device, or a workstation, mainframe, or other hardware device with strong computing power; it can also be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, where each server is functionally and hierarchically equivalent in the transaction chain, and each server can provide services to the outside world independently. Providing services independently can be understood as not requiring the assistance of other servers.

[0034] For example, a server can be multiple physical servers, each with independent hardware. Alternatively, a server can be multiple virtual servers deployed within the same hardware resource pool. Virtual server deployment methods include, but are not limited to, VMware, VirtualBox, and Virtual PC.

[0035] It is understood that server 101 also possesses other service capabilities and functions to complete the tasks described in the following embodiments. For example, server 101 also provides portal services, resource management services, and CI / CD services, etc.

[0036] Terminals include, but are not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Terminals may have different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and terminals in 5G networks or future evolved networks.

[0037] In the embodiments of this specification, display devices may also be installed on terminals such as the first terminal 1021, the second terminal 1022, and the third terminal 1023. These display devices can be various devices capable of display functions, such as cathode ray tube displays (CR), light-emitting diode displays (LED), electronic ink screens, liquid crystal displays (LCD), and plasma display panels (PDP). For example, a user can use the display device on the first terminal 1021 to send a request to the server 101 to generate test data. This request instructs the server 101 to generate test data. As another example, a user can use the first terminal 1021 to send policy information and multiple historical data points to the server 101, and can also send numerical values ​​for hierarchical variables or rule-based admission variables to the server 101.

[0038] Multiple terminals and multiple servers can communicate through communication links established by communication protocols. For example, the network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.

[0039] In one embodiment, such as Figure 2 The diagram shown is a flowchart illustrating a test data generation method provided in an embodiment of this specification. This method can be implemented using a computer program and can run on a test data generation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application.

[0040] Specifically, the test data generation method includes:

[0041] S102. Filter the strategy information and multiple historical data to obtain multiple hierarchical variables and multiple rule admission variables.

[0042] Strategy information can be understood as data and descriptions related to decision-making logic or action guidelines, and can include structured rules, model parameters, tags, approval process records, etc. For example, credit-related strategy information may include refusing to provide credit services to users with credit scores below 600; marketing-related strategy information may include pushing coupons to users who have made purchases in the past 3 months; and gaming-related strategy information may include preventing players with equipment scores below 25,000 from entering new dungeons.

[0043] For example, in the embodiments of this specification, before the strategy information related to financial scenarios is formally deployed to the real-time decision-making model (such as a risk control system or rule engine) in the production environment, strategy consistency verification is required. The specific verification steps involve constructing rich test data, submitting the test data to the decision-making model in the pre-production environment that includes the strategy information, obtaining the decision results corresponding to the test data, and implementing judgment logic that is completely consistent with the online strategy using SQL statements in the offline environment, comparing the consistency between the offline and online results. Only through rigorous consistency verification can it be guaranteed that the strategy's behavior is predictable, the risk is controllable, and the audit compliance is achieved after it is deployed to the decision-making model. Therefore, in one embodiment, this embodiment constructs test data for verifying strategy information and decision-making models related to financial scenarios; in another embodiment, this embodiment can also construct test data for verifying strategy information and decision-making models related to games, healthcare, education, marketing activities, etc., and this embodiment does not impose any limitations on this.

[0044] Historical data can be understood as data accumulated in the past that reflects user behavior, transaction records, system logs, or other activities. In this embodiment, historical data is related to strategy information. For example, historical data related to credit includes data such as user loan application records and repayment status generated when providing credit services to a large number of users based on strategy information; historical data related to marketing activities includes data such as user shopping frequency, average order value, and browsing behavior required to analyze user-matched coupons before pushing coupons to users based on strategy information; historical data related to games may include game players' account login time, device information, game data, etc.

[0045] The strategy information and historical data are filtered to obtain multiple stratified variables and multiple rule-based access variables. Stratified variables are used to divide the overall user base into different levels or groups, aiming to achieve differentiated management or analysis and help identify user groups with different risk levels, value levels, or behavioral patterns. For example, stratified variables include age, place of residence, number of overdue payments in the past six months, and monthly consumer finance transactions. Rule-based access variables can be understood as Boolean or categorical variables indicating whether a user meets certain preset rule conditions. They are used to determine whether a user is eligible to enter a certain process or enjoy a certain service. For example, rule-based access variables may include whether the user meets the age requirement (≥18 years old), has completed real-name authentication, has any overdue payment records in the past 30 days, and has a credit score greater than 700.

[0046] In the embodiments of this specification, multiple hierarchical variables can be selected by clustering or scoring behavioral indicators in historical data, and by mapping them in combination with labels in policy information (e.g., dividing levels according to credit limits). Furthermore, multiple rule-based admission variables can be selected by decomposing rules in policy information or by matching rules calculated from historical data.

[0047] S104. Assign values ​​to the stratified variables at least once to obtain at least one assigned stratified variable corresponding to each stratified variable, and combine multiple assigned stratified variables to obtain multiple initial test data.

[0048] The initial test data includes multiple pre-assigned hierarchical variables, which are obtained by assigning values ​​to different hierarchical variables. In other words, the hierarchical variables themselves are abstract, but by assigning values ​​to them, they become concrete descriptions. The initial test data is a relatively complete user profile formed by combining multiple pre-assigned hierarchical variables.

[0049] For example, the stratification variable is the user's risk level, and the corresponding value for this stratification variable can be low risk, medium risk, or high risk; the stratification variable is the salary level, and the corresponding value for this stratification variable can be low salary (<5k), medium salary (5k~15k), or high salary (15k); the stratification variable is the age group, and the corresponding value for this stratification variable can be 1-80; the stratification variable is the occupation type, and the corresponding value for this stratification variable can be white-collar, blue-collar, freelance, or student; the stratification variable is the credit score range, and the corresponding value for this stratification variable can be A (700+), B (600-699), C (<600), or a specific numerical value.

[0050] In the embodiments of this specification, multiple assigned hierarchical variables are combined to obtain multiple initial test data. The combination method can be a Cartesian product combination, i.e., a full combination. For example, one value is selected from each hierarchical variable to form an initial test data. For example, the initial test data could be a combination of low risk + low salary + 50 years old + student, or a combination of low risk + low salary + 30 years old + white-collar worker. In another embodiment, initial test data can also be obtained through orthogonal experimentation, i.e., instead of a full combination, representative combinations are selected. For example, only the combination of low risk + low salary + 30 years old + white-collar worker and the combination of high risk + high salary + 50 years old + high credit score of 700 are retained. In another embodiment, meaningful combinations can also be selected based on transaction priorities. For example, the initial test data corresponding to the combination of high salary + student is removed.

[0051] like Figure 3 As shown, Figure 3This is a schematic diagram of an initial test data generation process provided in an embodiment of this specification. In this embodiment, the selected stratified variables include at least stratified variable a, stratified variable b, stratified variable c, and stratified variable d. At least one assignment is performed on stratified variable a to obtain assigned stratified variables a1, a2, and a3. At least one assignment is performed on stratified variable b to obtain stratified variables b1 and b2. At least one assignment is performed on stratified variable c to obtain stratified variables c1, c2, c3, and c4. At least one assignment is performed on stratified variable d to obtain stratified variables d1, d2, and d3.

[0052] By combining the above-mentioned multiple assigned variables, at least the following initial test data are obtained: 201 includes assigned hierarchical variables a1, b1, and c1; 202 includes assigned hierarchical variables a2 and d1; and 203 includes assigned hierarchical variables a2, b3, c2, and d2.

[0053] Understandable Figure 3 The number of stratified variables, the values ​​assigned to each stratified variable, and the number and content of the initial test data obtained by combining them are for illustrative purposes only.

[0054] In one embodiment, strategy information and multiple historical data are filtered to obtain the correlation between the values ​​of at least two hierarchical variables; based on the correlation between the values ​​of at least two hierarchical variables, multiple hierarchical variables are assigned values ​​at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable; and multiple assigned hierarchical variables are combined to obtain multiple initial test data.

[0055] The association between multiple stratified variables can be derived by summarizing patterns in historical data. For example, there is no association between the stratified variable "student" and the stratified variable "high salary," so they cannot be combined to obtain initial test data. The association between multiple stratified variables can also be derived based on policy information settings. For example, if policy information includes users with fewer than 5 annual delinquencies and a monthly salary of 10k as having a high credit score, then there is an association between the stratified variables "annual delinquency count," "salary," and "credit score."

[0056] In this embodiment, logical relationships between multiple hierarchical variables are mined from strategy information and multiple historical data, and high-quality test samples are generated based on these logical relationships to improve the quality of the initial test data, making the initial test data more scientific and closer to the actual transaction scenario.

[0057] S106. Determine the accurate values ​​that conform to the admission rules and the rejection values ​​that do not conform to the admission rules for the rule admission variables. Add at least one rule admission variable to each initial test data and assign an accurate value or rejection value to the rule admission variable to obtain multiple test data.

[0058] The test data is used at least to train a decision-making model that uses policy information to make decisions.

[0059] The exact value indicates that the admission rule is satisfied when the rule's admission variable is that value, while the rejection value indicates that the admission rule is not satisfied when the rule's admission variable is that value. For example, if the admission rule is that the person is older than 18, the rejection value could be 10, and the exact value could be 20. Another example is if the admission rule is that the person must pass real-name authentication; the rejection value could be "verified_identity=0", and the exact value could be "verified_identity=1".

[0060] If only the rule admission variables are assigned corresponding exact and rejection values ​​and added to the initial test data to obtain multiple test data sets, each initial test data set must include at least one assigned rule admission variable. For example, for the initial test data set consisting of a combination of low salary and student, the rule admission variable is assigned the value verified_identity=1, indicating that the initial test data set has passed real-name authentication. Additionally, a rule admission variable verified_identity=0 is also added for this initial test data set.

[0061] like Figure 4 As shown, Figure 4 This is a schematic diagram of a test data generation process provided in an embodiment of this specification. In this embodiment, at least three rule-based admission variables are included: X, Y, and Z. Accurate values ​​X1 and X2 are assigned to rule-based admission variables X; accurate values ​​Y1 and Y2 are assigned to rule-based admission variables Y; and accurate values ​​Z1 and Z2 are assigned to rule-based admission variables Z. Based on... Figure 3 For the multiple initial test data obtained, add rule admission variables X and Y with accurate values ​​X1 and Y1 respectively for initial test data 201; add rule admission variables X with rejection values ​​X2 and Z with accurate values ​​Z1 respectively for initial test data 202; and add rule admission variables X with rejection values ​​X2, Y with rejection values ​​Y2 and Z with rejection values ​​Z2 respectively for initial test data 203.

[0062] Understandable Figure 4 The number of rule admission variables, the values ​​assigned to the rule admission variables, and the number and content of test data obtained by combining them are shown for illustrative purposes only.

[0063] This embodiment of the specification improves the accuracy of the obtained stratified variables and rule-based admission variables by filtering and analyzing policy information and multiple historical data, thus providing reliable data support for the subsequent construction of test data. Furthermore, by assigning values ​​to the stratified variables at least once, at least one assigned stratified variable is obtained for each stratified variable. These assigned stratified variables are combined to obtain multiple initial test data sets. The accurate values ​​that conform to the admission rules and the rejection values ​​that do not conform to the admission rules for the rule-based admission variables are determined. At least one rule-based admission variable is added to each initial test data set and an accurate or rejection value is assigned to it, resulting in multiple test data sets sufficient for training a decision-making model that uses policy information for decision-making. By assigning values ​​to multiple stratified variables multiple times and constructing initial test data in its entirety, this embodiment of the specification enriches the final test data, ensuring that the test data covers normal, abnormal, and extreme scenarios. Furthermore, the embodiments in this specification reduce the amount of test data generated by assigning accurate and rejection values ​​only to the rule admission variables and adding them to the initial test data. This allows for the focus on constructing test data that conforms to transaction logic, thereby improving the efficiency of test data generation and the overall quality of the test data.

[0064] In one embodiment, such as Figure 5 The diagram shown is a flowchart illustrating a test data generation method provided in an embodiment of this specification. This method can be implemented using a computer program and can run on a test data generation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application.

[0065] Specifically, the test data generation method includes:

[0066] S202. Filter the strategy information and multiple historical data to obtain multiple feature variables and the set of values ​​for the feature variables.

[0067] The process aligns and integrates strategy information and multiple historical data sets. Based on the structure and specific content of the strategy information and historical data, multiple feature variables and their value sets are selected. For example, features such as trend characteristics, volatility characteristics, technical indicators, and time indicators are extracted from historical data as feature variables to be selected. Strategy parameters and the number of strategy triggers are extracted from the strategy information as feature variables to be selected. Further analysis of these feature variables using feature extraction methods such as filtering, packaging, and embedding yields multiple feature variables and their value sets. For example, based on multiple historical data sets, the feature variable obtained is the return rate over the past 5 days, and the corresponding value set for this feature variable includes values ​​between 2% and 20%.

[0068] S204. Based on the type of each feature variable, the multiple feature variables are divided into multiple hierarchical variables and multiple rule-based entry variables.

[0069] Multiple feature variables are categorized based on their purpose and semantics, determining whether each feature variable is a stratified variable or a rule-based admission variable, and further defining the value sets for both stratified and rule-based admission variables. The value set for stratified variables is used to assign values ​​to them at least once, while the value set for rule-based admission variables is used to determine their exact and rejection values. For example, if a feature variable's content is net asset value, it is a stratified variable; if a feature variable's content is its historical maximum drawdown, and includes an admission rule prohibiting further investment if the drawdown exceeds 20%, then it is a rule-based admission variable.

[0070] S206. Assign values ​​to the stratified variables at least once to obtain at least one assigned stratified variable corresponding to each stratified variable, and combine multiple assigned stratified variables to obtain multiple initial test data.

[0071] Assign a value to the stratified variable at least once based on the set of values ​​corresponding to the stratified variable, and obtain at least one assigned stratified variable corresponding to each stratified variable. Combine multiple assigned stratified variables to obtain multiple initial test data.

[0072] In one embodiment, based on the set of values ​​for the hierarchical variables and at least one numerical value input by the user, the hierarchical variables are assigned a value at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and the multiple assigned hierarchical variables are combined to obtain multiple initial test data.

[0073] Users can input at least one numerical value for a specific hierarchical variable into the processor executing the test data generation method via a page or interface. This value can be a specific number or a range of values. Furthermore, based on the set of values ​​for the hierarchical variable extracted from multiple historical data and strategy information, and the at least one numerical value input by the user, the hierarchical variable is assigned a value at least once.

[0074] In this embodiment, the hierarchical variables can be assigned values ​​based on at least one numerical value input by the user, thereby increasing the richness of the hierarchical variable assignments and enabling the test data to cover more scenarios.

[0075] S208. Determine the accurate values ​​that conform to the admission rules and the rejection values ​​that do not conform to the admission rules for the rule admission variables. Add at least one rule admission variable to each initial test data and assign an accurate value or rejection value to the rule admission variable to obtain multiple test data.

[0076] See S106 above; it will not be repeated here.

[0077] This embodiment of the specification improves the accuracy of the acquired hierarchical variables and rule admission variables by filtering and analyzing strategy information and multiple historical data, thus providing reliable data support for subsequent test data construction. Furthermore, by assigning values ​​to multiple hierarchical variables multiple times and constructing initial test data from scratch, this embodiment enriches the final test data, ensuring it covers normal, abnormal, and extreme scenarios. Moreover, by assigning accurate and rejection values ​​only to the rule admission variables and adding them to the initial test data, this embodiment reduces the amount of test data generated, focusing on constructing test data that conforms to transaction logic, thereby improving both the efficiency and overall quality of test data generation.

[0078] based on Figure 5 Please refer to the embodiments shown below. Figure 6 The illustrated embodiment. Figure 6 This is a flowchart illustrating an embodiment of the present specification for obtaining a set of values ​​for a feature variable. S202 includes the following steps:

[0079] S202-2. Filter the strategy information and multiple historical data to obtain multiple feature variables.

[0080] Align and integrate strategy information and multiple historical data, and filter multiple feature variables based on the structure and specific content of the strategy information and multiple historical data.

[0081] S202-4. Analyze whether the value type of the characteristic variable is discrete or continuous. Based on the numerical acquisition method corresponding to different value types, select multiple values ​​of the characteristic variable from the strategy information and multiple historical data to obtain the value set of the characteristic variable.

[0082] Based on the set of values ​​for feature variables provided by multiple historical data and strategy information, it is determined whether the values ​​of the feature variables belong to a certain category and whether they are ordered, thereby determining whether the value type of the feature variable is discrete or continuous. For example, if the set of values ​​for a feature variable includes a finite number of categories, then the value type of the feature variable is discrete; if the set of values ​​for a feature variable includes a range of values ​​and the range includes infinitely subdivisible real numbers, then the value type of the feature variable is continuous.

[0083] Methods for obtaining the set of values ​​for discrete feature variables include enumeration, frequency statistics, and rule analysis. For continuous feature variables, the set of values ​​can be determined using extreme value methods, equal interval binning, Gaussian distribution fitting, and sliding window statistics.

[0084] In one embodiment, when the feature variable is discrete, all policy information and multiple historical data are recorded to obtain a set of feature variable values; when the feature variable is continuous, the continuous value range of the feature variable is obtained from the policy information and multiple historical data, and at least one key value in the continuous value range of the feature variable is recorded to obtain a set of feature variable values; wherein, the key value includes at least the boundary value of the continuous value range of the feature variable.

[0085] Specifically, when the feature variable is discrete, the policy information and multiple values ​​of the feature variable in multiple historical data are fully recorded and deduplicated. For example, by enumerating the values ​​corresponding to multiple admission conditions included in the policy information and recording the values ​​mentioned in multiple historical data, the set of values ​​corresponding to the feature variable is obtained.

[0086] When the feature variable is continuous, the upper and lower bounds of the continuous range of the feature variable are obtained from the strategy information and multiple historical data. There can be multiple continuous ranges of the feature variable, and at least one key value in the continuous range of the feature variable is recorded. The key value includes at least the boundary value (i.e., the upper and lower bounds) of the continuous range of the feature variable, and may also include the mean, median, threshold corresponding to the admission rule, benchmark value, etc., so as to obtain the set of values ​​of the feature variable.

[0087] In this embodiment, the value ranges corresponding to different types of feature variables are obtained based on different value-taking methods. For discrete feature variables, a full recording method is adopted to ensure that no legal state is missed, supporting complete enumeration testing. The value ranges of continuous feature variables are determined by focusing on boundary values ​​and key values. The value set of structured feature variables provides data support for the subsequent generation of test data.

[0088] In one embodiment, such as Figure 7 The diagram shown is a flowchart illustrating a test data generation method provided in an embodiment of this specification. This method can be implemented using a computer program and can run on a test data generation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application.

[0089] Specifically, the test data generation method includes:

[0090] S302. Filter the strategy information and multiple historical data to obtain multiple hierarchical variables and multiple rule admission variables.

[0091] See S102 above, which will not be repeated here.

[0092] S304. Assign values ​​to the stratified variables at least once to obtain at least one assigned stratified variable corresponding to each stratified variable.

[0093] Based on multiple values ​​within the range of values ​​of the hierarchical variable, a value can be assigned to the hierarchical variable at least once, resulting in at least one assigned hierarchical variable corresponding to each hierarchical variable.

[0094] S306. Calculate the Cartesian product of at least one assigned stratified variable corresponding to each of the multiple stratified variables to obtain multiple initial test data.

[0095] The Cartesian product represents all the possibilities of forming an ordered combination by taking one element from each of multiple sets. The multiple sets can be understood as sets consisting of at least one assigned hierarchical variable corresponding to each hierarchical variable. The result of each Cartesian product is a complete combination sample of assigned hierarchical variables, which serves as the initial test data.

[0096] S308. Based on the multiple admission rules included in the policy information and at least one rule admission variable used by the admission rules, determine the accurate value that conforms to the admission rule and the rejection value that does not conform to the admission rule corresponding to the rule admission variable.

[0097] The exact value indicates that the admission rule is satisfied when the rule's admission variable is that value, while the rejection value indicates that the admission rule is not satisfied when the rule's admission variable is that value. For example, if the admission rule is that the person is older than 18, the rejection value could be 10, and the exact value could be 20. Another example is if the admission rule is that the person must pass real-name authentication; the rejection value could be "verified_identity=0", and the exact value could be "verified_identity=1".

[0098] S310. Add at least one rule admission variable to each initial test data and assign an accurate value or a rejection value to the rule admission variable. Add different rule admission variables to each initial test data multiple times to obtain multiple test data.

[0099] If only the rule admission variables are assigned corresponding exact and rejection values ​​and added to the initial test data to obtain multiple test data sets, each initial test data set must include at least one assigned rule admission variable. For example, for the initial test data set consisting of a combination of low salary and student, the rule admission variable is assigned the value verified_identity=1, indicating that the initial test data set has passed real-name authentication. Additionally, a rule admission variable verified_identity=0 is also added for this initial test data set.

[0100] In one embodiment, based on policy information and multiple historical data, a set of values ​​for multiple rule admission variables is determined; based on multiple admission rules included in the policy information and at least one rule admission variable used by the admission rules, the set of values ​​for the rule admission variables is divided into a set of accurate values ​​that conform to the admission rules and a set of rejection values ​​that do not conform to the admission rules; based on at least one value in the set of accurate values ​​of the rule admission variables and at least one value in the set of rejection values, the accurate value and rejection value corresponding to the rule admission variables are determined respectively.

[0101] After obtaining the set of values ​​for the rule admission variables, based on the admission rules corresponding to the rule admission variables, the set of values ​​for the rule admission variables is divided into an accurate value set (containing only accurate values) and a rejection value set (containing only rejection values). Accurate and rejection values ​​corresponding to the rule admission variables are randomly selected based on at least one value from the accurate value set and at least one value from the rejection value set, respectively. Furthermore, at least one rule admission variable is added to each initial test data set and assigned an accurate or rejection value. Different rule admission variables are added multiple times to each initial test data set to obtain multiple test data sets.

[0102] This embodiment of the specification improves the accuracy of the acquired hierarchical variables and rule admission variables by filtering and analyzing strategy information and multiple historical data, thus providing reliable data support for subsequent test data construction. Furthermore, by assigning values ​​to multiple hierarchical variables multiple times and constructing initial test data from scratch, this embodiment enriches the final test data, ensuring it covers normal, abnormal, and extreme scenarios. Moreover, by assigning accurate and rejection values ​​only to the rule admission variables and adding them to the initial test data, this embodiment reduces the amount of test data generated, focusing on constructing test data that conforms to transaction logic, thereby improving both the efficiency and overall quality of test data generation.

[0103] The following are embodiments of the apparatus described in this specification, which can be used to execute the embodiments of the methods described in this specification. For details not disclosed in the apparatus embodiments of this specification, please refer to the embodiments of the methods described in this specification.

[0104] Please see Figure 8 This diagram illustrates the structure of a test data generation apparatus provided in an exemplary embodiment of this specification. The test data generation apparatus can be implemented as all or part of a device through software, hardware, or a combination of both. The test data generation apparatus includes a data filtering module 401, a first data generation module 402, and a second data generation module 403.

[0105] The data filtering module 401 is used to filter strategy information and multiple historical data to obtain multiple hierarchical variables and multiple rule admission variables;

[0106] The first data generation module 402 is used to assign values ​​to the hierarchical variables at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and to combine multiple assigned hierarchical variables to obtain multiple initial test data; wherein, the multiple assigned hierarchical variables included in the initial test data are obtained by assigning values ​​to different hierarchical variables;

[0107] The second data generation module 403 is used to determine the accurate value that conforms to the admission rule and the rejection value that does not conform to the admission rule corresponding to the rule admission variable, add at least one rule admission variable to each of the initial test data and assign an accurate value or rejection value to the rule admission variable to obtain multiple test data; wherein, the test data is used at least to train a decision model that makes decisions using the policy information.

[0108] In one or more embodiments, the data filtering module 401 includes:

[0109] The information filtering unit is used to filter strategy information and multiple historical data to obtain multiple feature variables and a set of values ​​for the feature variables.

[0110] The variable differentiation unit is used to differentiate the multiple feature variables into multiple stratified variables and multiple rule admission variables according to the type of each feature variable; wherein, the value set of the stratified variables is used to assign a value to the stratified variables at least once, and the value set of the rule admission variables is used to determine the exact value and rejection value of the rule admission variables.

[0111] In one or more embodiments, the information filtering unit includes:

[0112] The first filtering subunit is used to filter strategy information and multiple historical data to obtain multiple feature variables;

[0113] The second filtering subunit is used to analyze whether the value type of the feature variable is discrete or continuous, and according to the numerical acquisition method corresponding to different value types, to filter multiple values ​​of the feature variable from the strategy information and the multiple historical data to obtain the value set of the feature variable.

[0114] In one or more embodiments, the second screening subunit is specifically used for:

[0115] When the feature variable is discrete, the strategy information and multiple values ​​of the feature variable in the multiple historical data are fully recorded to obtain the set of values ​​of the feature variable;

[0116] When the feature variable is continuous, the continuous value range of the feature variable is obtained from the strategy information and the multiple historical data, and at least one key value in the continuous value range of the feature variable is recorded to obtain the value set of the feature variable; wherein, the key value includes at least the boundary value of the continuous value range of the feature variable.

[0117] In one or more embodiments, the test data generation apparatus further includes:

[0118] The correlation module is used to filter strategy information and multiple historical data to obtain the correlation relationship between the values ​​of at least two of the hierarchical variables;

[0119] The first data generation module 402 includes:

[0120] The associated data generation unit is used to assign values ​​to multiple hierarchical variables at least once according to the correlation between the values ​​of at least two hierarchical variables, to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and to combine multiple assigned hierarchical variables to obtain multiple initial test data.

[0121] In one or more embodiments, the first data generation module 402 includes:

[0122] The input data generation unit is used to assign a value to the hierarchical variable at least once based on the set of values ​​of the hierarchical variable and at least one numerical value input by the user, to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and to combine multiple assigned hierarchical variables to obtain multiple initial test data.

[0123] In one or more embodiments, the first data generation module 402 includes:

[0124] The first assignment unit is used to assign a value to the hierarchical variable at least once, so as to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable.

[0125] The first combination unit is used to perform Cartesian product calculation on at least one assigned hierarchical variable corresponding to each of the multiple hierarchical variables to obtain multiple initial test data.

[0126] In one or more embodiments, the second data generation module 403 includes:

[0127] The second assignment unit is used to determine the accurate value of the rule admission variable that conforms to the admission rule and the rejection value that does not conform to the admission rule, based on the multiple admission rules included in the strategy information and at least one rule admission variable used by the admission rule.

[0128] The second combination unit is used to add at least one rule admission variable to each of the initial test data and assign an accurate value or a rejection value to the rule admission variable, and to add different rule admission variables to each of the initial test data multiple times to obtain multiple test data.

[0129] In one or more embodiments, the test data generation apparatus includes:

[0130] The value-determining unit is used to determine a set of values ​​for multiple rule admission variables based on the strategy information and the multiple historical data.

[0131] The second assignment unit includes:

[0132] The first assignment subunit is used to divide the set of values ​​of the rule admission variable into a set of accurate values ​​that conform to the admission rule and a set of rejection values ​​that do not conform to the admission rule, based on the multiple admission rules included in the strategy information and at least one rule admission variable used by the admission rule.

[0133] The second assignment subunit is used to determine the accurate value and rejection value corresponding to the rule admission variable based on at least one value in the accurate value set and at least one value in the rejection value set of the rule admission variable, respectively.

[0134] This embodiment of the specification improves the accuracy of the obtained stratified variables and rule-based admission variables by filtering and analyzing policy information and multiple historical data, thus providing reliable data support for the subsequent construction of test data. Furthermore, by assigning values ​​to the stratified variables at least once, at least one assigned stratified variable is obtained for each stratified variable. These assigned stratified variables are combined to obtain multiple initial test data sets. The accurate values ​​that conform to the admission rules and the rejection values ​​that do not conform to the admission rules for the rule-based admission variables are determined. At least one rule-based admission variable is added to each initial test data set and an accurate or rejection value is assigned to it, resulting in multiple test data sets sufficient for training a decision-making model that uses policy information for decision-making. By assigning values ​​to multiple stratified variables multiple times and constructing initial test data in its entirety, this embodiment of the specification enriches the final test data, ensuring that the test data covers normal, abnormal, and extreme scenarios. Furthermore, the embodiments in this specification reduce the amount of test data generated by assigning accurate and rejection values ​​only to the rule admission variables and adding them to the initial test data. This allows for the focus on constructing test data that conforms to transaction logic, thereby improving the efficiency of test data generation and the overall quality of the test data.

[0135] It should be noted that the test data generation device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the test data generation method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the test data generation device and the test data generation method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0136] The example numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the examples.

[0137] This specification also provides a computer storage medium that can store multiple instructions adapted to be loaded and executed by a processor as described above. Figure 1 - Figure 8 The test data generation method of the illustrated embodiment can be found in the following documentation for a detailed execution process: Figure 1 - Figure 8 The specific details of the illustrated embodiments will not be elaborated here.

[0138] This specification also provides a computer program product that stores at least one instruction, which is loaded and executed by a processor as described above. Figure 1 - Figure 8 The test data generation method of the illustrated embodiment can be found in the following documentation for a detailed execution process: Figure 1 - Figure 8 The specific details of the illustrated embodiments will not be elaborated here.

[0139] Please see Figure 9 This document provides a schematic diagram of the structure of an electronic device as an embodiment of the present specification. Figure 9 As shown, the electronic device 500 may include: at least one processor 501, at least one network interface 504, user interface 503, memory 505, and at least one communication bus 502.

[0140] The communication bus 502 is used to enable communication between these components.

[0141] The user interface 503 may include a display screen and a camera. Optionally, the user interface 503 may also include a standard wired interface and a wireless interface.

[0142] The network interface 504 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0143] The processor 501 may include one or more processing cores. The processor 501 connects to various parts within the electronic device 500 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 505, and by calling data stored in the memory 505. Optionally, the processor 501 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 501 may integrate one or a combination of several of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 501 and may be implemented as a separate chip.

[0144] The memory 505 may include random access memory (RAM) or read-only memory. Optionally, the memory 505 may include a non-transitory computer-readable storage medium. The memory 505 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 505 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 505 may also be at least one storage device located remotely from the aforementioned processor 501. Figure 9 As shown, the memory 505, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a test data generation application.

[0145] exist Figure 9In the illustrated electronic device 500, the user interface 503 is mainly used to provide an input interface for the user and to acquire user input data; while the processor 501 can be used to call the test data stored in the memory 505 to generate an application program, and specifically perform the following operations:

[0146] By filtering strategy information and multiple historical data, multiple hierarchical variables and multiple rule admission variables are obtained;

[0147] The hierarchical variables are assigned values ​​at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable. The multiple assigned hierarchical variables are combined to obtain multiple initial test data. The multiple assigned hierarchical variables included in the initial test data are obtained by assigning values ​​to different hierarchical variables.

[0148] Determine the accurate values ​​that conform to the admission rule and the rejection values ​​that do not conform to the admission rule corresponding to the rule admission variable. Add at least one rule admission variable to each of the initial test data and assign an accurate value or a rejection value to the rule admission variable to obtain multiple test data. The test data is used at least to train a decision model that makes decisions using the policy information.

[0149] In one or more embodiments, processor 501 performs the filtering of policy information and multiple historical data to obtain multiple hierarchical variables and multiple rule admission variables, specifically executing:

[0150] By filtering strategy information and multiple historical data, multiple feature variables and the set of values ​​for the feature variables are obtained;

[0151] Based on the type of each feature variable, the multiple feature variables are divided into multiple stratified variables and multiple rule admission variables; wherein, the value set of the stratified variables is used to assign a value to the stratified variables at least once, and the value set of the rule admission variables is used to determine the exact value and rejection value of the rule admission variables.

[0152] In one or more embodiments, the processor 501 performs the filtering of policy information and multiple historical data to obtain multiple feature variables and a set of values ​​for the feature variables, specifically executing:

[0153] By filtering strategy information and multiple historical data, several feature variables are obtained;

[0154] The value type of the feature variable is analyzed to be discrete or continuous. Based on the numerical acquisition method corresponding to different value types, multiple values ​​of the feature variable are filtered from the strategy information and the multiple historical data to obtain the value set of the feature variable.

[0155] In one or more embodiments, the processor 501 executes the analysis to determine whether the value type of the feature variable is discrete or continuous, and according to the numerical acquisition method corresponding to different value types, filters multiple values ​​of the feature variable from the strategy information and the multiple historical data to obtain the value set of the feature variable. Specifically, the following is performed:

[0156] When the feature variable is discrete, the strategy information and multiple values ​​of the feature variable in the multiple historical data are fully recorded to obtain the set of values ​​of the feature variable;

[0157] When the feature variable is continuous, the continuous value range of the feature variable is obtained from the strategy information and the multiple historical data, and at least one key value in the continuous value range of the feature variable is recorded to obtain the value set of the feature variable; wherein, the key value includes at least the boundary value of the continuous value range of the feature variable.

[0158] In one or more embodiments, before the processor 501 performs the process of assigning values ​​to the hierarchical variables at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and combines the multiple assigned hierarchical variables to obtain multiple initial test data, it further performs the following:

[0159] By filtering the strategy information and multiple historical data, the correlation between the values ​​of at least two of the hierarchical variables is obtained;

[0160] Processor 501 performs the process of assigning values ​​to the hierarchical variables at least once, obtaining at least one assigned hierarchical variable corresponding to each hierarchical variable, and combining multiple assigned hierarchical variables to obtain multiple initial test data. Specifically, the process is as follows:

[0161] Based on the correlation between the values ​​of at least two of the hierarchical variables, each of the hierarchical variables is assigned a value at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable. The multiple assigned hierarchical variables are then combined to obtain multiple initial test data.

[0162] In one or more embodiments, the processor 501 performs the process of assigning values ​​to the hierarchical variables at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and combines multiple assigned hierarchical variables to obtain multiple initial test data. Specifically, the process is as follows:

[0163] Based on the set of values ​​of the hierarchical variables and at least one numerical value input by the user, the hierarchical variables are assigned a value at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable. Multiple assigned hierarchical variables are combined to obtain multiple initial test data.

[0164] In one or more embodiments, the processor 501 performs the process of assigning values ​​to the hierarchical variables at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and combines multiple assigned hierarchical variables to obtain multiple initial test data. Specifically, the process is as follows:

[0165] The hierarchical variables are assigned values ​​at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable.

[0166] Cartesian products are calculated for at least one assigned hierarchical variable corresponding to each of the multiple hierarchical variables to obtain multiple initial test data.

[0167] In one or more embodiments, processor 501 executes the process of determining the accurate value that conforms to the admission rule and the rejection value that does not conform to the admission rule corresponding to the rule admission variable, adding at least one rule admission variable to each of the initial test data and assigning an accurate value or a rejection value to the rule admission variable, thereby obtaining multiple test data. Specifically, the process involves:

[0168] Based on the multiple admission rules included in the policy information and at least one rule admission variable used by the admission rules, determine the accurate value of the rule admission variable that conforms to the admission rule and the rejection value that does not conform to the admission rule;

[0169] Add at least one rule admission variable to each of the initial test data and assign an accurate value or a rejection value to the rule admission variable, and add different rule admission variables to each of the initial test data multiple times to obtain multiple test data.

[0170] In one or more embodiments, before the processor 501 executes the determination of the exact value corresponding to the rule admission variable that conforms to the admission rule and the rejection value that does not conform to the admission rule based on the plurality of admission rules included in the policy information and at least one rule admission variable used by the admission rules, it further executes:

[0171] Based on the strategy information and the multiple historical data, determine a set of values ​​for multiple rule admission variables;

[0172] Processor 501 executes the process of determining, based on the multiple admission rules included in the policy information and at least one rule admission variable used by the admission rules, the exact value corresponding to the rule admission variable that conforms to the admission rule and the rejection value that does not conform to the admission rule, specifically:

[0173] Based on the multiple admission rules included in the policy information and at least one rule admission variable used by the admission rules, the set of values ​​of the rule admission variable is divided into a set of accurate values ​​that conform to the admission rules and a set of rejection values ​​that do not conform to the admission rules.

[0174] The accurate value and rejection value corresponding to the rule admission variable are determined based on at least one value from the set of accurate values ​​and at least one value from the set of rejection values, respectively.

[0175] This embodiment of the specification improves the accuracy of the obtained stratified variables and rule-based admission variables by filtering and analyzing policy information and multiple historical data, thus providing reliable data support for the subsequent construction of test data. Furthermore, by assigning values ​​to the stratified variables at least once, at least one assigned stratified variable is obtained for each stratified variable. These assigned stratified variables are combined to obtain multiple initial test data sets. The accurate values ​​that conform to the admission rules and the rejection values ​​that do not conform to the admission rules for the rule-based admission variables are determined. At least one rule-based admission variable is added to each initial test data set and an accurate or rejection value is assigned to it, resulting in multiple test data sets sufficient for training a decision-making model that uses policy information for decision-making. By assigning values ​​to multiple stratified variables multiple times and constructing initial test data in its entirety, this embodiment of the specification enriches the final test data, ensuring that the test data covers normal, abnormal, and extreme scenarios. Furthermore, the embodiments in this specification reduce the amount of test data generated by assigning accurate and rejection values ​​only to the rule admission variables and adding them to the initial test data. This allows for the focus on constructing test data that conforms to transaction logic, thereby improving the efficiency of test data generation and the overall quality of the test data.

[0176] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented. Each of the above methods can be executed by a computer program instructing related hardware. The program corresponding to each method can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The storage medium of the electronic device 500 can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.

[0177] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0178] The above-disclosed embodiments are merely preferred embodiments of this specification and should not be construed as limiting the scope of this specification. Therefore, any equivalent variations made in accordance with the claims of this specification shall still fall within the scope of this specification.

Claims

1. A method for generating test data, the method comprising: By filtering strategy information and multiple historical data, multiple hierarchical variables and multiple rule admission variables are obtained; The hierarchical variables are assigned values ​​at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable. The multiple assigned hierarchical variables are combined to obtain multiple initial test data. The multiple assigned hierarchical variables included in the initial test data are obtained by assigning values ​​to different hierarchical variables. Determine the accurate values ​​that conform to the admission rule and the rejection values ​​that do not conform to the admission rule corresponding to the rule admission variable. Add at least one rule admission variable to each of the initial test data and assign an accurate value or a rejection value to the rule admission variable to obtain multiple test data. The test data is used at least to train a decision model that makes decisions using the policy information.

2. The test data generation method according to claim 1, wherein filtering the strategy information and multiple historical data to obtain multiple hierarchical variables and multiple rule admission variables includes: By filtering strategy information and multiple historical data, multiple feature variables and the set of values ​​for the feature variables are obtained; Based on the type of each feature variable, the multiple feature variables are divided into multiple stratified variables and multiple rule admission variables; wherein, the value set of the stratified variables is used to assign a value to the stratified variables at least once, and the value set of the rule admission variables is used to determine the exact value and rejection value of the rule admission variables.

3. The test data generation method according to claim 2, wherein filtering the strategy information and multiple historical data to obtain multiple feature variables and a set of values ​​for the feature variables includes: By filtering strategy information and multiple historical data, several feature variables are obtained; The value type of the feature variable is analyzed to be discrete or continuous. Based on the numerical acquisition method corresponding to different value types, multiple values ​​of the feature variable are filtered from the strategy information and the multiple historical data to obtain the value set of the feature variable.

4. The test data generation method according to claim 3, wherein the analysis of the value type of the feature variable is discrete or continuous, and according to the numerical acquisition method corresponding to different value types, multiple values ​​of the feature variable are filtered from the strategy information and the multiple historical data to obtain the value set of the feature variable, including: When the feature variable is discrete, the strategy information and multiple values ​​of the feature variable in the multiple historical data are fully recorded to obtain the set of values ​​of the feature variable; When the feature variable is continuous, the continuous value range of the feature variable is obtained from the strategy information and the multiple historical data, and at least one key value in the continuous value range of the feature variable is recorded to obtain the value set of the feature variable; wherein, the key value includes at least the boundary value of the continuous value range of the feature variable.

5. The test data generation method according to claim 1, before assigning values ​​to the hierarchical variables at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and combining multiple assigned hierarchical variables to obtain multiple initial test data, further includes: By filtering the strategy information and multiple historical data, the correlation between the values ​​of at least two of the hierarchical variables is obtained; The process involves assigning values ​​to the hierarchical variables at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and combining multiple assigned hierarchical variables to obtain multiple initial test data, including: Based on the correlation between the values ​​of at least two of the hierarchical variables, each of the hierarchical variables is assigned a value at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable. The multiple assigned hierarchical variables are then combined to obtain multiple initial test data.

6. The test data generation method according to claim 2, wherein assigning values ​​to the hierarchical variables at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and combining the multiple assigned hierarchical variables to obtain multiple initial test data, includes: Based on the set of values ​​of the hierarchical variables and at least one numerical value input by the user, the hierarchical variables are assigned a value at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable. Multiple assigned hierarchical variables are combined to obtain multiple initial test data.

7. The test data generation method according to claim 1, wherein assigning values ​​to the hierarchical variables at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and combining the multiple assigned hierarchical variables to obtain multiple initial test data, includes: The hierarchical variables are assigned values ​​at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable. Cartesian products are calculated for at least one assigned hierarchical variable corresponding to each of the multiple hierarchical variables to obtain multiple initial test data.

8. The test data generation method according to claim 1, wherein determining the accurate value that conforms to the admission rule and the rejection value that does not conform to the admission rule corresponding to the rule admission variable, adding at least one rule admission variable to each of the initial test data and assigning an accurate value or a rejection value to the rule admission variable, to obtain multiple test data, includes: Based on the multiple admission rules included in the policy information and at least one rule admission variable used by the admission rules, determine the accurate value of the rule admission variable that conforms to the admission rule and the rejection value that does not conform to the admission rule; Add at least one rule admission variable to each of the initial test data and assign an accurate value or a rejection value to the rule admission variable, and add different rule admission variables to each of the initial test data multiple times to obtain multiple test data.

9. The test data generation method according to claim 8, before determining the accurate value of the rule admission variable that conforms to the admission rule and the rejection value that does not conform to the admission rule based on the multiple admission rules included in the strategy information and at least one rule admission variable used by the admission rule, the method further includes: Based on the strategy information and the multiple historical data, determine a set of values ​​for multiple rule admission variables; The step of determining the exact value of the rule admission variable that conforms to the admission rule and the rejection value that does not conform to the admission rule, based on the multiple admission rules included in the policy information and at least one rule admission variable used by the admission rules, includes: Based on the multiple admission rules included in the policy information and at least one rule admission variable used by the admission rules, the set of values ​​of the rule admission variable is divided into a set of accurate values ​​that conform to the admission rules and a set of rejection values ​​that do not conform to the admission rules. The accurate value and rejection value corresponding to the rule admission variable are determined based on at least one value from the set of accurate values ​​and at least one value from the set of rejection values, respectively.

10. A test data generation apparatus, the apparatus comprising: The data filtering module is used to filter strategy information and multiple historical data to obtain multiple hierarchical variables and multiple rule admission variables; The first data generation module is used to assign values ​​to the hierarchical variables at least once to obtain at least one assigned hierarchical variable corresponding to each hierarchical variable, and to combine multiple assigned hierarchical variables to obtain multiple initial test data; wherein, the multiple assigned hierarchical variables included in the initial test data are obtained by assigning values ​​to different hierarchical variables; The second data generation module is used to determine the accurate values ​​that conform to the admission rules and the rejection values ​​that do not conform to the admission rules corresponding to the rule admission variables, add at least one rule admission variable to each of the initial test data and assign an accurate value or a rejection value to the rule admission variable to obtain multiple test data; wherein, the test data is used at least to train a decision model that makes decisions using the policy information.

11. A computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method steps of any one of claims 1 to 9.

12. A computer program product storing a plurality of instructions adapted for loading by a processor and executing the method steps of any one of claims 1 to 9.

13. An electronic device, comprising: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed the method steps as claimed in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method for quickly generating decision test data

    CN116204417A

  • Multivariable decision processing method and device, storage medium and electronic equipment

    CN119576297A