Test data generation method and device
By acquiring and analyzing the characteristic rules and data rules of production data, and generating test data close to the characteristics of production data, the problem of large differences between test data and production data in the prior art is solved, and the effectiveness and privacy and security of test data are achieved.
Patent Information
- Application Number
- CN202110863087.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-07-29
AI Technical Summary
In the prior art, the test data generated by manual rules differs greatly from the actual production data, making it difficult to maintain the characteristics and complexity of the production data, and there is a risk of privacy leakage.
By obtaining production data, extracting multiple field features and performing statistics, statistical rules and data rules are obtained, and batches of test data are generated based on these rules to ensure that the test data is as close as possible to the production data.
Simplify and reduce manual intervention to avoid loss of data characteristics, generate test data close to production data, maintain the characteristics and complexity of production data, eliminate the hidden dangers of production data leakage, and regularly update and adjust the manufacturing strategy.
Smart Images

Figure CN113568949B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a test data generation method and device. Background Art
[0002] At present, online financial business is being carried out more and more frequently, and the corresponding business systems are becoming more mature and complex. Most of the customer transaction data in the financial industry, such as customer name, account number, transaction party account number, transaction type, product name, etc., are character data. Customers' financial transaction business data has high security requirements. Even if it is desensitized through data deformation, the original customer transaction information may be restored after being leaked due to the deformation algorithm. Therefore, using production data such as financial transaction business data as test data has the risk of privacy leakage.
[0003] As the number of business systems continues to increase, non-critical business systems often use manual rule-based number generation to prepare test data. However, manual rules often vary greatly due to different personal experiences and are difficult to keep close to the characteristics of production data.
[0004] To address the above problems, no effective solution has been proposed yet. Summary of the invention
[0005] The embodiments of this specification provide a test data generation method and device to solve the problem in the prior art that the test data generated by artificial rules is significantly different from the actual production data.
[0006] An embodiment of the present specification provides a test data generation method, comprising: acquiring production data; extracting multiple field features from the production data, and performing statistics on the extracted multiple field features to obtain statistical rules corresponding to the production data; based on the statistical rules, extracting data rules corresponding to the production data, wherein the data rules are used to characterize the dependency relationship between the numerical values of the multiple field features; and generating batches of test data according to the statistical rules and the data rules.
[0007] The embodiments of the present specification also provide a test data generating device, including: an acquisition module, used to acquire production data; an extraction module, used to extract multiple field features in the production data, and perform statistics on the extracted multiple field features to obtain statistical rules corresponding to the production data; an extraction module, used to extract data rules corresponding to the production data based on the statistical rules, wherein the data rules are used to characterize the dependency relationship between the numerical values of the multiple field features; and a generation module, used to generate batches of test data according to the statistical rules and the data rules.
[0008] An embodiment of the present specification also provides a computer device, including a processor and a memory for storing processor-executable instructions, wherein the processor implements the steps of the test data generating method described in any of the above embodiments when executing the instructions.
[0009] The embodiments of the present specification also provide a computer-readable storage medium on which computer instructions are stored. When the instructions are executed, the steps of the test data generating method described in any of the above embodiments are implemented.
[0010] In an embodiment of the present specification, a test data generation method is provided, wherein the server can obtain actual production data, extract multiple field features from the production data, and perform statistics on the extracted multiple field features to obtain statistical rules corresponding to the production data, and based on the statistical rules, extract data rules for characterizing the dependency relationship between the numerical values of the multiple field features, and then generate batches of test data according to the statistical rules and the data rules. In the above scheme, statistical characteristics and data characteristics are extracted according to production data to obtain statistical rules and data rules for data, which can simplify manual intervention and avoid data characteristic loss; batch test data is generated using the obtained statistical rules and data rules, which can make the generated test data as close to the production data as possible, can keep the characteristics and complexity of the production data as much as possible, avoid data distortion, and eliminate the hidden dangers of production data leakage. In addition, production data can be synchronized regularly to obtain the latest data characteristics, and the number-making strategy can be updated and adjusted in time to maintain the validity of the test data. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of this specification, constitute a part of this specification, and do not constitute a limitation of this specification. In the drawings:
[0012] Figure 1 An overall flow chart of a test data generation method in an embodiment of this specification is shown;
[0013] Figure 2 A schematic diagram of statistical rules and data rule extraction in an embodiment of this specification is shown;
[0014] Figure 3 A schematic diagram of a fully connected graph in an embodiment of this specification is shown;
[0015] Figure 4 A schematic diagram of an undirected connected graph in an embodiment of this specification is shown;
[0016] Figure 5 A schematic diagram of a directed connected graph in an embodiment of this specification is shown;
[0017] Figure 6A schematic diagram of the test data generation process in the embodiment of this specification is shown;
[0018] Figure 7 A schematic diagram of the test data rule verification process in an embodiment of this specification is shown;
[0019] Figure 8 A flow chart showing a method for generating test data in an embodiment of this specification is shown;
[0020] Fig. 9 A schematic diagram of a test data generating device in an embodiment of this specification is shown;
[0021] Fig.10 A schematic diagram of a computer device in an embodiment of the present specification is shown. DETAILED DESCRIPTION
[0022] The principles and spirit of this specification will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement this specification, and are not intended to limit the scope of this specification in any way. On the contrary, these embodiments are provided to make this specification more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.
[0023] Those skilled in the art will appreciate that the embodiments of this specification may be implemented as a system, device, method, or computer program product. Therefore, this specification may be implemented in the following forms: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0024] The embodiment of this specification provides a test data generation method. In a scenario example of this specification, the test data generation method can be applied to a number-making server. The number-making server can obtain business data from the database of the business server. Afterwards, the number-making server can extract multiple field features in the production data, and perform statistics on the extracted multiple field features to obtain statistical rules corresponding to the production data. The number-making server can extract data rules corresponding to the production data based on the statistical rules. The data rules can be used to characterize the dependency between the numerical values of the multiple field features. Afterwards, the number-making server can generate batches of test data based on the statistical rules and the data rules. The number-making server can send the generated batches of test data to the test server for software testing.
[0025] The above-mentioned number-generating server, business server and test server may be a single server, a server cluster including multiple servers, or a cloud server, which is not limited in this application.
[0026] Figure 1 This is the overall flow chart of the test data generation method. Figure 1 As shown, the overall flow chart includes production data input 1, rule extraction 2, statistical rule storage device 3, data rule storage device 4, data generation 5, rule verification 6 and algorithm storage device 7. Among them, rule extraction 2 can be combined with algorithm storage device 7 to complete the extraction of statistical rules and data rules. Statistical rule storage device 3 can save the extracted statistical rules. Data rule storage device 4 can save the extracted data rules. Data generation 5 is used to complete test data generation in combination with statistical rules and data rules. Rule verification 6 can be combined with statistical rule storage device 3 and algorithm storage device 7 to perform rule verification on the generated test data to obtain test data that meets the statistical rules.
[0027] like Figure 1 As shown, the system first extracts rules based on production data. The specific process is: the statistical value of production data in each field dimension can be obtained through the statistical algorithm in the algorithm storage device 7, the statistical caliber is distinguished according to the field type, the statistical rule construction is completed, and the statistical rule is stored in the statistical rule storage device 3. Afterwards, the numerical features of each field can be calculated according to the field type through the feature algorithm in the algorithm storage device 7, the data rule construction is completed, and the data rule is stored in the data rule storage device 4. Data generation 5 can realize the generation of test data. The specific process is: single data generation is completed through data rules, and multiple data generation is completed through statistical rules. Then, the generated test data is input into rule verification 6. Rule verification 6 can be combined with algorithm storage device 7 to recalculate the statistical values of each field. If the statistical value deviation of the test data is within the preset threshold range (such as 5%), it can be considered that the test data meets the number generation requirements in a macro sense. If the statistical value deviation of the test data is large, the rule extraction logic is adjusted and the test data generation is re-performed. By regularly inputting production data, the rule extraction algorithm logic can be automatically adjusted by this method to achieve synchronous update of the number generation process.
[0028] Figure 2 Schematic diagram of statistical rules and data rule extraction in the embodiment of this specification is shown. Figure 2 In the example, statistical rule extraction can be completed by field feature extraction 201 and field value statistics 202. Data rule extraction can be completed by field dependency extraction 203. The specific process is as follows:
[0029] (I) Statistical rule extraction
[0030] The field feature extraction can be completed according to the field type rules preset in the field rule storage device 204 in combination with the field value. The number of features calculated by the preset field type rules in combination with the field value needs to be set to an upper limit to avoid the scenario where uniformly distributed data has no regular features. The specific rules for extracting character features are as follows:
[0031] (1) Identification and enumeration
[0032] That is, the value in the field exists in the data dictionary and belongs to enumeration data. The characteristic of this type of field is the enumeration dictionary value. The calculation process is:
[0033] Step 1: Get the number of duplicate values of the field in the production data. If it is less than the preset threshold, proceed to step 2.
[0034] Step 2: Number the production data, take the data with odd serial numbers to calculate the number of values after deduplication, and take the data with even serial numbers to calculate the number of values after deduplication. If the two are equal, proceed to step 3.
[0035] Step 3: Compare the number of values obtained in step 1 and step 2. If they are equal, the field value is considered to be an enumeration.
[0036] For example, the field "sex" has 2 deduplicated values ("male" and "female" respectively) in 10,000 data records; the number of deduplicated values in 5,000 records with odd serial numbers is 2, and the number of deduplicated values in 5,000 records with even serial numbers is 2. This means that the field is an enumeration type, and the characteristics of the field are: "male" and "female".
[0037] (2) Determine the finite length
[0038] That is, the field is not of enumeration type, but the field length is enumerable.
[0039] The calculation process is:
[0040] Step 1: Get the number of values of the field length after deduplication in the production data. If it is less than the preset threshold, go to step 2.
[0041] Step 2: Number the production data, take the data with odd serial numbers to calculate the value of the field length after deduplication, and take the data with even serial numbers to calculate the value of the field length after deduplication. If the two are equal, go to step 3.
[0042] Step 3: Compare the values obtained in step 1 and step 2. If they are equal, the field value is considered to be of finite length.
[0043] For example, the field "phone" is not an enumeration type, but among 10,000 data records, the number of values after length deduplication is 3, and the values are [8, 11, 13]. The values after length deduplication of 5,000 records of data with odd numbers are [8, 11, 13], and the corresponding values of 5,000 records of data with even numbers are [8, 11, 13]. Therefore, the field is deemed to be of finite length, and the length characteristic of the field is: [8, 11, 13].
[0044] (3) Identify character composition
[0045] That is, identify the composition of the segment, including Chinese, English, numbers, mixed or special symbol data, to form field features. The calculation process is:
[0046] Step 1: Calculate the unicode value of the field value character by character and identify the field value type: Chinese, English, numbers, and special symbols.
[0047] Step 2: If the field contains only one unicode type, it is a single character type; otherwise, it is a mixed character type, and go to step 3.
[0048] Step 3: Calculate the proportion of each type of characters in step 2.
[0049] For example, the field composition and proportion of the field "mobile" is: numbers (100%); the field composition and proportion of the field "title" is: Chinese (95%), English (2%), numbers (3%)
[0050] (4) Calculate length distribution
[0051] That is, the probability distribution of the field length is calculated, which is used to control the character length when generating test data. The calculation process is:
[0052] Step 1: Calculate the mean and median of the field length. If the two are close (within 5%), the length distribution of the field is relatively even and is calculated as an average interval.
[0053] Step 2: Divide the data into two intervals based on the median, and calculate the mean and median after the division respectively. If the two are not close, repeat step 2 until they are close.
[0054] Step 3: Use the data interval divided in step 2 as the length distribution interval, take the average as the representative length, and calculate the length proportion of each interval.
[0055] For example, among 100 records, the length of the field "title" is: 7 (25 records), 8 (25 records), 10 (20 records), 12 (24 records), 15 (2 records), 18-20 (1 record each), 25 (1 record), then the calculated length distribution interval is shown in Table 1 below:
[0056] Table 1
[0057] length Proportion 9 94% 19 6%
[0058] The output of the above results is used as the field feature to input the field value statistics 202, calculate the occurrence probability of each field feature, and complete the statistical rule extraction. The data structure of the statistical rule is shown in Table 2 below. In Table 2, Field represents the field name as the unique identifier of the identified field; Features represents the feature item as the dimension of the field feature. Probability represents the value probability as the basis for the value of the field feature.
[0059] Table 2
[0060] Field Features Probility
[0061] Exemplarily, the data structure of the obtained statistical rules is shown in Table 3 below:
[0062] Table 3
[0063] Field Features Probility sex Enum[male, female] Male: 45%; Female: 55% mobile Type[Digital] Digital: 100% title Type[CN,Letter,Digital] CN:95%; EN:2%; Digital:3% title Length "9”:94%;"19”:6%
[0064] (II) Data rule extraction
[0065] Data rule extraction is mainly to calculate the dependency between field values and save the dependency as a rule. Take the purchase record of a fund product as an example. If the production data is as shown in Table 4, Sex represents gender, Income represents income, Fund risk represents user fund risk, and Mobile represents user mobile phone number:
[0066] Table 4
[0067] Sex(S) Income(I) Fund risk(FR) Mobile(M) Male H R3 13918276475 Female L R1 13518246343 Male M R1 17836542398 Male H R3 13965423654 Male M R3 13135468455
[0068] (1) Take the fields as nodes of an undirected connected graph and construct a completely connected graph, such as Figure 3 As shown. Figure 3 In a fully connected graph, any two nodes are connected.
[0069] (2) Referring to the above statistical rule extraction, the field feature recognition and probability calculation are completed. The calculated statistical rules are shown in Table 5 below:
[0070] Table 5
[0071] Field Features Probility Sex Enum[Male, Female] Male: 80%; Female: 20% Income Enum[H, M, L] H: 40%; M: 40%; L: 20% Fund risk Enum[R1, R3] R1: 40%; R3: 60% Mobile Type[Digital] Digital: 100% Mobile Length 11:100%
[0072] (3) Starting from any node, calculate the correlation of the eigenvalues with the adjacent nodes as the weight of the edge between the two nodes, including the following steps:
[0073] Step 1: Delete invalid edges
[0074] Calculate the probability of occurrence of the adjacent node features when a feature of the current node is true. If the probability of occurrence is close to the probability of the statistical rule, the two nodes are not related and the edge is eliminated. If the probability of occurrence changes, the two nodes are related and go to step 2.
[0075] For example, when Sex is Male, the characteristics of Income are [H: 50%; M: 50%], which has changed, indicating that the value of Sex affects the value of Income, and the edge E(S, I) is retained. When Mobile is a numeric type, the characteristics of Income are [H: 40%; M: 40%; L: 20%], which has not changed, indicating that Mobile will not affect the value of Income, and the edge E(M, I) is deleted. The undirected graph after the calculation is as follows Figure 4 shown.
[0076] Step 2: Calculate the correlation between nodes
[0077] The correlation between nodes is expressed as weight, and the weight w for a single feature is i,B The calculation formula is:
[0078]
[0079] Among them, w i,B represents the weight between the i-th feature of node A and node B, F i,B is the number of features of node B when the value of node A is the i-th feature, F i+1,B is the number of features of node B when the value of node A is the i+1th feature, P i,Bm is the probability of occurrence of the mth feature of node B when the value of node A is the i-th feature, P i+1,Bm is the probability of occurrence of the mth feature of node B when the value of node A is the i+1th feature, and M is the total number of features of node B.
[0080] The weight W of the edge between nodes A and B AB The calculation formula is:
[0081]
[0082] Among them, w i,BRepresents the weight between the i-th feature of node A and node B, and N is the total number of features of node A.
[0083] Starting from any node, calculate the weight of the directed edge to other nodes. After the weight calculation in the above example, a weighted directed connected graph is obtained, such as Figure 5 shown.
[0084] According to the shortest path algorithm device 205, the shortest path can be obtained as: S→I→FR, with a value of 1.58. The data rules on the path are recalculated, and the results are shown in Table 6 below:
[0085] Table 6
[0086]
[0087] Please refer to Figure 6 , shows a schematic diagram of the test data generation process in the embodiment of this specification. Single data generation 601 can be combined with the rules in the data rule storage device 603 and the statistical rule storage device 604 to complete the single data field generation. The specific strategy is:
[0088] (1) Find the first field of the record according to the data rule, and generate the value of the field according to the statistical rule. For example, the first node of the shortest path is S, and the probability that the field is Male is 80%. The results are shown in Table 7:
[0089] Table 7
[0090] Sex(S) Income(I) Fund risk(FR) Mobile(M) Male
[0091] (2) Complete the generation of relevant fields according to the data rules. For example, under the premise that Sex is Male, the value of Income is H: 50%; M: 50%. The results are shown in Table 8:
[0092] Table 8
[0093] Sex(S) Income(I) Fund risk(FR) Mobile(M) Male H
[0094] Under the premise that Income is H, the Fund risk value is R1: 0%; R3: 100%. The results are shown in Table 9:
[0095] Table 9
[0096] Sex(S) Income(I) Fund risk(FR) Mobile(M) Male H R3
[0097] (3) Complete the generation of the remaining fields according to the statistical rules. For example, the generation rule of the Mobile field is a number with a length of 11 digits. The results are shown in Table 10:
[0098] Table 10
[0099] Sex(S) Income(I) Fund risk(FR) Mobile(M) Male H R3 65424569853
[0100] The multiple data generation 602 can generate batch data according to the single data generation 601 and the statistical rules in the statistical rule storage device 604 to obtain batch test data.
[0101] Please refer to Figure 7 , shows a schematic diagram of the test data rule verification process in an embodiment of this specification. Figure 7 As shown, field characteristic extraction 701 can complete the extraction of test data field characteristics according to the field type rules preset by field rule storage device 704 and field values. This process is similar to the process of production data processing. The extracted field characteristics can be counted via field value statistics 702 to generate test data statistical rules. Afterwards, the generated test data statistical rules can be input into rule comparison 703. Rule comparison 703 can compare the production data statistical rules implemented by statistical rule storage device 705 with the test data statistical rules. The specific comparison strategy is as follows:
[0102] (1) Compare the feature rules of each field. If the feature rules of the test data contain the feature rules of the production data, the number generation requirement is met. Otherwise, there is a missing rule.
[0103] (2) Compare the probability distribution of data with the same feature rules and set a difference threshold. If the difference in data probability is within the threshold, the number generation requirement is met; otherwise, the number generation difference is too large.
[0104] (3) If the test data cannot pass the consistency comparison of statistical rules, it is necessary to regenerate the test data until it meets the rule verification.
[0105] The method in the above embodiment automatically extracts data rules by analyzing the characteristics of production data, and retains the characteristics of production data as much as possible by analyzing the data distribution and data correlation of the fields, avoiding the singleness of artificial rule number generation, and can dynamically track the latest data characteristics as the business continues to develop, keep the test data and production data synchronized and adjusted, and support the effective implementation of the test work. By extracting statistical characteristics and data characteristics based on production data, manual intervention can be simplified to avoid data feature loss. Using the above rules to generate data and evaluating the effectiveness of the generated results can effectively avoid the probability problem of number generation in some fields, which leads to difference amplification and data distortion. In addition, by regularly synchronizing production data, the latest data features can be obtained, and the number generation strategy can be adjusted more tightly in time to maintain the effectiveness of the test data.
[0106] Based on the above scenario example, this specification also provides a test data generation method. Figure 8A flow chart of a test data generating method in an embodiment of the present specification is shown. Although the present specification provides method operation steps or device structures as shown in the following embodiments or drawings, more or fewer operation steps or module units may be included in the method or device based on routine or no creative labor. In the steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure described in the embodiments of the present specification and shown in the drawings. When the method or module structure is applied to an actual device or terminal product, it can be connected according to the method or module structure shown in the embodiments or drawings for sequential execution or parallel execution (for example, a parallel processor or a multi-threaded processing environment, or even a distributed processing environment).
[0107] Specifically, Figure 8 As shown, a test data generation method provided in an embodiment of this specification may include the following steps:
[0108] Step S801, obtaining production data.
[0109] The method in the embodiments of this specification can be applied to a data generation server. The data generation server can obtain production data from a business server, or can obtain production data from a data server. The production data can be various business data, for example, it can include character data such as customer name, account number, transaction party account number, transaction type, product name, etc. Obtaining production data can include large batches of production data, and a piece of production data can include multiple fields.
[0110] Step S802: extract multiple field features from the production data, and perform statistics on the extracted multiple field features to obtain statistical rules corresponding to the production data.
[0111] The data generation server can extract multiple field features from the production data, and perform statistics on the extracted multiple field features to obtain statistical rules corresponding to the production data. For example, if the production data includes two fields, the features of the first field in the batch production data are extracted and statistically analyzed, and the features of the second field are extracted and statistically analyzed to obtain statistical rules corresponding to the production data. The statistical rules may include rules such as the name, type, and probability of the field features.
[0112] Step S803: extracting data rules corresponding to the production data based on the statistical rules, wherein the data rules are used to characterize the dependency relationship between the numerical values of the multiple field features.
[0113] After obtaining the statistical rules of production data, the data rules corresponding to the production data can be extracted based on the statistical rules. Among them, the data rules are used to characterize the dependency relationship between the numerical values of multiple field features. The dependency relationship between the numerical values of multiple field features refers to the influence of the value of one field feature on the value of another field feature. For example, when feature A takes a1, feature B can take b1 and b2, and when feature A takes a2, feature B can only take b1. For another example, when feature A takes a1, feature B can take b1 and b2, and the probabilities of taking values are 20% and 80% respectively; and when feature A takes a2, feature B can also take b1 and b2, but the probabilities of taking values are 30% and 70% respectively. In this case, it can be seen that there is a dependency relationship between the numerical value of feature A and the type and / or probability of taking values of feature B. The data rules are used to characterize this dependency relationship.
[0114] Step S804: Generate batches of test data according to the statistical rules and the data rules.
[0115] After obtaining the statistical rules and data rules of the production data, batches of test data can be generated according to the statistical rules and data rules.
[0116] In the above embodiment, statistical characteristics and data characteristics are extracted according to production data, and statistical rules and data rules of data are obtained, which can simplify manual intervention and avoid loss of data characteristics; batch test data is generated by using the obtained statistical rules and data rules, which can automatically generate test data and make the generated test data as close to production data as possible, which can maintain the characteristics and complexity of production data as much as possible, avoid data distortion, and eliminate the hidden dangers of production data leakage. In addition, production data can be synchronized regularly to obtain the latest data characteristics, and the number-making strategy can be updated and adjusted in time to maintain the validity of test data.
[0117] In some embodiments of the present specification, extracting multiple field features in the production data and performing statistics on the extracted multiple field features may include at least one of the following: determining whether a specified field in the production data is enumeration type data; determining whether a specified field in the production data is finite length data; identifying the character composition of the specified field in the production data; calculating the probability distribution of the field length of the specified field in the production data.
[0118] Specifically, multiple field features in the production data can be extracted. Among them, multiple field features mean that the production data includes multiple fields, and for each of the multiple fields, corresponding features can be extracted. It can be determined whether the specified field in the production data is enumerated type data. Among them, enumerated type data means that the value of the field exists in the data dictionary. It can be judged whether the specified field in the production data is limited length data. Limited length data means that the length of the field is enumerable. The character composition of the specified field in the production data can be identified. The character composition refers to the character type contained in the field, and the character type can include Chinese, English, numbers, mixed or special symbol data. The probability distribution of the field length of the specified field in the production data can be calculated. That is, the probability distribution of the field length of the field is calculated, which is used to control the character length when generating test data. In the above manner, multiple field features can be extracted to obtain statistical rules for production data.
[0119] In some embodiments of the present specification, determining whether the designated field in the production data is enumerated type data may include: reading the value of the designated field in the production data; removing duplicates from the value of the designated field to obtain a first number of the value of the designated field; when the first number is less than the preset number, numbering the production data, removing duplicates from the value of the designated field of the production data with an odd sequence number, obtaining a second number of the value of the designated field of the production data with an odd sequence number, removing duplicates from the value of the designated field of the production data with an even sequence number, obtaining a third number of the value of the designated field of the production data with an even sequence number; judging whether the first number, the second number, and the third number are equal; when judging that the first number, the second number, and the third number are equal, determining that the designated field of the production data is enumerated type data; using the value obtained after removing duplicates as the enumerated value, and calculating the proportion of the enumerated value. In the above manner, it is possible to determine whether the designated field of the production data is enumerated type data, and it is also possible to count the proportion of each enumerated value.
[0120] In some embodiments of the present specification, judging whether the designated field in the production data is limited length data may include: calculating the field length of the designated field of the production data and removing the calculated field length to obtain a first field length; numbering the production data, calculating the field length of the designated field of the production data with an odd sequence number and removing the duplicates to obtain a second field length, calculating the field length of the designated field of the production data with an even sequence number and removing the duplicates to obtain a third field length; determining whether the first field length, the second field length, and the third field length are the same; if it is determined that the first field length, the second field length, and the third field length are the same, determining that the designated field of the production data is limited length data; and calculating the proportion of each field length in the first field length. In the above manner, it is possible to judge whether the designated field of the production data is limited length data, and to count the proportion of each field length, so that the proportion of the field length of the generated test data is consistent with the production data.
[0121] In some embodiments of the present specification, identifying the character composition of the designated field in the production data may include: identifying the designated field of the production data character by character to obtain the character type of the designated field; if the designated field contains only one character type, the designated field is a single character type; if the designated field contains multiple character types, then the composition ratio of each character type in the designated field is calculated. In the above manner, the character type of the designated field in the production data can be identified, and the composition ratio of each character type can be calculated, so that the character composition of the field of the generated test data is consistent with the production data.
[0122] In some embodiments of the present specification, calculating the probability distribution of the field length of the specified field in the production data may include: calculating the mean and median of the field length of the specified field of the production data; when the difference between the mean and the median exceeds a preset range, dividing the production data into two intervals with the median as the dividing point, and calculating the mean and median of the field length of the specified field of the production data in each of the two intervals obtained after the division, repeating this step until the mean and median corresponding to the interval obtained by the division do not exceed the preset range; calculating the mean of the field length of the specified field of the production data in each of the multiple intervals obtained by the division, and using the mean corresponding to each interval as the representative field length corresponding to each interval; calculating the percentage of the number of production data in each interval to the total number of production data as the proportion corresponding to each interval. In the above manner, the probability distribution of each field of the production data can be calculated, which is convenient for controlling the field length distribution of the generated test data.
[0123] In some embodiments of the present specification, extracting data rules corresponding to the production data based on the statistical rules may include: constructing a completely connected graph, wherein any two nodes among a plurality of nodes in the completely connected graph are connected by an undirected edge, and each node among the plurality of nodes corresponds to each field among a plurality of fields in the production data; determining whether there is a correlation between adjacent nodes in the completely connected graph based on the statistical rules; in the case where there is no correlation between adjacent nodes, deleting invalid edges between the adjacent nodes in the completely connected graph to obtain a target undirected connected graph; using the statistical rules, calculating the weight of directed edges from each node to adjacent nodes among a plurality of nodes in the target undirected connected graph to obtain a target directed connected graph; determining a target path based on the target directed connected graph, wherein the target path is a unidirectional path traversing the plurality of nodes in the target directed connected graph; and calculating data rules on the target path, wherein the data rules on the target path include field names and field features corresponding to each node among a plurality of nodes on the target path, and numerical dependencies between the node and the previous node.
[0124] Specifically, the production data includes multiple fields, each field can correspond to a node, and a completely connected graph can be constructed. Any two nodes in the multiple nodes in the completely connected graph are connected by undirected edges. Based on statistical rules, it can be determined whether there is a correlation between adjacent nodes in the completely connected graph. For example, for a pair of adjacent nodes, it can be determined whether the value change of one node has an impact on the value or value probability distribution of another node. If there is an impact, it means that there is a correlation between the pair of adjacent nodes, and the undirected edge is retained. If there is no impact, it means that there is no correlation between the pair of adjacent nodes, and the undirected edge is deleted. After this operation, the target undirected connected graph can be obtained. After that, the weight of the directed edge from each node to the adjacent node in the multiple nodes in the target undirected connected graph can be calculated by using statistical rules to obtain the target directed connected graph. The target path can be determined based on the target directed connected graph. Among them, the target path is a unidirectional path that traverses the multiple nodes in the target directed connected graph. In one embodiment, the target path can be determined using the shortest path algorithm. After obtaining the target path, the data rules on the target path can be obtained. The data rules on the target path may include the field name, field characteristics, and the value dependency relationship between the node and the previous node of each node in the multiple nodes on the target path. In the above manner, the corresponding data rules can be generated based on the statistical rules of the production data, taking into account the value dependency relationship between the fields of the production data, and the accuracy of the production test data can be improved.
[0125] In some embodiments of the present specification, obtaining production data may include: regularly obtaining production data for a target time period; correspondingly, generating batches of test data according to the statistical rules and the data rules may include: generating batches of test data according to the statistical rules and data rules corresponding to the currently obtained production data.
[0126] Specifically, the data generation server can obtain production data from the business server regularly, for example, every preset time period, from the business server. The preset time period can be daily, weekly, monthly, quarterly, or annual, etc. Accordingly, after obtaining the latest production data, multiple field features in the latest production data currently obtained can be extracted, and the extracted multiple field features can be statistically analyzed to obtain the statistical rules corresponding to the currently obtained production data, and based on the statistical rules, the corresponding data rules can be extracted. Afterwards, batches of test data can be generated according to the statistical rules and data rules corresponding to the currently obtained production data. By regularly obtaining production data and regularly extracting the corresponding statistical rules and data rules, the rules can be dynamically updated to ensure the validity of the test data.
[0127] In some embodiments of the present specification, generating batches of test data according to the statistical rules and the data rules may include: extracting multiple field features from the test data, and performing statistics on the extracted multiple field features to obtain statistical rules corresponding to the test data; comparing the statistical rules corresponding to the test data with the statistical rules corresponding to the production data to determine whether the test data is valid.
[0128] Specifically, in order to further ensure the correctness of the test data, the test data is first verified before being used for testing. Multiple field features in the test data can be extracted, and the extracted multiple field features are counted to obtain statistical rules corresponding to the test data. The specific extraction and statistical method can refer to the implementation method in the aforementioned embodiment. After obtaining the statistical rules corresponding to the test data, the statistical rules corresponding to the test data are compared with the statistical rules of the production data to determine whether the generated test data is valid.
[0129] In one embodiment, the feature rules of each field can be compared. If the feature rules of the test data contain the feature rules of the production data, the numbering requirements are met, otherwise there is a missing rule. The data probability distribution of the same feature rules can be compared, and a difference threshold can be set. If the difference in data probability is within the threshold, the numbering requirements are met, otherwise the numbering difference is too large. If the test data cannot pass the consistency comparison of the statistical rules, the test data needs to be regenerated until the rule verification is met.
[0130] Based on the same inventive concept, a test data generating device is also provided in the embodiments of this specification, as described in the following embodiments. Since the principle of solving the problem by the test data generating device is similar to that of the test data generating method, the implementation of the test data generating device can refer to the implementation of the test data generating method, and the repeated parts are not repeated. As used below, the term "unit" or "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived. Fig. 9 is a structural block diagram of a test data generating device according to an embodiment of this specification, such as Fig. 9 As shown, it includes: an acquisition module 901, an extraction module 902, an extraction module 903 and a generation module 904. The structure is described below.
[0131] The acquisition module 901 is used to acquire production data.
[0132] The extraction module 902 is used to extract multiple field features from the production data, and perform statistics on the extracted multiple field features to obtain statistical rules corresponding to the production data.
[0133] The extraction module 903 is used to extract the data rules corresponding to the production data based on the statistical rules, wherein the data rules are used to characterize the dependency relationship between the numerical values of the multiple field features.
[0134] The generating module 904 is used to generate batches of test data according to the statistical rules and the data rules.
[0135] From the above description, it can be seen that the embodiments of this specification achieve the following technical effects: by extracting statistical characteristics and data characteristics based on production data, statistical rules and data rules of data are obtained, which can simplify manual intervention and avoid loss of data characteristics; using the obtained statistical rules and data rules to generate batch test data, the generated test data can be as close to the production data as possible, the characteristics and complexity of the production data can be maintained as much as possible, data distortion can be avoided, and the hidden dangers of production data leakage can be eliminated. In addition, production data can be synchronized regularly to obtain the latest data characteristics, and the number-making strategy can be updated and adjusted in a timely manner to maintain the validity of the test data.
[0136] This specification also provides a computer device, which can be found in Fig.10The computer device structure diagram of the test data generation method provided by the embodiment of this specification is shown, and the computer device may specifically include an input device 10, a processor 12, and a memory 14. Among them, the memory 14 is used to store processor executable instructions. When the processor 12 executes the instructions, the steps of the test data generation method described in any of the above embodiments are implemented.
[0137] In this embodiment, the input device may specifically be one of the main devices for information exchange between the user and the computer system. The input device may include a keyboard, a mouse, a camera, a scanner, a light pen, a handwriting input board, a voice input device, etc.; the input device is used to input the original data and the program for processing these numbers into the computer. The input device can also obtain and receive data transmitted from other modules, units, and devices. The processor can be implemented in any appropriate manner. For example, the processor can take the form of a computer-readable medium, a logic gate, a switch, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a programmable logic controller, and an embedded microcontroller, etc., such as a microprocessor or a processor and a computer-readable program code (such as software or firmware) that can be executed by the (micro) processor. The memory may specifically be a memory device used to store information in modern information technology. The memory may include multiple levels. In a digital system, anything that can store binary data can be a memory; in an integrated circuit, a circuit with a storage function without a physical form is also called a memory, such as a RAM, a FIFO, etc.; in a system, a storage device with a physical form is also called a memory, such as a memory stick, a TF card, etc.
[0138] In this embodiment, the functions and effects specifically realized by the computer device can be explained in comparison with other embodiments and will not be described in detail here.
[0139] A computer storage medium based on the test data generation method is also provided in the implementation manner of this specification, wherein the computer storage medium stores computer program instructions, and when the computer program instructions are executed, the steps of the test data generation method described in any of the above embodiments are implemented.
[0140] In this embodiment, the storage medium includes, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a cache, a hard disk (HDD) or a memory card. The memory may be used to store computer program instructions. The network communication unit may be an interface for network connection communication set in accordance with the standard specified by the communication protocol.
[0141] In this embodiment, the functions and effects specifically implemented by the program instructions stored in the computer storage medium can be explained in comparison with other embodiments and will not be repeated here.
[0142] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of this specification can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of this specification are not limited to any specific combination of hardware and software.
[0143] It should be understood that the above description is for illustration and not for limitation. Many embodiments and many applications beyond the examples provided will be apparent to those skilled in the art upon reading the above description. Therefore, the scope of this specification should not be determined with reference to the above description, but should be determined with reference to the preceding claims and the full scope of equivalents to which these claims belong.
[0144] The above description is only the preferred embodiment of this specification and is not intended to limit this specification. For those skilled in the art, the embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included in the protection scope of this specification.
Claims
1. A test data generation method, characterized in that: include: Obtain production data; Extracting multiple field features from the production data, and performing statistics on the extracted multiple field features to obtain statistical rules corresponding to the production data; Based on the statistical rules, extracting data rules corresponding to the production data, wherein the data rules are used to characterize the dependency relationship between the numerical values of the multiple field features; Generate batches of test data according to the statistical rules and the data rules; The extracting of multiple field features in the production data and performing statistics on the extracted multiple field features include at least one of the following: determining whether the specified field in the production data is enumerated type data; determining whether the specified field in the production data is limited length data; identifying the character composition of the specified field in the production data; calculating the probability distribution of the field length of the specified field in the production data; Wherein, based on the statistical rules, extracting data rules corresponding to the production data includes: constructing a completely connected graph, wherein any two nodes among a plurality of nodes in the completely connected graph are connected by an undirected edge, and each node among the plurality of nodes corresponds to each field among a plurality of fields in the production data; determining whether there is a correlation between adjacent nodes in the completely connected graph based on the statistical rules; in the case that there is no correlation between adjacent nodes, deleting invalid edges between the adjacent nodes in the completely connected graph to obtain a target undirected connected graph; using the statistical rules, calculating the weight of directed edges from each node to adjacent nodes among a plurality of nodes in the target undirected connected graph to obtain a target directed connected graph; determining a target path based on the target directed connected graph, wherein the target path is a unidirectional path traversing the plurality of nodes in the target directed connected graph; calculating data rules on the target path, wherein the data rules on the target path include field names and field features corresponding to each node among a plurality of nodes on the target path, and numerical dependencies between the node and the previous node.
2. The method according to claim 1, characterized in that Determining whether a specified field in the production data is enumeration type data includes: Reading the value of a specified field in the production data; Deduplication of the values of the specified field is performed to obtain a first number of the values of the specified field; In the case where the first number is less than the preset number, the production data are numbered, and the values of the designated fields of the production data with odd serial numbers are deduplicated to obtain a second number of the values of the designated fields of the production data with odd serial numbers, and the values of the designated fields of the production data with even serial numbers are deduplicated to obtain a third number of the values of the designated fields of the production data with even serial numbers; determining whether the first quantity, the second quantity, and the third quantity are equal; In the case where it is determined that the first quantity, the second quantity and the third quantity are equal, determining that the designated field of the production data is enumeration type data; The obtained value after deduplication is used as the enumeration value, and the proportion of the enumeration value is calculated.
3. The method according to claim 1, characterized in that Determining whether a specified field in the production data is limited-length data includes: Calculating the field length of the specified field of the production data and removing duplicates from the calculated field length to obtain a first field length; The production data are numbered, the field length of the designated field of the production data with an odd sequence number is calculated and duplicate removal is performed to obtain a second field length, and the field length of the designated field of the production data with an even sequence number is calculated and duplicate removal is performed to obtain a third field length; Determining whether the first field length, the second field length, and the third field length are the same; In the case where it is determined that the first field length, the second field length and the third field length are the same, determining that the designated field of the production data is limited-length data; Calculate the proportion of each field length in the first field length.
4. The method according to claim 1, characterized in that: Identifying the character composition of a specified field in the production data, including: Identify the designated field of the production data character by character to obtain the character type of the designated field; If the specified field contains only one character type, then the specified field is of a single character type; If the designated field includes multiple character types, the composition ratio of each character type in the designated field is calculated.
5. The method according to claim 1, characterized in that Calculating the probability distribution of the field length of the specified field in the production data, including: Calculate the average and median of the field lengths of the specified fields of the production data; In the case where the difference between the mean and the median exceeds a preset range, the production data is divided into two intervals with the median as a dividing point, and the mean and median of the field length of the specified field of the production data in each of the two intervals obtained after the division are calculated, and this step is repeated until the mean and median corresponding to the intervals obtained after the division do not exceed the preset range; Calculating the average field length of the designated field of the production data in each of the plurality of intervals obtained by segmentation, and using the average corresponding to each interval as the representative field length corresponding to each interval; The percentage of the amount of production data in each interval to the total amount of production data is calculated as the proportion corresponding to each interval.
6. The method according to claim 1, characterized in that Access production data, including: Regularly obtain production data for the target time period; Accordingly, batches of test data are generated according to the statistical rules and the data rules, including: Generate batches of test data based on the statistical rules and data rules corresponding to the currently acquired production data.
7. A test data generating device, characterized in that: include: An acquisition module is used to acquire production data; An extraction module, used to extract multiple field features from the production data, and perform statistics on the extracted multiple field features to obtain statistical rules corresponding to the production data; An extraction module, configured to extract data rules corresponding to the production data based on the statistical rules, wherein the data rules are used to characterize the dependency relationship between the numerical values of the plurality of field features; A generating module, used for generating batches of test data according to the statistical rules and the data rules; The extraction module is specifically used for at least one of the following: determining whether the specified field in the production data is enumerated type data; determining whether the specified field in the production data is limited length data; identifying the character composition of the specified field in the production data; calculating the probability distribution of the field length of the specified field in the production data; The extraction module is specifically used to: construct a completely connected graph, wherein any two nodes among the multiple nodes in the completely connected graph are connected by an undirected edge, and each node among the multiple nodes corresponds to each field among the multiple fields in the production data; based on the statistical rules, determine whether there is a correlation between adjacent nodes in the completely connected graph; when there is no correlation between adjacent nodes, delete the invalid edges between the adjacent nodes in the completely connected graph to obtain a target undirected connected graph; using the statistical rules, calculate the weight of the directed edges from each node to the adjacent nodes in the multiple nodes in the target undirected connected graph to obtain a target directed connected graph; determine the target path based on the target directed connected graph, wherein the target path is a unidirectional path traversing the multiple nodes in the target directed connected graph; calculate the data rules on the target path, wherein the data rules on the target path include the field name and field characteristics corresponding to each node among the multiple nodes on the target path, and the numerical dependency between the node and the previous node.
8. A computer device, characterized in that: The method comprises a processor and a memory for storing processor-executable instructions, wherein the processor implements the steps of the method according to any one of claims 1 to 6 when executing the instructions.
9. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instructions are executed, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Software application data generating device and method
CN104572122A
Data rule generation method and device
CN112199416A