Data processing method and computing device

By generating test samples related to the main data policy group, using neural network models or mapping tables to determine the data generation rules, adjusting the rules operation objects and quantity, the problems of low efficiency and poor accuracy in the existing testing methods are solved, and efficient and accurate main data policy group testing is achieved.

CN120296000APending Publication Date: 2025-07-11HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410038106.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing master data policy group testing method requires testing a large number of samples, resulting in low testing efficiency and high computational overhead, and inaccurate test results when the sampled samples are not related to the policy.

Method used

By generating test samples related to the main data policy group, using a small number of test samples to determine whether the main data policy group is correct, using neural network models or mapping tables to determine the data generation rules, adjust the operation objects and number of data generation rules, and generate test samples that conform to the actual business rules.

Benefits of technology

It improves testing efficiency and accuracy, reduces calculation overhead, and can quickly determine whether the master data policy group is correct. The generated test samples have strong correlation with the policy group.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296000A_ABST
    Figure CN120296000A_ABST
Patent Text Reader

Abstract

According to the data processing method provided by the embodiment of the invention, the test samples related to the main data strategy group can be generated, and whether the main data strategy group is correct or not can be judged by using a small number of test samples, so that the test efficiency is improved. The method comprises the following steps: after acquiring a main data parameter and a main data strategy group, a computing device determines a data generation rule corresponding to a strategy of the main data strategy group, acquires sampling data from main data according to the main data parameter, and then modifies a main data parameter value in a copy of the sampling data by using the data generation rule to obtain a test sample; and displaying the sampling data and the test sample. The embodiment of the invention further provides a computing device, a computing device cluster, a computer readable storage medium and a computer program product which can implement the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a data processing method and a computing device. Background Art

[0002] Master data is the basic data shared by the system, which includes data of core business entities (such as customers, partners, etc.). The master data policy group includes, but is not limited to, data cleaning policies.

[0003] Currently, a testing method for the master data policy group is generally as follows: receiving the master data policy group input by the user, processing the master data according to the master data policy group, when the processing result is different from the expected result, determining that the master data policy group is incorrect, modifying the policy of the master data policy group, and looping through the above steps; when the processing result is the same as the expected result, determining that the master data policy group is correct.

[0004] The above method requires testing a large number of master data samples, resulting in low testing efficiency and large computational overhead. Summary of the Invention

[0005] This application provides a data processing method that can generate test samples related to the master data policy group, and can determine whether the master data policy group is correct by using a small number of test samples, thereby improving the testing efficiency.

[0006] In a first aspect, a data processing method is provided. This method is executed by a computing device of a cloud service platform. After obtaining master data parameters and a master data policy group, it determines the data generation rules corresponding to the policies and obtains sampling data from the master data according to the master data parameters. Then, the computing device uses the data generation rules to modify the master data parameter values in the copy of the sampling data to obtain test samples, and displays the sampling data and the test samples. Among them, the master data policy group includes multiple policies, and the master data parameters are parameters of at least one policy.

[0007] Implementing in this way can pre-configure policies and data generation rules with a corresponding relationship, so that the data generation rules related to the master data policy group can be obtained. The test samples generated using this data generation rule have a strong correlation with the master data policy group. Therefore, it is possible to determine whether the master data policy group is correct by using a small number of test samples. Compared with the existing testing methods that use a large number of samples, it can reduce the computational overhead and improve the testing efficiency. In the existing testing methods, the sampled samples may not be relevant to the policies, resulting in inaccurate test results. The test samples of this application have a strong correlation with the master data policy group, and using the above test samples can improve the accuracy of testing the master data policy.

[0008] In the first possible implementation of the first aspect, the computing device determines that the data generation rule corresponding to the policy includes: the computing device inputs the function name included in the policy into the neural network model, and outputs the data generation rule corresponding to the policy through the neural network model. For policies that do not belong to the training samples, after inputting the function name of the policy into the neural network model, the neural network model can also output valid data generation rules, so it has good scalability.

[0009] In the second possible implementation of the first aspect, the computing device determines that the data generation rule corresponding to the policy includes: the computing device determines the data generation rule corresponding to the policy in the mapping table. The mapping table includes the corresponding relationship between the policy and the data generation rule, and the data generation rule can be quickly found through the mapping table.

[0010] In the third possible implementation of the first aspect, the computing device modifies the main data parameter value in the copy of the sampled data using the data generation rule to obtain a test sample, including: the computing device modifies the main data parameter value in the copy of the sampled data using the data generation rule to obtain a candidate test sample; when each main data parameter value of the candidate test sample is within the value range of the main data parameter, the computing device uses the candidate test sample as the test sample. This can remove samples that do not conform to the actual business, so that the main data parameter values of the test samples conform to the value range of the actual business data, thus providing valid test samples.

[0011] Combined with the third possible implementation of the first aspect, in the fourth possible implementation of the first aspect, the computing device adjusts the operation object of the data generation rule, and uses the adjusted data generation rule to modify the main data parameter value in the copy of the sampled data to obtain a candidate test sample. The operation object includes at least one of a character, a word, a data component string, or a numerical value.

[0012] In this implementation, the computing device can adjust the operation object according to the user input instruction, or can automatically adjust the operation object according to the obtained data component features. When the computing device displays the data generation rule, it is convenient for the user to preview the data generation rule, so as to judge whether the main data policy is correct or repeated, providing a method for quickly checking the main data policy. When the data generation rule is incorrect or does not meet the actual requirements, the main data policy group can be re-input or the data generation rule can be adjusted, so as to quickly correct the main data policy group or the data generation rule.

[0013] It should be understood that in addition to adjusting the operation object of the data generation rule, the number of operation objects or the data operation type of the data generation rule can also be adjusted, specifically, one or more of them can be adjusted.

[0014] Combined with the fourth possible implementation manner of the first aspect, in some possible implementation manners, the computing device inputs the main data parameters into the language model, and outputs the data component features corresponding to the main data parameters through the language model; adjusts the data component string of the data generation rule to the data component features. This can accurately adjust the data components of the main data parameters and improve the flexibility of adjustment.

[0015] Combined with the fourth possible implementation manner of the first aspect, in some possible implementation manners, the data processing method of the present application further includes: the computing device combines the data generation rules corresponding to all the main data parameter values; when all the main data parameter values of the candidate test sample are within the value range of the main data parameters, the computing device adds a positive example identifier to the combined data generation rule combination; when at least one of the main data parameter values of the candidate test sample is not within the value range of the main data parameters, the computing device adds a negative example identifier to the combined data generation rule combination, and then displays the data generation rule combination with the positive example identifier and / or the data generation rule combination with the negative example identifier. Checking positive and negative examples can determine which data generation rules are compliant and which operation rules are non-compliant, so as to judge in advance whether the main data policy group meets the requirements. Optionally, the number of test samples is the same as the number of data generation rule combinations.

[0016] In the fifth possible implementation manner of the first aspect, the data processing method of the present application further includes: after the computing device processes the copy of the sampling data into a reference test sample using the adjusted partial data generation rules, it displays the reference test sample. Checking the differences among the sampling data, the test sample, and the reference test sample can determine whether the data generation rules are correct.

[0017] In some possible implementation manners, the computing device tests the sampling data and the test sample using the main data policy group; when the test data is consistent with the expected data, the computing device determines that the main data policy group is correct; when the test data is different from the expected data, the computing device determines that the main data policy group is incorrect. Since the test sample has a strong correlation with the main data policy group, only a small number of test samples are needed to determine whether the main data policy group is correct.

[0018] In some possible implementation manners, the computing device modifies the policy of the main data policy group, determines the data generation rule corresponding to the modified policy, and then uses the data generation rule corresponding to the modified policy to modify the main data parameter value in the copy of the sampling data to obtain a modified test sample, and displays the sampling data and the modified test sample.

[0019] A second aspect provides a computing device, which includes an interaction module, a policy rule module, a sampling module, and a data generation module. The interaction module is used to obtain main data parameters and a main data policy group; the policy rule module is used to determine a data generation rule corresponding to the policy; the sampling module is used to obtain sampling data from the main data according to the main data parameters; the data generation module is used to modify the main data parameter values in a copy of the sampling data by using the data generation rule to obtain a test sample; and the interaction module is used to display the sampling data and the test sample.

[0020] In some possible implementation manners, the policy rule module is specifically configured to input the function name included in the policy into a neural network model, and output the data generation rule corresponding to the policy through the neural network model.

[0021] In some possible implementation manners, the policy rule module is specifically configured to determine the data generation rule corresponding to the policy in a mapping table.

[0022] In some possible implementation manners, the data generation module is specifically configured to modify the main data parameter values in a copy of the sampling data by using the data generation rule to obtain a candidate test sample; when each main data parameter value of the candidate test sample is within the value range of the main data parameters, the candidate test sample is used as the test sample.

[0023] In some possible implementation manners, the policy rule module is further used to adjust the operation object of the data generation rule, and the data generation module is specifically configured to modify the main data parameter values in a copy of the sampling data by using the adjusted data generation rule to obtain a candidate test sample.

[0024] In some possible implementation manners, the computing device further includes a feature module, and the feature module is used to input the main data parameters into a language model and output the data component features corresponding to the main data parameters through the language model; the policy rule module is specifically configured to adjust the data component string of the data generation rule to the data component features when the operation object of the data generation rule includes the data component string.

[0025] In some possible implementation manners, the policy rule module is further used to adjust the number of operation objects of the data generation rule.

[0026] In some possible implementation manners, the policy rule module is further configured to combine the data generation rules corresponding to all the master data parameter values according to the policy relationship; when all the master data parameter values of the candidate test sample are within the value range of the master data parameter, a positive example identifier is added to the combined data generation rule combination; when at least one master data parameter value of the candidate test sample is not within the value range of the master data parameter, a negative example identifier is added to the combined data generation rule combination; the interaction module is further configured to display the data generation rule combination with the positive example identifier and / or the data generation rule combination with the negative example identifier.

[0027] In some possible implementation manners, the data generation module is further configured to modify a copy of the sampled data using some of the adjusted data generation rules to obtain a reference test sample; the interaction module is further configured to display the reference test sample.

[0028] In some possible implementation manners, after the policy rule module determines the data generation rule corresponding to the policy, the interaction module is further configured to display the data generation rule.

[0029] In some possible implementation manners, the computing device further includes a testing module, and the testing module is configured to test the sampled data and the test sample using the master data policy group; when the test data is consistent with the expected data, it is determined that the master data policy group is correct; when the test data is different from the expected data, it is determined that the master data policy group is incorrect.

[0030] In some possible implementation manners, the policy rule module is configured to modify the policy of the master data policy group and determine the data generation rule corresponding to the modified policy; the data generation module is further configured to modify the master data parameter values in the copy of the sampled data using the data generation rule corresponding to the modified policy to obtain a modified test sample; the interaction module is further configured to display the sampled data and the modified test sample.

[0031] For the glossary of terms, the steps performed by each module, and the technical effects in the second aspect, reference may be made to the corresponding descriptions in the first aspect.

[0032] In a third aspect, a computing device cluster is provided, which includes at least one computing device, and each computing device includes a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method as described in the first aspect or any one of the possible implementation manners of the first aspect.

[0033] In a fourth aspect, a computer-readable storage medium is provided, which includes computer program instructions, and when the computer program instructions are executed by the computing device cluster, the computing device cluster executes the method as described in the first aspect or any one of the possible implementation manners of the first aspect.

[0034] A fifth aspect provides a computer program product comprising instructions, characterized in that when the instructions are run by a cluster of computing devices, the cluster of computing devices is caused to execute the method as described in the first aspect or any possible implementation manner of the first aspect. Description of the Drawings

[0035] Figure 1 A schematic diagram of a cloud service scenario in an embodiment of the present application;

[0036] Figure 2 A flowchart of testing a master data policy and executing a master data policy in an embodiment of the present application;

[0037] Figure 3 A flowchart of generating test samples in an embodiment of the present application;

[0038] Figure 4 A structural diagram of a computing device in an embodiment of the present application;

[0039] Figure 5 A schematic diagram of modifying a data generation rule and generating test samples in an embodiment of the present application;

[0040] Figure 6 Another structural diagram of a computing device in an embodiment of the present application;

[0041] Figure 7 Another structural diagram of a cluster of computing devices in an embodiment of the present application;

[0042] Figure 8 Another structural diagram of a cluster of computing devices in an embodiment of the present application. Detailed Embodiments

[0043] The data processing method of the present application can be applied to a cloud service system, which refers to a system that provides cloud services. The cloud service system may include one or more data centers, or may include some servers of a data center.

[0044] Referring to Figure 1 , in one example, the cloud service system includes a cloud management platform and multiple servers in a data center. The cloud management platform and the multiple servers are connected through an internal network of the data center. The client can be connected to the cloud management platform through the Internet.

[0045] Functions of the Cloud Management Platform: It provides access interfaces (such as interfaces or application programming interfaces). Tenants can operate the client to remotely access the access interface, register cloud accounts and passwords on the cloud management platform, and log in to the cloud management platform. After the cloud management platform authenticates the cloud accounts and passwords successfully, tenants can further pay on the cloud management platform to select and purchase virtual machines of specific specifications (processors, memory, disks). After the successful payment and purchase, the cloud management platform provides the remote login account password of the purchased virtual machine, and the client can remotely log in to the virtual machine and install and run the tenant's applications in the virtual machine.

[0046] Logical Function Division of the Cloud Management Platform: User Console, Computing Management Service, Network Management Service, Storage Management Service, Authentication Service, Image Management Service. The User Console provides interfaces or application programming interfaces to interact with tenants. The Computing Management Service is used to manage servers running virtual machines and containers as well as bare metal servers. The Network Management Service is used to manage network services (such as gateways, firewalls, etc.). The Storage Management Service is used to manage storage services (such as data bucket services). The Authentication Service is used to manage tenant account passwords. The Image Management Service is used to manage virtual machine images.

[0047] Functions of the Cloud Management Platform Client: Receive control plane commands sent by the cloud management platform and create and perform full-life cycle management on virtual machines on the server according to the control commands of the control plane. The cloud management platform client can be installed on user devices.

[0048] Therefore, tenants can create, manage, log in to, and operate virtual machines in the cloud data center through the cloud management platform. Among them, virtual machines can also be called elastic compute service (ECS) or elastic instances.

[0049] A virtual machine refers to a complete computer system with complete hardware system functions simulated by software and running in a completely isolated environment. What can be done on a server can all be achieved in a virtual machine. When creating a virtual machine on a server, part of the hard disk and memory capacity of the physical machine needs to be used as the hard disk and memory capacity of the virtual machine. Each virtual machine has an independent hard disk and operating system, and users of the virtual machine can operate the virtual machine just like using a server.

[0050] The server includes a hardware layer and a software layer. The hardware layer is the conventional configuration of the server. Among them, peripheral component interconnect (PCI) devices include network cards, graphics processing units (GPUs), offloading cards, and other devices that can be inserted into the PCI / PCIe slots of the server. The software layer includes the operating system installed and running on the server (which can be called the host operating system relative to the operating system of the virtual machine). The virtual machine manager (VMM) is set in the host operating system. The virtual machine manager is also called Hypervisor. The role of the virtual machine manager is to implement computing virtualization, network virtualization, and storage virtualization of the virtual machine and is responsible for managing the virtual machine.

[0051] Computing virtualization means providing part of the server's processor and memory to the virtual machine. Network virtualization means providing part of the functions of the network card (such as bandwidth) to the virtual machine. Storage virtualization means providing part of the disk to the virtual machine. The virtual machine manager can also implement logical isolation between different virtual machines and manage the virtual machine. For example, creating a virtual machine, simulating virtual hardware for the virtual machine according to the hardware layer (hardware simulation function), deleting the virtual machine, forwarding and / or processing network packets between all virtual machines running on this server (such as virtual machine 1 and virtual machine 2), or forwarding network packets between the virtual machine on this server and the external network (virtual switching function), and processing the I / O generated by the virtual machine.

[0052] The running environments (such as virtual machine applications, operating systems, and virtual hardware) in different virtual machines are completely isolated. To communicate between virtual machine 1 and virtual machine 2, it is necessary to forward network packets through the virtual manager. The tenant can remotely log in to the virtual machine and operate the installation, setting, and uninstallation of applications in the virtual machine operating system environment.

[0053] Due to the large variety of policies, when the user configures the master data policy, there may be a situation of inputting incorrect policies. In the existing testing methods, the sampled data may not meet the requirements of the policy, making it difficult to determine whether the policy rules are correct. For example, for the deduplication policy in the policy, when there are no duplicate names in the sampled data or an incorrect deduplication policy is input, the test result is 0 in both cases. Therefore, it is difficult to determine whether the sampled data does not meet the policy requirements or the deduplication policy is incorrect. To determine whether the policy takes effect, it is usually necessary to obtain a large amount of test data for testing. A large number of test samples are not relevant to the policy, resulting in too low test efficiency. To solve this problem, this application provides a method for quickly testing the master data policy. Refer to Figure 2 , in one example, the method includes the following steps:

[0054] S201. Configure the master data policy group. The computing device may display a configuration interface. After the user inputs the master data policy group in the configuration interface, the computing device configures the master data policy group.

[0055] S202. Generate test samples. After the computing device configures the master data policy group, it generates test samples related to the master data policy group.

[0056] S203. Test the master data policy group. Use the master data policy group to test the test samples to obtain test results.

[0057] S204. Determine whether the test results are correct. If so, execute S205; if not, execute S201. When the test results are consistent with the expected results, it is determined that the test results are correct; when the test results are inconsistent with the expected results, it is determined that the test results are incorrect.

[0058] S205. Process the master data using the master data policy group.

[0059] S206. Sample from the processing results. S206 is an optional step. The user can sample to view the processing results of the master data policy or view all the processing results.

[0060] The process of generating test samples is introduced in detail below. Refer to Figure 3 , an embodiment of the data processing method of this application includes the following steps:

[0061] S301. The computing device obtains the master data parameters and the master data policy group.

[0062] In this embodiment, the computing device may provide a page for inputting master data parameters and a page for inputting the master data policy group, or may provide a region for inputting master data parameters and a region for inputting the master data policy group on one page. After the user inputs the master data parameters and the master data policy group, the computing device obtains the master data parameters and the master data policy group. The master data policy group includes multiple policies, and the parameters of the policies are the master data parameters. Specifically, the policy includes a function, and the function parameters are the master data parameters. It should be understood that the parameters of multiple functions can be the same master data parameter.

[0063] S302. The computing device determines the data generation rules corresponding to the policies.

[0064] In an optional embodiment, S302 includes: determining the data generation rules corresponding to the policies in the mapping table. The mapping table includes the correspondence between the policies and the data generation rules, and the data generation rules can be quickly found through the mapping table.

[0065] In another alternative embodiment, S302 includes: the computing device inputs the function name of the policy into the neural network model, and outputs a data generation rule through the neural network model. A sample set including function names is obtained in advance, the function names are used as training samples, and the data generation rules corresponding to the function names are used as sample labels. After training the neural network model according to the sample set, the neural network model is deployed on the computing device. For policies not in the sample set, after inputting them into the neural network model, the neural network model can also output effective data generation rules, so it has good scalability.

[0066] A policy can correspond to multiple data generation rules, and the data generation rules are also called operators. The data generation rules include data operations, the number of operation objects, and the operation objects. The data operations include addition, deletion, and modification. The operation objects include one or more of characters, words, or numerical values. The number of operation objects can be set according to the actual situation. The initial value of the number of operation objects can be 1, 2, 3, 4, or other values, or it can be a random number within a value range, which is not limited in this application.

[0067] S303. The computing device obtains sampling data from the master data according to the master data parameters.

[0068] Specifically, when the master data is a table, sampling data can be obtained from the table according to the master data parameters. When the master data includes multiple tables, the computing device obtains sampling data from multiple tables according to the table names and the master data parameters, and the master data parameters are the column names. There is no fixed sequence between S302 and S303.

[0069] S304. The computing device uses the data generation rule to modify the value of the master data parameter in the copy of the sampling data to obtain a test sample. For each value of the master data parameter in the copy of the sampling data, the computing device can obtain the corresponding policy and data generation rule, and then use the data generation rule to modify the value of the master data parameter.

[0070] Optionally, S304 includes: the computing device uses the data generation rule to modify the value of the master data parameter in the copy of the sampling data to obtain a candidate test sample; when each value of the master data parameter of the candidate test sample is within the value range of the master data parameter, the computing device uses the candidate test sample as a test sample. This can ensure that the values of the test samples do not exceed the value range, thus meeting the actual business requirements. The value range of the master data parameter can be obtained from the knowledge base or set according to actual experience. When at least one value of the master data parameter of the candidate test sample is not within the value range of the master data parameter, it indicates that the data of the candidate test sample does not meet the requirements of the actual business rules, and the candidate test sample can be displayed to prompt that there is a problem with the data generation rule or the master data policy.

[0071] S305. The computing device displays the sampled data and the test samples.

[0072] The user can check the differences between the sampled data and the test samples, and based on these differences, can determine whether the master data policy is correct before testing the master policy combination.

[0073] In this embodiment, after pre-configuring the policies and data generation rules with corresponding relationships, the data generation rules related to the master data policy group can be obtained. The test samples generated using these data generation rules have a strong correlation with the master data policy group. Therefore, it is possible to determine whether the master data policy group is correct using a small number of test samples. Compared with the existing test methods that use a large number of samples, it can reduce the computational overhead and improve the test efficiency.

[0074] Secondly, in the existing test methods, the sampled samples may be irrelevant to the policy, resulting in inaccurate test results. The test samples in this embodiment have a strong correlation with the master data policy group, so it can improve the accuracy of testing the master data policy.

[0075] Thirdly, in this embodiment, the test samples can be automatically generated without user input, so it can improve the speed of obtaining test samples.

[0076] This application can also preview and modify the data generation rules to improve the diversity of test samples. The computing device can adjust the operation object according to the user input data, or can also adjust the operation object according to the characteristics of the automatically generated data components. In an alternative embodiment, the computing device adjusts the operation object of the data generation rule according to the user input data, and uses the adjusted data generation rule to modify the master data parameter values in the copy of the sampled data to obtain candidate test samples.

[0077] In this embodiment, the computing device can display the data generation rules for the user to preview the data generation rules and check whether the policy is correct before generating the test samples.

[0078] Specifically, the computing device can adjust the operation object of the data generation rule, and the operation object includes at least one of characters, words, data component strings, or numerical values. For example, adjusting the characters of the data generation rule to words, adjusting the words of the data generation rule to data component strings, and adjusting the data component strings of the data generation rule to characters.

[0079] Taking data operation modification as an example, the computing device modifies the main data parameter value in the copy of the sampled data using the adjusted data generation rule, including: when the operation object of the data generation rule includes characters, the computing device randomly generates characters and modifies the characters of the main data parameter value to the randomly generated characters; when the operation object of the data generation rule includes words, the computing device randomly generates words and modifies the words of the main data parameter value to the randomly generated words; when the operation object of the data generation rule includes data component strings, the computing device obtains related words (such as synonyms, near-synonyms, misspelled words or abbreviations) of the data component string and modifies the data component string of the main data parameter value to a synonym, near-synonym, misspelled word or abbreviation; when the operation object of the data generation rule includes numerical values, the computing device randomly generates numerical values and modifies the numerical values of the main data parameter value to the randomly generated numerical values.

[0080] It should be noted that the computing device can also adjust the number of operation objects. For example, it can adjust 1 character in the data generation rule to 2 characters, 1 word in the data generation rule to 2 words, or 1 character in the data generation rule to 1 - 3 characters. The computing device can also adjust the type of data operation of the operation object. For example, it can adjust "addition" in the data generation rule to "deletion". It should be understood that the above adjustments are only illustrative examples and can be adjusted according to the actual situation. This application does not make any limitations.

[0081] The method for adjusting the operation object according to the data component characteristics is introduced below. In an optional embodiment, the computing device inputs the main data parameter into a language model, and the language model outputs the data component characteristics corresponding to the main data parameter; adjusts the data component string of the data generation rule to the data component characteristics, and uses the adjusted data generation rule to modify the main data parameter value in the copy of the sampled data to obtain a candidate test sample.

[0082] In this embodiment, when the operation object is a data component string, the main data parameter can be input into the language model, and the language model outputs the data component characteristics; modify the data component strings corresponding to one or more data component characteristics in the main data parameter value. For example, input the name field into the language model, and the language model outputs {English first name} and {English last name}, and {English first name} and {English last name} are the data component characteristics corresponding to name. Modify the data component string corresponding to {English first name}, and / or modify the data component string corresponding to {English last name}.

[0083] Optionally, the data processing method of the present application further includes: the computing device combines the data generation rules corresponding to all the main data parameter values according to the policy relationship; when all the main data parameter values of the candidate test sample are within the value range of the main data parameter, the computing device adds a positive example identifier to the combined data generation rule combination; when at least one of the main data parameter values of the candidate test sample is not within the value range of the main data parameter, the computing device adds a negative example identifier to the combined data generation rule combination; the computing device displays the data generation rule combination with a positive example identifier and / or the data generation rule combination with a negative example identifier.

[0084] In this embodiment, the main data policy group may include multiple policies, such as a similarity policy, a merging policy, etc. The main data policy may also include a policy relationship, such as logical AND (&&), logical OR (||), logical NOT (!). By displaying positive and negative examples, it can be determined which data generation rules are compliant and which operation rules are non-compliant, so as to determine in advance whether the main data policy group meets the requirements. To facilitate the user to view the data generation rules, after the computing device generates or adjusts the data generation rules, it can display the data generation rules to facilitate the user to check whether the data generation rules are correct.

[0085] Optionally, the data processing method of the present application further includes: the computing device uses some of the adjusted data generation rules to modify the copy of the sampled data to obtain a reference test sample; displays the reference test sample. By checking the differences between the sampled data, the test sample, and the reference test sample, it can be seen whether some of the data generation rules are correct.

[0086] Optionally, the data processing method of the present application further includes: the computing device tests the sampled data and the test sample using the main data policy group; when the test data is consistent with the expected data, the computing device determines that the main data policy group is correct; when the test data is different from the expected data, the computing device determines that the main data policy group is incorrect.

[0087] In this embodiment, the expected data can be configured by the user. The sampled data, the test sample, and the main data policy are highly correlated. Therefore, using a small amount of sampled data and test samples can test whether the main data policy group is correct. Compared with the existing test method using a large amount of sampled data, it can improve the test efficiency and shorten the test duration.

[0088] Optionally, when the user adjusts the policy, the computing device can modify the policies of the main data policy group according to the policy input by the user; determine the data generation rules corresponding to the modified policy; use the data generation rules corresponding to the modified policy to modify the main data parameter values in the copy of the sampled data to obtain a modified test sample; display the sampled data and the modified test sample. In this way, the test sample can be dynamically displayed, and thus it can quickly check whether the modified main data policy is correct.

[0089] For ease of understanding, the data processing method of the present application is introduced below through an example:

[0090] In this example, the main data parameters input by the user include name, address, and update_time. The sampling data obtained by the computing device from the main data according to the main data parameters is shown in Table 1:

[0091] id name address update_time 001 David Smith 123 Jiangsu Rd Street 2023-01-01 002 Amy Brown 456 Main St. 2022-12-01

[0092] Table 1

[0093] The main data policy groups input by the user include: Policy 1 is Editsimilarity(name)>0.8, Policy 2 is Jaccardsimilarity(address)>0.6, and Policy 3 is Newest(update_time). Editsimilarity(name) calculates the similarity of the distance class, which is used to obtain the name similarity. Jaccardsimilarity(address) calculates the similarity of the set type, which is used to obtain the set similarity of the address. Newest(update_time) belongs to the merging policy of the address, which is used to select the latest address as the value of the main data record.

[0094] In the mapping table, the function names of each policy correspond to the data generation rules as shown in Table 2 or Table 3:

[0095] Rule Number Data Generation Rule Function Name of the Strategy 1 Randomly add several characters, randomly delete several characters or randomly modify several characters Edit similarity 2 Randomly add several words, randomly delete several words or randomly delete several words Jaccard similarity 3 Add the current date and historical dates Newest

[0096] Table 2

[0097]

[0098]

[0099] Table 3

[0100] When the product of the attribute value length * 0.2 is not an integer or the product of the number of random words * 0.4 is not an integer, the rounding operation can be performed on the attribute value length. For the data generation rules corresponding to each policy, the user can select one or more data generation rules.

[0101] In order to increase the diversity of sample data, the data generation rules can be adjusted. The corresponding relationship between the adjusted data generation rules and the function names of the policies is shown in Table 4:

[0102]

[0103] Table 4

[0104] Combine the data generation rules corresponding to the three fields or use the data generation rule corresponding to a single field as a combination to obtain the data generation rule combinations shown in Table 5:

[0105]

[0106] Table 5

[0107] Use the above data generation rule combinations 1 to 3 to modify the sample copy of Table 1, and the modification results are shown in Table 6:

[0108] id name address update_time 001 David Smith 123 Jiangsu Rd Street 2023-01-01 001-1 Dave Smith 123 Jiangsu Rd 2023-10-17 001-2 David Brown 123 Jiangsu Street 2023-10-17 001-3 David Smth 123 Jiangsu Rd Street 2023-01-01 002 Amy Brown 456 Main St. 2022-12-01 002-1 Army Brown 456 Main Street 2023-10-17 002-2 Amy Smith 123 Main St. 2023-10-17 002-3 Amy Bron 456 Main St. 2023-01-01

[0109] Table 6

[0110] If the modification results shown in Table 6 are all within the value range, then determine the data generation rule combination as a positive example rule combination. It should be noted that when the modification result is not within the value range, determine the data generation rule as a negative example rule or the data generation rule combination as a negative example rule combination. For example, the length of the random attribute value is equal to 100, and the data generation rule 2 is to randomly add 20 characters, randomly delete 20 characters, or randomly modify 20 characters. When the value range of the name length is [1 character, 15 characters], then after the data generation rule 2 modifies the name parameter value, the modification result exceeds the value range. Therefore, the data generation rule 2 is a negative example rule.

[0111] The following introduces the computing device of the present application. Refer to Figure 4 In one embodiment, the computing device 400 provided by the present application includes an interaction module 401, a policy rule module 402, a sampling module 403, a data generation module 404, a feature module 405, and a test module 406, where the feature module 405 and the test module 406 are optional modules. The interaction module 401 is used to obtain the main data parameters and the main data policy group. The policy rule module 402 is used to determine the data generation rule corresponding to the policy. The sampling module 403 is used to obtain sampling data from the main data according to the main data parameters. The data generation module 404 is used to modify the main data parameter values in the copy of the sampling data by using the data generation rule to obtain test samples. The interaction module 401 is used to display the sampling data and the test samples.

[0112] As an example of a software functional unit, the data generation module 404 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the data generation module 404 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.

[0113] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set up within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC, and the interconnection between VPCs is achieved through the communication gateway.

[0114] As an example of a hardware functional unit, the data generation module 404 may include at least one computing device, such as a server, etc. Alternatively, the data generation module 404 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0115] The multiple computing devices included in the data generation module 404 can be distributed in the same region or in different regions. The multiple computing devices included in the data generation module 404 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the data generation module 404 can be distributed in the same VPC or in multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0116] It should be noted that in other embodiments, the data generation module 404 can be used to execute Figure 2 or Figure 3 any step in the data processing method shown, and the interaction module 401 can be used to execute Figure 2 or Figure 3 any step in the data processing method shown, the policy rule module 402 can be used to execute Figure 2 or Figure 3 any step in the data processing method shown, the sampling module 403 can be used to execute Figure 2 or Figure 3 any step in the data processing method shown. The steps to be implemented by the interaction module 401, the policy rule module 402, the sampling module 403, and the data generation module 404 can be specified as needed. The entire function of the computing device 400 is implemented by separately implementing different steps in the data processing method through the interaction module 401, the policy rule module 402, the sampling module 403, and the data generation module 404. Figure 4 For the glossary of terms and technical effects in the embodiments shown, reference can be made to Figure 3 the corresponding descriptions in the embodiments shown.

[0117] In one embodiment, the feature module 405 is used to input the main data parameters from the interaction module 401 into the language model, and output the data component features corresponding to the main data parameters through the language model; the policy rule module 402 is specifically used to adjust the data component string of the data generation rule to the data component feature when the operation object of the data generation rule includes the data component string.

[0118] In another embodiment, the test module 406 is used to test the sampled data and the test samples using the main data policy group; when the test data is consistent with the expected data, it is determined that the main data policy group is correct; when the test data is different from the expected data, it is determined that the main data policy group is incorrect.

[0119] The process of adjusting the data generation rule and generating the test sample of the present application will be introduced below in conjunction with the computing device 400. Refer to Figure 5, in an alternative embodiment, the data table feed includes one or more master data parameters. The interaction module 401 receives the master data parameters and the master data policy group, and the policy rule module 402 determines the data generation rules corresponding to each policy according to the master data policy group. After obtaining the master data parameters, the feature module 405 inputs the master data parameters into the language model, and outputs the data component features corresponding to the master data parameters through the language model. When the data generation rule includes a data component string, the policy rule module 402 adjusts the data component string of the data generation rule to the data component feature. After obtaining the sampling data according to the master data parameters, the sampling module 403 generates a sampling data copy, and the data generation module 404 modifies the master data parameter value in the sampling data copy using the adjusted data generation rule to obtain a candidate test sample. After the data generation module 404 obtains the value range of the master data parameter value from the knowledge base, it determines whether the master data parameter value of the candidate test sample is within the value range. When all the master data parameter values of the candidate test sample are within the value range, the test sample is output.

[0120] This application also provides a computing device 600. As Figure 6 shown, the computing device 600 includes: a bus 602, a processor 604, a memory 606, and a communication interface 608. The processor 604, the memory 606, and the communication interface 608 communicate with each other through the bus 602. The computing device 600 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 600.

[0121] The bus 602 can be a PCI bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 only one line is shown here, but it does not mean that there is only one bus or one type of bus. The bus 604 can include a path for transmitting information between various components of the computing device 600 (for example, the memory 606, the processor 604, and the communication interface 608).

[0122] The processor 604 can include any one or more of a central processing unit (CPU), a GPU, a microprocessor (MP), or a digital signal processor (DSP), etc. The processor 604 can perform calculations or processing on data, such as metadata management, deduplication, data compression, virtualized storage space, and address translation, etc.

[0123] The memory 606 may include volatile memory, such as random access memory (RAM). The processor 604 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD). In some embodiments, executable program code is stored in the memory 606, and the processor 604 executes the executable program code to implement the functions of the foregoing interaction module 401, policy rule module 402, sampling module 403, data generation module 404, feature module 405, and test module 406 respectively, so as to implement the data processing method.

[0124] The communication interface 608 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 600 and other devices or a communication network.

[0125] An embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device 600 may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device 600 may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0126] As Figure 7 shown, the computing device cluster includes at least one computing device 600. The same instructions for executing the data processing method may be stored in the memory 606 of one or more computing devices 600 in the computing device cluster.

[0127] In some possible implementation manners, partial instructions for executing the data processing method may also be stored separately in the memory 606 of one or more computing devices 600 in the computing device cluster. In other words, a combination of one or more computing devices 600 may jointly execute the instructions for executing the data processing method.

[0128] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 8 shows a possible implementation manner. As Figure 8As shown, two computing devices 600A and 600B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation, the memory 606 in computing device 600A stores instructions for executing the functions of the interaction module 401. At the same time, the memory 606 in computing device 900B stores instructions for executing the functions of the policy rule module 402, the sampling module 403, the data generation module 404, the feature module 405, and the test module 406.

[0129] An embodiment of the present application also provides a computer program product containing instructions. The computer program product can be software or a program product that contains instructions and can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes the at least one computing device to execute a data processing method.

[0130] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute a data processing method. The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data processing method, characterized in that, The method is executed by a computing device of a cloud service platform, and the computing device is used to generate test samples of a master data policy group. The method includes: The computing device obtains master data parameters and a master data policy group. The master data policy group includes multiple policies, and the master data parameters are parameters of at least one of the policies. The computing device determines the data generation rule corresponding to the policy. The computing device obtains sampling data from the master data according to the master data parameters. The computing device uses the data generation rule to modify the master data parameter values in the copy of the sampling data to obtain test samples. The computing device displays the sampling data and the test samples.

2. The method according to claim 1, wherein The computing device determining the data generation rule corresponding to the policy includes: The computing device inputs the function name included in the policy into a neural network model, and outputs the data generation rule corresponding to the policy through the neural network model.

3. The method according to claim 1, characterized in that, The computing device determining the data generation rule corresponding to the policy includes: The computing device determines the data generation rule corresponding to the policy in a mapping table.

4. The method according to claim 1, characterized in that, The computing device using the data generation rule to modify the master data parameter values in the copy of the sampling data to obtain test samples includes: The computing device uses the data generation rule to modify the master data parameter values in the copy of the sampling data to obtain candidate test samples. When each master data parameter value of the candidate test sample is within the value range of the master data parameters, the computing device uses the candidate test sample as a test sample.

5. The method according to claim 4, wherein: The method further includes: the computing device adjusts the operation object of the data generation rule, and the operation object includes at least one of characters, words, data component strings, or numerical values. The computing device using the data generation rule to modify the master data parameter values in the copy of the sampling data to obtain candidate test samples includes: the computing device uses the adjusted data generation rule to modify the master data parameter values in the copy of the sampling data to obtain candidate test samples.

6. The method according to claim 5, wherein The operation object of the data generation rule includes a data component string. The method further includes: the computing device inputs the master data parameters into a language model, and outputs the data component features corresponding to the master data parameters through the language model. The computing device adjusting the operation object of the data generation rule includes: the computing device adjusts the data component string of the data generation rule to the data component feature.

7. The method according to claim 5, wherein Before the computing device uses the adjusted data generation rule to modify the master data parameter values in the copy of the sampling data, the method further includes: The computing device adjusts the number of operation objects of the data generation rule.

8. The method according to claim 4, wherein The master data policy group further includes a policy relationship, and the method further includes: The computing device combines the data generation rules corresponding to all the master data parameter values according to the policy relationship. When all the master data parameter values of the candidate test sample are within the value range of the master data parameters, the computing device adds a positive example identifier to the combined data generation rule combination. When at least one main data parameter value of the candidate test sample is not within the value range of the main data parameter, the computing device adds a negative example identifier to the data generation rule combination obtained by combination. The computing device displays the data generation rule combination with the positive example identifier and / or the data generation rule combination with the negative example identifier.

9. The method according to any one of claims 5 to 8, characterized in that The method further includes: The computing device uses the adjusted partial data generation rule to modify the copy of the sampled data to obtain a reference test sample. The computing device displays the reference test sample.

10. The method according to any one of claims 1 to 8, characterized in that The method further includes: The computing device uses the main data policy group to test the sampled data and the test sample. When the test data is consistent with the expected data, the computing device determines that the main data policy group is correct. When the test data is different from the expected data, the computing device determines that the main data policy group is incorrect.

11. The method according to any one of claims 1 to 8, characterized in that, The method further includes: The computing device modifies the policy of the main data policy group. The computing device determines the data generation rule corresponding to the modified policy. The computing device uses the data generation rule corresponding to the modified policy to modify the main data parameter value in the copy of the sampled data to obtain a modified test sample. The computing device displays the sampled data and the modified test sample.

12. A computing device, characterized in that, Includes: An interaction module for obtaining main data parameters and a main data policy group, the main data policy group includes multiple policies, and the main data parameter is a parameter of at least one of the policies. A policy rule module for determining the data generation rule corresponding to the policy. A sampling module for obtaining sampled data from the main data according to the main data parameter. A data generation module for using the data generation rule to modify the main data parameter value in the copy of the sampled data to obtain a test sample. An interaction module for displaying the sampled data and the test sample.

13. The device according to claim 12, characterized in that, The policy rule module is specifically configured to input the function name included in the policy into a neural network model, and output the data generation rule corresponding to the policy through the neural network model.

14. The device according to claim 12, characterized in that, The policy rule module is specifically configured to determine the data generation rule corresponding to the policy in a mapping table.

15. The device according to claim 12, characterized in that, The data generation module is specifically configured to use the data generation rule to modify the main data parameter value in the copy of the sampled data to obtain a candidate test sample; when each main data parameter value of the candidate test sample is within the value range of the main data parameter, the candidate test sample is used as a test sample.

16. The device according to claim 15, characterized in that, The policy rule module is further configured to adjust the operation object of the data generation rule, and the operation object includes at least one of characters, words, data component strings, or numerical values. The data generation module is specifically configured to use the adjusted data generation rule to modify the main data parameter value in the copy of the sampled data to obtain a candidate test sample.

17. The device according to claim 16, wherein, The device further includes: A feature module for inputting the main data parameter into a language model and outputting the data component feature corresponding to the main data parameter through the language model. Specifically, when the operation object of the data generation rule includes a data component string, the policy rule module is configured to adjust the data component string of the data generation rule to the data component feature.

18. The device according to claim 16, characterized in that, The policy rule module is further configured to adjust the number of operation objects of the data generation rule.

19. The apparatus according to claim 15, wherein the policy rule module is further configured to combine the data generation rules corresponding to all the main data parameter values according to a policy relationship; when all the main data parameter values of the candidate test sample are within the value range of the main data parameter, add a positive example identifier to the combined data generation rule combination; when at least one of the main data parameter values of the candidate test sample is not within the value range of the main data parameter, add a negative example identifier to the combined data generation rule combination; the interaction module is further configured to display the data generation rule combination with the positive example identifier and / or the data generation rule combination with the negative example identifier.

20. The apparatus according to any one of claims 16 to 19, wherein the data generation module is further configured to use the adjusted partial data generation rule to modify a copy of the sampled data to obtain a reference test sample; the interaction module is further configured to display the reference test sample.

21. The device according to any one of claims 12 to 19, characterized in that The apparatus further comprises: a test module, configured to test the sampled data and the test sample by using the main data policy group; when the test data is consistent with the expected data, determine that the main data policy group is correct; when the test data is different from the expected data, determine that the main data policy group is incorrect.

22. The apparatus according to any one of claims 12 to 19, wherein the policy rule module is further configured to modify the policy of the main data policy group and determine the data generation rule corresponding to the modified policy; the data generation module is further configured to use the data generation rule corresponding to the modified policy to modify the main data parameter value in the copy of the sampled data to obtain a modified test sample; the interaction module is further configured to display the sampled data and the modified test sample.

23. A cluster of computing devices, characterized in that, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 11.

24. A computer-readable storage medium, characterized in that, comprising computer program instructions, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 11.

25. A computer program product comprising instructions, characterized in that, When the instructions are run by a computing device cluster, the computing device cluster is caused to execute the method according to any one of claims 1 to 11.