Data testing method and device, computer equipment, storage medium and program product

By using pre-trained data to generate models and model prompt information to generate and synthesize test data, the problems of low efficiency and high cost of software testing are solved, and more efficient and accurate test results are achieved.

CN120216382APending Publication Date: 2025-06-27NEW H3C BIG DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510381159.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Software testing is inefficient, costly and error-prone.

Method used

By obtaining multiple pre-trained data, the model and model prompt information are generated, and the model is prompted to use these prompt information to guide the model to generate original test data, and the data is synthesized based on semantics to obtain the target test data. Then, execute the test cases corresponding to the target test data to generate the test results.

Benefits of technology

It improves the effectiveness and accuracy of the test, reduces the risk of manual intervention, and improves the scientificity and efficiency of data testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216382A_ABST
    Figure CN120216382A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of software testing, and discloses a data testing method and device, computer equipment, a storage medium and a program product. The method comprises the steps of obtaining a plurality of pre-trained data generation models and first model prompt information; utilizing the first model prompt information to guide each data generation model to generate corresponding original test data; based on the semantics of each piece of original test data, performing synthesis processing on each piece of original test data to obtain target test data; and executing the test case corresponding to the target test data, and generating a test result corresponding to the target test data. According to the technical scheme, automatic test data efficient synthesis based on multi-model collaborative generation and semantic fusion can be achieved, and the coverage rate and the test efficiency of test cases are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software testing, and in particular to a data testing method, device, computer equipment, storage medium and program product. Background Art

[0002] Software testing is a key step in the software development process, which aims to ensure that the software system meets the predetermined functional and performance requirements and reduce the risks of application failures, data loss and security vulnerabilities. At present, software testing mostly relies on manual operations, but manual testing has the limitations of being time-consuming, high labor costs, low efficiency and susceptible to human factors. Summary of the invention

[0003] In view of this, the present invention provides a data testing method, apparatus, computer equipment, storage medium and program product to solve the problems of low efficiency, high cost and easy error in software testing.

[0004] In a first aspect, the present invention provides a data testing method, comprising: obtaining a plurality of pre-trained data generation models and first model prompt information; using the first model prompt information to guide each data generation model to generate corresponding original test data; based on the semantics of each original test data, performing synthesis processing on each original test data to obtain target test data; executing a test case corresponding to the target test data to generate a test result corresponding to the target test data.

[0005] The data testing method provided by the embodiment of the present invention can generate diversified original test data with the help of the characteristics of different models by acquiring multiple pre-trained data generation models and model prompt information. Prompt information is used to guide generation, making data generation more targeted. The original test data can be synthesized and processed based on semantics to integrate advantages, making the target test data more comprehensive and accurate, thereby making the test results generated by executing the corresponding test cases more reliable, and comprehensively improving the effectiveness and accuracy of the test.

[0006] In an optional embodiment, the first model prompt information is used to guide each data generation model to generate corresponding original test data, including: obtaining the model order of each data generation model; using the first model prompt information to simultaneously guide each data generation model to generate corresponding original test data, and determining the data priority of each original test data based on the model order.

[0007] The data testing method provided by the embodiments of the present invention can organize and utilize the first model prompt information according to certain logic and rules by obtaining the model order of each data generation model, so as to orderly guide different models to generate corresponding original test data. This orderly generation method avoids the chaos and disorder of data generation, effectively improves the generation efficiency, makes the original test data generated by each model have coherence and traceability in sequence, and is conducive to subsequent data management, analysis and synthesis processing.

[0008] In an alternative embodiment, based on the semantics of each original test data, the synthesis process of each original test data is carried out to obtain the target test data, including: separating each original test data based on the field attributes of each original test data to obtain multiple field data corresponding to each original test data; obtaining the field verification rules matching each field data, and verifying each field data according to the field verification rules to obtain the verification results; based on the verification results and the semantics of the field data, carrying out the synthesis process of each original test data to obtain the target test data.

[0009] The data testing method provided by the embodiments of the present invention separates according to the field attributes of the original test data, accurately disassembles complex data into multiple field data, making data processing more targeted. Obtaining and verifying the verification rules matching each field data effectively guarantees the accuracy and reliability of the field data, and can timely discover and correct data errors. The synthesis process is carried out based on the verification results and the semantics of the field data, which not only ensures that the synthesized target test data has excellent quality, but also makes full use of the semantic information of the original data, so that the finally generated target test data can more comprehensively and accurately reflect the test requirements, greatly improving the scientificity and effectiveness of data testing.

[0010] In an alternative embodiment, based on the verification results and the semantics of the field data, the synthesis process of each original test data is carried out to obtain the target test data, including: if the verification results indicate that any one of the field data of the current original test data does not conform to the corresponding field verification rules, the field data that does not conform to the field verification rules is determined as abnormal data; based on the verification results of the remaining original test data other than the current original test data, determining the target field data matching the abnormal data; using the target field data to update the abnormal data to obtain the target test data corresponding to the current original test data.

[0011] The data testing method provided by the embodiments of the present invention can quickly locate the problem when the verification result shows that there is abnormal data that does not conform to the field verification rules, improving the accuracy of data error correction. By determining the matching target field data based on the verification results of the remaining original test data and using it to update the abnormal data, the effective correction of the original test data is achieved. This not only ensures the semantic coherence and logic of the target test data, but also improves the quality and reliability of the data, enabling the test data to more realistically simulate the actual situation.

[0012] In an alternative embodiment, before executing the test case corresponding to the target test data and generating the test result corresponding to the target test data, the method further includes: generating noise data for the target test data based on a preset noise parameter; injecting the noise data into the target test data to obtain the target test data with noise perturbation; and using the second model hint information to guide any data generation model to perform desensitization processing on the target test data with noise perturbation to obtain the desensitized target test data.

[0013] The data testing method provided by the embodiments of the present invention generates noise data based on a preset noise parameter and injects it into the target test data. By adding interference elements, the precise features of the original data are blurred, simulating the noise interference in the real environment while reducing the data sensitivity. Using the second model hint information to guide the data generation model to perform desensitization processing on the target test data that has been injected with noise and preliminarily desensitized achieves a double desensitization effect. This double desensitization mechanism greatly enhances the protection of sensitive information and avoids potential risks caused by data sensitivity.

[0014] In an alternative embodiment, if the test result indicates that the target test data passes the test of the test case, the target test data is determined as valid test data; and the parameter of each data generation model is adjusted using the valid test data to obtain the adjusted data generation model.

[0015] The data testing method provided by the embodiments of the present invention determines the target test data as valid test data when the test result shows that the target test data passes the test case, providing a clear determination basis for the validity of the data. Subsequently, using these valid test data to adjust the parameters of each data generation model enables the model to optimize its own parameters based on the data feedback that performs well in the actual test. This not only helps to improve the quality and accuracy of the data generated by the data generation model, making the data generated by it more likely to pass the test case in future tests, but also enables the model to better adapt to the requirements of the actual application scenario.

[0016] In an alternative embodiment, parameter adjustment is performed on each data generation model using valid test data to obtain an adjusted data generation model, including: distributing the valid test data to each data generation model in a load balancing manner to obtain model test data corresponding to each data generation model; and performing parameter adjustment on the corresponding data generation model using each model test data to obtain an adjusted data generation model.

[0017] The data testing method provided by the embodiments of the present invention distributes valid test data to each data generation model in a load balancing manner, which can ensure uniform data distribution and avoid the influence of excessive or insufficient data volume on a single model on the parameter adjustment effect, greatly improving the resource utilization efficiency. After each model obtains suitable model test data, parameter adjustment is performed based on this, making the adjustment process more targeted and scientific. This method not only optimizes the performance of the data generation model but also enhances the stability of the entire data testing and model optimization system, and can efficiently and accurately obtain an adjusted data generation model, laying a solid foundation for generating high-quality test data subsequently.

[0018] In an alternative embodiment, before guiding each data generation model to generate corresponding original test data using the first model prompt information, it further includes: obtaining a preset model configuration fixture; and invoking each data generation model based on the model configuration fixture to enable each data generation model to generate corresponding original test data.

[0019] The data testing method provided by the embodiments of the present invention provides a standardized tool and process for automatically invoking each data generation model by obtaining a preset model configuration fixture. This not only retains the stability and coherence of the original automated testing system but also improves the reliability and efficiency of the automated testing.

[0020] In a second aspect, the present invention provides a data testing device, including: a first acquisition module for acquiring a plurality of pre-trained data generation models and first model prompt information; a generation module for guiding each data generation model to generate corresponding original test data using the first model prompt information; a synthesis module for performing synthesis processing on each original test data based on the semantics of each original test data to obtain target test data; and a testing module for executing test cases corresponding to the target test data and generating test results corresponding to the target test data.

[0021] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the data testing method according to the first aspect or any corresponding embodiment thereof.

[0022] Fourthly, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to make a computer execute the data testing method according to the first aspect or any corresponding embodiment thereof.

[0023] Fifthly, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to make a computer execute the data testing method according to the first aspect or any corresponding embodiment thereof. Description of the Drawings

[0024] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is a schematic flowchart of the data testing method according to an embodiment of the present invention;

[0026] Figure 2 is a schematic flowchart of another data testing method according to an embodiment of the present invention;

[0027] Figure 3 is a schematic diagram of the synthesis processing of each original test data according to an embodiment of the present invention;

[0028] Figure 4 is a schematic flowchart of yet another data testing method according to an embodiment of the present invention;

[0029] Figure 5 is a schematic diagram of the secondary desensitization processing of the target test data according to an embodiment of the present invention;

[0030] Figure 6 is a structural block diagram of the data testing device according to an embodiment of the present invention;

[0031] Figure 7 is a schematic hardware structure diagram of the computer device according to an embodiment of the present invention. Detailed Embodiments

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0033] Software testing is a key step in software development, which aims to ensure that the software system meets the predetermined functional and performance requirements and reduce the risk of application failures, data loss or security vulnerabilities. Software testing is mainly divided into manual testing and automated testing. Manual testing has problems such as time-consuming, high labor costs, low efficiency and susceptibility to human factors, while automated testing can improve efficiency, accelerate the testing process, improve accuracy, and reduce costs, and is especially suitable for continuous integration and continuous delivery.

[0034] In automated testing, test data management is crucial. Correct and rich test data can significantly improve test efficiency and coverage, ensure the accuracy and uniqueness of test results, and help detect potential problems early, thereby improving software quality. Currently, automated testing is usually managed by directly reading the business database and storing the data in a CSV file. However, the problem with this approach is that the business data is relatively fixed and cannot meet the needs of automated testing for data diversity.

[0035] In view of this, the technical solution of the present invention obtains multiple trained data generation models and their prompt information, and uses these prompt information to guide the model to generate original test data. Subsequently, the original test data is synthesized according to the semantics to obtain the final target test data. Finally, the test results are generated by executing the test cases corresponding to the target test data, thereby effectively solving the problem of single test data and insufficient richness in automated testing, and significantly improving the test efficiency.

[0036] According to an embodiment of the present invention, a data testing method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0037] In this embodiment, a data testing method is provided, which can be used for computer equipment, such as notebook computers, desktop computers, etc. Figure 1 is a flow chart of a data testing method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0038] Step S101, obtaining multiple pre-trained data generation models and first model prompt information.

[0039] Multiple data generation models refer to large pre-trained generative AI models (such as GPT-4, models based on the Transformer architecture, etc.) that have the ability to perform natural language processing and data generation. Specifically, multiple data generation models can be obtained through open-source libraries or commercial application programming interfaces (APIs). For example, open-source libraries usually provide the code and weight files of a large number of pre-trained models (such as the transformers library), and users can install the library and directly call the models through the interface; users can also call cloud model services through APIs without the need for local deployment.

[0040] The first model prompt information is an instruction or input (such as a Prompt) that guides the data generation model to generate specific types of test data. For example, the first model prompt information can include data formats, field rules, or test scenario descriptions, such as "Generate test data containing name, mobile phone number, and address". Specifically, users can refer to the official design guidelines of the data generation model or existing test data generation templates in the open-source community to obtain the first model prompt information. Alternatively, users can also write natural language instructions (i.e., the first model prompt information) according to test requirements, such as "Generate 10 pieces of user information data, with fields including name, mobile phone number (in line with the Chinese format), and address".

[0041] Step S102: Use the first model prompt information to guide each data generation model to generate corresponding original test data.

[0042] The original test data is the preliminary test data generated by each data generation model according to the prompt information. Specifically, the first model prompt information is input into each data generation model, and the original test data is generated by calling the APIs of these models.

[0043] Step S103: Based on the semantics of each original test data, perform synthesis processing on each original test data to obtain the target test data.

[0044] The target test data is the final test data after semantic synthesis and optimization of the original test data. Specifically, semantic elements are extracted from each original test data, and abnormal data is identified through analysis. The abnormal data is repaired using the complementarity between the original test data. Finally, the processed original test data is fused and optimized to generate a high-quality data set with coherent semantics and conforming to business rules, that is, the target test data.

[0045] Step S104: Execute the test cases corresponding to the target test data to generate the test results corresponding to the target test data.

[0046] A test case is an automated test script or scenario designed based on target test data, used to verify the functionality or performance of a software system. For example, input the target test data and check whether the system can correctly process the user registration request. The test result is the output after executing the test case, used to determine whether the system under test meets the expectations. Specifically, an automated test script or scenario is designed based on the target test data. For example, for the user registration function, the user information in the target test data is used as input data to fill in the corresponding registration form fields of the test script. Then, the corresponding automated test tools are used to execute these test scripts. During the execution process, simulating real user operations, requests are sent to the software system under test. After the system receives the requests and processes the target test data, the automated test tool captures the response information of the system. According to the pre-set expected result criteria, the actual output of the system is compared with the expected output. If the system correctly processes the target test data and the output meets the expectations, the test result is determined to pass; if the system returns an error message or the output does not match the expectations, the test result is determined to fail.

[0047] The data testing method provided by the embodiments of the present invention can generate diverse original test data by obtaining multiple pre-trained data generation models and model prompt information, and utilize the characteristics of different models. Guided by the prompt information for generation, the data generation is more targeted. Based on semantic synthesis processing of the original test data, advantages can be integrated to make the target test data more comprehensive and accurate, thereby making the test results generated by executing its corresponding test cases more reliable and comprehensively improving the effectiveness and accuracy of testing.

[0048] In this embodiment, a data testing method is provided, which can be used in computer devices such as laptops and desktop computers. Figure 2 is a flowchart of the data testing method according to the embodiments of the present invention, as Figure 2 shown, and this process includes the following steps:

[0049] Step S201, obtain multiple pre-trained data generation models and the first model prompt information. For details, please refer to Figure 1 Step S101 of the embodiment shown, which will not be elaborated here.

[0050] Step S202, use the first model prompt information to guide each data generation model to generate corresponding original test data.

[0051] Specifically, the above step S202 includes:

[0052] Step S2021, obtain the model order of each data generation model.

[0053] The model order refers to the logical priority of multiple data generation models in data processing. For example, GPT-4 is used as the main model to generate data, and its outliers are detected in real time. Meanwhile, a model with a Transformer architecture is called as an auxiliary model to provide supplementary data to repair the outliers. Specifically, the model order of each data generation model can be defined through a configuration file or a rule engine (such as the main model: GPT-4, the auxiliary model: Transformer), or it can also be automatically generated through a dynamic algorithm (such as based on the historical accuracy of the model), which is not specifically limited here.

[0054] In step S2022, the first model prompt information is used to guide each data generation model to generate corresponding original test data simultaneously, and the data priority of each original test data is determined based on the model order.

[0055] All data generation models receive the same first model prompt information, such as "generate test data including name, ID number, and mobile phone number", so as to ensure that the generation targets of each model are consistent. On this basis, all data generation models carry out generation operations in parallel according to the first model prompt information to produce original test data. The generated original test data are sorted according to the above predefined model order. For example, the original test data generated by the main model is used as the basic output (the first priority), and the original test data generated by the auxiliary model is used as an alternative pool (the secondary priority).

[0056] The data testing method provided by the embodiments of the present invention can organize and utilize the first model prompt information according to certain logic and rules by obtaining the model order of each data generation model, so as to orderly guide different models to generate corresponding original test data. This orderly generation method avoids the chaos and disorder of data generation, effectively improves the generation efficiency, makes the original test data generated by each model have coherence and traceability in the sequence, and is conducive to subsequent data management, analysis, and synthesis processing.

[0057] In step S203, based on the semantics of each original test data, each original test data is synthetically processed to obtain target test data.

[0058] Specifically, the above step S203 includes:

[0059] In step S2031, based on the field attributes of each original test data, each original test data is separated to obtain multiple field data corresponding to each original test data.

[0060] Field attributes refer to the metadata information of each field in the original test data (such as "name", "ID number", "mobile phone number", etc.). For example, it can include data type (string, number, etc.), format requirements (such as regular expressions), whether it is required, value range, etc. Multiple field data refers to the independent data units obtained by disassembling the original test data according to field attributes. For example, an original test data "Zhang San, 110101199003077654, 13523064737" will be separated into three field data: name = Zhang San, ID number = 110101199003077654, mobile phone number = 13523064737. Specifically, according to the data type, format requirements or delimiters defined in the field attributes, the original test data is split into independent field data. For example, structured data (such as CSV) extracts fields by comma separation; semi-structured data (such as JSON) parses key-value pairs; unstructured data matches specific patterns through regular expressions (such as extracting mobile phone numbers and ID numbers). Field attributes define the boundary rules for each field (such as "name" is a string, "ID number" is 18 digits), ensuring that the splitting process conforms to business logic.

[0061] Step S2032: Obtain the field verification rules matching each field data, and perform data verification on each field data according to the field verification rules to obtain the verification results.

[0062] Field verification rules refer to the verification logic defined for each field, which is used to ensure the compliance of field data. For example: The mobile phone number must be 11 digits and comply with the operator number segment rules, and the ID number needs to be verified through the check digit calculation, etc. The verification result refers to the determination result after applying the verification rules to each field data. For example, it can be a boolean value (passed / failed) or a structured result containing error information. Specifically, the field verification rules are predefined in the configuration file, database or rule library. For example, it is required that the mobile phone number is 11 digits and the ID number needs to pass the check digit verification. During verification, the field data is compared with the rules one by one. For example, regular expressions are used to verify the format, algorithms are called to verify the logic (such as ID number check digit calculation), and the required nature or value range is checked, etc. The verification result is recorded as the passed / failed status and error details, providing a basis for subsequent synthesis.

[0063] Step S2033: Based on the verification results and the semantics of the field data, perform synthesis processing on each original test data to obtain the target test data.

[0064] Determine the primary model among multiple data generation models. If a field check fails for the original test data of the primary model, then supplement the field value that passes the check from the original test data of other models (such as the name of Model A + the ID number of Model B). Finally, combine the valid field data according to the business semantics to ensure that the synthesized target test data is logically consistent and complies with the business rules (such as integrating the name, compliant ID number, and valid mobile phone number).

[0065] In some alternative embodiments, step S2033 includes:

[0066] Step a1, if the check result indicates that any field data of the current original test data does not conform to the corresponding field check rule, then determine the field data that does not conform to the field check rule as abnormal data.

[0067] When a field data of the original test data (such as the ID number or mobile phone number) fails to pass the predefined field check rule (such as format, length, check digit, etc.), mark this field as abnormal data. For example, if the ID number field rule requires 18 digits, but the actual data is "1101011990030776552" (19 digits), this field is determined to be abnormal after the check fails.

[0068] Step a2, based on the check results of the remaining original test data other than the current original test data, determine the target field data that matches the abnormal data.

[0069] The remaining original test data refers to the original test data generated by other data generation models in the same generation process except for the current original test data being processed. For example, if a certain piece of data generated by the primary model has an abnormal field, then the remaining original test data is the batch of test data generated by the auxiliary model. Specifically, according to the predefined model order (such as primary model → auxiliary model 1 → auxiliary model 2), sequentially check the check results of the same field in other ranked models, and screen out the field data that passes the check and has a semantic match as the target field data. For example, when the data of a certain field in the original test data generated by the primary model does not conform to the corresponding check rule, sequentially check the check results of this field in the original test data of other ranked models (such as auxiliary model 1, auxiliary model 2) according to the model order. If the field check of the second ranked model (such as auxiliary model 1) still fails, then continue to search for the third ranked model (such as auxiliary model 2), and so on, until the field data of a certain model passes the check. For example, if the ID number field check of the primary model fails, and the ID number field of the auxiliary model 1 passes the check and has the correct format, then select the ID number field of the auxiliary model 1 as the replacement data.

[0070] Step a3: Update the abnormal data with the target field data to obtain the target test data corresponding to the current original test data.

[0071] Replace the abnormal field with the target field data and ensure that the replaced field is semantically consistent with the overall data. For example, the abnormal ID number "1101011990030776552" in the current original test data is replaced with the compliant value "110101199003077654" generated by Model 2, while retaining the original values of other fields (such as name and mobile phone number), finally forming a complete, compliant and logically consistent target test data.

[0072] For example, as Figure 3 shown, use the first model hint information to guide Model 1 (main model), Model 2 (auxiliary model) and Model 3 (auxiliary model) to generate the corresponding original test data. Perform data verification on the multiple field data corresponding to the original test data generated by Model 1, and determine that the ID number is abnormal and the mobile phone number is null. Model 1 needs to supplement the ID number and mobile phone number from the original test data of Model 2. However, when Model 1 uses the data of Model 2 to repair the ID number, it is found that the ID number of Model 2 is still abnormal. At this time, use the data of Model 3 to repair the ID number of Model 1. Model 1 uses the data of Model 2 to repair the null value to obtain the target test data.

[0073] In the above embodiment, when the verification result shows that there is abnormal data that does not conform to the field verification rules, the problem can be quickly located, improving the accuracy of data error correction. By determining the matching target field data based on the verification results of the remaining original test data and using it to update the abnormal data, the effective correction of the original test data is realized. This not only ensures the coherence and logic of the target test data in semantics, but also improves the quality and reliability of the data, enabling the test data to more realistically simulate the actual situation.

[0074] Step S204: Execute the test case corresponding to the target test data to generate the test result corresponding to the target test data. For details, please refer to Figure 1 Step S104 of the embodiment shown here, which will not be elaborated here.

[0075] The data testing method provided by the embodiments of the present invention separates according to the field attributes of the original test data, accurately disassembles complex data into multiple field data, making data processing more targeted. Obtain the verification rules matching each field data and verify them, effectively ensuring the accuracy and reliability of the field data, and can timely discover and correct data errors. Based on the verification results and the semantics of the field data for synthesis processing, this not only ensures that the synthesized target test data has excellent quality, but also makes full use of the semantic information of the original data, so that the finally generated target test data can more comprehensively and accurately reflect the test requirements, greatly improving the scientificity and effectiveness of data testing.

[0076] In this embodiment, a data testing method is provided, which can be used in computer devices such as laptops, desktop computers, etc. Figure 4 It is a flowchart of the data testing method according to the embodiments of the present invention, as Figure 4 shown, and this process includes the following steps:

[0077] Step S301, obtain multiple pre-trained data generation models and the first model prompt information. For details, please refer to Figure 2 Step S201 of the embodiment shown, which will not be elaborated here.

[0078] Step S302, obtain a preset model configuration fixture; based on the model configuration fixture, call each data generation model to enable each data generation model to generate corresponding original test data.

[0079] A model configuration fixture (fixture) is a predefined configuration tool or module used to automatically call and configure multiple data generation models before the execution of automated tests. The model configuration fixture ensures the correct loading and operation of each data generation model in the test environment through preset rules and parameters (such as model paths, hyperparameters, call order, etc.). Specifically, the model configuration fixture is loaded through a predefined configuration file or script, which contains the initialization parameters of the data generation model (such as model paths, call order, hyperparameters, etc.). Read the configuration information of the model configuration fixture, automatically call each data generation model (such as models based on GPT-4, Transformer architecture, etc.) to generate corresponding original test data and process the original test data (such as data desensitization processing described later).

[0080] The data testing method provided by the embodiments of the present invention provides a standardized tool and process for automatically calling each data generation model by obtaining a preset model configuration fixture. This not only retains the stability and coherence of the original automated test system, but also improves the reliability and efficiency of automated tests.

[0081] Step S303: Use the first model prompt information to guide each data generation model to generate corresponding original test data. For details, please refer to Figure 2 Step S202 of the embodiment shown, which will not be elaborated here.

[0082] Step S304: Based on the semantics of each original test data, perform synthesis processing on each original test data to obtain target test data. For details, please refer to Figure 2 Step S203 of the embodiment shown, which will not be elaborated here.

[0083] Step S305: Generate noise data for the target test data based on preset noise parameters; inject the noise data into the target test data to obtain the target test data after noise perturbation.

[0084] The preset noise parameters include a position parameter and a noise intensity parameter. When the position parameter is 0, it means that the value of the sensitive field can be added or subtracted by the noise data. The noise intensity parameter (epsilon) is used to control the intensity of the noise. For example, when epsilon = 0.1, the scale parameter of the Laplace distribution is 1 / epsilon = 10, and the generated noise range is larger, and the privacy protection effect is stronger; if epsilon = 1, the noise range is smaller, and the data retains more original information. The noise data refers to a random perturbation value generated by a specific algorithm (such as the Laplace noise generation function) for privacy protection of sensitive information. Specifically, call the add_noise_to_number function for the sensitive fields (such as mobile phone numbers and ID card numbers) in the target test data. For example, as Figure 5 shown, input the ID card number 410125198610304364, the function internally generates Laplace noise (such as noise = 4), adds the integer part of the noise to the integer converted from the original value (41012519861030436(4 + 4) = 410125198610304368), and finally returns the perturbed data "410125198610304368" in string form. Integrate all fields (sensitive fields have been perturbed, non-sensitive fields remain unchanged) to generate the target test data after noise perturbation.

[0085] Step S306: Use the second model prompt information to guide any data generation model to perform desensitization processing on the target test data after noise perturbation to obtain the desensitized target test data.

[0086] The second model prompt information refers to the instructions or input templates used to guide the data generation model to perform desensitization processing, which is used to guide the data generation model to perform secondary desensitization on the data after noise perturbation. Specifically, the prompt information is stored in the configuration file, database, or rule engine, and can be dynamically loaded through system calls. The target test data after noise perturbation is combined with the second model prompt information, and the data generation model (such as GPT-4) is called for secondary desensitization to obtain the desensitized target test data. For example, as Figure 5 shown, the data after noise perturbation is {Wang Wu; 99; 410125198610304368; 13786625654}, and the second model prompt information can be "Change the numbers in the 4th and 10th positions of the ID number". According to the desensitization needle in the data generation model, the positions of the numbers to be changed are determined, and the second model prompt information is used to guide the data generation model to perform desensitization processing on the target test data after noise perturbation, obtaining the desensitized target test data as {Wang Wu; 99; 410825198510304368; 13786965654}.

[0087] The data testing method provided by the embodiments of the present invention generates noise data based on preset noise parameters and injects it into the target test data. By adding interference elements, the precise features of the original data are blurred, simulating real environment noise interference while reducing data sensitivity. Using the second model prompt information to guide the data generation model to perform desensitization processing on the target test data that has been injected with noise and preliminarily desensitized achieves a double desensitization effect. This double desensitization mechanism greatly enhances the protection of sensitive information and avoids potential risks caused by data sensitivity.

[0088] Step S307, execute the test case corresponding to the target test data to generate the test result corresponding to the target test data. For details, please refer to Figure 2 step S204 of the embodiment shown, which will not be elaborated here.

[0089] Step S308, if the test result indicates that the target test data passes the test of the test case, then determine the target test data as valid test data.

[0090] Valid test data is high-quality data verified by test cases. Specifically, after executing the test case, if the target test data passes the test case, then mark the target test data as valid test data. For example, input the target test data (desensitized test data including name, mobile phone number, and address). If the target test data passes the test case of "user registration request", then mark the target test data as valid test data.

[0091] Step S309: Use the valid test data to adjust the parameters of each data generation model, and obtain the adjusted data generation model.

[0092] Use the valid test data as the training set and input it into each data generation model. Update the model parameters of each data generation model through the backpropagation algorithm to obtain the adjusted data generation model.

[0093] Specifically, the above Step S309 includes:

[0094] Step S3091: Use the load balancing method to distribute the valid test data to each data generation model, and obtain the model test data corresponding to each data generation model.

[0095] Intelligently distribute the verified valid test data to each data generation model instance through a load balancer (such as NGINX). Specifically, the distribution strategy can be round-robin to ensure the load balancing of each model instance and avoid resource idleness or overload. For example, assume there are 3 data generation model instances (Model A, B, and C). The load balancer divides a batch of valid test data into multiple subsets (i.e., model test data) and distributes the multiple subsets to Model A, B, and C in a round-robin manner.

[0096] Step S3092: Use each model test data to adjust the parameters of the corresponding data generation model, and obtain the adjusted data generation model.

[0097] Each data generation model instance uses the model test data allocated to it for training, and updates the model parameters (such as neural network weights) through the backpropagation algorithm. After the training is completed, each data generation model instance generates an updated parameter file to replace the original model parameters.

[0098] In the data testing method provided by the embodiments of the present invention, when the test result shows that the target test data passes the test case, it is determined as valid test data, providing a clear determination basis for the validity of the data. Subsequently, use these valid test data to adjust the parameters of each data generation model, enabling the model to optimize its own parameters based on the data feedback that performs well in the actual test. This not only helps to improve the quality and accuracy of the data generated by the data generation model, making the data generated more likely to pass the test case in future tests, but also enables the model to better adapt to the requirements of the actual application scenario.

[0099] In this embodiment, a data testing device is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0100] This embodiment provides a data testing device, as Figure 6 shown, including:

[0101] A first acquisition module 401, configured to acquire a plurality of pre-trained data generation models and first model prompt information;

[0102] A generation module 402, configured to use the first model prompt information to guide each data generation model to generate corresponding original test data;

[0103] A synthesis module 403, configured to perform synthesis processing on each original test data based on the semantics of each original test data to obtain target test data;

[0104] A test module 404, configured to execute a test case corresponding to the target test data and generate a test result corresponding to the target test data.

[0105] In some alternative implementation manners, the generation module 402 includes:

[0106] An acquisition sub-module, configured to acquire the model order of each data generation model;

[0107] A generation sub-module, configured to use the first model prompt information to simultaneously guide each data generation model to generate corresponding original test data, and determine the data priority of each original test data based on the model order.

[0108] In some alternative implementation manners, the synthesis module 403 includes:

[0109] A segmentation unit, configured to separate each original test data based on the field attributes of each original test data to obtain a plurality of field data corresponding to each original test data;

[0110] An acquisition unit, configured to acquire a field verification rule matching each field data, and perform data verification on each field data according to the field verification rule to obtain a verification result;

[0111] A synthesis unit, configured to perform synthesis processing on each original test data based on the verification result and the semantics of the field data to obtain target test data.

[0112] In some alternative implementation manners, the synthesis unit includes:

[0113] A first determination unit, configured to determine, if the verification result indicates that any field data of the current original test data does not conform to the corresponding field verification rule, the field data that does not conform to the field verification rule as abnormal data;

[0114] A second determination unit, configured to determine target field data matching the abnormal data based on the verification results of the remaining original test data other than the current original test data;

[0115] An update unit, configured to update the abnormal data with the target field data to obtain target test data corresponding to the current original test data.

[0116] In some alternative embodiments, the data testing device further includes:

[0117] A noise module, configured to generate noise data for the target test data based on preset noise parameters;

[0118] A perturbation module, configured to inject the noise data into the target test data to obtain the target test data after noise perturbation;

[0119] A second acquisition module, configured to use second model prompt information to guide any data generation model to desensitize the target test data after noise perturbation to obtain desensitized target test data.

[0120] In some alternative embodiments, the data testing device further includes:

[0121] A determination module, configured to determine the target test data as valid test data if the test result indicates that the target test data passes the test of the test case;

[0122] An adjustment module, configured to adjust the parameters of each data generation model with the valid test data to obtain an adjusted data generation model.

[0123] In some alternative embodiments, the adjustment module includes:

[0124] A distribution sub-module, configured to distribute the valid test data to each data generation model in a load balancing manner to obtain model test data corresponding to each data generation model;

[0125] An adjustment sub-module, configured to adjust the parameters of the corresponding data generation model with each model test data to obtain an adjusted data generation model.

[0126] In some alternative embodiments, the data testing device further includes:

[0127] A third acquisition module, configured to acquire a preset model configuration fixture;

[0128] A calling module, configured to call each data generation model based on the model configuration fixture, so that each data generation model generates corresponding original test data.

[0129] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.

[0130] The data testing device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0131] The data testing device provided by the embodiment of the present invention can generate diversified original test data by acquiring multiple pre-trained data generation models and the first model hint information, and utilize the characteristics of different models. The generation guided by the hint information makes the data generation more targeted. The original test data is synthesized and processed based on semantics, which can integrate advantages and make the target test data more comprehensive and accurate, so that the test results generated by executing its corresponding test cases are more reliable, comprehensively improving the effectiveness and accuracy of the test.

[0132] The embodiment of the present invention also provides a computer device having the above-mentioned Figure 6 shown data testing device.

[0133] Please refer to Figure 7 , Figure 7 is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As Figure 7 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 7 In

[0134] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device may be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.

[0135] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0136] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0137] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memories.

[0138] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 7 Taking the connection through the bus as an example.

[0139] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.

[0140] The computer device further includes a communication interface for the computer device to communicate with other devices or communication networks.

[0141] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0142] A part of the present invention can be applied as a computer program product, for example, computer program instructions, which when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should be able to understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.

[0143] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A data testing method, characterized in that: The method comprises: Obtain multiple pre-trained data generation models and first model prompt information; Using the first model prompt information to guide each of the data generation models to generate corresponding original test data; Based on the semantics of each of the original test data, synthesizing each of the original test data to obtain target test data; Execute the test case corresponding to the target test data and generate the test result corresponding to the target test data.

2. The method according to claim 1, characterized in that The using the first model prompt information to guide each of the data generation models to generate corresponding original test data includes: Obtaining a model order of each of the data generation models; The first model prompt information is used to simultaneously guide each of the data generation models to generate corresponding original test data, and the data priority of each of the original test data is determined based on the model sequence.

3. The method according to claim 1 or 2, characterized in that: Based on the semantics of each of the original test data, synthesizing each of the original test data to obtain target test data includes: Separating each of the original test data based on a field attribute of each of the original test data to obtain a plurality of field data corresponding to each of the original test data; Obtaining a field verification rule that matches each of the field data, performing data verification on each of the field data according to the field verification rule, and obtaining a verification result; Based on the verification result and the semantics of the field data, each of the original test data is synthesized to obtain the target test data.

4. The method according to claim 3, characterized in that The synthesizing and processing each of the original test data based on the verification result and the semantics of the field data to obtain the target test data includes: If the verification result indicates that any field data of the current original test data does not conform to the corresponding field verification rule, the field data that does not conform to the field verification rule is determined as abnormal data; Determine target field data matching the abnormal data based on the verification results of the remaining original test data other than the current original test data; The abnormal data is updated using the target field data to obtain target test data corresponding to the current original test data.

5. The method according to claim 1, characterized in that Before executing the test case corresponding to the target test data and generating the test result corresponding to the target test data, the method further includes: generating noise data for the target test data based on preset noise parameters; Injecting the noise data into the target test data to obtain the target test data after noise disturbance; The second model prompt information is used to guide any of the data generation models to perform desensitization processing on the target test data after the noise disturbance to obtain the desensitized target test data.

6. The method according to claim 1, characterized in that Also includes: If the test result indicates that the target test data passes the test of the test case, the target test data is determined as valid test data; The valid test data is used to adjust the parameters of each of the data generation models to obtain an adjusted data generation model.

7. The method according to claim 6, characterized in that The using the valid test data to adjust the parameters of each of the data generation models to obtain the adjusted data generation model includes: Distributing the valid test data to each of the data generation models in a load balancing manner to obtain model test data corresponding to each of the data generation models; Utilize each of the model test data to adjust the parameters of the corresponding data generation model to obtain an adjusted data generation model.

8. The method according to claim 1, characterized in that Before using the first model prompt information to guide each of the data generation models to generate corresponding original test data, the method further includes: Get the preset model configuration fixture; The data generation models are called based on the model configuration fixture, so that each data generation model generates corresponding original test data.

9. A data testing device, characterized in that: The device comprises: A first acquisition module is used to acquire a plurality of pre-trained data generation models and first model prompt information; A generation module, used for guiding each of the data generation models to generate corresponding original test data by using the first model prompt information; A synthesis module, used for synthesizing each of the original test data based on the semantics of each of the original test data to obtain target test data; The test module is used to execute the test cases corresponding to the target test data and generate the test results corresponding to the target test data.

10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data testing method according to any one of claims 1 to 8 by executing the computer instructions.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data testing method according to any one of claims 1 to 8.

12. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to cause a computer to execute the data testing method according to any one of claims 1 to 8.