Test data generation method, device, computer device, storage medium and product
By determining and adjusting the failure representation vector of the artificial intelligence software system and generating target test data, the problem of insufficient uncovering errors in the reliability test of the artificial intelligence software system in the prior art is solved, and a more accurate reliability evaluation is achieved.
Patent Information
- Application Number
- CN202410118457.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-01-26
AI Technical Summary
In third-party reliability testing, relying on the test set provided by developers may lead to excessive reliability estimation, while directly using domain data sets may lead to excessive reliability estimation and weak ability to uncover errors.
By determining the first failure characterization vector of the test set sample and the multiple second failure characterization vectors of the domain data set sample, a plurality of third failure characterization vectors similar to the first failure characterization vector are determined, and the third failure characterization vectors are adjusted based on the first failure characterization vector to generate target test data.
A test sample between the developer's test set and the domain data set was constructed, which effectively expanded the existing test set and improved the error-removing ability of artificial intelligence software system reliability testing.
Smart Images

Figure CN118227463B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a test data generation method, device, computer device, storage medium, and product. Background Art
[0002] With the rapid growth of the artificial intelligence industry. China's artificial intelligence industry is gradually evolving from the development stage of integrating artificial intelligence technology with typical application scenarios in various industries to the mature stage of efficient and industrial production, and the industry growth trend is rapid. The new features of artificial intelligence software systems pose challenges to reliability evaluation.
[0003] When conducting third-party reliability testing on artificial intelligence software systems and artificial intelligence components, relying on the test sets provided by developers may overestimate the reliability level because the model may overfit on the test case sets provided by developers. However, directly using domain data sets also has drawbacks. For example, domain data sets may contain objects that the model does not need to identify, scenarios that it will not encounter, etc., which may lead to an underestimation of the reliability level, and the ability of the artificial intelligence software system to detect errors using the above test data is weak. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a test data generation method, device, computer device, storage medium, and product that can improve the error detection ability.
[0005] In a first aspect, the present application provides a test data generation method, which is characterized in that the method includes:
[0006] Determine a first failure characterization vector of a test set sample and multiple second failure characterization vectors of domain data set samples;
[0007] From the multiple second failure characterization vectors, determine multiple third failure characterization vectors that are similar to the first failure characterization vector;
[0008] Based on the first failure characterization vector, perform adjustment processing on each third failure characterization vector to obtain multiple target test data.
[0009] In one of the embodiments, the above-mentioned performing adjustment processing on each third failure characterization vector based on the first failure characterization vector to obtain multiple target test data includes:
[0010] Respectively determine the number of different bits between each third failure characterization vector and the first failure characterization vector; the number of different bits is used to represent the difference number between the first failure characterization vector and the third failure characterization vector in the dimension of failure characterization elements;
[0011] For each third failure characterization vector, the third failure characterization vector is adjusted according to the pre-obtained variability threshold and different number of digits to obtain target test data.
[0012] In one embodiment, the above-mentioned adjusting and processing the third failure characterization vector according to the variability threshold and different number of digits to obtain target test data includes:
[0013] Calculate the product between the variability threshold and the different number of digits;
[0014] Determine the target adjustment position on the third failure characterization vector according to the product;
[0015] Adjust the data at the target adjustment position according to the first failure characterization vector to obtain target test data.
[0016] In one embodiment, the above-mentioned determining a plurality of third failure characterization vectors similar to the first failure characterization vector from a plurality of second failure characterization vectors includes:
[0017] Calculate the similarity between the first failure characterization vector and each second failure characterization vector respectively;
[0018] Determine the second failure characterization vector with a similarity greater than the preset similarity threshold as the third failure characterization vector.
[0019] In one embodiment, the above-mentioned determining the first failure characterization vector of the test set sample includes:
[0020] Obtain a plurality of test set samples;
[0021] Determine the data class characterization elements and environment class characterization elements corresponding to each test set sample;
[0022] When the data class characterization elements and environment class characterization elements corresponding to each test set sample meet the preset failure conditions, determine the first failure characterization vector.
[0023] In one embodiment, the above-mentioned determining the data class characterization elements and environment class characterization elements corresponding to each test set sample includes:
[0024] Use a panoramic object detection model to identify the environment class characterization elements corresponding to each test set sample;
[0025] Use an image class large model to identify the data class characterization elements corresponding to each test set sample.
[0026] In a second aspect, the present application also provides a test data generation device. The device includes:
[0027] A vector determination module, configured to determine a first failure characterization vector of a test set sample and multiple second failure characterization vectors of domain data set samples;
[0028] A similar vector determination module, configured to determine multiple third failure characterization vectors similar to the first failure characterization vector from the multiple second failure characterization vectors;
[0029] A vector adjustment module, configured to perform adjustment processing on each of the third failure characterization vectors based on the first failure characterization vector to obtain multiple target test data.
[0030] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:
[0031] Determine a first failure characterization vector of a test set sample and multiple second failure characterization vectors of domain data set samples;
[0032] Determine multiple third failure characterization vectors similar to the first failure characterization vector from the multiple second failure characterization vectors;
[0033] Based on the first failure characterization vector, perform adjustment processing on each of the third failure characterization vectors to obtain multiple target test data.
[0034] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0035] Determine a first failure characterization vector of a test set sample and multiple second failure characterization vectors of domain data set samples;
[0036] Determine multiple third failure characterization vectors similar to the first failure characterization vector from the multiple second failure characterization vectors;
[0037] Based on the first failure characterization vector, perform adjustment processing on each of the third failure characterization vectors to obtain multiple target test data.
[0038] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0039] Determine a first failure characterization vector of a test set sample and multiple second failure characterization vectors of domain data set samples;
[0040] Determine multiple third failure characterization vectors similar to the first failure characterization vector from the multiple second failure characterization vectors;
[0041] Based on the first failure characterization vector, adjust each third failure characterization vector respectively to obtain multiple target test data.
[0042] The above test data generation method, device, computer device, storage medium and product determine the first failure characterization vector of the test set samples and multiple second failure characterization vectors of the domain data set samples, determine multiple third failure characterization vectors similar to the first failure characterization vector from the multiple second failure characterization vectors, and based on the first failure characterization vector, adjust each third failure characterization vector respectively to obtain multiple target test data. The embodiments of the present application construct a class of test samples that can be understood and operated by testers and are between the test sets of developers and the domain data sets of test parties, which can effectively expand the existing test sets and improve the error detection ability of the reliability test of artificial intelligence software systems. Description of the Drawings
[0043] Figure 1 It is an application environment diagram of the test data generation method in an embodiment;
[0044] Figure 2 It is a flowchart of the test data generation method in an embodiment;
[0045] Figure 3 It is a flowchart of obtaining multiple target test data in an embodiment;
[0046] Figure 4 It is a flowchart of obtaining target test data in an embodiment;
[0047] Figure 5 It is a flowchart of determining multiple similar third failure characterization vectors in an embodiment;
[0048] Figure 6 It is a flowchart of determining the first failure characterization vector of the test set samples in an embodiment;
[0049] Figure 7 It is a schematic diagram of determining data type characterization elements and environment type characterization elements in an embodiment;
[0050] Figure 8 It is a structural block diagram of the test data generation device in an embodiment;
[0051] Figure 9 It is an internal structure diagram of a computer device in an embodiment. Detailed Description of the Invention
[0052] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application.
[0053] First, before specifically introducing the technical solutions of the embodiments of this application, the technical background on which the embodiments of this application are based will be introduced first.
[0054] The artificial intelligence industry has grown rapidly. China's artificial intelligence industry is gradually evolving from the development stage of integrating artificial intelligence technologies with typical application scenarios in various industries to the mature stage of efficient and industrialized production, with a rapid growth trend. The new features of artificial intelligence software systems pose challenges to reliability evaluation. An artificial intelligence software system is a computer software system that generates human intelligent behavior based on artificial intelligence technology and has received extensive attention and accelerated application in key fields such as healthcare, transportation, and finance. However, the core technology of artificial intelligence software systems - deep learning technology has the characteristics of a probabilistic system and a technology black box, and its technical principle is very different from that of traditional software. A series of new features have emerged in artificial intelligence software systems, such as strong adaptability and evolvability, weak generalization ability, weak interpretability, high uncertainty, and unclear failure mechanisms. It is easier to make unpredictable wrong decisions, and traditional software reliability analysis, testing, evaluation and other technologies are difficult to adapt to these new features. If the reliability of artificial intelligence software systems is not implemented properly, it will cause frequent software product failures, a significant increase in the total cost, and damage to the brand image in the lightest case, and may cause casualties, serious property losses and environmental damage in the most serious case. Therefore, it is necessary and urgent to carry out research on the reliability of artificial intelligence software systems.
[0055] The artificial intelligence core components in an artificial intelligence software system essentially belong to a probabilistic system, with high endogenous uncertainty, being sensitive to changes in the external environment, and prone to significant decline in prediction ability in new scenarios not included in the training set; the artificial intelligence components do not have a typical control structure, and the unclear failure mechanism makes it difficult to conduct a complete coverage test on the artificial intelligence software system; the test inputs of the core components of the artificial intelligence software system are mostly unstructured data, such as images, texts, voices, etc., and the process of generating reliability test data is not intuitive, making it difficult to determine whether it reflects the actual usage scenario; the above problems make it easier for artificial intelligence software systems to make unpredictable wrong decisions in rare scenarios compared with traditional software, and it is more difficult to eliminate software failures through prior testing.
[0056] The low-level requirements of artificial intelligence components are unclear, and the artificial intelligence software system is sensitive to the dynamically changing external environment, making it difficult to comprehensively test the reliability of the artificial intelligence software system in diverse scenarios. Research on the reliability testing technology of artificial intelligence software systems, construct a scenario profile of the artificial intelligence software system with time series changes, covering various typical scenarios such as normal conditions, out-of-distribution data, environmental disturbances, activation of failure modes, and adversarial attacks, and respectively adopt AIGC generation technology, data enhancement technology, adversarial attack technology, and out-of-distribution sample generation technology at the data layer based on the failure characterization law, and adopt statistical testing technology and software fault injection technology at the software layer to generate reliability test cases for various typical scenarios.
[0057] For the research on the artificial intelligence scenario profile construction technology, the operation profile of the artificial intelligence software system is represented in the form of Markov (decision process model). Different from the traditional software operation profile, it not only models the sequence relationship and operation probability between software operations, but also models the sequence relationship and operation probability between artificial intelligence component operations and software operations, and the artificial intelligence component operations can be further described by the artificial intelligence scenario profile. According to the different application scenarios of the artificial intelligence components to be tested, such as highways, tunnels, overpasses, etc. in autonomous driving, divide the application scenarios, and establish the following five types of preset scenarios under each scenario: test set scenario, out-of-distribution scenario, fault injection scenario, perturbation scenario, and adversarial scenario. The constructed artificial intelligence scenario profile changes with time, mainly reflected in the change of the proportion of different preset scenarios in different usage periods. For example, when the artificial intelligence software system is first put into use, the test set scenario can generally better summarize the actual usage scenario of the software; while after the software has been used for a long time, more out-of-distribution, failure mode, perturbation and other scenarios will be encountered.
[0058] The operation profile of the artificial intelligence software system based on the artificial intelligence scenario profile supports the following tests at the software layer: software statistical testing, which supports the generation of software reliability test cases based on random sampling;
[0059] Software failure mode testing, which supports injecting specified faults into the software to trigger failure modes and verifying whether the artificial intelligence software system correctly handles the specified failures.
[0060] Traditional software reliability testing includes more test types, such as model-based reliability testing, finite state machine-based reliability testing, software test coverage, etc. These software reliability test items are applicable to artificial intelligence software systems and will not be further studied in this application. The operation profile of the artificial intelligence software system based on the artificial intelligence scenario profile supports the following tests at the artificial intelligence component layer:
[0061] Test set scenario test: Conduct tests using the test set provided by the developer to verify the prediction ability of the artificial intelligence component under test in the preset usage scenarios;
[0062] Out-of-distribution scenario test: Use the domain dataset that the developer did not use during model training and testing, and adopt out-of-distribution data generation techniques to improve the similarity between the samples in the domain dataset and the samples in the test set provided by the developer, and verify the prediction ability of the artificial intelligence component under test in reasonable out-of-distribution scenarios;
[0063] Fault injection scenario test: Based on the test set provided by the developer, select applicable failure modes, and preferentially select failure modes that are prone to failure to generate new test data samples, and verify the prediction ability of the artificial intelligence component under test under specified fault injection;
[0064] Perturbation scenario test: Based on the test set provided by the developer, select applicable data perturbation methods, and preferentially select data perturbation methods that are prone to failure to generate new test data samples, and verify the prediction ability of the artificial intelligence component under test under specified data perturbation;
[0065] Adversarial scenario test: Based on the test set provided by the developer, select applicable adversarial attack methods, and preferentially select adversarial attack methods that are prone to failure to generate new test data samples, and verify the prediction ability of the artificial intelligence component under test under specified adversarial attacks.
[0066] The test set scenario test, perturbation scenario test, and adversarial scenario test are important components of the reliability test of artificial intelligence components. However, these test methods have been relatively well studied in artificial intelligence model testing and do not fall within the scope of the key research of this application.
[0067] When conducting third-party reliability tests on artificial intelligence software systems and artificial intelligence components, relying on the test set provided by the developer may overestimate the reliability level because various model architectures, different data preprocessing methods, etc. may be used during model training. And during the process of evaluating the quality of the model based on the test set provided by the developer, test set leakage may be introduced, resulting in overfitting of the model on the test case set provided by the developer. However, directly using the domain dataset also has drawbacks. For example, the domain dataset may contain objects that the model does not need to identify, scenarios that it will not encounter, etc. These pieces of information are often not described in detail in the software requirements and design documents but are implicit in the datasets used for model training and testing.
[0068] Based on this, this application provides a test data generation method, device, computer device, storage medium, and product, aiming to solve the above technical problems.
[0069] The test data generation method provided in the embodiments of this application can be applied to, for exampleFigure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed on the cloud or other network servers. The terminal 102 determines the first failure characterization vector of the test set sample and multiple second failure characterization vectors of the domain data set samples, determines multiple third failure characterization vectors similar to the first failure characterization vector from the multiple second failure characterization vectors, and based on the first failure characterization vector, adjusts each third failure characterization vector to obtain multiple target test data. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, and tablet computers. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0070] In an exemplary embodiment, as Figure 2 shown, the embodiment of the present application provides a test data generation method. Taking the method applied to the Figure 1 terminal 102 in it as an example for description, it includes the following S201 to S203. Among them:
[0071] S201, determine the first failure characterization vector of the test set sample and multiple second failure characterization vectors of the domain data set samples.
[0072] Among them, the test set sample can be the test set of the developer, and the domain data set is the domain data set of the test party.
[0073] In the embodiment of the present application, mainly test the test set samples in the out-of-distribution scenario. The terminal can select an appropriate number of test set samples according to requirements such as the number of test cases and time. For each selected test set sample, the panoramic object detection method and the large image model method can be used to identify the data type and environment type characterization elements, so as to construct the failure characterization vector. The first failure characterization vector can also be determined in advance, and the terminal can directly call the first failure characterization vector. The above method is also used to determine multiple second failure characterization vectors for the domain data set samples.
[0074] S202, determine multiple third failure characterization vectors similar to the first failure characterization vector from the multiple second failure characterization vectors.
[0075] In the embodiment of the present application, based on the first failure characterization vector and the second failure characterization vector obtained in the above embodiment, the terminal can calculate the similarity between the first failure characterization vector and the second failure characterization vector, or input both into the model to determine multiple third failure characterization vectors similar to the first failure characterization vector.
[0076] S203. Based on the first failure characterization vector, adjust each third failure characterization vector respectively to obtain multiple target test data.
[0077] In the embodiments of the present application, based on the third failure characterization vectors obtained from the above embodiments, for each third failure characterization vector, the terminal can compare the first failure characterization vector with the third failure characterization vector, and based on the first failure characterization vector, modify all or part of the values of the third failure characterization vector at the same bit positions to obtain the target test data.
[0078] The above test data generation method determines the first failure characterization vector of the test set samples and multiple second failure characterization vectors of the domain dataset samples, determines multiple third failure characterization vectors similar to the first failure characterization vector from the multiple second failure characterization vectors, and based on the first failure characterization vector, adjusts each third failure characterization vector respectively to obtain multiple target test data. In the embodiments of the present application, a type of test samples that can be understood and operated by testers and are between the test set of developers and the domain dataset of the test party are constructed, which can effectively expand the existing test set and improve the error discovery ability of the reliability test of the artificial intelligence software system.
[0079] In an exemplary embodiment, based on the above embodiments, please refer to Figure 3 , the embodiments of the present application relate to the process of adjusting each third failure characterization vector respectively based on the first failure characterization vector to obtain multiple target test data, including the following S301 to S302. Wherein:
[0080] S301. Determine the number of different bits between each third failure characterization vector and the first failure characterization vector respectively.
[0081] Among them, the number of different bits is used to represent the difference number between the first failure characterization vector and the third failure characterization vector in the dimension of failure characterization elements.
[0082] In the embodiments of the present application, based on the third failure characterization vectors and the first failure characterization vectors obtained from the above embodiments, the third failure characterization vectors and the first failure characterization vectors can be input into a comparison model to determine the number of different bits between each third failure characterization vector and the first failure characterization vector. It is also possible to compare the values of each bit of the third failure characterization vector and the first failure characterization vector one by one through traversal to determine the number of different bits between each third failure characterization vector and the first failure characterization vector.
[0083] S302. For each third failure characterization vector, adjust the third failure characterization vector according to the pre-obtained variability threshold and the number of different bits to obtain the target test data.
[0084] In the embodiments of the present application, based on the different number of bits obtained in the above embodiments, for each third failure characterization vector, the terminal first obtains a preset variability threshold, multiplies or divides the variability threshold by the pair of different number of bits, and searches for the corresponding number of bits in the third failure characterization vector according to the calculation result, and adjusts the values corresponding to the number of bits to obtain target test data.
[0085] In the embodiments of the present application, by pre-obtaining the variability threshold and different number of bits to adjust the values of the third failure characterization vector and the first failure characterization vector on different number of bits, a class of test samples that can be understood and operated by testers and are between the test set of developers and the domain data set of the test party can be constructed, which can effectively expand the existing test set and improve the error detection ability of the reliability test of the artificial intelligence software system.
[0086] In an exemplary embodiment, based on the above embodiments, please refer to Figure 4 , the embodiments of the present application relate to the process of adjusting the third failure characterization vector according to the variability threshold and different number of bits to obtain target test data, including the following S401 to S403. Among them:
[0087] S401, calculate the product between the variability threshold and the different number of bits.
[0088] In the embodiments of the present application, based on the variability threshold and different number of bits obtained in the above embodiments, the terminal can calculate the product between the variability threshold and the different number of bits. Specifically, first set the variability threshold M, and calculate that the different number of bits of the first failure characterization vector α and the third failure characterization vector βi of the test set sample is S. Then, randomly select S*M% bits different from the failure characterization vector βi of the similar domain data set sample on the first failure characterization vector α, and set them as the values of βi on this bit.
[0089] S402, determine the target adjustment position on the third failure characterization vector according to the product.
[0090] In the embodiments of the present application, based on the product obtained in the above embodiments, a position with the same value as the product value can be searched in the third failure characterization vector, and this position is determined as the target adjustment position. For example, when there are 50 different bits between α and βi, and the variability threshold M is 4%, then modification operations need to be performed on 2 randomly selected (50*4% = 2) bits of βi that are different from α.
[0091] S403, adjust the data at the target adjustment position according to the first failure characterization vector to obtain target test data.
[0092] In the embodiments of the present application, based on the target adjustment position obtained from the above embodiments, the terminal can adjust the data at the target adjustment position according to the first failure characterization vector to obtain target test data. Among them, the adjustment operation actually means generating target test data β'i that is the same as α in two failure characterization dimensions with βi as the blueprint.
[0093] In the embodiments of the present application, by calculating the product of the variability threshold and different digits, the number of digits that need to be changed can be calculated, so as to construct a class of test samples that can be understood and operated by testers and that are between the test sets of developers and the domain data sets of the testing party.
[0094] In an exemplary embodiment, based on the above embodiments, please refer to Figure 5 , the embodiments of the present application relate to the process of determining multiple third failure characterization vectors similar to the first failure characterization vector from multiple second failure characterization vectors, including the following S501 to S502. Among them:
[0095] S501, calculate the similarity between the first failure characterization vector and each second failure characterization vector respectively.
[0096] In the embodiments of the present application, the terminal first needs to calculate the cosine similarity between the failure characterization vector α of the test set sample and the failure characterization vector β of the domain data set sample.
[0097] S502, determine the second failure characterization vectors with similarity greater than the preset similarity threshold as the third failure characterization vectors.
[0098] In the embodiments of the present application, based on the similarity between the first failure characterization vector and each second failure characterization vector. First, set the similarity threshold T. When the cosine similarity between the failure characterization vector α of the test set sample and the failure characterization vector β of the domain data set sample is higher than the threshold T, then β and α are recorded as similar vectors. After retrieving the similarity of the failure characterization vectors, K similar domain data set sample failure characterization vectors (β1…βk) can be obtained.
[0099] In the embodiments of the present application, similar third failure characterization vectors can be calculated quickly and conveniently by calculating the similarity.
[0100] In an exemplary embodiment, based on the above embodiments, please refer to Figure 6 , the embodiments of the present application relate to the process of determining the first failure characterization vector of the test set sample, including the following S601 to S603. Among them:
[0101] S601, obtain multiple test set samples.
[0102] In the embodiments of the present application, the terminal can obtain the test set samples input by the user or obtain the test set samples from the database.
[0103] S602. Determine the data class characterization elements and environment class characterization elements corresponding to each test set sample.
[0104] In the embodiments of the present application, the data class characterization elements and environment class characterization elements corresponding to each test set sample can be determined through a pre-constructed model.
[0105] S603. When the data class characterization elements and environment class characterization elements corresponding to each test set sample meet the preset failure conditions, determine the first failure characterization vector.
[0106] In the embodiments of the present application, based on the data class characterization elements and environment class characterization elements corresponding to each test set sample obtained in the above embodiments, when both meet the failure conditions, the corresponding vector is determined as the first failure characterization vector.
[0107] In the embodiments of the present application, by comparing the data class characterization elements and environment class characterization elements corresponding to each test set sample with the failure conditions, the first failure characterization vector can be accurately obtained.
[0108] In an exemplary embodiment, based on the above embodiments, please refer to Figure 7 , the embodiments of the present application relate to the process of determining the data class characterization elements and environment class characterization elements corresponding to each test set sample, including the following S701 and S702. Wherein:
[0109] S701. Use a panoramic object detection model to identify the environment class characterization elements corresponding to each test set sample.
[0110] Among them, the panoramic object detection model can perform a two-dimensional display of panoramic information based on the panoramic image, or generate a corresponding floor plan or 3D model.
[0111] In the embodiments of the present application, for each test set sample, each test set sample is input into the panoramic object detection model to obtain all data class failure characterization elements.
[0112] S702. Use an image class large model to identify the data class characterization elements corresponding to each test set sample.
[0113] Among them, the image class large model refers to a model for image recognition.
[0114] In the embodiments of the present application, for each test set sample of the image class, each test set sample of the image class is input into the image class large model to obtain the data class characterization elements corresponding to each test set sample.
[0115] In an exemplary embodiment, based on the above embodiment, the method involved in the embodiments of the present application further includes the following steps:
[0116] Step 1: Obtain multiple test set samples; use a panoramic object detection model to identify the environmental characterization elements corresponding to each test set sample, and use an image large model to identify the data characterization elements corresponding to each test set sample. When the data characterization elements and environmental characterization elements corresponding to each test set sample meet the preset failure conditions, determine the first failure characterization vector.
[0117] Step 2: Obtain multiple domain dataset samples; use a panoramic object detection model to identify the environmental characterization elements corresponding to each domain dataset sample, and use an image large model to identify the data characterization elements corresponding to each domain dataset sample. When the data characterization elements and environmental characterization elements corresponding to each domain dataset sample meet the preset failure conditions, determine the second failure characterization vector.
[0118] Step 3: Calculate the similarity between the first failure characterization vector and each second failure characterization vector respectively, and determine the second failure characterization vectors with similarity greater than the preset similarity threshold as the third failure characterization vectors.
[0119] Step 4: Determine the number of different bits between each third failure characterization vector and the first failure characterization vector respectively; the number of different bits is used to represent the difference number between the first failure characterization vector and the third failure characterization vector in the dimension of failure characterization elements.
[0120] Step 5: For each third failure characterization vector, calculate the product of the mutation degree threshold and the number of different bits, determine the target adjustment position on the third failure characterization vector according to the product, and adjust the data at the target adjustment position according to the first failure characterization vector to obtain the target test data.
[0121] In the embodiments of the present application, when conducting third-party reliability testing on artificial intelligence software systems and artificial intelligence components, relying on the test sets provided by the developer may overestimate the reliability level because the model may overfit on the test case sets provided by the developer. However, directly using domain datasets also has drawbacks. For example, domain datasets may contain objects that the model does not need to identify, scenarios that it will not encounter, etc., which may lead to an underestimated reliability level. The embodiments of the present application construct a type of test samples that can be understood and operated by testers, which are between the developer's test sets and the tester's domain datasets, and can effectively expand the existing test sets and improve the error discovery ability of the reliability testing of artificial intelligence software systems.
[0122] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0123] Based on the same inventive concept, an embodiment of the present application further provides a test data generation device for implementing the above-mentioned test data generation method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the test data generation device provided below can refer to the limitations on the test data generation method in the above text, and will not be repeated here.
[0124] In one embodiment, as Figure 8 shown, a test data generation device 800 is provided, including:
[0125] A vector determination module 801, configured to determine a first failure characterization vector of a test set sample and a plurality of second failure characterization vectors of a domain data set sample;
[0126] A similar vector determination module 802, configured to determine a plurality of third failure characterization vectors similar to the first failure characterization vector from the plurality of second failure characterization vectors;
[0127] A vector adjustment module 803, configured to perform adjustment processing on each third failure characterization vector based on the first failure characterization vector to obtain a plurality of target test data.
[0128] In one of the embodiments, the above vector adjustment module 803 includes:
[0129] A digit determination unit, configured to respectively determine the number of different digits between each third failure characterization vector and the first failure characterization vector; the number of different digits is used to characterize the difference number between the first failure characterization vector and the third failure characterization vector in the dimension of failure characterization elements;
[0130] A data determination unit, configured to perform adjustment processing on each third failure characterization vector according to a pre-acquired variability threshold and the number of different digits to obtain target test data.
[0131] In one embodiment, the above data determination unit includes:
[0132] A product calculation sub-unit for calculating the product between the variability threshold and different digits;
[0133] A position determination sub-unit for determining the target adjustment position on the third failure characterization vector according to the product;
[0134] A data determination sub-unit for adjusting the data at the target adjustment position according to the first failure characterization vector to obtain the target test data.
[0135] In one embodiment, the above similarity vector determination module 802 includes:
[0136] A similarity determination unit for calculating the similarity between the first failure characterization vector and each second failure characterization vector respectively;
[0137] A similar vector determination unit for determining the second failure characterization vector with a similarity greater than the preset similarity threshold as the third failure characterization vector.
[0138] In one embodiment, the above vector determination module 801 includes:
[0139] A sample acquisition unit for acquiring multiple test set samples;
[0140] A feature determination unit for determining the data class characterization features and environment class characterization features corresponding to each test set sample;
[0141] A vector determination unit for determining the first failure characterization vector when the data class characterization features and environment class characterization features corresponding to each test set sample meet the preset failure conditions.
[0142] In one embodiment, the above feature determination unit includes:
[0143] An environment feature determination sub-unit for identifying the environment class characterization features corresponding to each test set sample by using a panoramic target detection model;
[0144] A data feature determination sub-unit for identifying the data class characterization features corresponding to each test set sample by using an image class large model.
[0145] Each module in the above test data generation device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0146] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structural diagram may be as shown in Figure 9 The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a test data generation method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0147] Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0148] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0149] Determine the first failure characterization vector of the test set sample and multiple second failure characterization vectors of the domain data set samples;
[0150] From the multiple second failure characterization vectors, determine multiple third failure characterization vectors similar to the first failure characterization vector;
[0151] Based on the first failure characterization vector, perform adjustment processing on each third failure characterization vector to obtain multiple target test data.
[0152] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0153] Determine the number of different bits between each third failure characterization vector and the first failure characterization vector respectively; the number of different bits is used to characterize the difference number between the first failure characterization vector and the third failure characterization vector in the dimension of failure characterization elements;
[0154] For each third failure characterization vector, perform adjustment processing on the third failure characterization vector according to the pre-acquired variability threshold and the number of different bits to obtain the target test data.
[0155] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0156] Calculate the product between the variability threshold and the number of different bits;
[0157] Determine the target adjustment position on the third failure characterization vector according to the product;
[0158] Adjust the data at the target adjustment position according to the first failure characterization vector to obtain the target test data.
[0159] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0160] Calculate the similarity between the first failure characterization vector and each second failure characterization vector respectively;
[0161] Determine the second failure characterization vectors with similarity greater than the preset similarity threshold as the third failure characterization vectors.
[0162] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0163] Obtain multiple test set samples;
[0164] Determine the data class characterization elements and environment class characterization elements corresponding to each test set sample;
[0165] When the data class characterization elements and environment class characterization elements corresponding to each test set sample meet the preset failure conditions, determine the first failure characterization vector.
[0166] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0167] Use a panoramic object detection model to identify the environment class characterization elements corresponding to each test set sample;
[0168] Use an image class large model to identify the data class characterization elements corresponding to each test set sample.
[0169] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0170] Determine a first failure characterization vector of a test set sample and multiple second failure characterization vectors of domain data set samples;
[0171] From the multiple second failure characterization vectors, determine multiple third failure characterization vectors that are similar to the first failure characterization vector;
[0172] Based on the first failure characterization vector, perform adjustment processing on each third failure characterization vector to obtain multiple target test data.
[0173] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0174] Respectively determine the number of different bits between each third failure characterization vector and the first failure characterization vector; the number of different bits is used to represent the number of differences between the first failure characterization vector and the third failure characterization vector in the dimension of failure characterization elements;
[0175] For each third failure characterization vector, perform adjustment processing on the third failure characterization vector according to a pre-acquired variability threshold and the number of different bits to obtain target test data.
[0176] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0177] Calculate the product between the variability threshold and the number of different bits;
[0178] According to the product, determine the target adjustment position on the third failure characterization vector;
[0179] Adjust the data at the target adjustment position according to the first failure characterization vector to obtain target test data.
[0180] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0181] Respectively calculate the similarity between the first failure characterization vector and each second failure characterization vector;
[0182] Determine the second failure characterization vectors with a similarity greater than a preset similarity threshold as the third failure characterization vectors.
[0183] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0184] Obtain multiple test set samples;
[0185] Determine the data - type representation elements and environment - type representation elements corresponding to each test - set sample;
[0186] When the data - type representation elements and environment - type representation elements corresponding to each test - set sample meet the preset failure conditions, determine the first failure representation vector.
[0187] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0188] Use a panoramic object - detection model to identify the environment - type representation elements corresponding to each test - set sample;
[0189] Use an image - type large model to identify the data - type representation elements corresponding to each test - set sample.
[0190] In one embodiment, a computer - program product is provided, including a computer program, which when executed by a processor implements the following steps:
[0191] Determine the first failure representation vector of the test - set sample and multiple second failure representation vectors of the domain - data - set samples;
[0192] From the multiple second failure representation vectors, determine multiple third failure representation vectors that are similar to the first failure representation vector;
[0193] Based on the first failure representation vector, perform adjustment processing on each third failure representation vector to obtain multiple target test data.
[0194] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0195] Respectively determine the number of different bits between each third failure representation vector and the first failure representation vector; the number of different bits is used to represent the difference number between the first failure representation vector and the third failure representation vector in the dimension of failure representation elements;
[0196] For each third failure representation vector, perform adjustment processing on the third failure representation vector according to the pre - obtained variability threshold and the number of different bits to obtain target test data.
[0197] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0198] Calculate the product of the variability threshold and the number of different bits;
[0199] According to the product, determine the target adjustment position on the third failure representation vector;
[0200] According to the first failure representation vector, adjust the data at the target adjustment position to obtain target test data.
[0201] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0202] Calculate the similarity between the first failure characterization vector and each second failure characterization vector respectively;
[0203] Determine the second failure characterization vectors with similarity greater than the preset similarity threshold as the third failure characterization vectors.
[0204] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0205] Obtain a plurality of test set samples;
[0206] Determine the data class characterization elements and environment class characterization elements corresponding to each test set sample;
[0207] When the data class characterization elements and environment class characterization elements corresponding to each test set sample meet the preset failure conditions, determine the first failure characterization vector.
[0208] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0209] Use a panoramic object detection model to identify the environment class characterization elements corresponding to each test set sample;
[0210] Use an image class large model to identify the data class characterization elements corresponding to each test set sample.
[0211] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0212] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.
[0213] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0214] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A test data generation method, characterized in that: The method comprises: Determining a first failure characterization vector of a test set sample from a developer and a plurality of second failure characterization vectors of a domain dataset sample; Determining, from the plurality of second failure characterization vectors, a plurality of third failure characterization vectors similar to the first failure characterization vector; respectively determining the number of different bits between each of the third failure characterization vectors and the first failure characterization vector; the number of different bits is used to represent the difference between the first failure characterization vector and the third failure characterization vector in the failure characterization element dimension; For each of the third failure characterization vectors, calculating the product between the pre-acquired variation threshold and the different number of bits; determining a target adjustment position on the third failure characterization vector according to the product; The data of the target adjustment position is adjusted according to the first failure characterization vector to obtain target test data.
2. The method according to claim 1, characterized in that The step of determining a plurality of third failure characterization vectors similar to the first failure characterization vector from the plurality of second failure characterization vectors comprises: respectively calculating similarities between the first failure characterization vector and each of the second failure characterization vectors; The second failure characterization vector having a similarity greater than a preset similarity threshold is determined as the third failure characterization vector.
3. The method according to claim 1, characterized in that The determining of a first failure characterization vector of a test set sample from a developer comprises: Get multiple test set samples; Determine the data type characterization element and the environment type characterization element corresponding to each of the test set samples; When the data characterization elements and the environment characterization elements corresponding to each of the test set samples meet preset failure conditions, the first failure characterization vector is determined.
4. The method according to claim 3, characterized in that The step of determining the data characterization elements and the environment characterization elements corresponding to each of the test set samples includes: Using a panoramic object detection model to identify the environmental representation elements corresponding to each of the test set samples; An image-based large model is used to identify the data class representation elements corresponding to each of the test set samples.
5. A test data generating device, characterized in that: The device comprises: A vector determination module, configured to determine a first failure characterization vector of a test set sample from a developer and a plurality of second failure characterization vectors of a domain data set sample; A similar vector determination module, configured to determine a plurality of third failure characterization vectors similar to the first failure characterization vector from the plurality of second failure characterization vectors; A vector adjustment module is used to respectively determine the different number of bits between each of the third failure characterization vectors and the first failure characterization vector; the different number of bits is used to characterize the difference between the first failure characterization vector and the third failure characterization vector in the failure characterization element dimension; for each of the third failure characterization vectors, calculate the product between the pre-acquired variation threshold and the different number of bits; determine the target adjustment position on the third failure characterization vector based on the product; adjust the data of the target adjustment position based on the first failure characterization vector to obtain target test data.
6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Detection model training method and device, computer equipment and storage medium
CN115396212A
Text category recognition method and apparatus, computer device, and medium
WO2023045184A1