Artificial Intelligence-Based Simulation Testing Method and Related Devices
Through the simulation test method based on artificial intelligence, the test data generation model is trained using historical test data, and the problems of human maintenance requirements and test data quality in the existing simulation test methods are solved, achieving higher test accuracy and reliability.
Patent Information
- Application Number
- CN202210952128.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-08-09
AI Technical Summary
The existing simulation testing methods require a lot of manpower to maintain, and due to differences in human configuration procedures, the quality of test data and the accuracy of test results cannot be guaranteed.
Using artificial intelligence-based simulation testing methods, we collect data from historical test records, encode and classify test data, train the test data to generate models, and automatically generate and transmit test data.
Improve the accuracy of simulated tests, reduce manpower maintenance requirements, and ensure the reliability of test results.
Smart Images

Figure CN115237802B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to an artificial intelligence-based simulation test method, device, electronic device, and storage medium. Background Art
[0002] With the development of information technology, more and more enterprises have an increasing demand for cross-system program performance testing. In some cross-system performance testing projects, due to objective factors such as limited test hardware resources and great difficulty in coordinating multiple systems, it is impossible to build a complete test environment to complete the test work. Therefore, software programs are usually used to simulate the functions of other systems to complete the test. This test method is usually called stub testing and is also called simulation testing.
[0003] Currently, simulation test programs are usually configured manually according to test requirements. However, this method requires a large amount of manpower for maintenance after the test program is configured, and due to differences in manually configured programs, the quality of test data cannot be guaranteed, and thus the accuracy of test results cannot be guaranteed. Summary of the Invention
[0004] In view of the above, it is necessary to provide an artificial intelligence-based simulation test method and related devices to solve the technical problem of how to improve the accuracy of simulation testing. Among them, the related devices include an artificial intelligence-based simulation test device, an electronic device, and a storage medium.
[0005] An embodiment of the present application provides an artificial intelligence-based simulation test method, and the method includes:
[0006] Collect historical test requirements and historical test data from historical test records, and perform encoding processing on the historical test requirements to obtain encoded data corresponding to each historical test requirement;
[0007] Classify each historical test requirement to obtain the category of each historical test requirement, and divide the encoded data and the historical test data into multiple training data sets according to the category, and the category corresponds to the training data set one by one;
[0008] Train a test data generation model corresponding to each training data set according to each training data set respectively;
[0009] Query the encoded data and communication protocol to be evaluated corresponding to the test requirement to be evaluated;
[0010] Classify the test requirement to be evaluated to obtain the category corresponding to the test requirement to be evaluated, and select a target model from the multiple test data generation models according to the category;
[0011] Input the to-be-evaluated encoded data into the target model multiple times to generate multiple batches of test data, and transmit the multiple batches of test data to a preset data receiver according to the communication protocol for multiple simulation tests.
[0012] In some embodiments, the encoding process for the historical test requirements to obtain the encoded data corresponding to each historical test requirement includes:
[0013] Perform word segmentation on the historical test requirements to obtain multiple words;
[0014] Encode each word according to a preset text encoding algorithm to obtain the encoding vector corresponding to each word, and associate the encoding vector with the word one by one as a word corpus;
[0015] Use the encoding vectors of all words corresponding to each historical test requirement as the encoded data corresponding to the historical test requirement.
[0016] In some embodiments, the classification of each historical test requirement to obtain the category of each historical test requirement, and the division of the encoded data and the historical test data into multiple training data sets according to the category includes:
[0017] Input the encoded data into a preset requirement classification model to obtain the category corresponding to each historical test requirement, and the category includes at least "credit investigation", "transaction", and "information query";
[0018] Use the historical test data corresponding to the historical test requirement as label data;
[0019] Use the encoded data corresponding to the historical test requirement as sample data, and associate the sample data with the label data one by one as training data;
[0020] Attribute the training data corresponding to the historical test requirements with the same category to the same training data set to obtain multiple training data sets, and the training data sets correspond to the categories one by one.
[0021] In some embodiments, the training of each test data generation model corresponding to each training data set respectively includes:
[0022] Construct an initial generation model, and the initial generation model includes an encoder and a generator;
[0023] For each of the training data sets, if the category corresponding to the training data set is not "credit investigation", then use the training data set to train the initial generation model, calculate the loss value of the initial generation model according to a preset loss function, and continuously update the parameters in the initial generation model until the loss value no longer changes, so as to obtain the first test data generation model corresponding to the training data set whose category is not "credit investigation".
[0024] If the category corresponding to the training data set is "credit investigation", then use the sample data in the training data set as keys and the label data as values to construct key-value pairs, use all the key-value pairs as the second test data generation model, and use all the first test data generation models and the second test data generation model together as the test data generation model.
[0025] In some embodiments, querying the to-be-evaluated encoded data and communication protocol corresponding to the to-be-evaluated test requirement includes:[[]]
[0026] Performing word segmentation on the to-be-evaluated test requirement to obtain a plurality of to-be-evaluated words;
[0027] Querying the encoded data corresponding to each of the to-be-evaluated words from the vocabulary corpus as the to-be-evaluated encoded data;
[0028] Querying the communication protocol corresponding to the to-be-evaluated test requirement, where the communication protocol is used to characterize the protocol for the to-be-evaluated test requirement to receive test data.
[0029] In some embodiments, classifying the to-be-evaluated test requirement to obtain the category corresponding to the to-be-evaluated test requirement, and selecting a target model from the plurality of test data generation models according to the category includes:[[]]
[0030] Inputting the to-be-evaluated encoded data into the preset requirement classification model to obtain the category corresponding to the to-be-evaluated requirement;
[0031] Traversing the test data generation models in sequence, comparing the category of the test data generation model with the category of the to-be-evaluated test requirement. If the category of the test data generation model is the same as the category of the to-be-evaluated test requirement, then use the test data generation model as the target model;
[0032] If the category of the test data generation model is different from the category of the to-be-evaluated test requirement, then continue to traverse until the target model is obtained and then stop traversing.
[0033] In some embodiments, the test data generation model is stored in a preset server. The steps of repeatedly inputting the to-be-evaluated encoded data into the target model to generate multiple batches of test data and transmitting the multiple batches of test data to a preset data receiver according to the communication protocol for multiple simulation tests include:
[0034] a. Continuously query the resource occupancy rate of the preset server;
[0035] b. When the resource occupancy rate is less than a preset occupancy rate threshold, input the to-be-evaluated encoded data into the target model to obtain test data. When the resource occupancy rate is not less than the preset occupancy rate threshold, stop executing the target model;
[0036] c. Repeat steps a and b to obtain multiple batches of test data, and transmit the test data to a preset data receiver according to the communication protocol for multiple simulation tests until the number of repetitions is not less than a preset repetition threshold, then stop repeating.
[0037] An embodiment of the present application further provides an artificial intelligence-based simulation test device, which includes:
[0038] An encoding unit, configured to collect historical test requirements and historical test data from historical test records, and perform encoding processing on the historical test requirements to obtain encoded data corresponding to each historical test requirement;
[0039] A classification unit, configured to classify each historical test requirement to obtain the category of each historical test requirement, and divide the encoded data and the historical test data into multiple training data sets according to the category, and the category corresponds to the training data set one by one;
[0040] A training unit, configured to train a test data generation model corresponding to each training data set according to each training data set;
[0041] A query unit, configured to query the to-be-evaluated encoded data and the communication protocol corresponding to the to-be-evaluated test requirement;
[0042] A selection unit, configured to classify the to-be-evaluated test requirement to obtain the category of the to-be-evaluated test requirement, and select a target model from the multiple test data generation models according to the category;
[0043] A test unit, configured to repeatedly input the to-be-evaluated encoded data into the target model to generate multiple batches of test data, and transmit the multiple batches of test data to a preset data receiver according to the communication protocol for multiple simulation tests.
[0044] An embodiment of the present application further provides an electronic device, which includes:
[0045] a memory for storing computer-readable instructions; and
[0046] a processor for executing the computer-readable instructions stored in the memory to implement the artificial intelligence-based simulation test method.
[0047] An embodiment of the present application further provides a computer-readable storage medium, in which computer-readable instructions are stored, and the computer-readable instructions are executed by a processor in an electronic device to implement the artificial intelligence-based simulation test method.
[0048] The above artificial intelligence-based simulation test method encodes a large number of historical test requirements to obtain a large amount of encoded data, classifies the encoded data to obtain the category of each historical test requirement, and divides the encoded data and historical test data into multiple training data sets according to the category of the historical test requirement, and uses each training data set to train a targeted test data generation model, which can provide model guidance for the test task to be evaluated, thereby improving the accuracy of data testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a flowchart of a preferred embodiment of an artificial intelligence-based simulation test method involved in the present application.
[0050] Figure 2 is a functional module diagram of a preferred embodiment of an artificial intelligence-based simulation test device involved in the present application.
[0051] Figure 3 is a schematic structural diagram of an electronic device of a preferred embodiment of an artificial intelligence-based simulation test method involved in the present application.
[0052] Figure 4 is a schematic structural diagram of an initial generation model involved in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to more clearly understand the purpose, features, and advantages of the present application, the present application will be described in detail below with reference to the drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other. Many specific details are set forth in the following description in order to provide a thorough understanding of the present application. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0054] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used in the description of this application herein are for the purpose of describing specific embodiments only and are not intended to limit this application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0056] An embodiment of this application provides an artificial intelligence-based simulation test method, which can be applied to one or more electronic devices. The electronic device is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0057] The electronic device may be any electronic product capable of human-computer interaction with a user. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.
[0058] The electronic device may further include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing (Cloud Computing).
[0059] The network where the electronic device is located includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.
[0060] Such as Figure 1As shown, it is a flowchart of a preferred embodiment of the simulation test method based on artificial intelligence in this application. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted.
[0061] S10. Collect historical test requirements and historical test data from historical test records, and perform encoding processing on the historical test requirements to obtain encoded data corresponding to each historical test requirement.
[0062] In an optional embodiment, the performing encoding processing on the historical test requirements to obtain encoded data corresponding to each historical test requirement includes:
[0063] Perform word segmentation processing on the historical test requirements to obtain a plurality of words;
[0064] Encode each word according to a preset text encoding algorithm to obtain an encoding vector corresponding to each word, and correspond the encoding vector to the word one by one to form a word corpus;
[0065] Use the encoding vectors of all words corresponding to each historical test requirement as the encoded data corresponding to the historical test requirement.
[0066] In this optional embodiment, the historical test requirements and the historical test data are in one-to-one correspondence. Exemplarily, when the historical test requirement is "scrape user credit data", the historical test data is user credit data; when the historical test requirement is "test the payment interface of an e-commerce platform", the test data is payment data.
[0067] In this optional embodiment, the preset text encoding algorithm can be an existing text encoding algorithm such as the GloVe algorithm (Global Vector algorithm), Skip-Gram algorithm, CBOW algorithm (Continuous Bag Of Words algorithm), etc. This application does not make any limitations in this regard.
[0068] In this optional embodiment, the encoding vectors corresponding to each word can be combined according to the arrangement order of all words in the historical test requirement to be used as the encoded data corresponding to the historical test requirement.
[0069] In this way, a large amount of encoded data is obtained by performing word segmentation and encoding on a large number of historical test requirements, providing data support for subsequent construction of a training dataset.
[0070] S11. Classify each historical test requirement to obtain the category of each historical test requirement, and divide the encoded data and the historical test data into multiple training datasets according to the category, with the category corresponding to the training dataset one by one.
[0071] In an optional embodiment, classifying each of the historical test requirements to obtain the category of each historical test requirement, and dividing the encoded data and the historical test data into multiple training data sets according to the category, including:
[0072] Inputting the encoded data into a preset requirement classification model to obtain the category corresponding to each historical test requirement, where the category at least includes "credit investigation", "transaction", and "information query";
[0073] Taking the historical test data corresponding to the historical test requirement as label data;
[0074] Taking the encoded data corresponding to the historical test requirement as sample data, and taking the sample data and the label data in one-to-one correspondence as training data;
[0075] Grouping the training data corresponding to the historical test requirements with the same category into the same training data set to obtain multiple training data sets, where the training data sets correspond to the categories one-to-one.
[0076] In this optional embodiment, the preset requirement classification model may be an existing classification model such as an XGBoost model (Extreme Gradient Boost), a LightGBM model (Light Gradient Boost Machine), a GBDT (Gradient Boost Decision Tree), or a random forest model. The present application does not limit this.
[0077] The category at least includes "credit investigation", "transaction", and "information query". When the category of the historical test requirement is "credit investigation", the historical test requirement needs to receive "credit investigation data" to test the system; when the category of the historical test requirement is "transaction", the historical test requirement needs to receive "transaction data" to test the system; when the category of the historical test requirement is "information query", the historical test requirement needs to receive "user information" to test the system.
[0078] In this way, by constructing multiple groups of training data with the historical test data and the encoded data in one-to-one correspondence, and dividing the training data into multiple training data sets according to the category of the historical requirements, it is ensured that a unique test data generation model can be trained using the training data set corresponding to each category subsequently, thereby improving the fit between the test data and the test requirements.
[0079] S12. Training a test data generation model corresponding to each training data set according to each training data set respectively.
[0080] In an optional embodiment, the generating the test data generation model corresponding to each of the training data sets respectively according to each of the training data sets includes:
[0081] Constructing an initial generation model, where the initial generation model includes an encoder and a generator;
[0082] For each of the training data sets, if the category corresponding to the training data set is not "credit investigation", then use the training data set to train the initial generation model, calculate the loss value of the initial generation model according to a preset loss function, and continuously update the parameters in the initial generation model until the loss value no longer changes, so as to obtain the first test data generation model corresponding to the training data set whose category is not "credit investigation";
[0083] If the category corresponding to the training data set is "credit investigation", then use the sample data in the training data set as keys and the label data as values to construct key-value pairs, use all the key-value pairs as the second test data generation model, and use all the first test data generation models and the second test data generation model together as the test data generation model.
[0084] In this optional embodiment, the initial generation model includes an encoder and a generator. Both the encoder and the generator can be existing neural network structures such as an LSTM model (Long Short Term Memory), an RNN model (Recurrent Neural Network), a GRU model (Gate Recurrent Unit), etc. This application does not make any limitations in this regard. Both the encoder and the generator include multiple neurons, and both the encoder and the generator are formed by connecting the multiple neurons in series. Taking the RNN model as an example, as Figure 4 shown is the structural schematic diagram of the initial generation model.
[0085] In this optional embodiment, taking any one of the training data sets whose category is not "credit investigation" as an example, the input of the encoder is the sample data in the training data set, and the output of the encoder is the feature vector corresponding to the sample data; the input of the generator is the feature vector, and the output of the generator is the virtual test data corresponding to the sample data.
[0086] In this optional embodiment, the virtual test data and the label data corresponding to the sample data may be input into a preset loss function to calculate the loss value of the initial generation model. The preset loss function may be an existing loss function such as a root mean square error function, a least squares function, or an Euclidean distance function. This application does not limit this.
[0087] In this optional embodiment, the parameters of the initial generation model may be continuously updated according to the gradient descent method until the loss value of the initial generation model no longer changes. Then, the update of the parameters of the initial generation model is stopped, and a first test data generation model corresponding to the training data set is obtained.
[0088] In this optional embodiment, if the category corresponding to the training data set is "credit investigation", it indicates that the test data required for the test requirement corresponding to the training data set is a pre-organized credit report. Then, the sample data in the training data set may be used as the key, and the label data in the training data set may be used as the value to construct a key-value pair, and the key-value pair is used as the second test data generation model.
[0089] In this way, the initial generation model is trained using the training data sets corresponding to each category respectively, and test data generation models corresponding to each training data set are obtained, so that test data generation models that meet the historical test requirements of each category can be obtained, and further the accuracy of subsequent test data generation can be improved.
[0090] S13. Query the to-be-evaluated encoded data and communication protocol corresponding to the to-be-evaluated test requirement.
[0091] In an optional embodiment, the querying the to-be-evaluated encoded data and communication protocol corresponding to the to-be-evaluated test requirement includes:
[0092] Performing word segmentation on the to-be-evaluated test requirement to obtain a plurality of to-be-evaluated words;
[0093] Querying the encoded data corresponding to each of the to-be-evaluated words from the vocabulary corpus as the to-be-evaluated encoded data;
[0094] Querying the communication protocol corresponding to the to-be-evaluated test requirement, where the communication protocol is used to characterize the protocol for the to-be-evaluated test requirement to receive test data.
[0095] In this optional embodiment, the to-be-evaluated test requirement may be segmented into words using the preset word segmentation tool to obtain a plurality of to-be-evaluated words, and the encoded vectors corresponding to each of the to-be-evaluated words are queried from the vocabulary corpus, and the encoded vectors corresponding to the to-be-evaluated words are combined according to the arrangement order of the to-be-evaluated words in the to-be-evaluated test requirement as the to-be-evaluated encoded data.
[0096] In this optional embodiment, the communication protocol refers to the data transmission rules preset in the requirements to be evaluated. The communication protocol may be existing communication protocols such as TCP / IP (Transmission Control Protocol / Internet Protocol), SPP protocol (Sequenced Packet Protocol), NetBEUI (NetBios Enhanced User Interface), etc. This application does not limit this.
[0097] In this way, by classifying the test requirements to be evaluated to determine the test data generation model corresponding to the test requirements to be evaluated, the baffle test model corresponding to the test data to be evaluated can obtain the baffle test mode corresponding to the test data.
[0098] S14. Classify the test requirements to be evaluated to obtain the category corresponding to the test requirements to be evaluated, and select a target model from the multiple test data generation models according to the category.
[0099] In an optional embodiment, the step of classifying the test requirements to be evaluated to obtain the category corresponding to the test requirements to be evaluated, and selecting a target model from the multiple test data generation models according to the category includes:
[0100] Input the encoded data to be evaluated into the preset requirement classification model to obtain the category corresponding to the requirement to be evaluated;
[0101] Traverse the test data generation models in sequence, compare the category of the test data generation model with the category of the test requirements to be evaluated. If the category of the test data generation model is the same as the category of the test requirements to be evaluated, then use the test data generation model as the target model;
[0102] If the category of the test data generation model is different from the category of the test requirements to be evaluated, continue to traverse until the target model is obtained and then stop traversing.
[0103] In this optional embodiment, the category corresponding to the test requirements to be evaluated at least includes "credit investigation", "transaction", and "information query".
[0104] In this way, by selecting the corresponding test data generation model according to the category of the test requirements to be evaluated, a model basis is provided for subsequent multiple simulation tests. Since the test data is generated by the test data generation model, it is more inclined to the historical test data of this category, and its accuracy is higher compared to artificially constructing test data.
[0105] S15. Input the to-be-evaluated encoded data into the target model multiple times to generate multiple batches of test data, and transmit the multiple batches of test data to a preset data receiver according to the communication protocol for multiple simulation tests.
[0106] In an optional embodiment, the test data generation model is stored in a preset server. The steps of inputting the to-be-evaluated encoded data into the target model multiple times to generate multiple batches of test data and transmitting the multiple batches of test data to a preset data receiver according to the communication protocol for multiple simulation tests include:
[0107] a. Continuously query the resource occupancy rate of the preset server;
[0108] b. When the resource occupancy rate is less than a preset occupancy rate threshold, input the to-be-evaluated encoded data into the target model to obtain test data. When the resource occupancy rate is not less than the preset occupancy rate threshold, stop executing the target model;
[0109] c. Repeat steps a and b to obtain multiple batches of test data, and transmit the test data to a preset data receiver according to the communication protocol for multiple simulation tests until the number of repetitions is not less than a preset repetition threshold, then stop repeating.
[0110] In this optional embodiment, the resource occupancy rate of the preset server can be continuously queried. The resource occupancy rate of the preset server at least includes the CPU occupancy rate. The preset occupancy rate threshold can be 30%, 40%, 50%, etc., and this application does not make any limitations in this regard.
[0111] In this optional embodiment, when the resource occupancy rate is less than the preset occupancy rate threshold, it indicates that the resource occupancy rate of the preset server is relatively low. Then, the multiple test data generation operations have a relatively small impact on the load of the preset server. Therefore, the to-be-evaluated encoded data can be input into the target model to obtain test data.
[0112] The preset repetition threshold can be 3 times, 4 times, 5 times, etc., and this application does not make any limitations in this regard.
[0113] In this optional embodiment, the step of inputting the to-be-evaluated encoded data into the target model to obtain test data as described in step b includes:
[0114] When the category of the to-be-evaluated encoded data is not "credit investigation", the target model is the first test data generation model. Each to-be-evaluated encoded vector in the to-be-evaluated encoded data can be sequentially input into each neuron of the encoder of the target model, and the output of the encoder of the target model is the intermediate vector corresponding to the to-be-evaluated encoded data. Further, the intermediate vector can be input into the generator of the target model, and the output of each neuron in the generator of the target model is a test vector. All the test vectors output by the neurons in the generator are combined according to the arrangement order of the neurons in the target model to obtain test data.
[0115] When the category of the to-be-evaluated encoded data is "credit investigation", the target model is the second test data generation model. The to-be-evaluated encoded data can be sequentially compared with each key in the target model. When the to-be-evaluated encoded data is the same as the key, the value corresponding to the key is used as test data.
[0116] In this way, by generating test data when the server resource occupancy rate is relatively small, multiple batches of test data can be generated on the premise of maintaining server stability for multiple simulation tests, thereby improving the accuracy of the simulation tests.
[0117] The above artificial intelligence-based simulation test method encodes a large number of historical test requirements to obtain a large amount of encoded data, classifies the encoded data to obtain the category of each historical test requirement, and divides the encoded data and historical test data into multiple training data sets according to the category of the historical test requirement. Each training data set is used to train a targeted test data generation model, which can provide model guidance for the to-be-evaluated test task, thereby improving the accuracy of data testing.
[0118] As Figure 2 shown, it is a functional module diagram of a preferred embodiment of the artificial intelligence-based simulation test device provided by an embodiment of the present application. The artificial intelligence-based simulation test device 11 includes an encoding unit 110, a classification unit 111, a training unit 112, a query unit 113, a selection unit 114, and a test unit 115. The module / unit referred to in the present application means a series of computer program segments that can be executed by a processor 13 and can complete fixed functions, and are stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0119] In an alternative embodiment, the encoding unit 110 is configured to collect historical test requirements and historical test data from historical test records, and perform encoding processing on the historical test requirements to obtain encoded data corresponding to each historical test requirement.
[0120] In an optional embodiment, the encoding process for the historical test requirements to obtain the encoding data corresponding to each of the historical test requirements includes:
[0121] Performing word segmentation on the historical test requirements to obtain a plurality of words;
[0122] Encoding each of the words according to a preset text encoding algorithm to obtain an encoding vector corresponding to each of the words, and corresponding the encoding vector to the word one by one to serve as a word corpus;
[0123] Taking the encoding vectors of all the words corresponding to each of the historical test requirements as the encoding data corresponding to the historical test requirements.
[0124] In this optional embodiment, the historical test requirements and the historical test data are in one-to-one correspondence. Exemplarily, when the historical test requirement is "scraping user credit investigation data", the historical test data is user credit investigation data; when the historical test requirement is "testing the payment interface of an e-commerce platform", the test data is payment data.
[0125] In this optional embodiment, the preset text encoding algorithm can be an existing text encoding algorithm such as the GloVe algorithm (Global Vector algorithm), Skip-Gram algorithm, CBOW algorithm (Continuous Bag Of Words algorithm), etc., and this application does not make any limitations in this regard.
[0126] In this optional embodiment, the encoding vectors corresponding to each of the words can be combined according to the arrangement order of all the words in the historical test requirements to serve as the encoding data corresponding to the historical test requirements.
[0127] In an optional embodiment, the classification unit 111 is used to classify each of the historical test requirements to obtain the category of each of the historical test requirements, and divide the encoding data and the historical test data into a plurality of training data sets according to the category, and the category is in one-to-one correspondence with the training data set.
[0128] In an optional embodiment, the process of classifying each of the historical test requirements to obtain the category of each of the historical test requirements, and dividing the encoding data and the historical test data into a plurality of training data sets includes:
[0129] Inputting the encoding data into a preset requirement classification model to obtain the category corresponding to each of the historical test requirements, and the category at least includes "credit investigation", "transaction", and "information query";
[0130] Use the historical test data corresponding to the historical test requirements as labeled data;
[0131] Use the encoded data corresponding to the historical test requirements as sample data, and pair the sample data with the labeled data one by one as training data;
[0132] Group the training data corresponding to the historical test requirements with the same category into the same training data set to obtain multiple training data sets, and the training data sets correspond to the categories one by one.
[0133] In this alternative embodiment, the preset requirement classification model can be an existing classification model such as an XGBoost model (Extreme Gradient Boost), a LightGBM model (Light Gradient Boost Machine), a GBDT (Gradient Boost Decision Tree), a random forest model, etc. This application does not limit this.
[0134] The categories at least include "credit investigation", "transaction", and "information query". When the category of the historical test requirement is "credit investigation", the historical test requirement needs to receive "credit investigation data" to test the system; when the category of the historical test requirement is "transaction", the historical test requirement needs to receive "transaction data" to test the system; when the category of the historical test requirement is "information query", the historical test requirement needs to receive "user information" to test the system.
[0135] In an alternative embodiment, the training unit 112 is used to train a test data generation model corresponding to each training data set respectively according to each training data set.
[0136] In an alternative embodiment, the training of the test data generation model corresponding to each training data set respectively according to each training data set includes:
[0137] Construct an initial generation model, where the initial generation model includes an encoder and a generator;
[0138] For each training data set, if the category corresponding to the training data set is not "credit investigation", then use the training data set to train the initial generation model, calculate the loss value of the initial generation model according to a preset loss function, and continuously update the parameters in the initial generation model until the loss value no longer changes, and obtain the first test data generation model corresponding to the training data set whose category is not "credit investigation";
[0139] If the category corresponding to the training data set is "credit investigation", the sample data in the training data set is used as the key, and the label data is used as the value to construct key-value pairs. All the key-value pairs are used to generate a model as the second test data, and all the first test data generation models and the second test data generation model are unified as the test data generation model.
[0140] In this optional embodiment, the initial generation model includes an encoder and a generator. Both the encoder and the generator can be existing neural network structures such as an LSTM model (Long Short Term Memory), an RNN model (Recurrent Neural Network), or a GRU model (Gate Recurrent Unit). This application does not limit this. Both the encoder and the generator contain multiple neurons, and both the encoder and the generator are formed by connecting the multiple neurons in series. Taking the RNN model as an example, as Figure 4 shown in the structural schematic diagram of the initial generation model.
[0141] In this optional embodiment, taking any training data set whose category is not "credit investigation" as an example, the input of the encoder is the sample data in the training data set, and the output of the encoder is the feature vector corresponding to the sample data; the input of the generator is the feature vector, and the output of the generator is the virtual test data corresponding to the sample data.
[0142] In this optional embodiment, the virtual test data and the label data corresponding to the sample data can be input into a preset loss function to calculate the loss value of the initial generation model. The preset loss function can be an existing loss function such as a root mean square error function, a least squares function, or an Euclidean distance function. This application does not limit this.
[0143] In this optional embodiment, the parameters of the initial generation model can be continuously updated according to the gradient descent method until the loss value of the initial generation model no longer changes. Then, the update of the parameters of the initial generation model is stopped, and a first test data generation model corresponding to the training data set is obtained.
[0144] In this optional embodiment, if the category corresponding to the training data set is "credit investigation", it indicates that the test data required for the test requirements corresponding to this training data set is a pre-organized credit report. Then, the sample data in the training data set can be used as the key, and the label data in the training data set can be used as the value to construct key-value pairs. The key-value pairs are used as the second test data generation model.
[0145] In an optional embodiment, the query unit 113 is configured to query the to-be-evaluated coding data and communication protocol corresponding to the to-be-evaluated test requirement.
[0146] In an optional embodiment, the querying the to-be-evaluated coding data and communication protocol corresponding to the to-be-evaluated test requirement includes:
[0147] Performing word segmentation on the to-be-evaluated test requirement to obtain a plurality of to-be-evaluated words;
[0148] Querying, from the vocabulary corpus, the coding data corresponding to each of the to-be-evaluated words as the to-be-evaluated coding data;
[0149] Querying the communication protocol corresponding to the to-be-evaluated test requirement, where the communication protocol is used to characterize the protocol for receiving test data of the to-be-evaluated test requirement.
[0150] In this optional embodiment, the to-be-evaluated test requirement can be segmented according to the preset word segmentation tool to obtain a plurality of to-be-evaluated words, and the coding vectors corresponding to each of the to-be-evaluated words are queried from the vocabulary corpus, and the coding vectors corresponding to the to-be-evaluated words are combined according to the arrangement order of the to-be-evaluated words in the to-be-evaluated test requirement as the to-be-evaluated coding data.
[0151] In this optional embodiment, the communication protocol refers to the data transmission rules preset in the to-be-evaluated requirement, and the communication protocol can be existing communication protocols such as TCP / IP (Transmission Control Protocol / Internet Protocol), SPP protocol (Sequenced Packet Protocol), NetBEUI (NetBios Enhanced User Interface), etc., and the present application does not limit this.
[0152] In an optional embodiment, the selection unit 114 is configured to classify the to-be-evaluated test requirement to obtain the category corresponding to the to-be-evaluated test requirement, and select a target model from the multiple test data generation models according to the category.
[0153] In an optional embodiment, the classifying the to-be-evaluated test requirement to obtain the category corresponding to the to-be-evaluated test requirement, and selecting a target model from the multiple test data generation models according to the category includes:
[0154] Inputting the to-be-evaluated coding data into the preset requirement classification model to obtain the category corresponding to the to-be-evaluated requirement;
[0155] Traverse the test data generation models in sequence, compare the category of the test data generation model with the category of the test requirements to be evaluated. If the category of the test data generation model is the same as the category of the test requirements to be evaluated, then use the test data generation model as the target model;
[0156] If the category of the test data generation model is different from the category of the test requirements to be evaluated, continue to traverse until the target model is obtained and then stop traversing.
[0157] In this optional embodiment, the categories corresponding to the test requirements to be evaluated at least include "credit investigation", "transaction", and "information query".
[0158] In an optional embodiment, the test unit 115 is configured to input the test encoding data to be evaluated into the target model multiple times to generate multiple batches of test data, and transmit the multiple batches of test data to a preset data receiver according to the communication protocol for multiple simulation tests.
[0159] In an optional embodiment, the test data generation model is stored in a preset server. The process of inputting the test encoding data to be evaluated into the target model multiple times to generate multiple batches of test data, and transmitting the multiple batches of test data to a preset data receiver according to the communication protocol for multiple simulation tests includes:
[0160] a. Continuously query the resource occupancy rate of the preset server;
[0161] b. When the resource occupancy rate is less than a preset occupancy rate threshold, input the test encoding data to be evaluated into the target model to obtain test data. When the resource occupancy rate is not less than the preset occupancy rate threshold, stop executing the target model;
[0162] c. Repeat steps a and b to obtain multiple batches of test data, and transmit the test data to a preset data receiver according to the communication protocol for multiple simulation tests until the number of repetitions is not less than a preset repetition threshold and then stop repeating.
[0163] In this optional embodiment, the resource occupancy rate of the preset server can be continuously queried. The resource occupancy rate of the preset server at least includes the CPU occupancy rate. The preset occupancy rate threshold can be 30%, 40%, 50%, etc., and the present application does not make any limitations in this regard.
[0164] In this alternative embodiment, when the resource occupancy rate is less than a preset occupancy rate threshold, it indicates that the resource occupancy rate of the preset server is relatively low. Therefore, the load impact of multiple test data generation operations on the preset server is relatively small. Thus, the encoding data to be evaluated can be input into the target model to obtain test data.
[0165] The preset repetition threshold can be 3 times, 4 times, 5 times, etc., and this application does not limit it.
[0166] In this alternative embodiment, as described in step b, inputting the encoding data to be evaluated into the target model to obtain test data includes:
[0167] When the category of the encoding data to be evaluated is not "credit investigation", the target model is the first test data generation model. Each encoding vector to be evaluated in the encoding data to be evaluated can be sequentially input into each neuron of the encoder of the target model. The output of the encoder of the target model is the intermediate vector corresponding to the encoding data to be evaluated. Further, the intermediate vector can be input into the generator of the target model. The output of each neuron in the generator of the target model is a test vector. All the test vectors output by the neurons in the generator are combined according to the arrangement order of the neurons in the target model to obtain test data.
[0168] When the category of the encoding data to be evaluated is "credit investigation", the target model is the second test data generation model. The encoding data to be evaluated can be sequentially compared with each key in the target model. When the encoding data to be evaluated is the same as the key, the value corresponding to the key is used as test data.
[0169] The above artificial intelligence-based simulation test method encodes a large number of historical test requirements to obtain a large amount of encoding data, classifies the encoding data to obtain the category of each historical test requirement, and divides the encoding data and historical test data into multiple training data sets according to the category of the historical test requirement. Each training data set is used to train a targeted test data generation model, which can provide a model guidance for the test task to be evaluated, thereby improving the accuracy of data testing.
[0170] As Figure 3 shown, it is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 1 includes a memory 12 and a processor 13. The memory 12 is used to store computer-readable instructions, and the processor 13 is used to execute the computer-readable instructions stored in the memory to implement the artificial intelligence-based simulation test method of any of the above embodiments.
[0171] In an alternative embodiment, the electronic device 1 further includes a bus, and a computer program stored in the memory 12 and executable on the processor 13, such as an artificial intelligence-based simulation test program.
[0172] Figure 3 Only the electronic device 1 with components 12-13 is shown. Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have a different component arrangement.
[0173] In combination with Figure 1 , the memory 12 in the electronic device 1 stores a plurality of computer-readable instructions to implement an artificial intelligence-based simulation test method, and the processor 13 can execute the plurality of instructions to implement:
[0174] Collect historical test requirements and historical test data from historical test records, and perform encoding processing on the historical test requirements to obtain encoded data corresponding to each historical test requirement;
[0175] Classify each historical test requirement to obtain the category of each historical test requirement, and divide the encoded data and the historical test data into multiple training data sets according to the category, with the category corresponding to the training data set one by one;
[0176] Train a test data generation model corresponding to each training data set respectively according to each training data set;
[0177] Query the to-be-evaluated encoded data and communication protocol corresponding to the to-be-evaluated test requirement;
[0178] Classify the to-be-evaluated test requirement to obtain the category corresponding to the to-be-evaluated test requirement, and select a target model from the multiple test data generation models according to the category;
[0179] Input the to-be-evaluated encoded data into the target model multiple times to generate multiple batches of test data, and transmit the multiple batches of test data to a preset data recipient according to the communication protocol for multiple simulation tests.
[0180] Specifically, for the specific implementation method of the above instructions by the processor 13, reference can be made to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.
[0181] Among them, the memory 12 includes at least one type of readable storage medium, which can be non-volatile or volatile. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the memory 12 can be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 12 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the electronic device 1. The memory 12 can be used not only to store application software installed in the electronic device 1 and various types of data, such as the code of an artificial intelligence-based simulation test program, etc., but also to temporarily store data that has been output or will be output.
[0182] In some embodiments, the processor 13 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions, including the combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the electronic device 1, connecting various components of the entire electronic device 1 through various interfaces and lines. By running or executing programs or modules stored in the memory 12 (such as executing an artificial intelligence-based simulation test program, etc.), and calling data stored in the memory 12, it performs various functions of the electronic device 1 and processes data.
[0183] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned embodiments of various artificial intelligence-based simulation test methods, such as Figure 1 the steps shown.
[0184] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into an encoding unit 110, a classification unit 111, a training unit 112, a query unit 113, a selection unit 114, and a testing unit 115.
[0185] The integrated units implemented in the form of software function modules as described above may be stored in a computer-readable storage medium. The above-mentioned software function modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute a part of the artificial intelligence-based simulation testing method described in various embodiments of the present application.
[0186] If the modules / units integrated in the electronic device 1 are implemented in the form of software function units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented.
[0187] Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory, and other memories, etc.
[0188] Furthermore, the computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the blockchain node, etc.
[0189] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.
[0190] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, in Figure 3 only one arrow is used to represent it, but it does not mean that there is only one bus or one type of bus. The bus is set to achieve connection and communication between the memory 12 and at least one processor 13, etc.
[0191] The embodiments of this application also provide a computer-readable storage medium (not shown in the figure). Computer-readable instructions are stored in the computer-readable storage medium, and the computer-readable instructions are executed by a processor in an electronic device to implement the artificial intelligence-based simulation test method described in any of the above embodiments.
[0192] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.
[0193] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0194] In addition, the functional modules in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0195] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular form does not exclude the plural form. A plurality of units or devices described in the specification can also be implemented by one unit or device through software or hardware. Terms such as first, second, etc. are used to denote names and do not denote any particular order.
[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. An artificial intelligence-based simulation test method, characterized in that, the method includes: Collect historical test requirements and historical test data from historical test records, and perform encoding processing on the historical test requirements to obtain encoding data corresponding to each historical test requirement, including: performing word segmentation on the historical test requirements to obtain multiple words; encoding each word according to a preset text encoding algorithm to obtain an encoding vector corresponding to each word, and corresponding the encoding vector with the word one by one as a word corpus; using the encoding vectors of all words corresponding to each historical test requirement as the encoding data corresponding to the historical test requirement; Classify each historical test requirement to obtain the category of each historical test requirement, and divide the encoding data and the historical test data into multiple training data sets according to the category, with the category corresponding to the training data set one by one; Train a test data generation model corresponding to each training data set according to each training data set respectively; Query the to-be-evaluated encoding data and communication protocol corresponding to the to-be-evaluated test requirement, including: performing word segmentation on the to-be-evaluated test requirement to obtain multiple to-be-evaluated words; querying the encoding data corresponding to each to-be-evaluated word from the word corpus as the to-be-evaluated encoding data; querying the communication protocol corresponding to the to-be-evaluated test requirement, where the communication protocol is used to characterize the protocol for the to-be-evaluated test requirement to receive test data; Classify the to-be-evaluated test requirement to obtain the category corresponding to the to-be-evaluated test requirement, and select a target model from multiple test data generation models according to the category; Input the to-be-evaluated encoding data into the target model multiple times to generate multiple batches of test data, and transmit the multiple batches of test data to a preset data receiver according to the communication protocol for multiple simulation tests.
2. The artificial intelligence-based simulation test method according to claim 1, characterized in that, the step of classifying each historical test requirement to obtain the category of each historical test requirement, and dividing the encoding data and the historical test data into multiple training data sets according to the category includes: Inputting the encoding data into a preset requirement classification model to obtain the category corresponding to each historical test requirement, where the category at least includes "credit investigation", "transaction", and "information query"; Using the historical test data corresponding to the historical test requirement as label data; Using the encoding data corresponding to the historical test requirement as sample data, and corresponding the sample data with the label data one by one as training data; Grouping the training data corresponding to historical test requirements with the same category into the same training data set to obtain multiple training data sets, with the training data set corresponding to the category one by one.
3. The artificial intelligence-based simulation test method according to claim 2, characterized in that, the step of training a test data generation model corresponding to each training data set according to each training data set respectively includes: Construct an initial generation model, where the initial generation model includes an encoder and a generator; For each of the training data sets, if the category corresponding to the training data set is not "credit investigation", then use the training data set to train the initial generation model, calculate the loss value of the initial generation model according to a preset loss function, and continuously update the parameters in the initial generation model until the loss value no longer changes, and obtain the first test data generation model corresponding to the training data set whose category is not "credit investigation"; If the category corresponding to the training data set is "credit investigation", then use the sample data in the training data set as keys and the label data as values to construct key-value pairs, use all the key-value pairs as the second test data generation model, and use all the first test data generation models and the second test data generation model as the test data generation model uniformly.
4. The artificial intelligence-based simulation test method according to claim 1, wherein, The step of classifying the to-be-evaluated test requirement to obtain the category corresponding to the to-be-evaluated test requirement, and selecting a target model from multiple test data generation models according to the category, includes: Input the to-be-evaluated encoded data into the preset requirement classification model to obtain the category corresponding to the to-be-evaluated test requirement; Traverse the test data generation models in sequence, compare the category of the test data generation model with the category of the to-be-evaluated test requirement. If the category of the test data generation model is the same as the category of the to-be-evaluated test requirement, then use the test data generation model as the target model; If the category of the test data generation model is different from the category of the to-be-evaluated test requirement, continue to traverse until the target model is obtained and then stop traversing.
5. The artificial intelligence-based simulation test method according to claim 1, wherein, The test data generation model is stored in a server. The step of inputting the to-be-evaluated encoded data into the target model multiple times to generate multiple batches of test data, and transmitting the multiple batches of test data to a preset data receiver according to the communication protocol for multiple simulation tests, includes: a. Continuously query the resource occupancy rate of the server; b. When the resource occupancy rate is less than a preset occupancy rate threshold, input the to-be-evaluated encoded data into the target model to obtain test data. When the resource occupancy rate is not less than the preset occupancy rate threshold, stop executing the target model; c. Repeat steps a and b to obtain multiple batches of test data, and transmit the test data to a preset data receiver according to the communication protocol for multiple simulation tests until the number of repetitions is not less than a preset repetition threshold and then stop repeating.
6. An artificial intelligence-based simulation test device, the device includes units for implementing the method according to any one of claims 1 to 5, wherein, The device includes: Coding unit, configured to collect historical test requirements and historical test data from historical test records, and perform coding processing on the historical test requirements to obtain coding data corresponding to each historical test requirement; Classification unit, configured to classify each historical test requirement to obtain the category of each historical test requirement, and divide the coding data and the historical test data into multiple training data sets according to the category, where the category corresponds to the training data set one by one; Training unit, configured to train a test data generation model corresponding to each training data set respectively according to each training data set; Query unit, configured to query the to-be-evaluated coding data and communication protocol corresponding to the to-be-evaluated test requirement; Selection unit, configured to classify the to-be-evaluated test requirement to obtain the category corresponding to the to-be-evaluated test requirement, and select a target model from multiple test data generation models according to the category; Testing unit, configured to input the to-be-evaluated coding data into the target model multiple times to generate multiple batches of test data, and transmit the multiple batches of test data to a preset data receiver according to the communication protocol for multiple simulation tests.
7. An electronic device, characterized in that, the electronic device includes: a memory, storing computer-readable instructions; and a processor, executing the computer-readable instructions stored in the memory to implement the artificial intelligence-based simulation test method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: computer-readable instructions are stored in the computer-readable storage medium, and the computer-readable instructions are executed by a processor in an electronic device to implement the artificial intelligence-based simulation test method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Performance test method and related equipment
CN110618922A
Method and device for test data generation based on data decision, and computer equipment
CN111176990A