Artificial Intelligence-Based Data System Testing Method and Related Devices

By conducting separate tests on different modules of the data system, calculating the weight of module indicators and building comprehensive indicators, the problem of inaccurate test results in the existing test methods is solved, and a more accurate and comprehensive data system performance evaluation is achieved.

CN114860588BActive Publication Date: 2025-06-20CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210451130.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-06-20
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

The existing data system testing methods lack consideration of the importance of the performance of different modules, resulting in inaccurate test results.

Method used

Using artificial intelligence-based data system testing method, the weights of different functional module indicators are calculated by individually testing different modules of the data system, and comprehensive data system indicators are constructed based on these weights to evaluate the performance of the data system.

Benefits of technology

It improves the accuracy of data system test results, takes into account the correlation between test results of different modules, and provides a more comprehensive performance evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114860588B_ABST
    Figure CN114860588B_ABST
Patent Text Reader

Abstract

The present application proposes an artificial intelligence-based data system testing method, apparatus, electronic device, and storage medium. The artificial intelligence-based data system testing method includes: extracting test data from an open-source database according to a preset program; testing the data system based on the test data to obtain test metrics, where the test metrics include a processing accuracy metric and a screening accuracy metric; calculating the eigenvalues of the processing accuracy metric and the screening accuracy metric to obtain the importance of the metrics; inputting the importance of the metrics and the test metrics into a custom integration model to obtain an integration result and using it as an updated metric; and evaluating the data system based on the updated metric to obtain a test result. This method takes into account the correlation between the test results of different programs in the data system and can improve the accuracy of the test results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, electronic device and storage medium for testing a data system based on artificial intelligence. Background Art

[0002] With the development of digital business, the importance of data systems in the enterprise operation process has become increasingly prominent. The performance of a data system is a relatively key dimension for evaluating a data system, and the performance of the data system mainly includes the accuracy of data processing and the accuracy of data screening, etc.

[0003] Existing data system testing methods are usually module-specific targeted testing, lacking consideration of the importance of the performance of different modules, resulting in inaccurate test results. Summary of the Invention

[0004] In view of the above, it is necessary to provide a method and related devices for testing a data system based on artificial intelligence to solve the technical problem of how to improve the accuracy of the test results of a data system, wherein the related devices include a data system testing device, an electronic device and a storage medium based on artificial intelligence.

[0005] An embodiment of this application provides a method for testing a data system based on artificial intelligence, and the method includes:

[0006] Extract test data from an open-source database according to a preset program, and the test data is formatted data;

[0007] Test the data system based on the test data to obtain test metrics, where the test metrics include a processing accuracy metric and a screening accuracy metric. The processing accuracy metric is used to characterize the accuracy of the data system in data processing, and the screening accuracy metric is used to characterize the accuracy of the data system in data screening;

[0008] Calculate the characteristic values of the processing accuracy metric and the screening accuracy metric to obtain the metric importance, and the metric importance is used to characterize the importance degree of the screening accuracy and the processing accuracy;

[0009] Input the metric importance and the test metrics into a custom integration model to obtain an integration result and use it as an updated metric, and the updated metric is used to characterize the accuracy of the data system in performing data screening and data processing link test tasks;

[0010] Evaluate the data system based on the updated metric to obtain a test result.

[0011] In the above data system testing method, different modules of the data system are separately tested to obtain a set of functional module metrics, the weights of different functional module metrics are calculated based on the set of functional module metrics, and finally a comprehensive metric of the data system is constructed based on the weights to evaluate the performance of the data system. The relevance between the test results of different modules is considered, thereby improving the accuracy of the test results.

[0012] In some embodiments, the data system is tested based on the test data to obtain test metrics, and the test metrics include a processing accuracy metric and a screening accuracy metric, including:

[0013] The test data is input into the data system to obtain processing result data and screening result data;

[0014] The processing result data and the screening result data are respectively calculated to obtain the processing accuracy metric and the screening accuracy metric.

[0015] In this way, the test data is input into the data system to obtain result data, and the test metrics are further constructed based on the result data. The test metrics can characterize the performance of the data system, can provide data support for the subsequent evaluation of the data system, and thus improve the accuracy of the subsequent evaluation results.

[0016] In some embodiments, the step of respectively calculating the processing result data and the screening result data to obtain the processing accuracy metric and the screening accuracy metric includes:

[0017] The distance data of the processing result data is calculated according to a preset distance measurement method, and the distance data is used to characterize the similarity degree of the processing result data;

[0018] The variance of the distance data is calculated multiple times to obtain the processing accuracy metric;

[0019] The variance of the screening result data is calculated multiple times to obtain a second variance set;

[0020] The aggregation degree of the second variance set is calculated according to a custom aggregation degree measurement method to obtain the screening accuracy metric. The higher the value of the screening accuracy metric, the more complete the data screening function of the data system.

[0021] In this way, the processing accuracy metric is obtained by measuring the distance between the processing result data, and the screening accuracy metric is obtained by calculating the variance of the screening result data. The processing accuracy metric and the screening accuracy metric can characterize the aggregation degree of the test result data, thereby characterizing the performance of the data system, and further improving the accuracy of the subsequent evaluation results.

[0022] In some embodiments, calculating the characteristic values of the processing accuracy index and the screening accuracy index to obtain the index importance includes:

[0023] Calculating the variance and covariance of the processing accuracy index and the screening accuracy index to construct a covariance matrix;

[0024] Calculating the eigenvalues of the covariance matrix as the index importance, and the eigenvalues include processing eigenvalues and screening eigenvalues.

[0025] In this way, the index importance is obtained by constructing the covariance matrix of the processing accuracy index and the screening accuracy index, providing data support for subsequent establishment of a custom integration model, thereby improving the accuracy of the test index.

[0026] In some embodiments, before inputting the index importance and the test index into a custom integration model to obtain an integration result and using it as an updated index, the method further includes:

[0027] Calculating the index importance according to a preset standardization method to obtain normalized index importance, and the normalized index importance includes normalized processing eigenvalues and normalized screening eigenvalues.

[0028] In this way, normalizing the index importance based on the custom standardization method can improve the confidence level of the index importance, and further improve the accuracy of the updated index.

[0029] In some embodiments, the custom integration model satisfies the relational expression:

[0030] T = α1·Mean1 + α2·Mean2

[0031] Wherein, T represents the updated index, and the updated index is used to characterize the accuracy of the data system when performing data screening and data processing link test tasks; β1 represents the normalized processing eigenvalue, used to characterize the importance of processing accuracy; β2 represents the normalized screening eigenvalue, used to characterize the importance of screening accuracy; Mean1 represents the average value of the processing accuracy index, and Mean2 represents the average value of the screening accuracy index.

[0032] In this way, by establishing an integration model, the test index can be integrated according to the normalized index importance to obtain an updated index, which is convenient for subsequent evaluation of the data system.

[0033] In some embodiments, evaluating the data system based on the updated index to obtain a test result includes:

[0034] Evaluate the data system based on the update metrics to obtain an evaluation result;

[0035] Compare the evaluation result with a preset threshold to obtain a test result.

[0036] In this way, a test result is obtained based on the preset threshold and the update metrics of the data system. The preset threshold can be set by developers based on enterprise requirements, so the way to obtain the test result is relatively flexible.

[0037] An embodiment of the present application further provides a data system testing device based on artificial intelligence. The device includes:

[0038] An acquisition unit, configured to extract test data from an open-source database according to a preset program, and the test data is formatted data;

[0039] A test unit, configured to test the data system based on the test data to obtain test metrics. The test metrics include a processing accuracy metric and a screening accuracy metric. The processing accuracy metric is used to characterize the accuracy of data processing by the data system, and the screening accuracy metric is used to characterize the accuracy of data screening by the data system;

[0040] A calculation unit, configured to calculate the eigenvalues of the processing accuracy metric and the screening accuracy metric to obtain metric importance, and the metric importance is used to characterize the importance degree of the screening accuracy and the processing accuracy;

[0041] An integration unit, configured to input the metric importance and the test metrics into a custom integration model to obtain an integration result and use it as an update metric. The update metric is used to characterize the accuracy of the data system when performing data screening and data processing link testing tasks;

[0042] An evaluation unit, configured to evaluate the data system based on the update metric to obtain a test result.

[0043] An embodiment of the present application further provides an electronic device. The electronic device includes:

[0044] A memory, storing computer-readable instructions; and

[0045] A processor, configured to execute the computer-readable instructions stored in the memory to implement the data system testing method based on artificial intelligence.

[0046] An embodiment of the present application further provides a computer-readable storage medium. Computer-readable instructions are stored in the computer-readable storage medium, and the computer-readable instructions are executed by a processor in an electronic device to implement the data system testing method based on artificial intelligence. Description of the Drawings

[0047] Figure 1 It is a flowchart of a preferred embodiment of the method for testing an artificial intelligence-based data system involved in the present application.

[0048] Figure 2 It is a flowchart of a preferred embodiment of testing the data system based on the test data to obtain test metrics involved in the present application.

[0049] Figure 3 It is a flowchart of a preferred embodiment of calculating the correlation between the processing accuracy metric and the screening accuracy metric to obtain the metric importance involved in the present application.

[0050] Figure 4 It is a functional block diagram of a preferred embodiment of an artificial intelligence-based data system testing apparatus involved in the present application.

[0051] Figure 5 It is a schematic structural diagram of an electronic device of a preferred embodiment of the method for testing an artificial intelligence-based data system involved in the present application.

[0052] Figure 6 It is a schematic diagram of the test data involved in the present application. Detailed implementation manners

[0053] In order to more clearly understand the purpose, features, and advantages of the present application, the present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other. Many specific details are set forth in the following description in order to fully understand the present application. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0054] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality" means two or more, unless otherwise specifically defined.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used in the description of the present application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0056] An embodiment of the present application provides an artificial intelligence-based data system testing method, which can be applied to one or more electronic devices. The electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0057] The electronic device can be any electronic product that can perform human-computer interaction with users. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.

[0058] The electronic device may further include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.

[0059] The network where the electronic device is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.

[0060] As Figure 1 shown, it is a flowchart of a preferred embodiment of the artificial intelligence-based data system testing method of the present application. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted.

[0061] S10, extract test data from an open-source database according to a preset program, and the test data is formatted data.

[0062] In this optional embodiment, the open-source database is a set of integrated and shareable data organized and stored in a certain structure. Exemplarily, the preset database can be a MySQL database, and the MySQL database is a relational database.

[0063] In this alternative embodiment, the purpose of obtaining the test data is to test a data system, which is a software system that provides data for an actually operable storage, maintenance, and application system, and is an aggregate of a storage medium, processing objects, and a management system.

[0064] In this alternative embodiment, the data system can receive any formatted data to output data processing results. Therefore, the test data can be any formatted data, and the data in any open-source database can be used as test data to test the data system. This solution does not make any limitations in this regard. Figure 6 It is a schematic diagram of the test data.

[0065] In this alternative embodiment, the preset program can be run on a computer to obtain test data. Exemplarily, the form of the preset program can be "select 'table' from 'database'", where 'table' represents the test data and 'database' represents the name of the database where the test data is located. The test data represented by 'table' is in the form of a table, where each row is a piece of data and each column is a field of the data. Exemplarily, 'table' can have n rows and m columns, where n can be 10,000 and m can be 5. Therefore, the set of field names included in 'table' can be {field 1, field 2, field 3, field 4, field 5}.

[0066] In this alternative embodiment, the data table represented by 'table' is used as the test data.

[0067] In this way, a test data set is obtained based on a preset database, providing data support for the subsequent test link of the data system and making the test results more accurate.

[0068] S11. Test the data system based on the test data to obtain test metrics. The test metrics include a processing accuracy metric and a screening accuracy metric. The processing accuracy metric is used to characterize the accuracy of the data system in data processing, and the screening accuracy metric is used to characterize the accuracy of the data system in data screening.

[0069] Please refer to Figure 2 , in an alternative embodiment, the testing of the data system based on the test data to obtain test metrics, where the test metrics include a processing accuracy metric and a screening accuracy metric, and the processing accuracy metric is used to characterize the accuracy of the data system in data processing, and the screening accuracy metric is used to characterize the accuracy of the data system in data screening includes:

[0070] S111. Input the test data into the data system to obtain processed result data and screened result data.

[0071] In this alternative embodiment, the data system includes a data processing program and a data screening program. The test data can be input into the data processing program to obtain processed result data, and the test data can be input into the data screening program to obtain screening result data.

[0072] In this alternative embodiment, the data processing program is a collection of multiple data operation programs. Each data operation program can be a piece of SQL script. The data operation programs can include a product program, a quotient program, a sum program, and a difference program. The form of the product program can be "select a * b as 'product' from table", and its function is to calculate the product of two fields of each piece of data in the test data; the form of the quotient program can be "select c / d as 'quotient' from table", and its function is to calculate the quotient of two fields of each piece of data in the test data; the form of the sum program can be "select e + f as'sum' from table", and its function is to calculate the sum of two fields of each piece of data in the test data; the form of the difference program can be "select g - h as 'difference' from table", and its function is to calculate the difference of two fields of each piece of data in the test data.

[0073] In this alternative embodiment, the specific implementation steps of inputting the test data into the data processing program to obtain processed result data are as follows:

[0074] Input the test data into the product program to obtain the product P of the test data. The P contains 10,000 products. In the product program, a can represent field 1 in the table, and b can represent field 2 in the table;

[0075] Input the test data into the quotient program to obtain the quotient value Q of the test data. The Q contains 10,000 quotients. In the quotient program, c can represent field 2 in the table, and d can represent field 3 in the table;

[0076] Input the test data into the sum program to obtain the sum value S of the test data. The S contains 10,000 sum values. In the sum program, e can represent field 3 in the table, and f can represent field 4 in the table;

[0077] Input the test data into the difference program to obtain the difference value D of the test data. The D contains 10,000 differences. In the difference program, g can represent field 4 in the table, and h can represent field 5 in the table;

[0078] Arranging the P, Q, S, and D column by column can obtain a matrix z with 10,000 rows and 4 columns.

[0079] In this alternative embodiment, the step of obtaining the processed result data can be repeated n times to obtain n sets of test results. Exemplarily, n can be 20. Each time the step of obtaining the processed result data is implemented, a corresponding matrix can be obtained. Therefore, 20 matrices of 10,000 * 4 can be obtained. Denote the set of the 20 matrices as Z, which can be expressed as Z = {z 1 , z 2 , …, z 20}.

[0080] In this alternative embodiment, the data screening program can be a section of SQL script, whose function is to screen out the data that meets the preset screening conditions. Its form can be "select data from table where 'condition'", where data represents the field where the return value of the screening program is located. Exemplarily, data = "field 4", table represents the test data, and condition represents the preset screening conditions.

[0081] In this alternative embodiment, the preset screening condition can be "ROWNUM = 3", which means to screen out the data in the 3rd row and 4th column of the test data. The specific implementation steps of inputting the test data into the data screening program to obtain the screened result data are to run the data screening program to obtain the return value that meets the preset screening conditions, and the return value is the data in the 3rd row and 4th column of the test data.

[0082] In this alternative embodiment, the method of obtaining the screened result data can be repeated m times to obtain the screened result data. Exemplarily, m can be 20. Therefore, the screened result data can be a data set {data1, data2, …, data 20}, where the subscript of each data represents the order of the screened result.

[0083] In this alternative embodiment, the matrix set Z = {z 1 , z 2 , …, z 20} can be used as the processed result data, and the data set {data1, data2, …, data 20} can be used as the screened result data.

[0084] S112. Calculate the processed result data and the screened result data respectively to obtain the processing accuracy index and the screening accuracy index.

[0085] In this optional embodiment, the distance data of the processing result data can be calculated according to a preset distance measurement method, and the distance data is used to characterize the similarity degree of the processing result data; then the variance of the example data is calculated to obtain a processing accuracy index. Taking z 1 and z 2 in the matrix set Z as an example, the implementation steps of the preset distance measurement method are as follows:

[0086] Calculate the cosine distance between the first column data of z 1 and the first column data of z 2 based on the cosine distance formula and denote it as Sd 11,21 . The cosine distance formula is:

[0087]

[0088] where A represents the first column vector of the z 1 ; B represents the first column vector of the z 2 ; Sd 11,21 represents the cosine distance between A and B; A i represents the value of the i-th dimension in vector A; B i represents the value of the i-th dimension in vector B; k represents the dimension of vector A and vector B;

[0089] Calculate the corresponding cosine distances of the remaining three column data in the z 1 and z 2 respectively based on the above cosine distance formula and denote them as Sd 12,22 , Sd 13,23 and Sd 14,24 ;

[0090] Calculate the mean value of the Sd 11,21 , Sd 12,22 , Sd 13,23 and Sd 14,24 and denote it as Sd. The mean value represented by Sd is used to characterize the similarity degree between the matrix z 1 and z 2 . The larger the value of Sd, the lower the similarity degree between the matrix z 1 and z 2 .

[0091] In this optional embodiment, since the matrix set Z contains 20 matrices, there are 190 distance data between pairwise matrices, denoted as Sd u ={Sd1, Sd2, …, Sd 190}, and the Sd u is the distance data of the processing result data.

[0092] In this alternative embodiment, the variance of the distance data can be calculated to obtain a machining accuracy index. The calculation formula for the variance of the distance data is as follows:

[0093]

[0094] where s1 represents the variance of the 190 distance data, and Sd i represents the i-th distance data in the Sd u The higher the value of s1, the higher the similarity between the machining result data, and thus the higher the machining accuracy.

[0095] In this alternative embodiment, the method of obtaining the variance s1 can be repeated 20 times to obtain a first variance set s1_z = {s11, s12,..., s1 20}, with the aim of testing the data processing program multiple times to simulate the working conditions of the data system in a high-load environment and testing the stability of the data processing program in the data system. The first variance set s1_z can be used as the machining accuracy index.

[0096] In this alternative embodiment, a second variance set of the screening result data can be calculated, and then the aggregation degree of the data in the second variance set can be calculated according to a custom aggregation degree measurement method to obtain a screening accuracy index. The calculation method for the variance of the screening result data is as follows:

[0097]

[0098] where s2 represents the variance of the screening result data. The larger the value of s2, the lower the aggregation degree of the screening data, the greater the difference between each screening result data, and the lower the screening accuracy; data i represents the i-th data in the screening result data.

[0099] In this alternative embodiment, the method of obtaining s2 can be repeated 20 times to obtain a second variance set s2_z = {s21, s22,..., s2 20}, with the aim of testing the screening program multiple times to simulate the working conditions of the data system in a high-load environment and testing the stability of the program in the data system.

[0100] In this alternative embodiment, the aggregation degree of the data in the second variance set can be calculated according to a custom aggregation degree measurement method to obtain a screening accuracy index. The specific implementation steps of the custom aggregation degree measurement method are as follows:

[0101] Calculate the kurtosis of the data in the second variance set. The calculation formula for the kurtosis of the second variance set is as follows:

[0102]

[0103] Wherein, K represents the kurtosis of the data in the second variance set, and the kurtosis represented by K is used to characterize the aggregation degree of the data in the second variance set. The larger the value of K, the higher the aggregation degree of the data in the second variance set. The value range of K is [1, ∞); s2 j represents the j-th data in the second variance set, μ represents the mean of the second variance set, and σ represents the standard deviation of the second variance set;

[0104] Update the second variance set based on the kurtosis of the data in the second variance set to obtain a screening accuracy index, and the update method satisfies the following relational expression:

[0105]

[0106] Wherein, s2 e represents the e-th variance data in the second variance set, K represents the kurtosis of the data in the second variance set, and E e represents the e-th variance data after update. Since there are 20 variance data in the second variance set, there are 20 variance data after update, denoted as {E 1 , E 2 , …, E 20}.

[0107] In this optional embodiment, the updated variance data {E 1 , E 2 , …, E 20} can be used as the screening accuracy index.

[0108] In this way, inputting the test data into the application program of the data system to obtain result data, and further constructing the test index based on the result data can provide data support for calculating the comprehensive index of the data system, thereby improving the accuracy of the subsequent evaluation result.

[0109] S12. Calculate the characteristic values of the processing accuracy index and the screening accuracy index to obtain the index importance, and the index importance is used to characterize the importance degrees of the screening accuracy and the processing accuracy.

[0110] Please refer to Figure 3 , in an optional embodiment, calculating the characteristic values of the processing accuracy index and the screening accuracy index to obtain the index importance, and the index importance is used to characterize the importance degrees of the screening accuracy and the processing accuracy includes:

[0111] S121. Calculate the variances and covariance of the processing accuracy index and the screening accuracy index to construct a covariance matrix.

[0112] In this alternative embodiment, the covariance matrix is constructed by calculating the covariance between the processing accuracy index and the screening accuracy index, taking the covariance as the elements in the lower left and upper right corners of the covariance matrix, and taking the variances of the processing accuracy index and the screening accuracy index as the elements in the upper left and lower right corners of the covariance matrix. The formula for calculating the covariance is:

[0113]

[0114] where Cov represents the covariance between the processing accuracy index and the screening accuracy index, s1 i represents the i-th processing accuracy index, and E i represents the i-th screening accuracy index.

[0115] S122. Calculate the eigenvalues of the covariance matrix as the index importance, where the eigenvalues include processing eigenvalues and screening eigenvalues.

[0116] In this alternative embodiment, the covariance matrix can be denoted as M, and the eigenvalues of the covariance matrix M are denoted as α. The condition that the eigenvalue α needs to satisfy is |α·I - Z| = 0, where I is an identity matrix with all internal elements being 1 and the dimension of I is the same as that of M. Since M has two dimensions, the eigenvalues α1 and α2 of the matrix M can be obtained according to this condition. Among them, if the row where α1 is located is the same as the row where the variance of the processing accuracy is located, then the value of α1 is used to characterize the importance degree of the processing accuracy; if the row where α2 is located is the same as the row where the variance of the screening accuracy is located, then the value of α2 is used to characterize the importance degree of the screening accuracy.

[0117] In this alternative embodiment, the eigenvalues α1 and α2 can be used as the index importance.

[0118] In this way, the index importance is obtained by constructing the covariance matrix of the processing accuracy index and the screening accuracy index, providing data support for the subsequent establishment of a custom integration model, thereby being able to improve the accuracy of the test index.

[0119] S13. Input the index importance and the test index into a custom integration model to obtain an integration result and use it as an updated index, where the updated index is used to characterize the accuracy of the data system when performing data screening and data processing link test tasks.

[0120] In this alternative embodiment, before obtaining the integration result by inputting the metric importance and the test metric into the custom integration model and using it as the updated metric, the method further includes:

[0121] Calculating the metric importance according to a preset standardization method to obtain the normalized metric importance, where the normalized metric importance includes the normalized processing eigenvalue and the normalized screening eigenvalue.

[0122] In this alternative embodiment, the preset standardization method may be the maximization method. The specific implementation of the maximization method is to denote the maximum value of the eigenvalues α1 and α2 as α max , and calculating the normalized metric importance based on the eigenvalues α1 and α2 and the maximum value α max . The specific calculation method is as follows:

[0123]

[0124]

[0125] where β1 represents the normalized processing eigenvalue, and the larger the value of β1, the higher the importance of the processing accuracy; β2 represents the normalized screening eigenvalue, and the larger the value of β2, the higher the importance of the screening accuracy; α1 represents the processing eigenvalue; α2 represents the screening eigenvalue; α max represents the maximum value of α1 and α2.

[0126] In this alternative embodiment, a custom integration model can be constructed based on the normalized metric importance. The custom integration model is used to integrate the processing accuracy metric and the screening accuracy metric, and the custom integration model satisfies the following formula:

[0127]

[0128] where T represents the updated metric, which is used to characterize the accuracy of the data system in performing data screening and data processing link test tasks; β1 represents the normalized processing eigenvalue, which is used to characterize the importance of the processing accuracy; β2 represents the normalized screening eigenvalue, which is used to characterize the importance of the screening accuracy; Mean1 represents the average value of the processing accuracy metric, and Mean2 represents the average value of the screening accuracy metric.

[0129] In this alternative embodiment, the value of T can be used as the updated metric of the data system.

[0130] In this way, by normalizing the importance of indicators, the error caused by the difference in the order of magnitude of the importance of indicators can be reduced, and a custom integration model is established based on the normalized importance of indicators to integrate the processing accuracy indicator and the screening accuracy indicator, and finally an updated indicator is obtained to characterize the accuracy of the data system processing link, which is convenient for subsequent evaluation of the data system.

[0131] S14, evaluate the data system based on the updated indicator to obtain a test result.

[0132] In this optional embodiment, the data system can be evaluated based on a preset threshold and the updated indicator of the data system to obtain a test result. The tester can set the threshold based on test experience. The preset threshold can be denoted as t. Exemplarily, t can be 0.9. If the updated indicator T of the data system is greater than t, it can be determined that the data system is qualified; otherwise, the data system is unqualified. The test result can be denoted as R, and the calculation method of R is as follows:

[0133]

[0134] In this optional embodiment, the result represented by R can be used as the test result.

[0135] In this way, the test result is obtained by comparing the preset threshold with the updated indicator of the data system. The preset threshold can be formulated by the developer based on enterprise requirements. Therefore, the method for obtaining the test result is relatively flexible.

[0136] The above-mentioned data system testing method based on artificial intelligence separately tests different programs of the data system to obtain test indicators, calculates the weights of different programs based on the test indicators, constructs an updated indicator of the data system based on the weights to evaluate the performance of the data system, and considers the importance of the test results of different programs. Compared with the existing testing methods, the accuracy of the test results is higher.

[0137] As Figure 4 shown, it is a functional module diagram of a preferred embodiment of the data system testing device based on artificial intelligence provided by an embodiment of the present application. The data system testing device 11 based on artificial intelligence includes an acquisition unit 110, a testing unit 111, a calculation unit 112, an integration unit 113, and an evaluation unit 114. The module / unit mentioned in the present application refers to a series of computer program segments that can be executed by a processor 13 and can complete fixed functions, and are stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0138] In an optional embodiment, the acquisition unit 110 is used to extract test data from the test database according to a preset program.

[0139] In this optional embodiment, the open-source database is a collection of integrated and sharable data organized in a certain structure. Exemplarily, the preset database can be a MySQL database, which is a relational database.

[0140] In this optional embodiment, the purpose of obtaining the test data is to test the data system, which is a software system that provides data for an actually operable storage, maintenance, and application system, and is an aggregate of a storage medium, processing objects, and a management system.

[0141] In this optional embodiment, the data system can receive any formatted data to output data processing results. Therefore, the test data can be any formatted data, and the data in any open-source database can be used as test data to test the data system. This solution does not make any limitations in this regard. Figure 6 It is a schematic diagram of the test data.

[0142] In this optional embodiment, the preset program can be run on a computer to obtain test data. Exemplarily, the form of the preset program can be "select 'table' from 'database'", where 'table' represents the test data and 'database' represents the name of the database where the test data is located. The test data represented by 'table' is in the form of a table, where each row is a piece of data and each column is a field of the data. Exemplarily, 'table' can have n rows and m columns, where n can be 10,000 and m can be 5. Therefore, the set of field names included in 'table' can be {field 1, field 2, field 3, field 4, field 5}.

[0143] In this optional embodiment, the data table represented by 'table' is used as the test data.

[0144] In an optional embodiment, the test unit 111 is used to test the data system based on the test data to obtain test metrics, which include a processing accuracy metric and a screening accuracy metric. The processing accuracy metric is used to characterize the accuracy of the data system in data processing, and the screening accuracy metric is used to characterize the accuracy of the data system in data screening.

[0145] In this optional embodiment, the testing of the data system based on the test data to obtain test metrics, which include a processing accuracy metric and a screening accuracy metric, and the processing accuracy metric is used to characterize the accuracy of the data system in data processing, and the screening accuracy metric is used to characterize the accuracy of the data system in data screening, includes:

[0146] Input the test data into the data system to obtain processing result data and screening result data;

[0147] Calculate the processing result data and the screening result data respectively to obtain the processing accuracy index and the screening accuracy index.

[0148] In this alternative embodiment, the data system includes a data processing program and a data screening program. The test data can be input into the data processing program to obtain processing result data, and the test data can be input into the data screening program to obtain screening result data.

[0149] In this alternative embodiment, the data system includes a data processing program and a data screening program. The test data can be input into the data processing program to obtain processing result data, and the test data can be input into the data screening program to obtain screening result data.

[0150] In this alternative embodiment, the data processing program is a collection of multiple data operation programs. Each data operation program can be a piece of SQL script. The data operation programs can include a product program, a quotient program, a sum program, and a difference program. The form of the product program can be "select a * b as 'product' from table", and its function is to calculate the product of two fields of each piece of data in the test data; the form of the quotient program can be "select c / d as 'quotient' from table", and its function is to calculate the quotient of two fields of each piece of data in the test data; the form of the sum program can be "select e + f as'sum' from table", and its function is to calculate the sum of two fields of each piece of data in the test data; the form of the difference program can be "select g - h as 'difference' from table", and its function is to calculate the difference of two fields of each piece of data in the test data.

[0151] In this alternative embodiment, the specific implementation steps of inputting the test data into the data processing program to obtain processing result data are as follows:

[0152] Input the test data into the product program to obtain the product P of the test data. The P contains 10,000 products. In the product program, a can represent field 1 in the table, and b can represent field 2 in the table;

[0153] Input the test data into the quotient program to obtain the quotient value Q of the test data. The Q contains 10,000 quotients. In the quotient program, c can represent field 2 in the table, and d can represent field 3 in the table;

[0154] Input the test data into the summation program to obtain the sum value S of the test data. S contains 10,000 sum values. In the summation program, e can represent field 3 in the table, and f can represent field 4 in the table;

[0155] Input the test data into the difference program to obtain the difference value D of the test data. D contains 10,000 difference values. In the difference program, g can represent field 4 in the table, and h can represent field 5 in the table;

[0156] Arranging P, Q, S, and D column by column can obtain a matrix z with 10,000 rows and 4 columns.

[0157] In this alternative embodiment, the step of obtaining the processed result data can be repeated n times to obtain n groups of test results. Exemplarily, n can be 20. Each time the step of obtaining the processed result data is implemented, a corresponding matrix can be obtained. Therefore, 20 matrices of 10,000 * 4 can be obtained. Denote the set of the 20 matrices as Z, which can be expressed as Z = {z 1 , z 2 , …, z 20}.

[0158] In this alternative embodiment, the data screening program can be a section of SQL script. Its function is to screen out the data that meets the preset screening conditions. Its form can be "select data from table where 'condition'", where data represents the field where the return value of the screening program is located. Exemplarily, data = "field 4", table represents the test data, and condition represents the preset screening conditions.

[0159] In this alternative embodiment, the preset screening condition can be "ROWNUM = 3", which means to screen out the data in the 3rd row and 4th column of the test data. The specific implementation steps of inputting the test data into the data screening program to obtain the screened result data are to run the data screening program to obtain the return value that meets the preset screening conditions, and the return value is the data in the 3rd row and 4th column of the test data.

[0160] In this alternative embodiment, the method of obtaining the screened result data can be repeated m times to obtain the screened result data. Exemplarily, m can be 20. Therefore, the screened result data can be the data set {data1, data2, …, data 20}, where the subscript of each data represents the order of the screened result.

[0161] In this alternative embodiment, the matrix set Z = {z 1 , z 2 , …, z 20} can be used as the processed result data, and the data set {data1, data2, …, data 20} can be used as the screening result data.

[0162] In this alternative embodiment, the distance data of the processed result data can be calculated according to a preset distance metric method, and the distance data is used to characterize the similarity degree of the processed result data; then the variance of the example data is calculated to obtain a processing accuracy index. Taking z 1 and z 2 in the matrix set Z as an example, the implementation steps of the preset distance metric method are as follows:

[0163] Calculate the cosine distance between the first column data of z 1 and the first column data of z 2 based on the cosine distance formula and denote it as Sd 11,21 . The cosine distance formula is:

[0164]

[0165] where A represents the first column vector of the z 1 ; B represents the first column vector of the z 2 , Sd 11,21 represents the cosine distance between A and B, A i represents the value of the i-th dimension in vector A, B i represents the value of the i-th dimension in vector B, and k represents the dimension of vector A and vector B;

[0166] Calculate the corresponding cosine distances of the remaining three column data in z 1 and z 2 respectively based on the above cosine distance formula and denote them as Sd 12,22 , Sd 13,23 and Sd 14,24 ;

[0167] Calculate the mean value of Sd 11,21 , Sd 12,22 , Sd 13,23 and Sd 14,24 and denote it as Sd. The mean value represented by Sd is used to characterize the similarity degree between the matrix z 1 and z 2 . The larger the value of Sd, the lower the similarity between the matrix z 1 and z 1 .

[0168] In this alternative embodiment, since the matrix set Z contains 20 matrices, there are 190 distance data between every two matrices, denoted as Sd u ={Sd1, Sd2, …, Sd 190}, and the Sd u is the distance data of the processing result data.

[0169] In this alternative embodiment, the variance of the distance data can be calculated to obtain a processing accuracy index. The calculation formula for the variance of the distance data is:

[0170]

[0171] where s1 represents the variance of the 190 distance data, Sd i represents the i-th distance data in the Sd u . The higher the value of s1, the higher the similarity between the processing result data, and thus the higher the processing accuracy.

[0172] In this alternative embodiment, the method of obtaining the variance s1 can be repeated 20 times to obtain the first variance set s1_z = {s11, s12, …, s1 20}. The purpose is to test the data processing program multiple times to simulate the working conditions of the data system in a high-load environment, so as to test the stability of the data processing program in the data system. The first variance set s1_z can be used as the processing accuracy index.

[0173] In this alternative embodiment, the second variance set of the screening result data can be calculated multiple times, and then the aggregation degree of the data in the second variance set can be calculated according to a custom aggregation degree measurement method to obtain a screening accuracy index. The calculation method for the variance of the screening result data is:

[0174]

[0175] where s2 represents the variance of the screening result data. The larger the value of s2, the lower the aggregation degree of the screening data, the greater the difference between each screening result data, and the lower the screening accuracy; data i represents the i-th data in the screening result data.

[0176] In this alternative embodiment, the method of obtaining s2 can be repeated 20 times to obtain the second variance set s2_z = {s21, s22, …, s2 20}. The purpose is to test the screening program multiple times to simulate the working conditions of the data system in a high-load environment, so as to test the stability of the program in the data system.

[0177] In this optional embodiment, the aggregation degree of the data in the second variance set may be calculated according to a custom aggregation degree measurement method to obtain a screening accuracy index. The specific implementation steps of the custom aggregation degree measurement method are as follows:

[0178] Calculate the kurtosis of the data in the second variance set. The calculation formula for the kurtosis of the second variance set is:

[0179]

[0180] where K represents the kurtosis of the data in the second variance set. The kurtosis represented by K is used to characterize the aggregation degree of the data in the second variance set. The larger the value of K, the higher the aggregation degree of the data in the second variance set. The value range of K is [1, ∞); s2 j represents the j-th data in the second variance set, μ represents the mean of the second variance set, and σ represents the standard deviation of the second variance set;

[0181] Update the second variance set based on the kurtosis of the data in the second variance set to obtain a screening accuracy index. The update method satisfies the following relationship:

[0182]

[0183] where s2 l represents the l-th variance data in the second variance set, K represents the kurtosis of the data in the second variance set, and E l represents the l-th variance data after update. Since there are 20 variance data in the second variance set, there are 20 updated variance data, denoted as {E 1 , E 2 , …, E 20}.

[0184] In this optional embodiment, the updated variance data {E 1 , E 2 , …, E 20} may be used as the screening accuracy index.

[0185] In an optional embodiment, the calculation unit 112 is configured to calculate the eigenvalue of the processing accuracy index and the screening accuracy index to obtain an index importance, and the index importance is used to characterize the importance degree of the screening accuracy and the processing accuracy.

[0186] In this alternative embodiment, calculating eigenvalues of the processing accuracy index and the screening accuracy index to obtain the index importance, where the index importance is used to characterize the importance degrees of the screening accuracy and the processing accuracy, includes:

[0187] Calculating the variances and covariances of the processing accuracy index and the screening accuracy index to construct a covariance matrix;

[0188] Calculating the eigenvalues of the covariance matrix as the index importance, where the index importance includes a processing eigenvalue and a screening eigenvalue.

[0189] In this alternative embodiment, the covariance matrix is constructed by calculating the covariance between the processing accuracy index and the screening accuracy index, taking the covariance as the elements in the lower left corner and the upper right corner of the covariance matrix, and taking the variances of the processing accuracy index and the screening accuracy index as the elements in the upper left corner and the lower right corner of the covariance matrix. The formula for calculating the covariance is:

[0190]

[0191] where Cov represents the covariance between the processing accuracy index and the screening accuracy index, s1 i represents the i-th processing accuracy index, and E i represents the i-th screening accuracy index.

[0192] In this alternative embodiment, the covariance matrix can be denoted as M, and the eigenvalues of the covariance matrix M are denoted as α. The condition that the eigenvalue α needs to satisfy is |α·I - Z| = 0, where I is an identity matrix with all internal elements being 1 and the dimension of I is the same as that of M. Since M has two dimensions, the eigenvalues α1 and α2 of the matrix M can be obtained according to this condition. Among them, if the row where α1 is located is the same as the row where the variance of the processing accuracy is located, then the value of α1 is used to characterize the importance degree of the processing accuracy; if the row where α2 is located is the same as the row where the variance of the screening accuracy is located, then the value of α2 is used to characterize the importance degree of the screening accuracy.

[0193] In this alternative embodiment, the eigenvalues α1 and α2 can be used as the index importance.

[0194] In an alternative embodiment, the calculation unit 113 is configured to input the index importance and the test index into a custom integration model to obtain an integration result and use it as an updated index, where the updated index is used to characterize the accuracy of the data system when performing data screening and data processing link test tasks.

[0195] In this alternative embodiment, before obtaining the integration result by inputting the importance of the metrics and the test metrics into the custom integration model and using it as the updated metric, the method further includes:

[0196] Calculating the importance of the metrics according to a preset standardization method to obtain the normalized importance of the metrics, where the normalized importance of the metrics includes the normalized processing eigenvalue and the normalized screening eigenvalue.

[0197] In this alternative embodiment, the importance of the metrics can be calculated according to a preset standardization method to obtain the normalized importance of the metrics. The preset standardization method can be the maximization method. The specific implementation of the maximization method is to denote the maximum value of the eigenvalues α1 and α2 as α max , and calculating the normalized importance of the metrics based on the eigenvalues α1 and α2 and the maximum value α max . The specific calculation method is as follows:

[0198]

[0199]

[0200] where β1 represents the normalized processing eigenvalue, and the larger the value of β1, the higher the importance of the processing accuracy; β2 represents the normalized screening eigenvalue, and the larger the value of β2, the higher the importance of the screening accuracy; α1 represents the processing eigenvalue; α2 represents the screening eigenvalue; α max represents the maximum value of α1 and α2.

[0201] In this alternative embodiment, a custom integration model can be constructed based on the normalized importance of the metrics. The custom integration model is used to integrate the processing accuracy metric and the screening accuracy metric, and the custom integration model satisfies the following formula:

[0202]

[0203] where T represents the updated metric, which is used to characterize the accuracy of the data system when performing data screening and data processing link test tasks; β1 represents the normalized processing eigenvalue, which is used to characterize the importance of the processing accuracy; β2 represents the normalized screening eigenvalue, which is used to characterize the importance of the screening accuracy; Mean1 represents the average value of the processing accuracy metric, and Mean2 represents the average value of the screening accuracy metric.

[0204] In this alternative embodiment, T can be used as the updated metric of the data system.

[0205] In an alternative embodiment, the evaluation unit 114 is configured to evaluate the data system based on the update metric to obtain a test result.

[0206] In this alternative embodiment, the data system may be evaluated based on a preset threshold and the update metric of the data system to obtain a test result. The tester may set the threshold based on testing experience. The preset threshold may be denoted as t. Exemplarily, t may be 0.9. If the update metric T of the data system is greater than t, it may be determined that the data system is qualified; otherwise, the data system is unqualified. The test result may be denoted as R, and the calculation method of R is as follows:

[0207]

[0208] In this alternative embodiment, the result represented by R may be used as the test result.

[0209] The above-mentioned artificial intelligence-based data system testing method separately tests different programs of the data system to obtain test metrics, calculates the weights of different programs based on the test metrics, constructs a data system update metric based on the weights to evaluate the performance of the data system, and takes into account the importance of the test results of different programs. Compared with existing testing methods, the accuracy of the test results is higher.

[0210] As Figure 5 shown, it is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 1 includes a memory 12 and a processor 13. The memory 12 is used to store computer-readable instructions, and the processor 13 is configured to execute the computer-readable instructions stored in the memory to implement the artificial intelligence-based data system testing method of any of the above embodiments.

[0211] In an alternative embodiment, the electronic device 1 further includes a bus and a computer program stored in the memory 12 and executable on the processor 13, such as an artificial intelligence-based data system testing program.

[0212] Figure 5 Only the electronic device 1 with components 12-13 is shown. Those skilled in the art can understand that Figure 5 the shown structure does not limit the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have a different component arrangement.

[0213] Combined with Figure 1 , the memory 12 in the electronic device 1 stores multiple computer-readable instructions to implement an artificial intelligence-based data system testing method, and the processor 13 can execute multiple instructions to implement:

[0214] Extract test data from an open-source database according to a preset program, where the test data is formatted data;

[0215] Test the data system based on the test data to obtain test metrics, where the test metrics include a processing accuracy metric and a screening accuracy metric. The processing accuracy metric is used to characterize the accuracy of the data system in data processing, and the screening accuracy metric is used to characterize the accuracy of the data system in data screening;

[0216] Calculate the eigenvalues of the processing accuracy metric and the screening accuracy metric to obtain the metric importance, where the metric importance is used to characterize the importance degree of the screening accuracy and the processing accuracy;

[0217] Input the metric importance and the test metrics into a custom integration model to obtain an integration result and use it as an updated metric, where the updated metric is used to characterize the accuracy of the data system when performing data screening and data processing link test tasks;

[0218] Evaluate the data system based on the updated metric to obtain a test result.

[0219] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0220] Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 1 and does not constitute a limitation on the electronic device 1. The electronic device 1 can be either a bus structure or a star structure. The electronic device 1 can also include more or fewer other hardware or software than shown in the figure, or different component arrangements. For example, the electronic device 1 can also include input / output devices, network access devices, etc.

[0221] It should be noted that the electronic device 1 is only an example. Other existing or future possible electronic products that can be adapted to this application should also be included in the protection scope of this application and are included herein by reference.

[0222] Among them, the memory 12 includes at least one type of readable storage medium, which can be non-volatile or volatile. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 12 can be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 12 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the electronic device 1. The memory 12 can be used not only to store application software installed in the electronic device 1 and various types of data, such as the code of the data system test program based on artificial intelligence, etc., but also to temporarily store the data that has been output or will be output.

[0223] In some embodiments, the processor 13 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged, including the combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the electronic device 1, connecting all components of the entire electronic device 1 through various interfaces and circuits. By running or executing the programs or modules stored in the memory 12 (such as executing the data system test program based on artificial intelligence, etc.), and calling the data stored in the memory 12, it can execute various functions of the electronic device 1 and process data.

[0224] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned embodiments of various data system test methods based on artificial intelligence, such as Figure 1 the steps shown.

[0225] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into an acquisition unit 110, a test unit 111, a calculation unit 112, an integration unit 114, and an evaluation unit 115.

[0226] The integrated units implemented in the form of software function modules as described above may be stored in a computer-readable storage medium. The above-mentioned software function modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute a part of the method for testing an artificial intelligence-based data system according to various embodiments of the present application.

[0227] If the modules / units integrated in the electronic device 1 are implemented in the form of software function units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present application, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned method embodiments may be implemented.

[0228] Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory, and other memories, etc.

[0229] Furthermore, the computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the blockchain node, etc.

[0230] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.

[0231] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, in Figure 5 it is only represented by one arrow, but it does not mean that there is only one bus or one type of bus. The bus is arranged to realize the connection and communication between the memory 12 and at least one processor 13, etc.

[0232] The embodiment of this application also provides a computer-readable storage medium (not shown in the figure). The computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions are executed by a processor in an electronic device to implement the method for testing an artificial-intelligence-based data system described in any of the above embodiments.

[0233] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0234] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0235] In addition, in each embodiment of this application, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0236] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular form does not exclude the plural form. A plurality of units or devices described in the specification can also be implemented by one unit or device through software or hardware. Terms such as first and second are used to denote names and do not denote any particular order.

[0237] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. An artificial intelligence-based data system testing method, characterized in that, The method includes: Extracting test data from an open-source database according to a preset program, where the test data is formatted data; Testing the data system based on the test data to obtain test metrics, where the test metrics include a processing accuracy metric and a screening accuracy metric. The processing accuracy metric is used to characterize the accuracy of data processing by the data system, and the screening accuracy metric is used to characterize the accuracy of data screening by the data system. Among them, testing the data system based on the test data to obtain test metrics, where the test metrics include a processing accuracy metric and a screening accuracy metric includes: inputting the test data into the data system to obtain processed result data and screening result data; calculating distance data of the processed result data according to a preset distance measurement method, where the distance data is used to characterize the similarity degree of the processed result data; calculating the variance of the distance data multiple times to obtain a processing accuracy metric; calculating the variance of the screening result data multiple times to obtain a second variance set; calculating the aggregation degree of the second variance set according to a custom aggregation degree measurement method to obtain a screening accuracy metric, and the higher the value of the screening accuracy metric, the more complete the data screening function of the data system. Among them, the method for determining the aggregation degree includes: , where K represents the kurtosis of the data in the second variance set, and the kurtosis represented by K is used to characterize the aggregation degree of the data in the second variance set. The larger the value of K, the higher the aggregation degree of the data in the second variance set, and the value range of K is [1, ∞); represents the j-th data in the second variance set, represents the mean of the second variance set, represents the standard deviation of the second variance set; Calculating the eigenvalues of the processing accuracy index and the screening accuracy index to obtain the index importance, which is used to characterize the importance levels of the screening accuracy and the processing accuracy; Input the importance of the metrics and the test metrics into a custom integration model to obtain an integration result, which is used as an updated metric. The updated metric is used to characterize the accuracy of the data system when performing data screening and data processing link test tasks. Among them, the custom integration model satisfies the relational expression: , where T represents the updated metric, and the updated metric is used to characterize the accuracy of the data system when performing data screening and data processing link test tasks; represents the normalized processing eigenvalue, which is used to characterize the importance of processing accuracy; represents the normalized screening eigenvalue, which is used to characterize the importance of screening accuracy; Mean1 represents the average value of the processing accuracy metric, and Mean2 represents the average value of the screening accuracy metric; Evaluating the data system based on the updated index to obtain a test result.

2. The artificial intelligence-based data system testing method according to claim 1, characterized in that, The calculating the eigenvalues of the processing accuracy index and the screening accuracy index to obtain the index importance includes: Calculating the variance and covariance of the processing accuracy index and the screening accuracy index to construct a covariance matrix; Calculating the eigenvalues of the covariance matrix as the index importance, where the eigenvalues include processing eigenvalues and screening eigenvalues.

3. The artificial intelligence-based data system testing method according to claim 2, characterized in that, Before inputting the index importance and the test index into a custom integration model to obtain an integration result and using it as the updated index, the method further includes: Calculating the index importance according to a preset standardization method to obtain the normalized index importance, where the normalized index importance includes the normalized processing eigenvalues and the normalized screening eigenvalues.

4. The artificial intelligence-based data system testing method according to claim 1, characterized in that, The evaluating the data system based on the updated index to obtain a test result includes: Evaluating the data system based on the updated index to obtain an evaluation result; Comparing the evaluation result with a preset threshold to obtain a test result.

5. An artificial intelligence-based data system testing device, characterized in that, The apparatus is used to implement the method according to any one of claims 1 to 4, and the apparatus includes: An acquisition unit, configured to extract test data from an open-source database according to a preset program, where the test data is formatted data; A test unit, configured to test the data system based on the test data to obtain test indexes, where the test indexes include a processing accuracy index and a screening accuracy index, the processing accuracy index is used to characterize the accuracy of the data system in data processing, and the screening accuracy index is used to characterize the accuracy of the data system in data screening; A calculation unit, configured to calculate the eigenvalues of the processing accuracy index and the screening accuracy index to obtain the index importance, which is used to characterize the importance levels of the screening accuracy and the processing accuracy; An integration unit, configured to input the index importance and the test index into a custom integration model to obtain an integration result and use it as the updated index, where the updated index is used to characterize the accuracy of the data system in performing the data screening and data processing link test tasks; An evaluation unit, configured to evaluate the data system based on the updated index to obtain a test result.

6. An electronic device, characterized in that, The electronic device includes: A memory, storing computer-readable instructions; and A processor, executing the computer-readable instructions stored in the memory to implement the artificial intelligence-based data system test method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: Computer-readable instructions are stored in the computer-readable storage medium, and the computer-readable instructions are executed by a processor in the electronic device to implement the artificial intelligence-based data system test method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Artificial intelligence prediction evaluation method and device, storage medium and electronic equipment

    CN110826908A

  • Microservice failure modeling and testing

    US10684940B1