Data checking management method and system based on artificial intelligence technology and medium
Through the data verification management method based on artificial intelligence, and using technical means such as verification models and allocation factor calculation, the problem of insufficient initiative and forward-looking data verification in the existing technology has been solved, and higher data verification accuracy and efficiency have been achieved.
Patent Information
- Application Number
- CN202510502662.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Existing data verification methods have poor initiative and prospects when they find errors that have not occurred, making it difficult to effectively judge whether the data is abnormal.
Using a data verification management method based on artificial intelligence technology, we use historical samples and data samples to be detected, conduct preliminary verification and data encoding, establish a verification model, and use allocation factor calculation and screening methods to screen elements participating in the verification, and divide them into calculation groups to improve the accuracy and initiative of data verification.
Through the interrelation between data, check the detection data, actively discover errors in new data, improve the accuracy of verification results, and have good initiative and forward-looking.
Smart Images

Figure CN120030202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data verification technology, and specifically to a data verification management method, system and medium based on artificial intelligence technology. Background Art
[0002] Data verification refers to the use of manual verification, program comparison and other means to conduct a comprehensive and detailed review of various types of data, aiming to confirm the accuracy, completeness, consistency and reliability of the data. It plays a key role in various fields, can accurately eliminate erroneous, missing and contradictory data, ensure data quality, lay a solid foundation for corporate strategic planning, government policy formulation, analysis and judgment of scientific research projects and other activities, and help make scientific, reasonable and practical decisions.
[0003] When conducting data verification, it is necessary to use a verification method to simplify the verification process, such as the verification method and storage medium for customer basic file data based on machine learning with patent publication number CN118964597A, which includes obtaining the internal data and external database of the enterprise; judging whether the unified social credit code in the internal data exists in the external database; the present invention determines whether there is the internal data of the enterprise in the external database by judging whether there is a unified social credit code for the internal data of the enterprise in the external database; if the unified social credit code exists in the external database, the external data corresponding to the unified social credit code in the external database is further verified with the internal data of the enterprise for text information; if the verification is consistent, it means that the internal and external data of the enterprise are consistent, otherwise they are inconsistent; adopting this method can effectively improve the accuracy of verification on the one hand, and on the other hand, it can also effectively improve the efficiency of verification.
[0004] Verifying newly acquired or newly registered data based on data in the database can avoid database contamination and improve the data quality of the database. However, when conducting the verification, the above method determines whether the data is abnormal by comparing the items of external data with internal data. This judgment method can only discover problems that are inconsistent with known rules or other system data. It is difficult to judge errors that have not occurred, and its initiative and foresight are poor. Summary of the invention
[0005] The purpose of the present invention is to provide a data verification management method, system and medium based on artificial intelligence technology to solve the problems raised in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solution: a data verification management method based on artificial intelligence technology, comprising:
[0007] Obtaining an original database with multiple historical samples and data samples to be tested;
[0008] Conduct preliminary verification on the data samples to be tested to improve the data reliability of the samples to be tested;
[0009] The non-numerical data in the historical samples and the samples to be tested are encoded by using a data encoding method, and are edited into a historical sample matrix and a sample matrix to be tested respectively;
[0010] Through the historical sample processing method, multiple historical sample matrices are randomly distributed into several groups, and the average matrix of each group is calculated to establish an average matrix data set, thereby reducing the number of samples in the data set and the amount of calculation;
[0011] The allocation factors between elements in the historical sample matrix are calculated by the allocation factor calculation method, the elements involved in the verification are screened out according to the screening method, and the elements involved in the verification are divided into several calculation groups, and the elements that can affect the verification results of the samples to be tested are screened and classified to improve the accuracy of data verification;
[0012] A verification model is established based on the screened elements in the historical database, and the elements of different categories in the sample to be tested are brought into the verification model for data verification and the verification results are output. Through the mutual correlation between the data, the data to be tested is verified to determine whether the data is abnormal, and errors in new data are actively discovered to improve the accuracy of the verification results.
[0013] Preferably, the allocation factor calculation method includes:
[0014] Group different elements and number the elements in the same group in order according to their numerical values;
[0015] Select two elements that need to be calculated and calculate the difference between the sequence numbers according to the formula, specifically:
[0016]
[0017] in Indicates the difference between the sequence number of the Xth element and the sequence number of the Yth element in sample i. and represent the Xth and Yth elements in sample i respectively, and They represent the sequence number of the Xth element in sample i and the sequence number of the Yth element in sample i respectively;
[0018] Then, the distribution factor between the two elements is calculated based on the difference between the multiple sequence numbers, specifically:
[0019]
[0020] in represents the distribution factor between the Xth element and the Yth element, represents the total number of samples, Represents the square of the difference between the sequence number of the Xth element and the sequence number of the Yth element in sample i.
[0021] Preferably, the allocation factor calculation method includes:
[0022] For each element, the elements of each historical sample matrix are sequentially numbered according to the numerical value;
[0023] When there are identical values in the data, the sequence numbers of the identical values are the average of their sorted positions;
[0024] The distribution factor between two elements is calculated based on the difference between multiple sequence numbers and the average value, specifically:
[0025]
[0026] in represents the distribution factor between the Xth element and the Yth element, and represent the Xth and Yth elements in sample i respectively, and They represent the sequence number of the Xth element in sample i and the sequence number of the Yth element in sample i, respectively. and are the average values of the sequence numbers of the Xth and Yth elements, Indicates the total number of samples.
[0027] Preferably, the screening method comprises:
[0028] Comparing the allocation factor with a plurality of preset level factors of different sizes, dividing the allocation factor into levels, and setting a preset threshold for each level of the division;
[0029] It is determined whether the allocation factors at the same level have crossover elements, and the elements corresponding to the multiple allocation factors with crossover elements are grouped as one calculation group, and the elements corresponding to the allocation factors without crossover elements are grouped into another calculation group.
[0030] Preferably, the establishment of the verification model is specifically as follows:
[0031] Select a calculation group, calculate the mean and standard deviation of each element in the target calculation group, and combine the mean values to obtain the average feature matrix ;
[0032] Calculate the mean of multiple allocation factors in the target calculation group as the correlation coefficient between the elements;
[0033] Then calculate the covariance matrix between every two elements according to the formula , specifically:
[0034]
[0035] in represents the data at the location with coordinates (X, Y) in the covariance matrix, Indicates the correlation coefficient between the Xth element and the Yth element. In the same calculation group, the correlation coefficient between every two elements is the same. represents the standard deviation of the Xth element, represents the standard deviation of the Yth element;
[0036] The verification model is established according to the probability density function, specifically:
[0037]
[0038] in Represents the sample matrix to be detected The abnormal prediction value, the value of which reflects the sample matrix to be detected The abnormal probability of represents the average feature matrix, Represents the sample matrix to be detected and the average feature matrix The number of elements of Represents the covariance matrix The determinant of Represents the covariance matrix The inverse matrix of Represents conversion, converting a column vector into a row vector;
[0039] The abnormal prediction value is compared with the prediction threshold. When the abnormal prediction value is greater than the prediction threshold, the data is judged to be normal, otherwise the data is judged to be abnormal.
[0040] Preferably, the establishment of the verification model is specifically as follows:
[0041] S1: Select a calculation group, convert each sample into a calculation group sample matrix, randomly select several calculation group sample matrices, and use their data as the central sample matrices of several communities;
[0042] S2: Calculate the vector gap between each calculation group sample matrix and each center sample matrix according to the formula, specifically:
[0043]
[0044] in Indicates the vector gap between the calculation group sample matrix i and the center sample matrix, represents the number of elements of the sample matrix of each calculation group in the target calculation group, It means calculating the value of the jth element in the sample matrix i of the group. Represents the value of the jth element of the center sample matrix;
[0045] S3: assigning the sample matrix of the calculation group and the central sample with the smallest vector gap to the same community according to the vector gap;
[0046] S4: Calculate the average matrix of each community and use the average matrix as the new central sample matrix;
[0047] S5: Repeat steps S2-S4 until the central sample matrix no longer changes;
[0048] S6: Calculate and compare the vector gap between the sample matrix to be detected and each central sample matrix, and compare the smallest vector gap with the gap threshold. When the vector gap is greater than the gap threshold, the data is judged to be abnormal, otherwise the data is judged to be normal.
[0049] Preferably, the preliminary verification is specifically:
[0050] Back up the data samples to be tested;
[0051] Check whether the data to be tested has missing values. If the data sample to be tested has missing values, prompt relevant personnel to conduct manual verification;
[0052] Formulate several logical judgment rules to detect logical errors in the data samples to be tested, and obtain change records when relevant personnel make changes to abnormal data;
[0053] When abnormal data and change records accumulate to a set level, the records will be sorted and sent to relevant personnel to facilitate the staff to sort and add or subtract logical judgment rules to improve the accuracy of preprocessing.
[0054] Preferably, the data encoding method comprises:
[0055] Establish a feature category table and encode the feature categories;
[0056] Extract text features from text data, classify and obtain codes according to set categories, mark and prompt text data with non-single classification and text data without categories, and make manual judgments by relevant personnel;
[0057] For image data, similarity detection is performed for each category. When the similarity ratio between the category with the highest similarity and the category with the second-highest similarity is higher than or equal to the ratio threshold, it is determined that the image data belongs to the category with the highest similarity and the encoding is obtained. When the similarity ratio between the category with the highest similarity and the category with the second-highest similarity is lower than the ratio threshold, a prompt is given for the image data and it is manually determined by relevant personnel.
[0058] Compared with the prior art, the beneficial effects of the present invention are:
[0059] A verification model is established through the selected elements. The elements of different classifications in the sample to be detected are respectively brought into the verification model for data verification and the verification results are output. Through the mutual association between the data, the data to be detected is verified to determine whether the data is abnormal, and errors in the new data are actively discovered, improving the accuracy of the verification results, and having good initiative and foresight.
[0060] Meanwhile, the distribution factors between the elements in the historical sample matrix are calculated through the distribution factor calculation method, and then the elements participating in the verification are selected according to the selection method, and the elements participating in the verification are divided into several calculation groups. The elements that can affect the verification results of the sample to be detected are selected and classified, and the distribution factor can reflect the influence of different calculation groups in data verification, further optimizing the accuracy of the verification results. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a schematic diagram of the overall process of the data verification management method of the present invention;
[0062] Figure 2 It is a schematic diagram of the process of the distribution factor calculation method in the first embodiment of the present invention;
[0063] Figure 3 It is a schematic diagram of the process of the distribution factor calculation method in the second embodiment of the present invention;
[0064] Figure 4 It is a schematic diagram of the process of the selection method in the present invention;
[0065] Figure 5 It is a schematic diagram of the process of establishing the verification model in the present invention;
[0066] Figure 6 It is a schematic diagram of the process of establishing the verification model in the third embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0068] In the present application, for ease of understanding, the method steps used do not need to be executed in the order of the steps in this embodiment during actual operation. In other embodiments, these steps may be performed simultaneously or in a different order.
[0069] Embodiment 1:
[0070] Verifying newly acquired or newly registered data based on the data in the database can avoid database contamination and improve the data quality of the database. Through the mutual correlation between the data, the data to be tested can be verified to determine whether the data is abnormal, and errors in new data can be proactively discovered to improve the accuracy of the verification results. It has good initiative and foresight.
[0071] like Figure 1 , Figure 2 , Figure 4 and Figure 5 As shown, the present invention provides a technical solution: a data verification management method based on artificial intelligence technology, comprising:
[0072] Obtaining an original database with multiple historical samples and data samples to be tested;
[0073] Conduct preliminary verification on the data samples to be tested to improve the data reliability of the samples to be tested;
[0074] The non-numerical data in the historical samples and the samples to be tested are encoded by using a data encoding method, and are edited into a historical sample matrix and a sample matrix to be tested respectively;
[0075] Through the historical sample processing method, multiple historical sample matrices are randomly distributed into several groups, and the average matrix of each group is calculated to establish an average matrix data set, thereby reducing the number of samples in the data set and the amount of calculation;
[0076] The allocation factors between elements in the historical sample matrix are calculated by the allocation factor calculation method, the elements involved in the verification are screened out according to the screening method, and the elements involved in the verification are divided into several calculation groups, and the elements that can affect the verification results of the samples to be tested are screened and classified to improve the accuracy of data verification;
[0077] A verification model is established based on the screened elements in the historical database, and the elements of different categories in the sample to be tested are brought into the verification model for data verification and the verification results are output. Through the mutual correlation between the data, the data to be tested is verified to determine whether the data is abnormal, and errors in new data are actively discovered to improve the accuracy of the verification results.
[0078] It should be noted that when verifying, if there is missing or difficult-to-read data in the original database, some samples in the original database can be abandoned, or supplemented using the mean, median, etc. to improve the quality of the original database. The specific supplement method can be selected according to the data type and will not be elaborated here.
[0079] Preliminary verification, specifically:
[0080] Back up the data samples to be tested;
[0081] Check whether the data to be tested has missing values. If the data sample to be tested has missing values, prompt relevant personnel to conduct manual verification;
[0082] Formulate several logical judgment rules to detect logical errors in the data samples to be tested, and obtain change records when relevant personnel make changes to abnormal data;
[0083] When abnormal data and change records accumulate to a set level, the records will be sorted and sent to relevant personnel to facilitate the staff to sort and add or subtract logical judgment rules to improve the accuracy of preprocessing.
[0084] It should be noted that the detection of missing values and the formulation of logical rules are both existing technologies and will not be elaborated here. For example, an upper limit is set for age, and the occurrence time of an event must be earlier than the end time. When new data is input, whenever the new data is judged to be abnormal data, the data is recorded, and the changes made to the data by relevant personnel are recorded, and then batch sorting is performed, which helps relevant personnel to set new logical rules or delete erroneous logical rules, which can improve the effect of preliminary verification and thus improve the overall verification efficiency.
[0085] Data encoding methods include:
[0086] Establish a feature category table and encode the feature categories;
[0087] Extract text features from text data, classify and obtain codes according to set categories, mark and prompt text data with non-single classification and text data without categories, and make manual judgments by relevant personnel;
[0088] For image data, a similarity check is performed on each category. When the similarity ratio between the category with the highest similarity and the category with the second highest similarity is higher than or equal to the ratio threshold, the image data is determined to belong to the category with the highest similarity and the encoding is obtained. When the similarity ratio between the category with the highest similarity and the category with the second highest similarity is lower than the ratio threshold, the image data is prompted and manually judged by relevant personnel.
[0089] It should be noted that for text data, a feature category table is established. For example, in the police system, the case type is coded: theft is coded as 1, robbery is coded as 2, violent crime is coded as 3, and fraud is coded as 4. The value range can also be changed according to actual needs, for example, theft is coded as 10, robbery is coded as 20, violent crime is coded as 30, and fraud is coded as 40. If the case type in the new data is endangering public safety, the category that appears in the feature category table can mark the data, and relevant personnel can change the case type of the new data or modify the feature category table. For image data, a representative image can be selected from each category in the feature category table, and then the images in the new data are compared for similarity (the specific similarity comparison method is an existing technology such as perceptual hashing, convolutional neural network, etc., which will not be repeated here). For example, the similarities between some images in the new data and the categories in the feature category table are 90%, 20%, 3%, 4%, 1%, and 0, respectively, and the set ratio threshold is 4. At this time, the ratio of 90% to 20% is 4.5, which is greater than the ratio threshold. The number of the category with a similarity of 90% can be used as the number of the image data in the new data.
[0090] like Figure 2 As shown, the allocation factor calculation method includes:
[0091] Group different elements and number the elements in the same group in order according to their numerical values;
[0092] Select two elements that need to be calculated and calculate the difference between the sequence numbers according to the formula, specifically:
[0093]
[0094] in Indicates the difference between the sequence number of the Xth element and the sequence number of the Yth element in sample i. and represent the Xth and Yth elements in sample i respectively, and They represent the sequence number of the Xth element in sample i and the sequence number of the Yth element in sample i respectively;
[0095] Then, the distribution factor between the two elements is calculated based on the difference between multiple sequential labels, the difference is squared, the influence of positive and negative numbers on the data is taken out, and then the sum of several difference squares is calculated. This sum reflects the overall degree of difference in the levels of the two variables. The larger the sum of squared differences, the more obvious the difference in the levels of the two variables and the weaker the correlation. The sum of squared level differences is normalized and used as the denominator. The weight of the sum of squared level differences is further adjusted so that the correlation coefficient can reasonably take values within the range of [−1,1], thereby accurately reflecting the degree of correlation between the two variables. Specifically:
[0096]
[0097] in represents the distribution factor between the Xth element and the Yth element, represents the total number of samples, Represents the square of the difference between the sequence number of the Xth element and the sequence number of the Yth element in sample i.
[0098] It should be noted that for the convenience of calculation, the simulated data is as follows
[0099] Assume that we monitor the operation data of fire protection facilities in a large shopping mall. The data includes four key indicators: fire protection water pressure (MPa), fire protection equipment temperature (℃), smoke detector signal strength (dimensionless), and fire door status (closed as 0 and open as 1). The following is the historical data of the past 5 regular inspections, as shown in Table 1:
[0100] Table 1: Raw data
[0101] sample Fire water pressure Equipment temperature Signal Strength Fire door status 1 0.80 31 20 0 2 0.75 28 18 0 3 0.90 32 22 0 4 0.70 26 16 0 5 0.85 30 20 0
[0102] Taking fire water pressure and equipment temperature as an example, the allocation factor is calculated as follows:
[0103] First, the data of fire water pressure and equipment temperature are numbered as (3, 2, 5, 1, 4) and (4, 2, 5, 1, 3) respectively, and then the calculation is done. , respectively (-1, 0, 0, 0, 1), then according to the formula =0.9, which is close to 1, indicating a strong hierarchical order between the two. The calculation method for the other elements is the same, so I will not go into details here.
[0104] like Figure 4 As shown, the screening methods include:
[0105] Compare the allocation factor with a plurality of preset level factors of different sizes, divide the allocation factor into levels, and set a preset threshold for each level of the division;
[0106] It is determined whether the allocation factors at the same level have crossover elements, and the elements corresponding to the multiple allocation factors with crossover elements are grouped as one calculation group, and the elements corresponding to the allocation factors without crossover elements are grouped into another calculation group.
[0107] It should be noted that, in order to facilitate calculation, the simulation data is set as follows:
[0108] There are several historical samples in the original database, and the samples contain 4 elements. After calculation, the distribution factor between the first element and the second element is 0.22, the distribution factor between the first element and the third element is 0.63, the distribution factor between the first element and the fourth element is 0.11, the distribution factor between the second element and the third element is 0.12, the distribution factor between the second element and the fourth element is 0.14, and the distribution factor between the third element and the fourth element is 0.32. The set level factors are 0.2 and 0.5. The distribution factor less than 0.2 is distributed at the first level, the distribution factor greater than or equal to 0.2 and less than 0.5 is distributed at the second level, and the distribution factor greater than or equal to 0.5 is distributed at the third level.
[0109] Taking the second level as an example, at this time, in the second level, the distribution factor between the second element and the third element is 0.12, the distribution factor between the second element and the fourth element is 0.14, and the distribution factor between the first element and the fourth element is 0.11. The two distribution factors of 0.14 and 0.12 have an intersecting element (the second element), and the two distribution factors of 0.14 and 0.11 have an intersecting element (the fourth element). Therefore, the elements corresponding to the three distribution factors of 0.14, 0.12 and 0.11 are a calculation group, that is, the second element, the third element and the fourth element are all a calculation group, that is, there is only one calculation group in the second level, and other levels also use the same method to divide the calculation groups, which will not be repeated here.
[0110] like Figure 5 As shown, a verification model is established, specifically:
[0111] Select a calculation group, calculate the mean and standard deviation of each element in the target calculation group, and combine the mean values to obtain the average feature matrix ;
[0112] Calculate the mean of multiple allocation factors in the target calculation group as the correlation coefficient between the elements;
[0113] Then calculate the covariance matrix between every two elements according to the formula , specifically:
[0114]
[0115] in represents the data at the location with coordinates (X, Y) in the covariance matrix, Indicates the correlation coefficient between the Xth element and the Yth element. In the same calculation group, the correlation coefficient between every two elements is the same, and when X and Y are equal, =1, represents the standard deviation of the Xth element, represents the standard deviation of the Yth element;
[0116] The verification model is established according to the probability density function, specifically:
[0117]
[0118] in Represents the sample matrix to be detected The abnormal prediction value, the value of which reflects the sample matrix to be detected The abnormal probability of represents the average feature matrix, Represents the sample matrix to be detected and the average feature matrix The number of elements of Represents the covariance matrix The determinant of Represents the covariance matrix The inverse matrix of Represents conversion, converting a column vector into a row vector;
[0119] The abnormal prediction value is compared with the prediction threshold. When the abnormal prediction value is greater than the prediction threshold, the data is judged to be normal, otherwise the data is judged to be abnormal.
[0120] It should be noted that, for the convenience of calculation, it is assumed that the data of a calculation group is as follows:
[0121] Assume that the fire protection facilities maintenance situation of an industrial park is evaluated. In one of the calculation groups, the three elements are: fire protection equipment inspection frequency (times / month), fire protection training duration (hours / quarter), and fire drill participation rate (%). The preset threshold corresponding to this calculation group is 0.001, as shown in Table 2:
[0122] Table 2: Calculation group data 1
[0123] sample Check frequency Training duration Exercise participation rate 1 8 10 85 2 5 6 70 3 12 15 90 4 3 4 60 5 10 12 88 6 6 8 75
[0124] The mean of the distribution factor of the calculation group is calculated to be about 0.82 (the calculation method is the same as the above distribution factor calculation method, which will not be demonstrated); the standard deviations of the three elements are calculated to be 3.33, 4.04, and 11.75 respectively; and the value of each element in the covariance matrix is calculated according to the formula. For example, the calculated result is 3.33×3.33×1≈11.09. Continue to calculate the values of the remaining positions in the covariance matrix to obtain the covariance = .
[0125] Assume that the sample matrix to be detected is = ,Will and Bring in It can be calculated in ≈3.63×10 −10 ,at this time It is smaller than the preset threshold, so the data to be detected is determined to be abnormal.
[0126] Embodiment 2:
[0127] In the first embodiment, when calculating the allocation factor, it is necessary to first sequentially label the elements of each historical sample matrix. However, when the same numerical value appears in the elements, the labeling order is difficult to determine, and it is difficult to ensure the accuracy of the calculation of the allocation factor. This embodiment provides another allocation factor calculation method to improve the calculation accuracy of the allocation factor when the same numerical value appears in the elements.
[0128] like Figure 3 As shown, the allocation factor calculation method includes:
[0129] For each element, the elements of each historical sample matrix are sequentially numbered according to the numerical value;
[0130] When there are identical values in the data, the sequence numbers of the identical values are the average of their sorted positions;
[0131] The distribution factor between two elements is calculated based on the difference between multiple sequential labels and the average value, specifically:
[0132]
[0133] in represents the distribution factor between the Xth element and the Yth element, and represent the Xth and Yth elements in sample i respectively, and They represent the sequence number of the Xth element in sample i and the sequence number of the Yth element in sample i, respectively. and are the average values of the sequence numbers of the Xth and Yth elements, Indicates the total number of samples.
[0134] It should be noted that, in order to facilitate calculation, the simulation data is set as follows:
[0135] Suppose that the data of two elements in different samples are (3, 5, 5, 8, 10) and (4, 6, 6, 9, 11) respectively.
[0136] First, we perform sequential labeling, and for the sequential labels with the same value, we take the average of their sorted positions. The results are (1, 2.5, 2.5, 4, 5) and (1, 2.5, 2.5, 4, 5), respectively. and ,get = =3.
[0137] according to It can be calculated =1.
[0138] When elements with the same value appear, the distribution factor can be calculated to improve the calculation accuracy of the distribution factor.
[0139] Embodiment three:
[0140] In one embodiment, a high probability density function is used to establish a verification model, which requires that the data arrangement in the original database conforms to the Gaussian distribution. When the data volume is large and the data dimension is large, it is difficult to determine the data distribution form. The verification result of this verification model will be biased and has great limitations in use. This embodiment provides another verification model to improve the verification accuracy when it is difficult to determine the distribution rules of the original data.
[0141] like Figure 6 As shown, a verification model is established, specifically:
[0142] S1: Select a calculation group, convert each sample into a calculation group sample matrix, randomly select several calculation group sample matrices, and use their data as the central sample matrices of several communities;
[0143] S2: Calculate the vector gap between each calculation group sample matrix and each center sample matrix according to the formula, specifically:
[0144]
[0145] in Indicates the vector gap between the calculation group sample matrix i and the center sample matrix, represents the number of elements of the sample matrix of each calculation group in the target calculation group, It means calculating the value of the jth element in the sample matrix i of the group. Represents the value of the jth element of the center sample matrix;
[0146] S3: assigning the sample matrix of the calculation group and the central sample with the smallest vector gap to the same community according to the vector gap;
[0147] S4: Calculate the average matrix of each community and use the average matrix as the new central sample matrix;
[0148] S5: Repeat steps S2-S4 until the central sample matrix no longer changes;
[0149] S6: Calculate and compare the vector gap between the sample matrix to be detected and each central sample matrix, and compare the smallest vector gap with the gap threshold. When the vector gap is greater than the gap threshold, the data is judged to be abnormal, otherwise the data is judged to be normal.
[0150] It should be noted that in order to facilitate calculation, the data is set as follows:
[0151] Assume that the police have collected criminal case data in a certain area over a period of time. One of the calculation groups contains information such as the latitude and longitude of the crime scene, the time of the crime, the type of case (such as theft, robbery, violent crime, fraud, etc.), and the number of suspects. For ease of processing, the case type is coded: theft is coded as 1, robbery is coded as 2, violent crime is coded as 3, and fraud is coded as 4. The calculation group data is shown in Table 3:
[0152] Table 3: Calculation group data 2
[0153] Case Number longitude latitude Type Encoding Number of staff 1 116.38 39.95 1 1 2 116.40 39.98 1 2 3 116.35 39.92 1 1 4 116.50 40.00 2 3 5 116.55 40.05 2 4 6 116.48 39.98 2 3 7 116.20 39.80 3 1 8 116.18 39.75 3 1 9 116.25 39.85 3 2
[0154] Randomly select the data of 3 samples as the central sample matrix. Assume that the selected central sample matrix is: [116.38, 39.95, 1, 1] (corresponding to case 1), [116.50, 40.00, 2, 3] (corresponding to case 4), and [116.20, 39.80, 3, 1] (corresponding to case 7).
[0155] Calculate the vector gap between the group sample matrix and each center sample matrix, for example, for case 2 ([116.40,39.98,1,2]):
[0156] The vector gap with the first center sample matrix [116.38, 39.95, 1, 1] is: |116.40−116.38|+|39.98−39.95|+|1−1|+|2−1|=0.02+0.03+0+1=1.05;
[0157] The vector gap with the second center sample matrix [116.50,40.00,2,3] is: |116.40−116.50|+|39.98−40.00|+|1−2|+|2−3|=0.1+0.02+1+1=2.12;
[0158] The vector gap with the third center sample matrix [116.20, 39.80, 3, 1] is: |116.40−116.20|+|39.98−39.80|+|1−3|+|2−1|=0.2+0.18+2+1=3.38;
[0159] Because the vector gap between case 2 and the first central sample matrix is the smallest, case 2 belongs to the first community.
[0160] Repeat this process to assign all data points to their corresponding clusters:
[0161] The first community: [116.38,39.95,1,1], [116.40,39.98,1,2], [116.35,39.92,1,1];
[0162] The second community: [116.50,40.00,2,3], [116.55,40.05,2,4], [116.48,39.98,2,3];
[0163] The third community: [116.20,39.80,3,1], [116.18,39.75,3,1], [116.25,39.85,3,2].
[0164] The average matrix of each community is calculated as the new central sample matrix. The new central sample matrices of the three communities are [116.38, 39.95, 1, 1.33], [116.51, 40.01, 2, 3.33], and [116.21, 39.80, 3, 1.33]. Then, each case is assigned to a community. In this embodiment, the central sample matrix calculated subsequently does not change, so the above assignment result is the final assignment result.
[0165] Suppose a new case occurs, the longitude and latitude of the crime scene are [117.00,40.50], the case type code is 4 (fraud), and the number of suspects is 5.
[0166] Calculate the vector gap from new data to each center sample matrix:
[0167] The vector gap to the first center sample matrix [116.38, 39.95, 1, 1.33] is: |117.00−116.38|+|40.50−39.95|+|4−1|+|5−1.33|=0.62+0.55+3+3.67=7.84;
[0168] The vector gap to the second center sample matrix [116.51, 40.01, 2, 3.33] is: |117.00−116.51|+|40.50−40.01|+|4−2|+|5−3.33|=0.49+0.49+2+1.67=4.65;
[0169] The vector gap to the third center sample matrix [116.21, 39.80, 3, 1.33] is: |117.00−116.21|+|40.50−39.80|+|4−3|+|5−1.33|=0.79+0.7+1+3.67=6.16;
[0170] Assuming that a gap threshold of 3.5 is set according to the actual situation, the vector gap between the new data and the nearest second central sample matrix is 4.65, which is greater than the gap threshold, so it can be judged that the data of this new case is abnormal data. Since this case is a fraud type, it appears in a location different from the regular distribution of other fraud cases, and the number of suspects is also quite different from that of other clusters, it may indicate the activities of new fraud gangs or other special circumstances, and the police need to conduct further in-depth investigations.
[0171] It is difficult to determine the distribution status of the above-mentioned simulation data, and it is difficult to perform accurate data verification using the verification model in Example 1. By allocating and continuously optimizing the communities, the accuracy of data verification is improved when it is difficult to determine the distribution rules of the original data.
[0172] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is limited by the attached embodiments and their equivalents.
Claims
1. Data verification management method based on artificial intelligence technology, including: An original database with multiple historical samples and a data sample to be detected is obtained, wherein: Conduct preliminary verification on the data samples to be tested to improve the data reliability of the samples to be tested; The non-numerical data in the historical samples and the samples to be tested are encoded by using a data encoding method, and are edited into a historical sample matrix and a sample matrix to be tested respectively; Through the historical sample processing method, multiple historical sample matrices are randomly distributed into several groups, and the average matrix of each group is calculated to establish an average matrix data set, thereby reducing the number of samples in the data set and the amount of calculation; The allocation factors between elements in the historical sample matrix are calculated by the allocation factor calculation method, the elements involved in the verification are screened out according to the screening method, and the elements involved in the verification are divided into several calculation groups, and the elements that can affect the verification results of the samples to be tested are screened and classified to improve the accuracy of data verification; A verification model is established based on the screened elements in the historical database, and the elements of different categories in the sample to be tested are brought into the verification model for data verification and the verification results are output. Through the mutual correlation between the data, the data to be tested is verified to determine whether the data is abnormal, and errors in new data are actively discovered to improve the accuracy of the verification results.
2. The data verification management method based on artificial intelligence technology according to claim 1 is characterized in that: The allocation factor calculation method includes: Group different elements and number the elements in the same group in order according to their numerical values; Select two elements that need to be calculated and calculate the difference between the sequence numbers according to the formula, specifically: ; in Indicates the difference between the sequence number of the Xth element and the sequence number of the Yth element in sample i. and represent the Xth and Yth elements in sample i respectively, and They represent the sequence number of the Xth element in sample i and the sequence number of the Yth element in sample i respectively; Then, the distribution factor between the two elements is calculated based on the difference between the multiple sequence numbers, specifically: ; in represents the distribution factor between the Xth element and the Yth element, represents the total number of samples, Represents the square of the difference between the sequence number of the Xth element and the sequence number of the Yth element in sample i.
3. The data verification management method based on artificial intelligence technology according to claim 1 is characterized in that: The allocation factor calculation method includes: For each element, the elements of each historical sample matrix are sequentially numbered according to the numerical value; When there are identical values in the data, the sequence numbers of the identical values are the average of their sorted positions; The distribution factor between two elements is calculated based on the difference between multiple sequence numbers and the average value, specifically: ; in represents the distribution factor between the Xth element and the Yth element, and represent the Xth and Yth elements in sample i respectively, and They represent the sequence number of the Xth element in sample i and the sequence number of the Yth element in sample i, respectively. and are the average values of the sequence numbers of the Xth and Yth elements, Indicates the total number of samples.
4. The data verification management method based on artificial intelligence technology according to claim 1 is characterized in that: The screening method comprises: Compare the allocation factor with a plurality of preset level factors of different sizes, divide the allocation factor into levels, and set a preset threshold for each level of the division; It is determined whether the allocation factors at the same level have crossover elements, and the elements corresponding to the multiple allocation factors with crossover elements are grouped as one calculation group, and the elements corresponding to the allocation factors without crossover elements are grouped into another calculation group.
5. The data verification management method based on artificial intelligence technology according to claim 1 is characterized in that: The establishment of the verification model is specifically as follows: Select a calculation group, calculate the mean and standard deviation of each element in the target calculation group, and combine the mean values to obtain the average feature matrix ; Calculate the mean of multiple allocation factors in the target calculation group as the correlation coefficient between the elements; Then calculate the covariance matrix between every two elements according to the formula , specifically: ; in represents the data at the location with coordinates (X, Y) in the covariance matrix, Indicates the correlation coefficient between the Xth element and the Yth element. In the same calculation group, the correlation coefficient between every two elements is the same. represents the standard deviation of the Xth element, represents the standard deviation of the Yth element; The verification model is established according to the probability density function, specifically: ; in Represents the sample matrix to be detected The abnormal prediction value, the value of which reflects the sample matrix to be detected The probability of abnormality is represents the average feature matrix, Represents the sample matrix to be detected and the average feature matrix The number of elements of Represents the covariance matrix The determinant of Represents the covariance matrix The inverse matrix of Represents conversion, converting a column vector into a row vector; The abnormal prediction value is compared with the prediction threshold. When the abnormal prediction value is greater than the prediction threshold, the data is judged to be normal, otherwise the data is judged to be abnormal.
6. The data verification management method based on artificial intelligence technology according to claim 1 is characterized in that: The establishment of the verification model is specifically as follows: S1: Select a calculation group, convert each sample into a calculation group sample matrix, randomly select several calculation group sample matrices, and use their data as the central sample matrices of several communities; S2: Calculate the vector gap between each calculation group sample matrix and each center sample matrix according to the formula, specifically: ; in Indicates the vector gap between the calculation group sample matrix i and the center sample matrix, represents the number of elements of the sample matrix of each calculation group in the target calculation group, It means calculating the value of the jth element in the sample matrix i of the group. Represents the value of the jth element of the center sample matrix; S3: assigning the sample matrix of the calculation group and the central sample with the smallest vector gap to the same community according to the vector gap; S4: Calculate the average matrix of each community and use the average matrix as the new central sample matrix; S5: Repeat steps S2-S4 until the central sample matrix no longer changes; S6: Calculate and compare the vector gap between the sample matrix to be detected and each central sample matrix, and compare the smallest vector gap with the gap threshold. When the vector gap is greater than the gap threshold, the data is judged to be abnormal, otherwise the data is judged to be normal.
7. The data verification management method based on artificial intelligence technology according to claim 1 is characterized in that: The preliminary verification is specifically: Back up the data samples to be tested; Check whether the data to be tested has missing values. If the data sample to be tested has missing values, prompt relevant personnel to conduct manual verification; Formulate several logical judgment rules to detect logical errors in the data samples to be tested, and obtain change records when relevant personnel make changes to abnormal data; When abnormal data and change records accumulate to a set level, the records will be sorted and sent to relevant personnel to facilitate the staff to sort and add or subtract logical judgment rules to improve the accuracy of preprocessing.
8. The data verification management method based on artificial intelligence technology according to claim 1 is characterized in that: The data encoding method comprises: Establish a feature category table and encode the feature categories; Extract text features from text data, classify and obtain codes according to set categories, mark and prompt text data with non-single classification and text data without categories, and make manual judgments by relevant personnel; For image data, a similarity check is performed on each category. When the similarity ratio between the category with the highest similarity and the category with the second highest similarity is higher than or equal to the ratio threshold, the image data is determined to belong to the category with the highest similarity and the encoding is obtained. When the similarity ratio between the category with the highest similarity and the category with the second highest similarity is lower than the ratio threshold, the image data is prompted and manually judged by relevant personnel.
9. Data verification management system based on artificial intelligence technology, characterized by: include: Data collection module: used to obtain the original database with multiple historical samples and the data samples to be tested; Preprocessing module: perform preliminary verification on the samples to be tested to improve the data reliability of the samples to be tested; Data processing module: Use data encoding method to encode the non-numerical data in historical samples and samples to be tested, and edit them into historical sample matrices and samples to be tested matrices respectively; Use historical sample processing method to randomly distribute multiple historical sample matrices into several groups, and calculate the average matrix of each group to establish an average matrix data set, reduce the number of samples in the data set, and reduce the amount of calculation; The allocation factors between elements in the historical sample matrix are calculated by the allocation factor calculation method, and the elements involved in the verification are screened out according to the screening method, and the elements involved in the verification are classified. The elements that can affect the verification results of the samples to be tested are screened and classified to improve the accuracy of data verification; the verification model is established based on the screened elements in the historical database; Data output module: The elements of different categories in the sample to be tested are brought into the verification model for data verification and the verification results are output to determine whether the data is abnormal.
10. A computer storage medium, characterized in that: The computer storage medium stores computer program instructions, which, when executed by a processor, implement the data verification management method based on artificial intelligence technology as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Customer basic archive data checking method based on machine learning and storage medium
CN118964597A
Automatic labeling method based on edge end traffic audio and video synchronization samples
CN112100435A
Large-scale power data anomaly detection method and detection system
CN115577756A
Fault early warning method for drainage system based on multi-source heterogeneous data
CN118506534A
System and method for modeling and quantifying regulatory capital, key risk indicators, probability of default, exposure at default, loss given default, liquidity ratios, and value at risk, within the areas of asset liability management, credit risk, market risk, operational risk, and liquidity risk for banks
US20150088783A1
Cited By
Intelligent auditing system for authentication process of enterprise management system
CN121119375A