Data Verification Management Method, System and Medium Based on Artificial Intelligence Technology
Through the data verification management method based on artificial intelligence, data encoding, allocation factor calculation and verification model are used to solve the problem of lack of initiative and forward-looking data verification in the existing technology, and efficient and accurate data anomaly detection is achieved.
Patent Information
- Application Number
- CN202510502662.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Existing data verification methods lack initiative and forward-looking in judging data abnormalities, making it difficult to find unseen errors, making it difficult to guarantee the quality of the database.
Using data verification and management methods based on artificial intelligence technology, data coding, allocation factor calculation and screening are carried out by obtaining historical samples and samples to be detected, verification is established, and data abnormalities are actively discovered.
It improves the accuracy and efficiency of data verification, has good initiative and forward-lookingness, can actively discover errors in new data, and optimize the verification results.
Smart Images

Figure CN120030202B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data verification, and specifically to a data verification management method, system and medium based on artificial intelligence technology. Background Art
[0002] Data verification refers to using means such as manual verification and program comparison to conduct a comprehensive and detailed review of various types of data, aiming to confirm the accuracy, integrity, consistency and reliability of the data. It plays a key role in various fields, can accurately eliminate incorrect, missing and contradictory data, ensure data quality, lay a solid foundation for activities such as the strategic planning of enterprises, the policy formulation of governments, and the analysis and judgment of scientific research projects, and help make scientific, reasonable and practical decisions.
[0003] When conducting data verification, it is necessary to use verification methods to simplify the verification process. For example, the verification method and storage medium for customer basic file data based on machine learning with the patent publication number CN118964597A includes obtaining the internal data of an enterprise and an external database; judging whether there is a unified social credit code in the internal data in the external database; the present invention determines whether there is the internal data of the enterprise in the external database by judging whether there is a unified social credit code of the internal data of the enterprise in the external database. If the unified social credit code exists in the external database, further conduct text information verification on the external data corresponding to the unified social credit code in the external database and the internal data of the enterprise. If the verification is consistent, it means that the internal and external data of the enterprise are consistent, otherwise they are inconsistent; adopting this method can effectively improve the accuracy of verification on the one hand and effectively improve the verification efficiency on the other hand.
[0004] Verifying newly obtained or newly registered data according to the data in the database can avoid database pollution and improve the data quality of the database. However, when the above method conducts verification, by comparing the items of external data and internal data to judge whether the data is abnormal, this judgment method can only find problems inconsistent with known rules or other system data, and it is difficult to judge unappeared errors, with poor initiative and foresight. Summary of the Invention
[0005] The purpose of the present invention is to provide a data verification management method, system and medium based on artificial intelligence technology to solve the problems raised in the above background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A data verification management method based on artificial intelligence technology, including:
[0007] Obtain an original database with multiple historical samples and a data sample to be detected;
[0008] Conduct a preliminary verification on the data samples to be detected to improve the data reliability of the samples to be detected;
[0009] Use a data encoding method to encode the non-numerical data in the historical samples and the samples to be detected, and respectively edit them into a historical sample matrix and a sample matrix to be detected;
[0010] Randomly allocate multiple historical sample matrices into several groups through a historical sample processing method, calculate the average matrix of each group, establish an average matrix data set, reduce the number of samples in the data set, and reduce the calculation amount;
[0011] Calculate the distribution factors between the elements in the historical sample matrix through a distribution factor calculation method, screen out the elements participating in the verification according to the screening method, divide the elements participating in the verification into several calculation groups, screen and classify the elements that can affect the verification results of the samples to be detected, and improve the accuracy of data verification;
[0012] Establish a verification model according to the elements screened in the historical database, respectively bring the elements of different classifications in the samples to be detected into the verification model for data verification and output the verification results, verify the data to be detected through the mutual association between the data, judge whether the data is abnormal, actively discover errors in the new data, and improve the accuracy of the verification results.
[0013] Preferably, the distribution factor calculation method includes:
[0014] Group according to different elements, and sequentially number the elements in the same group according to the numerical size;
[0015] Select two elements that need to be calculated, and calculate the difference between the sequential numbers according to the formula. Specifically:
[0016]
[0017] Among them represents the difference between the sequential numbers of the Xth element and the Yth element in sample i, and respectively represent the Xth element and the Yth element in sample i, and respectively represent the sequential number of the Xth element in sample i and the sequential number of the Yth element in sample i;
[0018] Then calculate the distribution factor between the two elements according to the differences between multiple sequential numbers. Specifically:
[0019]
[0020] Among them Represents the distribution factor between the Xth element and the Yth element. Represents the total number of samples. Represents the square of the difference between the sequence numbers of the Xth element and the Yth element in sample i.
[0021] Preferably, the method for calculating the distribution factor includes:
[0022] For each element, the elements of each historical sample matrix are sequentially numbered according to the numerical size.
[0023] When there are the same numerical values in the data, the sequence numbers of the same numerical values take the average of their sorted positions.
[0024] Calculate the distribution factor between two elements according to the difference between multiple sequence numbers and the average value, specifically:
[0025]
[0026] Where Represents the distribution factor between the Xth element and the Yth element. And Respectively represent the Xth element and the Yth element in sample i. And Respectively represent the sequence number of the Xth element in sample i and the sequence number of the Yth element in sample i. And Are respectively the average values of the sequence numbers of the Xth element and the Yth element. Represents the total number of samples.
[0027] Preferably, the screening method includes:
[0028] Compare the distribution factor with multiple preset rank factors of different sizes, classify the distribution factor into ranks, and set a preset threshold for each classified rank.
[0029] Judge whether the distribution factors in the same rank have cross elements, take the elements corresponding to the multiple distribution factors with cross elements as a calculation group, and divide the elements corresponding to the distribution factors without cross elements into another calculation group.
[0030] Preferably, the establishment of the verification model is specifically:
[0031] Select a calculation group, calculate the average value and standard deviation of each element in the target calculation group, and combine the average values to obtain the average feature matrix ;
[0032] Calculate the mean value of multiple distribution factors in the target calculation group as the correlation coefficient between elements.
[0033] Then calculate the covariance matrix between every two elements according to the formula , specifically as follows:
[0034]
[0035] where represents the data at the position of coordinates (X, Y) in the covariance matrix, represents the correlation coefficient between the Xth element and the Yth element. In the same calculation group, the correlation coefficients between every two elements are the same, represents the standard deviation of the Xth element, represents the standard deviation of the Yth element;
[0036] Establish a verification model according to the probability density function, specifically as follows:
[0037]
[0038] where represents the abnormal prediction value for the matrix of the sample to be detected. The magnitude of its value reflects the level of the abnormal probability of the matrix of the sample to be detected, represents the average feature matrix, represents the matrix of the sample to be detected and the average feature matrix in terms of the number of elements, represents the determinant of the covariance matrix , represents the inverse matrix of the covariance matrix , represents the conversion of converting a column vector into a row vector;
[0039] Compare the abnormal prediction value with the prediction threshold. When the abnormal prediction value is greater than the prediction threshold, the data is determined to be normal; otherwise, the data is determined to be abnormal.
[0040] Preferably, the establishment of the verification model is specifically as follows:
[0041] S1: Select a calculation group, convert each sample into a calculation group sample matrix, randomly select several of the calculation group sample matrices, and use their data as the central sample matrices of several communities;
[0042] S2: Calculate the vector gap between each calculation group sample matrix and each central sample matrix according to the formula, specifically as follows:
[0043]
[0044] where Indicates the vector gap between the sample matrix i of the calculation group and the central sample matrix, Indicates the number of elements of each sample matrix in the target calculation group, Indicates the value of the j-th element in the sample matrix i of the calculation group, Indicates the value of the j-th element of the central sample matrix;
[0045] S3: According to the vector gap, the sample matrix of the calculation group and the central sample with the smallest vector gap are assigned to the same community respectively;
[0046] S4: Calculate the average matrix of each community and use the average matrix as the new central sample matrix;
[0047] S5: Repeat steps S2 - S4 until the central sample matrix no longer changes;
[0048] S6: Calculate the vector gap between the sample matrix to be detected and each central sample matrix and make a comparison. Compare the smallest vector gap with the gap threshold. When the vector gap is greater than the gap threshold, it is determined that the data is abnormal, otherwise it is determined that the data is normal.
[0049] Preferably, the preliminary verification is specifically:
[0050] Back up the data sample to be detected;
[0051] Detect whether the data to be detected has missing values. When the data sample to be detected has missing values, prompt relevant personnel for manual verification;
[0052] Formulate several logical judgment rules, detect logical errors in the data sample to be detected, and obtain the change record when relevant personnel change the abnormal data;
[0053] When the abnormal data and change records accumulate to a set quantity level, organize the records and send them to relevant personnel, which is convenient for the staff to organize and increase or decrease the logical judgment rules to improve the accuracy of preprocessing.
[0054] Preferably, the data coding method includes:
[0055] Establish a feature category table and encode the feature categories;
[0056] Extract text features from text data, classify according to the set categories and obtain the encoding. Mark and prompt the text data with non-single classification and text data without categories, and make manual judgments by relevant personnel;
[0057] For image data, similarity detection is performed for each category. When the similarity ratio between the category with the highest similarity and the category with the second highest similarity is higher than or equal to the ratio threshold, it is determined that the image data belongs to the category with the highest similarity and the encoding is obtained. When the similarity ratio between the category with the highest similarity and the category with the second highest similarity is lower than the ratio threshold, a prompt is given for the image data and it is manually determined by relevant personnel.
[0058] Compared with the prior art, the beneficial effects of the present invention are:
[0059] A verification model is established through the selected elements. The elements of different classifications in the sample to be detected are respectively brought into the verification model for data verification and the verification results are output. Through the mutual association between the data, the data to be detected is verified to determine whether the data is abnormal, actively discover errors in new data, improve the accuracy of the verification results, and have good initiative and foresight.
[0060] At the same time, the distribution factors between the elements in the historical sample matrix are calculated through the distribution factor calculation method, and then the elements participating in the verification are selected according to the selection method and divided into several calculation groups. The elements that can affect the verification results of the sample to be detected are selected and classified, and the distribution factors can reflect the influence of different calculation groups in data verification, further optimizing the accuracy of the verification results. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a schematic diagram of the overall process of the data verification management method of the present invention;
[0062] Figure 2 It is a schematic diagram of the process of the distribution factor calculation method in the first embodiment of the present invention;
[0063] Figure 3 It is a schematic diagram of the process of the distribution factor calculation method in the second embodiment of the present invention;
[0064] Figure 4 It is a schematic diagram of the process of the selection method in the present invention;
[0065] Figure 5 It is a schematic diagram of the process of establishing the verification model in the present invention;
[0066] Figure 6 It is a schematic diagram of the process of establishing the verification model in the third embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0067] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0068] In this application, for the convenience of understanding, the method steps used do not necessarily need to be executed in the order of the steps in this embodiment during actual operation. In some other embodiments, these steps can be carried out synchronously or the order can be changed.
[0069] Embodiment 1:
[0070] Verifying newly acquired or newly registered data according to the data in the database can avoid database pollution, improve the data quality of the database, and through the mutual association between data, verify the data to be detected, judge whether the data is abnormal, actively discover errors in the new data, improve the accuracy of the verification results, and have good initiative and foresight.
[0071] As Figure 1 、 Figure 2 、 Figure 4 and Figure 5 shown, the present invention provides a technical solution: a data verification management method based on artificial intelligence technology, including:
[0072] Obtain an original database with multiple historical samples and a data sample to be detected;
[0073] Conduct a preliminary verification on the data sample to be detected to improve the data reliability of the sample to be detected;
[0074] Use a data encoding method to encode the non-numerical data in the historical samples and the data sample to be detected, and respectively compile them into a historical sample matrix and a data sample matrix to be detected;
[0075] Randomly assign multiple historical sample matrices into several groups through a historical sample processing method, calculate the average matrix of each group, and establish an average matrix data set to reduce the sample quantity of the data set and reduce the calculation amount;
[0076] Calculate the distribution factors between the elements in the historical sample matrix through a distribution factor calculation method, screen out the elements participating in the verification according to a screening method, and divide the elements participating in the verification into several calculation groups, and screen and classify the elements that can affect the verification results of the data sample to be detected to improve the accuracy of data verification;
[0077] Establish a verification model based on the screened elements in the historical database, bring the elements of different classifications in the sample to be detected into the verification model for data verification respectively, and output the verification results. Through the mutual association between data, verify the data to be detected, judge whether the data is abnormal, actively discover errors in new data, and improve the accuracy of the verification results.
[0078] It should be noted that when there are missing or difficult-to-read data in the original database during verification, some samples in the original database can be abandoned, or methods such as mean and median can be used for supplementation to improve the quality of the original database. The specific supplementation method can be selected according to the data type and will not be elaborated here.
[0079] Preliminary verification, specifically:
[0080] Back up the data sample to be detected;
[0081] Detect whether the data to be detected has missing values. When the data sample to be detected has missing values, prompt relevant personnel for manual verification;
[0082] Formulate several logical judgment rules, detect logical errors in the data sample to be detected, and obtain the change records when relevant personnel change the abnormal data;
[0083] When the abnormal data and change records accumulate to a set quantity level, organize the records and send them to relevant personnel, which is convenient for the staff to organize and increase or decrease the logical judgment rules, and improve the accuracy of preprocessing.
[0084] It should be noted that the detection of missing values and the formulation of logical rules are both existing technologies and will not be elaborated here. For example, setting an upper limit for age, setting that the occurrence time of an event must be earlier than the end time, etc. as judgment rules. When inputting new data, whenever the new data is determined to be abnormal data, record the data and the changes made by relevant personnel to the data, and then conduct batch sorting, which helps relevant personnel set new logical rules or delete incorrect logical rules, can improve the effect of preliminary verification, and thus improve the overall verification efficiency.
[0085] The data coding method includes:
[0086] Establish a feature category table and encode the feature categories;
[0087] Extract text features from text data, classify according to the set categories and obtain the encoding. For text data with non-single classifications and text data without categories, mark and prompt, and conduct manual judgment by relevant personnel;
[0088] For image data, similarity detection is performed for each category. When the similarity ratio between the category with the highest similarity and the category with the second-highest similarity is higher than or equal to the ratio threshold, it is determined that the image data belongs to the category with the highest similarity and the encoding is obtained. When the similarity ratio between the category with the highest similarity and the category with the second-highest similarity is lower than the ratio threshold, a prompt is given for the image data, and it is manually determined by relevant personnel.
[0089] It should be noted that for text data, a feature category table is established. For example, in a policing system, the case types are encoded as follows: theft is encoded as 1, robbery is encoded as 2, violent crime is encoded as 3, and fraud is encoded as 4. According to actual needs, the value range can also be changed. For example, theft can be encoded as 10, robbery as 20, violent crime as 30, and fraud as 40. If the case type in the new data is endangering public safety and appears in the feature category table, the data can be marked, and relevant personnel can change the case type of the new data or modify the feature category table. For picture data, a representative picture can be selected for each category in the feature category table, and then the pictures in the new data are compared for similarity (the specific similarity comparison method is an existing technology such as perceptual hashing, convolutional neural network, etc., which will not be elaborated here). For example, the similarities between the pictures in the new data and the categories in the feature category table are 90%, 20%, 3%, 4%, 1%, and 0 respectively, and the set ratio threshold is 4. At this time, the ratio of 90% to 20% is 4.5, which is greater than the ratio threshold. Then, the number of the category with a similarity of 90% can be used as the number of the image data in the new data.
[0090] As Figure 2 shown, the allocation factor calculation method includes:
[0091] Group by different elements and sequentially number the elements in the same group according to the numerical size;
[0092] Select two elements that need to be calculated, and calculate the difference between the sequential numbers according to the formula. Specifically:
[0093]
[0094] where represents the difference between the sequential numbers of the Xth element and the Yth element in sample i, and respectively represent the Xth element and the Yth element in sample i, and respectively represent the sequential number of the Xth element in sample i and the sequential number of the Yth element in sample i;
[0095] Calculate the distribution factor between two elements based on the difference between multiple sequential labels. Square the difference to eliminate the influence of positive and negative numbers on the data, and then calculate the sum of the squares of several differences. This summation term reflects the overall degree of the rank difference between two variables. The larger the sum of the squared differences, the more obvious the rank difference between the two variables and the weaker the correlation. Normalize the sum of the squared rank differences as the denominator to further adjust the weight of the sum of the squared rank differences, enabling the correlation coefficient to be reasonably valued within the range of [−1,1], thereby accurately reflecting the degree of rank correlation between two variables. Specifically:
[0096]
[0097] Among them represents the distribution factor between the Xth element and the Yth element, represents the total number of samples, represents the square of the difference between the sequential label of the Xth element and the sequential label of the Yth element in sample i.
[0098] It should be noted that for the convenience of calculation, the following simulated data is used
[0099] Suppose we monitor the operation data of the fire protection facilities in a large shopping mall. The data includes four key indicators: fire pressure (MPa), temperature of fire protection equipment (℃), signal strength of smoke detectors (dimensionless), and status of fire doors (0 for closed and 1 for open). The following is the historical data of the past 5 regular inspections, as shown in Table 1:
[0100] Table 1: Original data
[0101] Sample Fire pressure Equipment temperature Signal strength Fire door status 1 0.80 31 20 0 2 0.75 28 18 0 3 0.90 32 22 0 4 0.70 26 16 0 5 0.85 30 20 0
[0102] Taking the fire pressure and equipment temperature as an example, calculate the distribution factor as follows:
[0103] First, number the data of the fire pressure and equipment temperature as (3, 2, 5, 1, 4) and (4, 2, 5, 1, 3) respectively, and then calculate , which are (-1, 0, 0, 0, 1) respectively. At this time, calculate according to the formula = 0.9. This value is relatively close to 1, indicating a strong rank order between the two. The calculation method for the remaining elements is the same and will not be elaborated here.
[0104] As Figure 4 shown, the screening method includes:
[0105] Compare the distribution factor with multiple preset rank factors of different sizes, divide the distribution factor into ranks, and set a preset threshold for each divided rank;
[0106] Determine whether the distribution factors at the same level have cross elements, and take the elements corresponding to multiple distribution factors with cross elements as a calculation group, and divide the elements corresponding to the distribution factors without cross elements into another calculation group.
[0107] It should be noted that for the convenience of calculation, the following simulated data is set:
[0108] There are several historical samples in the original database, and the samples contain 4 elements. After calculation, the distribution factor between the first element and the second element is 0.22, the distribution factor between the first element and the third element is 0.63, the distribution factor between the first element and the fourth element is 0.11, the distribution factor between the second element and the third element is 0.12, the distribution factor between the second element and the fourth element is 0.14, and the distribution factor between the third element and the fourth element is 0.32. The set level factors are 0.2 and 0.5. The distribution factors less than 0.2 are assigned to the first level, the distribution factors greater than or equal to 0.2 and less than 0.5 are assigned to the second level, and the distribution factors greater than or equal to 0.5 are assigned to the third level.
[0109] Taking the second level as an example, at this time in the second level, the distribution factor between the second element and the third element is 0.12, the distribution factor between the second element and the fourth element is 0.14, and the distribution factor between the first element and the fourth element is 0.11. Among them, the two distribution factors 0.14 and 0.12 have cross elements (the second element), and the two distribution factors 0.14 and 0.11 have cross elements (the fourth element). Therefore, the elements corresponding to the three distribution factors 0.14, 0.12, and 0.11 are a calculation group, that is, the second element, the third element, and the fourth element are all in one calculation group, that is, there is only one calculation group in the second level. The same method is used for the calculation group division in other levels and will not be elaborated here.
[0110] As Figure 5 shown, a verification model is established, specifically:
[0111] Select a calculation group, calculate the average value and standard deviation of each element in the target calculation group, and combine the average values to obtain the average feature matrix ;
[0112] Calculate the mean value of multiple distribution factors in the target calculation group as the correlation coefficient between elements;
[0113] Then calculate the covariance matrix between every two elements according to the formula , specifically:
[0114]
[0115] wherein represents the data at the position with coordinates (X, Y) in the covariance matrix, represents the correlation coefficient between the Xth element and the Yth element. In the same calculation group, the correlation coefficient between every two elements is the same, and when X and Y are equal, = 1, represents the standard deviation of the Xth element, represents the standard deviation of the Yth element;
[0116] A verification model is established according to the probability density function, specifically:
[0117]
[0118] wherein represents the abnormal prediction value of the sample matrix to be detected , and the magnitude of its value reflects the level of the abnormal probability of the sample matrix to be detected . represents the average feature matrix, represents the sample matrix to be detected and the average feature matrix represents the number of elements, represents the determinant of the covariance matrix , represents the inverse matrix of the covariance matrix , represents the conversion of converting a column vector into a row vector;
[0119] The abnormal prediction value is compared with the prediction threshold. When the abnormal prediction value is greater than the prediction threshold, the data is determined to be normal, otherwise the data is determined to be abnormal.
[0120] It should be noted that for the convenience of calculation, assume that the data of a calculation group is as follows:
[0121] Suppose the maintenance situation of the fire protection facilities in an industrial park is evaluated. In one calculation group, the three elements are respectively: the fire equipment inspection frequency (times / month), the fire training duration (hours / quarter), and the fire drill participation rate (%). And the preset threshold corresponding to this calculation group is 0.001, as shown in Table 2:
[0122] Table 2: Calculation Group Data One
[0123] Sample Inspection frequency Training duration Drill participation rate 1 8 10 85 2 5 6 70 3 12 15 90 4 3 4 60 5 10 12 88 6 6 8 75
[0124] The mean of the distribution factors of this calculation group is approximately 0.82 after calculation (the calculation method is the same as the above-mentioned distribution factor calculation method and will not be demonstrated here); the standard deviations of the three elements are 3.33, 4.04, and 11.75 respectively; then calculate the values of each element in the covariance matrix according to the formula. Taking as an example, the calculation result is 3.33×3.33×1≈11.09. Continue to calculate the values in the remaining positions of the covariance matrix to obtain the covariance = .
[0125] Suppose the sample matrix to be detected = . Substitute and into to calculate ≈3.63×10 −10 . At this time, is less than the preset threshold, so it is determined that the data to be detected is abnormal.
[0126] Example 2:
[0127] In Example 1, when calculating the distribution factor, it is necessary to first sequentially label the elements of each historical sample matrix. However, when there are the same numerical values in the elements, it is difficult to determine the labeling order, and it is difficult to ensure the accuracy of the distribution factor calculation. This example provides another distribution factor calculation method to improve the calculation accuracy of the distribution factor when there are the same numerical values in the elements.
[0128] As Figure 3 shown, the distribution factor calculation method includes:
[0129] For each type of element, sequentially label the elements of each historical sample matrix according to the numerical size;
[0130] When there are the same numerical values in the data, the sequential labels of the same numerical values take the average of their sorted positions;
[0131] Calculate the distribution factor between two elements according to the difference between multiple sequential labels and the average value. Specifically:
[0132]
[0133] Among them represents the distribution factor between the Xth element and the Yth element, and respectively represent the Xth element and the Yth element in sample i, and respectively represent the sequential label of the Xth element in sample i and the sequential label of the Yth element in sample i, and are the average values of the sequence numbers of the Xth element and the Yth element respectively, indicating the total number of samples.
[0134] It should be noted that for the convenience of calculation, the simulated data is set as follows:
[0135] Suppose the data of the two elements in different samples are (3, 5, 5, 8, 10) and (4, 6, 6, 9, 11) respectively.
[0136] First, perform sequence numbering, and take the average of the positions after sorting for the sequence numbers of the same values. The results are (1, 2.5, 2.5, 4, 5) and (1, 2.5, 2.5, 4, 5). Calculate and respectively, and obtain = = 3.
[0137] According to it can be calculated that = 1.
[0138] It can calculate the distribution factor when there are elements with the same value, improving the calculation accuracy of the distribution factor.
[0139] Example 3:
[0140] In Example 1, a high-probability density function is used to establish a verification model, which requires the data arrangement in the original database to conform to the Gaussian distribution. When the data volume is large and the data dimension is large, it is difficult to determine the data distribution form, and the verification result of this verification model will deviate, and the use limitation is relatively large. This example provides another verification model to improve the verification accuracy when it is difficult to determine the original data distribution rule.
[0141] As Figure 6 shown, establish a verification model, specifically:
[0142] S1: Select a calculation group, convert each sample into a calculation group sample matrix, randomly select several of the calculation group sample matrices, and use their data as the central sample matrices of several communities;
[0143] S2: Calculate the vector difference between each calculation group sample matrix and each central sample matrix according to the formula, specifically:
[0144]
[0145] where represents the vector difference between the calculation group sample matrix i and the central sample matrix, represents the number of elements of each calculation group sample matrix in the target calculation group, represents the numerical value of the j-th element in the calculation group sample matrix i, represents the numerical value of the j-th element in the central sample matrix;
[0146] S3: According to the vector gap, assign the calculation group sample matrix and the central sample with the smallest vector gap to the same community respectively;
[0147] S4: Calculate the average matrix of each community and use the average matrix as the new central sample matrix;
[0148] S5: Repeat steps S2 - S4 until the central sample matrix no longer changes;
[0149] S6: Calculate the vector gap between the sample matrix to be detected and each central sample matrix and make a comparison. Compare the smallest vector gap with the gap threshold. When the vector gap is greater than the gap threshold, it is determined that the data is abnormal; otherwise, the data is determined to be normal.
[0150] It should be noted that for the convenience of calculation, the following data is set:
[0151] Suppose the police have collected the crime case data in a certain area for a period of time. The records in one calculation group include the longitude and latitude of the crime location, the crime time, the type of crime (such as theft, robbery, violent crime, fraud, etc.), the number of suspects, etc. For the convenience of processing, the types of crimes are encoded: theft is encoded as 1, robbery is encoded as 2, violent crime is encoded as 3, and fraud is encoded as 4. The calculation group data is shown in Table 3:
[0152] Table 3: Calculation Group Data Two
[0153] Case number Longitude Latitude Type code Number of personnel 1 116.38 39.95 1 1 2 116.40 39.98 1 2 3 116.35 39.92 1 1 4 116.50 40.00 2 3 5 116.55 40.05 2 4 6 116.48 39.98 2 3 7 116.20 39.80 3 1 8 116.18 39.75 3 1 9 116.25 39.85 3 2
[0154] Randomly select the data of 3 samples as the central sample matrix. Suppose the selected central sample matrix is: [116.38, 39.95, 1, 1] (corresponding to case 1), [116.50, 40.00, 2, 3] (corresponding to case 4), [116.20, 39.80, 3, 1] (corresponding to case 7).
[0155] Calculate the vector gap between the calculation group sample matrix and each central sample matrix. For example, for case 2 ([116.40, 39.98, 1, 2]):
[0156] The vector gap with the first central sample matrix [116.38, 39.95, 1, 1] is: ∣116.40 - 116.38∣ + ∣39.98 - 39.95∣ + ∣1 - 1∣ + ∣2 - 1∣ = 0.02 + 0.03 + 0 + 1 = 1.05;
[0157] The vector difference from the second central sample matrix [116.50, 40.00, 2, 3] is: ∣116.40 - 116.50∣ + ∣39.98 - 40.00∣ + ∣1 - 2∣ + ∣2 - 3∣ = 0.1 + 0.02 + 1 + 1 = 2.12;
[0158] The vector difference from the third central sample matrix [116.20, 39.80, 3, 1] is: ∣116.40 - 116.20∣ + ∣39.98 - 39.80∣ + ∣1 - 3∣ + ∣2 - 1∣ = 0.2 + 0.18 + 2 + 1 = 3.38;
[0159] Since the vector difference of Case 2 from the first central sample matrix is the smallest, Case 2 belongs to the first community.
[0160] Repeat this process to assign all data points to the corresponding communities:
[0161] The first community: [116.38, 39.95, 1, 1], [116.40, 39.98, 1, 2], [116.35, 39.92, 1, 1];
[0162] The second community: [116.50, 40.00, 2, 3], [116.55, 40.05, 2, 4], [116.48, 39.98, 2, 3];
[0163] The third community: [116.20, 39.80, 3, 1], [116.18, 39.75, 3, 1], [116.25, 39.85, 3, 2].
[0164] Calculate the average matrix of each community as the new central sample matrix. The new central sample matrices of the three communities are [116.38, 39.95, 1, 1.33], [116.51, 40.01, 2, 3.33], [116.21, 39.80, 3, 1.33]. Then continue to assign communities to each case. In this embodiment, the central sample matrix calculated subsequently does not change, so the above assignment result is the final assignment result.
[0165] Suppose a new case appears, the longitude and latitude of the crime scene is [117.00, 40.50], the case type code is 4 (fraud), and the number of suspects is 5.
[0166] Calculate the vector difference of the new data from each central sample matrix:
[0167] The vector gap to the first central sample matrix [116.38, 39.95, 1, 1.33] is: ∣117.00−116.38∣+∣40.50−39.95∣+∣4−1∣+∣5−1.33∣ = 0.62 + 0.55 + 3 + 3.67 = 7.84;
[0168] The vector gap to the second central sample matrix [116.51, 40.01, 2, 3.33] is: ∣117.00−116.51∣+∣40.50−40.01∣+∣4−2∣+∣5−3.33∣ = 0.49 + 0.49 + 2 + 1.67 = 4.65;
[0169] The vector gap to the third central sample matrix [116.21, 39.80, 3, 1.33] is: ∣117.00−116.21∣+∣40.50−39.80∣+∣4−3∣+∣5−1.33∣ = 0.79 + 0.7 + 1 + 3.67 = 6.16;
[0170] Assume that a gap threshold of 3.5 is set according to the actual situation. The vector gap between the new data and the nearest second central sample matrix is 4.65, which is greater than the gap threshold. Therefore, it can be determined that the data of this new case is abnormal data. Since this case is of the fraud type, it appears at a location different from the normal distribution of other fraud cases, and the number of suspects also varies greatly from that of other clusters, which may indicate the activities of a new fraud gang or other special circumstances, and the police need to conduct further in-depth investigations.
[0171] It is difficult to determine the distribution state of the above simulation data, and it is difficult for the verification model in the first embodiment to conduct accurate data verification. By the allocation and continuous optimization of the community, the accuracy of data verification is improved when it is difficult to determine the distribution rule of the original data.
[0172] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended embodiments and their equivalents.
Claims
1. A data verification and management method based on artificial intelligence technology, comprising: Obtain the original database with multiple historical samples and the data samples to be detected. Both the original database and the data samples to be detected contain text data, image data, and numerical data. It is characterized in that: Conduct a preliminary verification on the data samples to be detected to improve the data reliability of the samples to be detected; Use a data encoding method to encode the non-numerical data in the historical samples and the samples to be detected, and respectively edit them into a historical sample matrix and a sample matrix to be detected; Randomly allocate multiple historical sample matrices into several groups through a historical sample processing method, calculate the average matrix of each group, and establish an average matrix data set to reduce the number of samples in the data set and reduce the computational amount; Calculate the distribution factors between the elements in the historical sample matrix through a distribution factor calculation method, screen out the elements participating in the verification according to the screening method, and divide the elements participating in the verification into several calculation groups, and screen and classify the elements that can affect the verification result of the sample to be detected to improve the accuracy of data verification; Establish a verification model according to the elements screened in the historical database, bring the elements of different classifications in the sample to be detected into the verification model respectively for data verification and output the verification result. Through the mutual association between the data, verify the data to be detected, judge whether the data is abnormal, actively discover errors in the new data, and improve the accuracy of the verification result; The screening method includes: Compare the distribution factor with multiple preset rank factors of different sizes, divide the distribution factor into ranks, and set a preset threshold for each divided rank; Judge whether the distribution factors in the same rank have cross elements, take the elements corresponding to the multiple distribution factors with cross elements as a calculation group, and divide the elements corresponding to the distribution factors without cross elements into another calculation group; The establishment of the verification model is specifically: Select a calculation group, calculate the mean and standard deviation of each element in the target calculation group, combine the means, and obtain the average feature matrix ; Calculate the mean value of multiple distribution factors in the target calculation group as the correlation coefficient between the elements; Then calculate the covariance matrix between every two elements according to the formula , specifically as follows: ; wherein represents the data at the position with coordinates (X, Y) in the covariance matrix, represents the correlation coefficient between the Xth element and the Yth element. In the same calculation group, the correlation coefficients between every two elements are the same, represents the standard deviation of the Xth element, represents the standard deviation of the Yth element; Establish a verification model according to the probability density function, specifically: ; Among them represents the abnormal prediction value of the sample matrix to be detected , and the magnitude of its value reflects the abnormal probability of the sample matrix to be detected . represents the average feature matrix represents the sample matrix to be detected and the average feature matrix in terms of the number of elements represents the determinant of the covariance matrix . represents the inverse matrix of the covariance matrix . represents a transformation that converts a column vector into a row vector; Compare the abnormal prediction value with the prediction threshold. When the abnormal prediction value is greater than the prediction threshold, it is determined that the data is normal, otherwise it is determined that the data is abnormal.
2. The data verification management method based on artificial intelligence technology according to claim 1, characterized in that: The distribution factor calculation method includes: Group according to different elements, and sequentially number the elements in the same group according to the numerical size; Select two elements that need to be calculated, and calculate the difference between the sequential numbers according to the formula, specifically: ; where represents the difference between the sequence numbers of the X-th element and the Y-th element in sample i, and represent the X-th element and the Y-th element in sample i respectively, and represent the sequence number of the X-th element and the sequence number of the Y-th element in sample i respectively; Then calculate the distribution factor between the two elements according to the differences between multiple sequential numbers, specifically: ; wherein represents the distribution factor between the Xth element and the Yth element, represents the total number of samples, represents the square of the difference between the sequence numbers of the Xth element and the Yth element in sample i.
3. The data verification management method based on artificial intelligence technology according to claim 1, wherein: The distribution factor calculation method includes: For each element, sequentially number the elements of each historical sample matrix according to the numerical size; When there are the same numerical values in the data, the sequential numbers of the same numerical values take the average value of their sorted positions; Calculate the distribution factor between the two elements according to the differences between multiple sequential numbers and the average value, specifically: ; where represents the distribution factor between the Xth element and the Yth element, and respectively represent the Xth element and the Yth element in sample i, and respectively represent the sequence number of the Xth element and the sequence number of the Yth element in sample i, and are respectively the average values of the sequence numbers of the Xth element and the Yth element, represents the total number of samples.
4. The data verification and management method based on artificial intelligence technology according to claim 1, characterized in that: The establishment of the verification model is specifically: S1: Select a calculation group, convert each sample into a calculation group sample matrix, randomly select several calculation group sample matrices among them, and use their data as the central sample matrices of several communities; S2: Calculate the vector gap between each calculation group sample matrix and each central sample matrix according to the formula, specifically as follows: ; Among them represents the vector gap between the calculation group sample matrix i and the central sample matrix, represents the number of elements of each calculation group sample matrix in the target calculation group, represents the value of the j-th element in the calculation group sample matrix i, represents the value of the j-th element of the central sample matrix; S3: Assign the calculation group sample matrices and the central sample matrix with the smallest vector gap to the same community according to the vector gap; S4: Calculate the average matrix of each community and use the average matrix as the new central sample matrix; S5: Repeat steps S2 - S4 until the central sample matrix no longer changes; S6: Calculate and compare the vector gap between the sample matrix to be detected and each central sample matrix, compare the smallest vector gap with the gap threshold. When the vector gap is greater than the gap threshold, it is determined that the data is abnormal, otherwise it is determined that the data is normal.
5. The data verification management method based on artificial intelligence technology according to claim 1, characterized in that: The preliminary verification is specifically as follows: Back up the data samples to be detected; Detect whether the data to be detected has missing values. When the data sample to be detected has missing values, prompt the relevant personnel for manual verification; Formulate several logical judgment rules, detect logical errors in the data samples to be detected, and obtain the change records when the relevant personnel modify the abnormal data; When the abnormal data and change records accumulate to a set quantity level, organize the records and send them to the relevant personnel to facilitate the staff to organize and increase or decrease the logical judgment rules to improve the accuracy of preprocessing.
6. The data verification management method based on artificial intelligence technology according to claim 1, characterized in that: The data encoding method includes: Establish a feature category table and encode the feature categories; Extract text features from text data, classify according to the set categories and obtain the encoding. For text data with non - single classifications and text data without categories, mark and prompt, and make a manual determination by the relevant personnel; For image data, perform similarity detection for each category. When the similarity ratio between the category with the highest similarity and the category with the second - highest similarity is higher than or equal to the ratio threshold, determine that the image data belongs to the category with the highest similarity and obtain the encoding. When the similarity ratio between the category with the highest similarity and the category with the second - highest similarity is lower than the ratio threshold, prompt the image data and make a manual determination by the relevant personnel.
7. A data verification management system based on artificial intelligence technology, characterized in that: Use the data verification and management method based on artificial intelligence technology described in any one of claims 1 - 6, including: Data collection module: used to obtain the original database with multiple historical samples and the data samples to be detected; Pre - processing module: perform preliminary verification on the samples to be detected to improve the data reliability of the samples to be detected; Data processing module: encode the non - numerical data in the historical samples and the samples to be detected using the data encoding method, and respectively edit them into a historical sample matrix and a sample matrix to be detected; randomly assign multiple historical sample matrices into several groups through the historical sample processing method, calculate the average matrix of each group, establish an average matrix data set, reduce the number of samples in the data set, and reduce the calculation amount; calculate the distribution factors between the elements in the historical sample matrix through the distribution factor calculation method, screen out the elements participating in the verification according to the screening method, classify the elements participating in the verification, screen and classify the elements that can affect the verification result of the sample to be detected, and improve the accuracy of data verification; establish a verification model according to the elements screened in the historical database. Data output module: Elements of different classifications in the sample to be detected are respectively input into the verification model for data verification and the verification results are output to determine whether the data is abnormal.
8. A computer storage medium, characterized in that: The computer storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the data verification management method based on artificial intelligence technology described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Customer basic archive data checking method based on machine learning and storage medium
CN118964597A
Automatic labeling method based on edge end traffic audio and video synchronization samples
CN112100435A
Large-scale power data anomaly detection method and detection system
CN115577756A