A data check rule optimization method, device and medium

CN120892472BActive Publication Date: 2026-09-25SHANDONG INSPUR DIGITAL BUSINESS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510895140.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2026-09-25
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

但仍存在以下技术问题:首先是复杂数据处理适应性低的问题,由于行政类数据字段之间业务逻辑关系复杂,但现有模型只支持建模线性关系数据,且规则树越深,易导致过拟合,对数据微小变化敏感,则导致数据校验处理不稳定;并且,无法支持实时更新校验规则及优化思考,且缺乏人工反馈闭环处理模式,补报期时,无法启用或停用某些特殊校验规则,会降低数据征集质量,造成迟报漏报情况的发生

Benefits of technology

[0044]通过过自动化的全局与局部异常检测生成初始规则库,替代传统人工定义规则,显著降低维护成本,避免因人工疏漏导致的误判漏检。通过挖掘字段关联规则优化初始规则库,有效解决现有模型对行政类复杂数据非线性关系建模不足的问题,提升对字段间复杂业务逻辑的适应性,降低过拟合风险,保障校验稳定性。引入人工反馈闭环机制,支持用户对规则库进行维护与优化,结合校验数据之间的数据分布触发规则库的动态更新策略,实现规则库随业务变化进行实时迭代。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892472B_ABST
    Figure CN120892472B_ABST
Patent Text Reader

Abstract

The application discloses a data verification rule optimization method, device and medium, the method comprising: obtaining an administrative data set, performing global anomaly detection and local anomaly detection on the administrative data set to determine global rule feature information and local rule feature information corresponding to abnormal data in the administrative data set; constructing an initial rule library corresponding to the administrative data set according to the global rule feature information and the local rule feature information; determining association rules between data sample fields in the administrative data set to perform association optimization on the initial rule library; receiving artificial feedback information of a user for the initial rule library, optimizing the initial rule library through the artificial feedback information to generate an optimized rule library; performing data verification on to-be-verified administrative data according to the optimized rule library, and iteratively optimizing the optimized rule library according to data distribution between obtained verification data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data quality management technology, specifically to a data verification rule optimization method, device, and medium. Background Technology

[0002] Traditional data validation rules (such as regular expressions, date format checks, dictionary checks, range checks, business logic rules, consistency checks, and integrity checks) rely on manual definition, resulting in high maintenance costs and a tendency to produce false positives, missed detections, or incomplete coverage of data anomalies. Currently, administrative datasets change with business needs; if data validation rules are not updated promptly, it will lead to a large number of missed erroneous data points and an increase in late-reported data, thus reducing the quality of data assessment.

[0003] While existing simple machine learning models have improved the speed of data verification and processing without requiring complex data preprocessing, they still have the following technical problems: First, they have low adaptability to complex data processing. Due to the complex business logic relationships between administrative data fields, existing models only support modeling linear relationship data. Furthermore, the deeper the rule tree, the more prone it is to overfitting. They are also sensitive to small changes in data, leading to instability in data verification and processing. In addition, they cannot support real-time updates of verification rules and optimization considerations, and they lack a closed-loop processing mode for human feedback. During the supplementary reporting period, it is impossible to enable or disable certain special verification rules, which will reduce the quality of data collection and cause late or missed reporting. Summary of the Invention

[0004] To address the aforementioned issues, this application proposes a data validation rule optimization method, comprising:

[0005] Obtain an administrative dataset, and perform global and local anomaly detection on the administrative dataset to determine the global and local rule feature information corresponding to the abnormal data in the administrative dataset.

[0006] Based on the global rule feature information and the local rule feature information, an initial rule base corresponding to the administrative dataset is constructed;

[0007] Determine the association rules between the fields of each data sample in the administrative dataset in order to optimize the association of the initial rule base;

[0008] Receive user feedback on the initial rule base, and optimize the initial rule base based on the feedback to generate an optimized rule base;

[0009] The administrative data to be verified is validated based on the optimized rule base, and the optimized rule base is iteratively optimized based on the data distribution among the obtained validation data.

[0010] In one implementation of this application, global anomaly detection and local anomaly detection are performed on the administrative dataset to determine the global rule feature information and local rule feature information corresponding to the abnormal data in the administrative dataset, specifically including:

[0011] The administrative dataset is subjected to global anomaly detection using the isolated forest detection algorithm. A first subsample is randomly extracted from the administrative dataset, an isolated tree corresponding to the first subsample is constructed, and the abnormal data in the administrative dataset and the global rule feature information corresponding to the abnormal data are determined based on the average path length of each data point in the isolated tree.

[0012] The local outlier factor algorithm is used to detect local anomalies in the administrative dataset. Several second sub-samples are extracted according to preset rule features. Based on the local outlier factor value corresponding to each data point in the second sub-sample, the local rule feature information corresponding to the abnormal data in the administrative dataset is determined.

[0013] In one implementation of this application, constructing the isolated tree corresponding to the first sub-sample specifically includes:

[0014] Perform a feature segmentation operation on the first subsample. The feature segmentation operation includes: determining a rule feature and a separation value corresponding to the rule feature; segmenting the first subsample according to the separation value to obtain a sample subset after segmentation and a tree node corresponding to the sample subset.

[0015] For the sample subset, the feature segmentation operation is repeated until all data points in the first subsample are isolated or the tree node reaches the maximum depth of the isolated tree, thus obtaining the isolated tree corresponding to the first subsample.

[0016] In one implementation of this application, the abnormal data in the administrative dataset and the global rule feature information corresponding to the abnormal data are determined based on the average path length of each data point in the isolated tree, specifically including:

[0017] The anomaly score corresponding to each data point is determined based on the average path length of each data point in the isolated tree.

[0018] Based on whether the abnormal score is within a preset abnormal score range, abnormal data in the administrative dataset is determined, and based on the path node of the abnormal data in the isolated tree, the global rule feature information corresponding to the abnormal data is determined.

[0019] In one implementation of this application, determining the local rule feature information corresponding to the abnormal data in the administrative dataset based on the local outlier factor values ​​corresponding to each data point in the second subsample specifically includes:

[0020] The reachability distance for each data point in the second subsample is calculated using the following formula:

[0021] reach-dist k (p,o)=max(k-distance(o),distance(p,o))

[0022] Where o represents a neighboring point of data point p, k-distance(o) represents the k-distance of neighboring point o, and distance(p,o) represents the distance between data point p and neighboring point o;

[0023] Calculate the local reachability density corresponding to the data point based on the reachability distance;

[0024] The local outlier factor value corresponding to the data point is determined based on the mean local reachability density of the neighboring points of the data point and the local reachability density.

[0025] Data points whose local outlier factor values ​​are greater than a preset threshold are considered as abnormal data in the administrative dataset, and local rule feature information corresponding to the abnormal data is determined according to the rule features.

[0026] In one implementation of this application, calculating the local reachability density corresponding to the data point based on the reachability distance specifically includes:

[0027] The local reachability density corresponding to the data point is calculated using the following formula:

[0028]

[0029] Where, N k (p) represents the number of neighboring points of data point p;

[0030] The local outlier factor value of the data point is determined using the following formula, based on the mean local reachability density of the data point's neighboring points and the local reachability density itself. Specifically, this includes:

[0031]

[0032] in, This represents the mean local reachability density corresponding to the neighboring point o.

[0033] In one implementation of this application, the optimization rule base is iteratively optimized based on the data distribution among the obtained verification data, specifically including:

[0034] Calculate the cumulative distribution function value and the theoretical distribution function value corresponding to the verification data;

[0035] Determine whether the maximum distance between the cumulative distribution function and the theoretical distribution function exceeds a preset value. If so, adopt a preset rule update strategy to iteratively optimize the optimization rule base.

[0036] In one implementation of this application, the manual feedback information includes maintenance processing information, deactivation or archiving information, data verification range setting information, interface call configuration information, database table judgment configuration information, combined processing configuration information, and permanent effect configuration information.

[0037] This application provides a data verification rule optimization device, the device comprising:

[0038] At least one processor;

[0039] And, a memory communicatively connected to the at least one processor;

[0040] The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform a data verification rule optimization method as described in any of the preceding claims.

[0041] This application provides a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured as follows:

[0042] One of the data validation rule optimization methods described above.

[0043] The data validation rule optimization method proposed in this application can bring the following beneficial effects:

[0044] An initial rule base is generated through automated global and local anomaly detection, replacing traditional manual rule definition. This significantly reduces maintenance costs and avoids misjudgments and missed detections due to human oversight. The initial rule base is optimized by mining field association rules, effectively addressing the shortcomings of existing models in modeling nonlinear relationships in complex administrative data. This improves adaptability to complex business logic between fields, reduces overfitting risk, and ensures validation stability. A closed-loop feedback mechanism is introduced, allowing users to maintain and optimize the rule base. The dynamic update strategy of the rule base is triggered by the data distribution among validation data, enabling real-time iteration of the rule base as business changes occur. Attached Figure Description

[0045] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0046] Figure 1 A flowchart illustrating a data validation rule optimization method provided in an embodiment of this application;

[0047] Figure 2 A schematic diagram illustrating the generation process of an initial rule base provided in an embodiment of this application;

[0048] Figure 3 A schematic diagram illustrating the optimization process of an initial rule base provided in an embodiment of this application;

[0049] Figure 4 This is a schematic diagram of the structure of a data verification rule optimization device provided in an embodiment of this application. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0051] Administrative datasets involve complex business logic, requiring robust data validation rules and mechanisms to ensure data quality and improve data credibility. However, administrative data collection methods are diverse, including front-end processing, file transfer, individual data entry, and APIs. Furthermore, the relationships between fields in administrative data are complex, involving prerequisite relationships, dependencies, and mutual exclusions.

[0052] For the aforementioned business scenarios, ensuring data validity is paramount. Currently, issues such as underreporting, delayed reporting, false reporting, and duplicate data exist, failing to provide effective data references for the credit infrastructure industry. As data compliance requirements become increasingly stringent, a comprehensive verification rule system is needed to control the quality of collected data. Simultaneously, the verification rule base needs to be updated and iterated in real time to improve data compliance rates and reduce underreporting and delayed reporting rates.

[0053] Existing data validation rules rely on manual definition, resulting in high maintenance costs and a tendency to produce false positives, false negatives, or failure to cover data anomalies. While existing simple machine learning models (such as decision trees) have improved data validation processing speed and eliminated the need for complex data preprocessing, they still suffer from limitations such as low adaptability to complex data processing, lack of dynamic validation rule update mechanisms, inadequate anomaly monitoring, and insufficient flexibility.

[0054] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0055] like Figure 1As shown in the embodiment of this application, a data verification rule optimization method includes:

[0056] S101: Obtain the administrative dataset, perform global anomaly detection and local anomaly detection on the administrative dataset, and determine the global rule feature information and local rule feature information corresponding to the abnormal data in the administrative dataset.

[0057] Machine learning technology is used to automatically detect anomalies in the collected administrative datasets. Reinforcement learning is applied to optimize the rules by combining administrative data business tags, improving accuracy. The decision logic of the machine learning model is transformed into interpretable verification rules, and anomaly rules are periodically detected, automatically triggering retraining and optimization of the rule model to build a complete verification rule learning engine. Specifically, the Isolation Forest (iForest) detection algorithm is used to analyze and process global anomaly data, and the Local Outlier Factor (LOF) algorithm is used to analyze local anomaly data to generate an initial rule base. Then, the FP-Growth algorithm is used to mine association rules between fields to optimize the initial rule base. Finally, a human feedback mechanism is used to refine the initial rule base. Therefore, for the collected administrative datasets, global and local anomaly detection are performed to determine the global and local rule feature information corresponding to anomalies in the administrative dataset.

[0058] In one embodiment, an isolated forest detection algorithm is used to perform global anomaly detection on the administrative dataset. This involves randomly selecting a first subsample from the administrative dataset, constructing an isolated tree corresponding to the first subsample, and determining the abnormal data in the administrative dataset and the global rule feature information corresponding to the abnormal data based on the average path length of each data point in the isolated tree.

[0059] Specifically, when constructing an isolated tree, a feature segmentation operation needs to be performed on the first subsample. This operation includes: determining the regular features and their corresponding separator values; segmenting the first subsample according to the separator values ​​to obtain a subset of samples and the corresponding tree nodes. For each subset, the feature segmentation operation is repeated until all data points in the first subsample are isolated or the tree nodes reach the maximum depth of the isolated tree, thus obtaining the isolated tree corresponding to the first subsample.

[0060] For example, in the first segmentation, the "Unified Social Credit Code" is selected as the rule feature, with the separator being "whether the 8th digit is a number," dividing the data into two sample subsets: "numerical" and "non-numerical." In the second segmentation, for the "non-numerical" subset, "penalty date" is selected as the rule feature, with the separator being "2025-01-01," dividing it into "≤2025-01-01" and ">2025-01-01." This process is repeated until a subset has only one data point remaining (isolated) or the preset tree depth limit is reached. At this point, the isolated tree corresponding to the administrative dataset is obtained.

[0061] After obtaining the isolated tree, irrelevant features are removed to improve performance on high-dimensional data. For each data point in the isolated tree, its corresponding average path length is calculated, and the score corresponding to the data point is determined based on the average path length using the following formula:

[0062]

[0063] Where x represents a data point, n represents the sample size of the data point, h(x) represents the average path length, E(h(x)) represents the expected value of the average path length of data point x in all isolated trees, and c(n) represents the standardization factor of the path length.

[0064] After calculating the anomaly score, abnormal data in the administrative dataset is identified based on whether the score falls within a preset range. If the score falls within the range, the data point is considered abnormal. Generally, anomaly scores less than 0.5 are considered normal, while scores greater than 0.5 are considered potential anomalies. Once anomaly data is identified, the global rule features corresponding to it can be determined based on its path nodes in the isolation tree. These rule features and their corresponding anomaly conditions constitute the validation rules to be considered when validating this type of data.

[0065] For local anomalies that the Isolation Forest detection algorithm may fail to detect, Loophole F (Local Outlier Factor) is used for detection. Based on the Local Outlier Factor algorithm, local anomaly detection is performed on the administrative dataset by quantifying the local density difference between a data point and its neighbors. Several second subsamples are extracted according to predefined rule features. The local outlier factor value is calculated by comparing the local density of a data point in the second subsample with its neighbors, thus determining the local rule feature information corresponding to the anomalous data in the administrative dataset. The larger the local outlier factor value, the higher the probability of an anomaly.

[0066] Specifically, after extracting the second subsample, for each data point p in the sample, its k nearest neighbors are calculated. The k-distance reflects the local density of data point p; the larger the distance, the lower the density. Based on the local reachability density, the local outlier value corresponding to the data point can be further calculated.

[0067] The reachability distance for each data point in the second subsample is calculated using the following formula:

[0068] reach-dist k (p,o)=max(k-distance(o),distance(p,o))

[0069] Where o represents a neighboring point of data point p, k-distance(o) represents the k-distance of neighboring point o, and distance(p,o) represents the distance between data point p and neighboring point o;

[0070] Based on the reachability distance, the local reachability density corresponding to the data points is calculated, which can be obtained using the following formula:

[0071]

[0072] Where, N k (p) represents the number of neighboring points of data point p.

[0073] After calculating the local reachability density corresponding to data point p, it is also necessary to calculate the mean local reachability density among its neighboring points o. Then, based on the mean and local reachability densities of the data point's neighboring points, the local outlier factor value corresponding to the data point can be determined. Specifically, this can be calculated using the following formula:

[0074]

[0075] in, This represents the mean local reachability density corresponding to the neighboring point o.

[0076] After calculating the local outlier factor values, data points with local outlier factor values ​​greater than a preset threshold are considered outliers in the administrative dataset. The rule features selected when extracting the second subsample, along with the outlier conditions for the outliers, constitute the local rule feature information corresponding to those outliers.

[0077] S102: Construct the initial rule base corresponding to the administrative dataset based on the global rule feature information and the local rule feature information.

[0078] The determined global and local rule features are integrated to form a complete set of rule features, i.e., the initial rule base. Based on these rules, the initial rule base contains a series of rules generated based on possible anomalies in the dataset.

[0079] S103: Determine the association rules between the fields of each data sample in the administrative dataset in order to optimize the association of the initial rule base.

[0080] The initial rule base obtained above needs to be optimized by mining the association rules between fields using the FP-Growth algorithm. These association rules reveal the inherent relationships and interdependencies between different fields in the dataset. Applying these rules to the initial rule base can verify and supplement the rules in the initial rule base, making it more accurately reflect the true situation of the data, thereby improving the effectiveness and practicality of the rule base.

[0081] Specifically, the administrative dataset is scanned to count the support of items and filter out infrequent items. Frequent items are then sorted in descending order of support, and each transaction is inserted into an FP-tree. Next, frequent itemsets are recursively mined from the bottom of the FP-tree header, constructing a conditional pattern set and a conditional FP-tree. An initial validation rule base is formed based on these frequent itemsets and stored as interpretable rule information, such as JSON or SQL expression format. Finally, the initial rule base is optimized through rule filtering and integration, rule validation and optimization, and human feedback and iteration. This improves the accuracy of data validation, enhances the comprehensiveness of the rule base, increases data processing efficiency, and reduces manual maintenance costs.

[0082] S104: Receive user feedback on the initial rule base, optimize the initial rule base based on the feedback, and generate an optimized rule base.

[0083] For the initial rule base generated by the above steps, this application embodiment provides a human-computer interactive visual editing interface, which allows for manual intervention, iterative optimization of rules, and improvement of data fault tolerance. Upon receiving user feedback on the initial rule base, the server optimizes the initial rule base based on this feedback, generating an optimized rule base. This feedback includes maintenance processing information, deactivation or archiving information, data validation range settings, interface call configuration information, database table judgment configuration information, combined processing configuration information, and permanent effect configuration information.

[0084] The maintenance and processing information is used to maintain the initial rule base, allowing for additions, modifications, and deletions. Manual processing logic can be added to the initial rule logic to accommodate complex business logic. For example, the criteria for judging late reports need to be validated according to different scenarios.

[0085] The information on deactivation or archiving is used for deactivation or archiving. Based on feedback of concealed, false, and problematic data, if a rule is determined to be an abnormal rule, it can be archived. If it is a national supplementary reporting period, certain rules can be temporarily deactivated to increase the amount of data collected.

[0086] The data validation scope setting information is used to determine the data validation scope of the administrative dataset. For example, if it only applies to a certain type of data information such as administrative licenses and administrative penalties, the corresponding data type range can be set; otherwise, it applies to all administrative data by default.

[0087] The API call configuration information supports configuring corresponding third-party API validation processing, and allows configuration of corresponding input and output parameter information, as well as setting the corresponding return result processing.

[0088] The database table judgment configuration information supports configuring the corresponding database table judgment logic, such as the matching and verification of the administrative counterpart's name and unified social credit code. It allows configuration of the corresponding data source and resource table information, and setting the corresponding query fields and return result processing logic.

[0089] Configuration information can be combined with corresponding validation rules for configuration processing. For example, if it is a document number field, it is necessary to perform validation for simplified penalty data determination, mandatory field validation, compliance validation, etc.

[0090] The permanently effective configuration information ensures that the rules for manual feedback intervention remain valid indefinitely. In other words, when the machine learning model is rerun, the rules for manual feedback intervention will not be cleared, and manual intervention is unnecessary. Furthermore, the unsupervised learning model can perform reinforcement learning based on the rule information from manual feedback, optimizing the current rule base system.

[0091] S105: Perform data verification on the administrative data to be verified based on the optimized rule base, and iteratively optimize the optimized rule base based on the data distribution among the obtained verification data.

[0092] Based on the aforementioned optimized rule base, data verification is performed on the administrative data to be verified, and all issues, issues to be confirmed, concealed, false, and omitted data are collected and fed back to the rule learning engine, triggering the dynamic update mechanism of the unsupervised learning model. Combined with the data distribution of the verification data, the verification rule base is dynamically iterated and optimized to adapt to changes in data distribution or evolution of business scenarios.

[0093] In one embodiment, this application uses the KS statistical test to detect data distribution shifts and trigger a dynamic update mechanism. That is, if the current data distribution is found to be significantly different from the historical baseline, a dynamic update to the optimization rule base will be triggered. The KS test is a non-parametric test method based on the comparison of empirical distribution functions. It determines whether the data distributions are consistent by calculating the maximum vertical distance (D statistic) between the cumulative distribution function of the observed sample and the cumulative distribution function of the theoretical distribution. The larger the D value, the greater the difference between the two distributions.

[0094] Specifically, based on the actual observed verification data, the cumulative distribution function value is calculated, i.e., the distribution of the sample data. Based on historical benchmark data, the corresponding theoretical distribution function value is calculated.

[0095] For example, for a certain rule feature, the corresponding verification rule is: determine whether the administrative data to be verified is less than the verification value x corresponding to the rule feature. The cumulative distribution function value can be expressed as the ratio between the number of data with values ​​less than x in the administrative data sample to be verified and the total sample size. The theoretical distribution function value is expressed as the probability that the data in the historical benchmark data is less than x.

[0096] After calculating the cumulative distribution function and the theoretical distribution function, the maximum distance between them is calculated to see if it exceeds a preset value. The maximum distance is used to quantify the most significant deviation between the current distribution and the historical benchmark. If it exceeds the preset value, it indicates that the current data distribution differs significantly from the historical benchmark distribution. In this case, a preset rule update strategy needs to be adopted to iteratively optimize the rule base.

[0097] The formula for calculating the D statistic is as follows:

[0098]

[0099] Where Fn(x) is the empirical distribution function, used to calculate the cumulative distribution function value, and F(x) is the theoretical distribution function, used to calculate the theoretical distribution function value. This indicates taking the upper bound (maximum distance) for all x.

[0100] Depending on the type of anomaly, different rule update strategies can be selected for update processing. The processing strategies are as follows:

[0101] Rule logic expansion strategy: If a new abnormal data pattern is discovered, the FP-Growth algorithm is used to associate rule mining to generate a new rule branch and iterate the current verification rule base system.

[0102] Rule weight adjustment strategy: If an old rule has a high false positive rate, dynamically weight it based on the rule's accuracy / recall rate to reduce the false positive rate.

[0103] Rule elimination strategy: If a rule has a coverage rate of 0, it is determined to be a long-term invalid rule, and this inefficient rule can be deactivated or archived.

[0104] After the rules are dynamically updated and optimized, manual intervention and review are supported. Based on the aforementioned dynamic update triggering mechanism, different update strategies are automatically selected for processing, and a manual visual review interface is supported to ensure the completeness of rule updates and improve system robustness.

[0105] Figure 2 This application provides a schematic diagram of an initial rule base generation process, as illustrated in the embodiments of this application. Figure 2 As shown, after collecting the administrative dataset, samples are extracted using the Isolation Forest detection algorithm, and an isolation tree corresponding to each sample is constructed based on rule features and separator values. Anomaly scores for each data point in the isolation tree are calculated to identify anomalous data and their corresponding global rule features. Simultaneously, subsamples are extracted using the Local Outlier Factor (LOF) algorithm. The reachability distance for each data point is determined by calculating its k-distance, and the Locator Rank (LOF) value is calculated by using the local reachability density of each data point and the mean of the local reachability densities of its neighbors. The LOF value is used to filter out anomalous data, thereby determining the corresponding local rule features. An initial rule base is constructed based on the global and local rule features.

[0106] Figure 3 This application provides a schematic diagram of an optimization process for an initial rule base, as shown in the embodiments below. Figure 3 As shown, after constructing the initial rule base, the FP-Growth algorithm is used to optimize it. Similarly, based on a human feedback mechanism, the initial rule base can also be optimized accordingly. During data validation, feedback from abnormal data can determine whether to trigger an automatic update mechanism for the optimized rule base. If the automatic update mechanism is triggered, a preset rule update strategy is used to iteratively optimize the optimized rule base.

[0107] The above are embodiments of the methods proposed in this application. Based on the same idea, some embodiments of this application also provide devices and non-volatile computer storage media corresponding to the above methods.

[0108] Figure 4 This is a schematic diagram of a data verification rule optimization device provided in an embodiment of this application. Figure 4 As shown, it includes:

[0109] At least one processor; and,

[0110] At least one processor-communication-connected memory; wherein,

[0111] The memory stores instructions that can be executed by at least one processor, which enables the at least one processor to perform a data verification rule optimization method as described in any of the preceding items.

[0112] This application provides a non-volatile computer storage medium storing computer-executable instructions, which are configured as follows:

[0113] One of the data validation rule optimization methods described above.

[0114] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the method embodiments.

[0115] The devices and media provided in this application are one-to-one with the methods. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0116] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0117] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0118] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0119] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0120] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0121] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0122] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0123] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0124] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for optimizing data validation rules, characterized in that, The method includes: Obtain an administrative dataset, and perform global and local anomaly detection on the administrative dataset to determine the global and local rule feature information corresponding to the abnormal data in the administrative dataset. Based on the global rule feature information and the local rule feature information, an initial rule base corresponding to the administrative dataset is constructed; The association rules between fields of each data sample in the administrative dataset are determined to optimize the association of the initial rule base; wherein, the association rules between fields are mined using the FP-Growth algorithm to optimize the association of the initial rule base. The system receives user feedback on the initial rule base and optimizes it to generate an optimized rule base. The user feedback includes maintenance processing information, deactivation or archiving information, data verification range setting information, interface call configuration information, database table judgment configuration information, combined processing configuration information, and permanent effect configuration information. The administrative data to be verified is verified according to the optimized rule base, and the optimized rule base is iteratively optimized according to the data distribution among the verified data. Global and local anomaly detection are performed on the administrative dataset to determine the global and local rule feature information corresponding to anomalous data in the administrative dataset, specifically including: The administrative dataset is subjected to global anomaly detection using the isolated forest detection algorithm. A first subsample is randomly extracted from the administrative dataset, an isolated tree corresponding to the first subsample is constructed, and the abnormal data in the administrative dataset and the global rule feature information corresponding to the abnormal data are determined based on the average path length of each data point in the isolated tree. The local outlier factor algorithm is used to detect local anomalies in the administrative dataset. Several second sub-samples are extracted according to preset rule features. Based on the local outlier factor value corresponding to each data point in the second sub-sample, the local rule feature information corresponding to the abnormal data in the administrative dataset is determined. Based on the data distribution among the obtained verification data, the optimization rule base is iteratively optimized, specifically including: Calculate the cumulative distribution function value and the theoretical distribution function value corresponding to the verification data; Determine whether the maximum distance between the cumulative distribution function and the theoretical distribution function exceeds a preset value. If so, adopt a preset rule update strategy to iteratively optimize the optimization rule base.

2. The data validation rule optimization method according to claim 1, characterized in that, Constructing the isolation tree corresponding to the first sub-sample specifically includes: Perform a feature segmentation operation on the first subsample. The feature segmentation operation includes: determining a rule feature and a separation value corresponding to the rule feature; segmenting the first subsample according to the separation value to obtain a sample subset after segmentation and a tree node corresponding to the sample subset. For the sample subset, the feature segmentation operation is repeated until all data points in the first subsample are isolated or the tree node reaches the maximum depth of the isolated tree, thus obtaining the isolated tree corresponding to the first subsample.

3. The data validation rule optimization method according to claim 1, characterized in that, Based on the average path length of each data point in the isolated tree, the abnormal data in the administrative dataset and the corresponding global rule feature information are determined, specifically including: The anomaly score corresponding to each data point is determined based on the average path length of each data point in the isolated tree. Based on whether the abnormal score is within a preset abnormal score range, abnormal data in the administrative dataset is determined, and based on the path node of the abnormal data in the isolated tree, the global rule feature information corresponding to the abnormal data is determined.

4. The data validation rule optimization method according to claim 1, characterized in that, Based on the local outlier factor values ​​corresponding to each data point in the second subsample, the local rule feature information corresponding to the abnormal data in the administrative dataset is determined, specifically including: The reachability distance for each data point in the second subsample is calculated using the following formula: Where o represents a neighboring point of data point p, This represents the k-distance from neighboring point o. This represents the distance between data point p and its neighboring point o; Calculate the local reachability density corresponding to the data point based on the reachability distance; The local outlier factor value corresponding to the data point is determined based on the mean local reachability density of the neighboring points of the data point and the local reachability density. Data points whose local outlier factor values ​​are greater than a preset threshold are considered as abnormal data in the administrative dataset, and local rule feature information corresponding to the abnormal data is determined according to the rule features.

5. The data validation rule optimization method according to claim 4, characterized in that, Based on the reachability distance, the local reachability density corresponding to the data point is calculated, specifically including: The local reachability density corresponding to the data point is calculated using the following formula: in, This represents the number of neighboring points of data point p; The local outlier factor value of the data point is determined using the following formula, based on the mean local reachability density of the data point's neighboring points and the local reachability density itself. Specifically, this includes: in, This represents the mean local reachability density corresponding to the neighboring point o.

6. A data verification rule optimization device, characterized in that, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform a data verification rule optimization method as described in any one of claims 1-5.

7. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are set as follows: A data verification rule optimization method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Automatic data verification method and device in fire-fighting business

    CN112307086A

  • Power grid data quality abnormity diagnosis method based on data analysis

    CN117786587A