Data quality analysis system and method for application authorization platform
By using causal correlation analysis and self-healing operations, the problem of locating and repairing abnormal authorized data was solved, improving the system's adaptability and detection efficiency, and ensuring the stability of the authorization chain and user experience.
Patent Information
- Application Number
- CN202610992341.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-08-25
AI Technical Summary
In complex multi-tenant, multi-application environments, existing technologies cannot effectively identify the causal transmission relationship between abnormal authorized data, leading to invalid permission configurations and a decline in user experience. Furthermore, the detection rules are fixed and cannot be adaptively optimized.
By analyzing causal relationships in three dimensions—timeliness, accuracy, and consistency—and combining lag time vectors and anomaly propagation paths, the root cause can be accurately located and self-healing operations can be performed, dynamically optimizing the detection process.
It enables accurate location and rapid repair of authorized link faults, improves the system's adaptability and detection efficiency, and reduces the false positive rate.
Smart Images

Figure CN122633668A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data quality management technology, specifically to a data quality analysis system and method for use in an application authorization platform. Background Technology
[0002] As enterprises deepen their digital transformation, application authorization platforms have become a core infrastructure for ensuring the secure operation of information systems. The authorization platform is responsible for managing the mapping relationship between user identities and application permissions, and its data quality directly determines the accuracy of access control and the security of system operation. However, in complex multi-tenant, multi-application environments, problems such as data synchronization delays leading to delayed permission activation, incorrect field formats causing permission configuration failures, and broken associations resulting in broken authorization chains frequently occur, seriously affecting the reliability of the authorization platform and user experience.
[0003] In existing technologies, quality inspection of authorized data typically employs an independent approach, checking each dimension separately. For example, the timeliness, accuracy, and consistency of the data are inspected separately, with each dimension's results output independently and unrelated. When anomalies are detected, operations personnel need to manually analyze the correlations between anomalies across different dimensions to trace the root cause of the problem. This fragmented inspection and manual analysis model, where each dimension is inspected in isolation, fails to identify the causal transmission relationships between anomalies. For instance, data synchronization delays may lead to errors in field content, causing a break in the data chain, but existing technologies struggle to establish such multi-dimensional correlation analysis. Anomaly handling remains at the level of single-point repair, lacking precise identification of the root cause, resulting in the recurrence of the same problem. Furthermore, fixed detection rules and judgment thresholds cannot be dynamically optimized based on historical processing results, and the system lacks adaptability. Summary of the Invention
[0004] The purpose of this invention is to provide a data quality analysis system and method for application licensing platforms to solve the problems raised in the prior art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a data quality analysis method for an application authorization platform, the method comprising: S100. Perform timeliness detection on authorized data, compare the data timestamp with the current system time and calculate the time difference, determine the lagging data according to the dynamic timeliness judgment threshold, and generate and output a timeliness anomaly list and a lagging time vector. S200. Take the timeliness anomaly list and the lag time vector as input data, use the lag data as the only verification object, use the lag time vector as the time series constraint, perform accuracy verification, generate accuracy anomaly markers, and establish a correlation between the accuracy anomaly markers and the corresponding lag data. S300: Based on accuracy anomaly markers and lag time vectors, consistency verification is performed using the complete operational logic of the authorized link as the verification benchmark. The link breakage caused by data errors is analyzed, the specific type and severity of the link breakage are identified, and the breakage node and anomaly propagation path are recorded. S400: Based on the causal relationship rules between the three dimensions of timeliness, accuracy and consistency, combined with the abnormal transmission path and the lag time vector, the detected abnormalities are analyzed in sequence and the data dependencies are traced. The interference of secondary abnormalities and irrelevant abnormalities is eliminated, the initial abnormal dimension and corresponding data node that triggered the chain reaction are located as the root cause, and the corresponding self-healing operation is triggered according to the type of root cause and the historical repair success rate. S500, re-execute timeliness detection, accuracy verification and consistency verification, update the timeliness anomaly list, accuracy anomaly markers and breakpoint records, and iteratively execute root cause determination and self-healing operations based on the update results, eliminating abnormal data in each dimension one by one until all dimension detection results are qualified; S600 independently calculates and stores the repair success rate of each closed-loop process, and dynamically updates the timeliness judgment threshold and causal association rule weights based on the repair success rate to achieve adaptive optimization of judgment conditions and association rules, and adjusts the detection order or execution frequency of timeliness detection, accuracy verification and consistency verification to continuously adapt the detection process to the system operation status.
[0006] According to the above scheme, step S100 includes: S110. Perform timeliness detection on the authorized data, extract the timestamp information carried by the authorized data, compare the timestamp information with the current system time, and calculate the time difference between the timestamp information and the current system time to characterize the degree of lag of the authorized data. S120. A dynamic timeliness judgment threshold is generated based on the time difference and historical repair success rate to determine whether the authorized data is lagging data. All authorized data records that are determined to be lagging data are collected, and after removing duplicate data entries and supplementing the basic identification information of the data, a timeliness anomaly list is generated. The dynamic timeliness judgment threshold is a critical value that is dynamically adjusted according to the historical repair success rate and is used to determine whether the data exceeds the allowable lag range. S130. Construct a lag time vector based on the time difference corresponding to each lag data, and output the timeliness anomaly list and the lag time vector; the lag time vector records the lag duration of each lag data in vector form, providing time sequence characteristics for accuracy verification and consistency verification.
[0007] According to the above scheme, step S200 includes: S210. Obtain the timeliness anomaly list and the lag time vector, and extract the lag data recorded therein as the data to be verified. S220. Combining the time-series characteristics reflected by the lag time vector, perform field-level accuracy verification on each piece of data to be verified; the field-level accuracy verification includes checking whether the format of the data field conforms to the specification, whether the enumeration value is within the legal range, and whether the foreign key points to an existing primary key record; S230. Generate corresponding accuracy anomaly markers for the data to be verified that contain anomalies based on the verification results; the accuracy anomaly markers include the location of the anomaly field and anomaly type information; S240. Establish a correspondence between the accuracy anomaly marker and the corresponding data to be verified and the lag time vector, and store them together, so that each lag data carries time series characteristics and accuracy anomaly marker.
[0008] According to the above scheme, step S300 includes: S310. Obtain all lagged data carrying accuracy anomaly markers and lag time vectors, and determine the authorization link and corresponding data node to which each lagged data belongs; the authorization link is an authorization path composed of user identity, role permissions, and application module data nodes in a hierarchical relationship; S320. For each authorized link, traverse its data nodes and check the integrity and timing consistency of the association between each data node by combining the lag time vector; the integrity check refers to verifying whether foreign key references exist between nodes and whether the dependency relationship is valid; the timing consistency check refers to verifying whether the data update time of the node conforms to the logical order of upstream and downstream. S330. When it is detected that a data node is carrying an accuracy anomaly flag and its lag time vector exceeds the allowable range, causing its association with an upstream or downstream node to be interrupted, it is determined that a link break has occurred at that location, and the abnormal propagation path is recorded; the abnormal propagation path records the complete transmission process from the accuracy anomaly node to the link break location. S340. Record the authorized link identifier of the link break, the data node information corresponding to the break location, the accuracy anomaly marker type that caused the break, and the anomaly propagation path.
[0009] According to the above scheme, step S400 includes: S410. Obtain the timeliness anomaly list, accuracy anomaly markers, link break node records, and anomaly propagation paths, and extract all the anomaly data contained therein and their corresponding anomaly dimension identifiers; the anomaly dimension identifiers are used to distinguish whether the anomaly belongs to the timeliness, accuracy, or consistency dimension. S420. According to the causal association rule base, the extracted abnormal data are sorted in chronological order of occurrence, and the triggering paths between each abnormality are traced based on the dependencies between data nodes and the abnormal transmission paths. The causal association rule base contains causal transmission rules for timeliness abnormalities leading to accuracy abnormalities and accuracy abnormalities leading to consistency abnormalities. The rule weights in the causal association rule base are dynamically adjusted by the historical repair success rate. The rule weights reflect the credibility of each causal rule in root cause localization. The higher the weight, the more reliable the transmission relationship described by the rule. S430. When the tracing results show that the abnormal data of a certain abnormal dimension precedes other abnormalities in time sequence, and serves as the data dependency basis for subsequent abnormalities and the abnormal propagation path points to this dimension, the abnormal dimension and its corresponding data node are located as the root cause; the root cause is the initial source that triggers the entire abnormal chain reaction. S440. Based on the anomaly dimension type to which the root cause belongs and the current dynamic timeliness judgment threshold, match the corresponding self-healing operation instruction from the self-healing strategy set and execute it; the self-healing strategy set includes the correspondence between different anomaly dimension types and self-healing operation instructions; the self-healing operation instruction includes specific repair actions such as re-pulling data, updating configuration, and rolling back the status.
[0010] According to the above scheme, step S500 includes: S510. After the self-healing operation is completed, obtain the authorized data after self-healing repair and update the lag time vector synchronously; the update of the lag time vector is to recalculate the lag duration based on the repaired data timestamp. S520. Re-perform timeliness checks, accuracy checks, and consistency checks on the repaired authorized data; S530. Obtain the timeliness anomaly list, accuracy anomaly marker, link break node record and anomaly propagation path generated after re-verification; S540. Determine whether there is any abnormal data in any dimension after re-verification. If there is abnormal data, perform root cause localization and self-healing operations again in combination with the updated lag time vector. S550. Repeat steps S520 to S540 until the results of the timeliness, accuracy and consistency tests are all qualified after re-verification.
[0011] According to the above scheme, step S600 includes: S610. After each round of closed-loop process, obtain the result data of root cause localization and self-healing operation, abnormal transmission path and lag time vector distribution of this round, calculate the repair success rate of timeliness dimension, accuracy dimension and consistency dimension and store it; the repair success rate is statistically analyzed by dimension to evaluate the repair effect of each dimension. S620. Retrieve historically stored data on repair success rates, anomaly propagation paths, and lag time vectors for each dimension, and comprehensively analyze the repair stability and anomaly triggering patterns for each dimension. The comprehensive analysis includes statistically analyzing the changing trends of repair success rates for each dimension, identifying high-frequency anomaly propagation paths, and analyzing the distribution characteristics of lag time vectors. S630. Dynamically update the timeliness judgment threshold and the time sequence weight of the causal association rule base based on the comprehensive analysis results, and adjust the execution order and execution frequency of each dimension detection in the next round of closed-loop process; the dynamic update enables the judgment threshold and rule weight to adaptively optimize with historical data; the adjustment of the detection order and execution frequency enables the detection resources to focus on dimensions with low repair success rate, thereby improving the overall detection efficiency; the adjusted detection order and execution frequency need to be recorded and filed for traceability in subsequent process optimization.
[0012] A data quality analysis system for application authorization platforms, comprising a multi-dimensional detection module, a root cause localization module, a self-healing execution module, a closed-loop verification module, and an adaptive optimization module; The multi-dimensional detection module is used to sequentially perform timeliness detection, accuracy verification, and consistency verification on authorized data, and generate a timeliness anomaly list, lag time vector, accuracy anomaly marker, link break node, and anomaly propagation path; The root cause localization module is used to trace the path of an anomaly and locate the root cause based on the causal relationship rules of timeliness, accuracy and consistency, combined with the anomaly propagation path and the lag time vector. The self-healing execution module is used to match and execute the corresponding self-healing operation instructions from the self-healing strategy set based on the anomaly dimension type to which the root cause belongs, the current dynamic timeliness judgment threshold, and the historical repair success rate. The closed-loop verification module is used to re-execute multidimensional detection after self-healing is completed, iteratively performing root cause localization and self-healing until all dimensions of detection are qualified. The adaptive optimization module is used to count and store the repair success rate of each closed loop, dynamically update the timeliness judgment threshold and the weight of causal association rules based on the repair success rate, and adjust the detection order and execution frequency.
[0013] According to the above scheme, the multi-dimensional detection module includes a timeliness detection unit, an accuracy verification unit, and a consistency verification unit; The timeliness detection unit is used to extract the timestamp information of authorized data, calculate the time difference, identify lagging data according to the dynamic timeliness judgment threshold, and generate a timeliness anomaly list and a lag time vector. The accuracy verification unit is used to perform field-level accuracy verification on lagged data by combining the lag time vector, and to generate and associate and store accuracy anomaly markers. The consistency verification unit is used to combine accuracy anomaly markers with lag time vectors to verify authorized links, identify link breaks, and record abnormal propagation paths.
[0014] According to the above scheme, the root cause localization module includes a time series analysis unit and a causal tracing unit; The time series analysis unit is used to sort the occurrence time of abnormal data in various dimensions and determine the order of abnormalities by combining the lag time vector. The causal tracing unit is used to trace the dependencies between anomalies based on the causal association rule base and the anomaly propagation path, and to locate the initial anomaly dimension that triggered the chain reaction and its corresponding data node as the root cause.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. By using time lag vectors and abnormal propagation paths, abnormal time sequence sorting and causal tracing are achieved, accurately locating the root cause node that triggers the chain reaction, improving the efficiency of abnormal location, and suppressing the spread of authorized link failures from the source. 2. This invention dynamically updates the timeliness judgment threshold and causal association rule weights based on the historical repair success rate, and adaptively adjusts the detection order and execution frequency, so that the system can continuously adapt to business changes and maintain high recognition accuracy and low false judgment rate in the long term. 3. This invention deeply couples timeliness, accuracy, and consistency detection, unifies input, output, and data interfaces, improves detection coverage, reduces the complexity of inter-module coupling, and is compatible with existing authorized platform architectures. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the steps of the data quality analysis method for an application authorization platform according to the present invention. Figure 2 This is a schematic diagram of the data quality analysis system of the present invention used in an application authorization platform. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example: Figures 1-2 As shown, the present invention provides a technical solution for a data quality analysis method for an application authorization platform, the method comprising: S100. Perform timeliness detection on authorized data, compare the data timestamp with the current system time and calculate the time difference, determine the lagging data according to the dynamic timeliness judgment threshold, and generate and output a timeliness anomaly list and a lagging time vector. S200. Take the timeliness anomaly list and the lag time vector as input data, use the lag data as the only verification object, use the lag time vector as the time series constraint, perform accuracy verification, generate accuracy anomaly markers, and establish a correlation between the accuracy anomaly markers and the corresponding lag data. S300: Based on accuracy anomaly markers and lag time vectors, consistency verification is performed using the complete operational logic of the authorized link as the verification benchmark. The link breakage caused by data errors is analyzed, the specific type and severity of the link breakage are identified, and the breakage node and anomaly propagation path are recorded. S400: Based on the causal relationship rules between the three dimensions of timeliness, accuracy and consistency, combined with the abnormal transmission path and the lag time vector, the detected abnormalities are analyzed in sequence and the data dependencies are traced. The interference of secondary abnormalities and irrelevant abnormalities is eliminated, the initial abnormal dimension and corresponding data node that triggered the chain reaction are located as the root cause, and the corresponding self-healing operation is triggered according to the type of root cause and the historical repair success rate. S500, re-execute timeliness detection, accuracy verification and consistency verification, update the timeliness anomaly list, accuracy anomaly markers and breakpoint records, and iteratively execute root cause determination and self-healing operations based on the update results, eliminating abnormal data in each dimension one by one until all dimension detection results are qualified; S600 independently calculates and stores the repair success rate of each closed-loop process, and dynamically updates the timeliness judgment threshold and causal association rule weights based on the repair success rate to achieve adaptive optimization of judgment conditions and association rules, and adjusts the detection order or execution frequency of timeliness detection, accuracy verification and consistency verification to continuously adapt the detection process to the system operation status.
[0019] Specifically, the authorization platform manages several authorization data records. Each record contains fields such as user identity identifier, role identifier, permission identifier, and timestamp. In this embodiment, the authorization data table contains 5 records: Record ID: R01, User ID: U001, Role ID: R001, Permission ID: P001, Timestamp: 2025-03-15 10:00:00; Record ID: R02, User ID: U002, Role ID: R002, Permission ID: P002, Timestamp: 2025-03-14 09:00:00; Record ID: R03, User ID: U003, Role ID: R003, Permission ID: P003, Timestamp: 2025-03-10 08:00:00; Record ID: R04, User ID: U004, Role ID: R004, Permission ID: P004, Timestamp: 2025-03-01 07:00:00; Record ID: R05, User ID: U005, Role ID: R005, Permission ID: P005, Timestamp: 2025-02-28 06:00:00; Specifically, step S100 includes: S110. Perform timeliness detection on the authorized data, extract the timestamp information carried by the authorized data, compare the timestamp information with the current system time, and calculate the time difference between the timestamp information and the current system time to characterize the degree of lag of the authorized data. Specifically, in this embodiment, the current system time is set to 2025-03-15 12:00:00; the initial value of the historical repair success rate is 0, and the initial value of the dynamic timeliness judgment threshold is set to 72 hours; the timestamp of each record is extracted, and the time difference with the current system time is calculated to obtain the time difference values: R01: 2 hours; R02: 27 hours; R03: 124 hours; R04: 341 hours; R05: 354 hours; this is only an example and is not a limitation. S120. A dynamic timeliness judgment threshold is generated based on the time difference and historical repair success rate to determine whether the authorized data is lagging data. All authorized data records that are determined to be lagging data are collected, and after removing duplicate data entries and supplementing the basic identification information of the data, a timeliness anomaly list is generated. The dynamic timeliness judgment threshold is a critical value that is dynamically adjusted according to the historical repair success rate and is used to determine whether the data exceeds the allowable lag range. Specifically, delayed data is determined based on the time difference and the dynamic timeliness judgment threshold; records with a time difference of more than 72 hours are delayed data, so R03, R04, and R05 are judged as delayed data; these records are compiled to generate a timeliness anomaly list, including R03, R04, and R05. S130. Construct a lag time vector based on the time difference corresponding to each lag data, and output the timeliness anomaly list and the lag time vector; the lag time vector records the lag duration of each lag data in vector form, providing time-series characteristics for accuracy verification and consistency verification; Specifically, a lag time vector is constructed based on the time difference of each lagged data point; the lagged data are in the order of R03, R04 and R05, and the corresponding lag time vector is V_lag=[124,341,354]; the timeliness anomaly list and the lag time vector are output.
[0020] Specifically, step S200 includes: S210. Obtain the timeliness anomaly list and the lag time vector, and extract the lag data recorded therein as the data to be verified. Specifically, obtain the timeliness anomaly list and the lag time vector, and extract the lag data R03, R04 and R05 as the data to be verified; S220. Combining the time-series characteristics reflected by the lag time vector, perform field-level accuracy verification on each piece of data to be verified; the field-level accuracy verification includes checking whether the format of the data field conforms to the specification, whether the enumeration value is within the legal range, and whether the foreign key points to an existing primary key record; Specifically, field-level accuracy is verified by combining the lag time vector, checking the field format, enumeration value validity, and foreign key existence of each record. In this embodiment, the check found that: the role ID R003 of R03 does not exist in the role table, which is an accuracy anomaly; the data of R04 is normal; the data of R05 is normal; this is only an example and is not a limitation. S230. Generate corresponding accuracy anomaly markers for the data to be verified that contain anomalies based on the verification results; the accuracy anomaly markers include the location of the anomaly field and anomaly type information; Specifically, an accuracy anomaly marker is generated for R03 with anomalies, and it is marked as a missing foreign key; S240. Establish a correspondence between the accuracy anomaly marker and the corresponding data to be verified and the lag time vector, and store them together, so that each piece of lag data carries time series characteristics and accuracy anomaly marker. Specifically, the accuracy anomaly marker is associated with R03 and its lag time of 124 hours and stored together; at this time, the lag data R03, which carries the time series feature 124 and the accuracy anomaly marker foreign key, is missing; R04 and R05 only carry the time series feature and have no accuracy anomaly marker.
[0021] Specifically, step S300 includes: S310. Obtain all lagged data carrying accuracy anomaly markers and lag time vectors, and determine the authorization link and corresponding data node to which each lagged data belongs; the authorization link is an authorization path composed of user identity, role permissions, and application module data nodes in a hierarchical relationship; Specifically, obtain all lagging data marked with an accuracy anomaly, i.e., R03; determine the authorization chain to which R03 belongs; the authorization chain to which R03 belongs is: User U003 - Role R003 - Permission P003, where R03 records the association between users and roles, and the association between roles and permissions is stored in another table; the data nodes of R03 are user-role associations; this is only an example and is not a limitation. S320. For each authorized link, traverse its data nodes and check the integrity and timing consistency of the association between each data node by combining the lag time vector; the integrity check refers to verifying whether foreign key references exist between nodes and whether the dependency relationship is valid; the timing consistency check refers to verifying whether the data update time of the node conforms to the logical order of upstream and downstream. Specifically, the authorization chain was traversed to check the integrity and timing consistency of the relationships between data nodes. The check revealed that role R003 did not exist in the role table, causing the relationship between the user-role association node and the role node to be interrupted. Based on the lag time vector, the lag time of R03 was 124 hours, while the most recent update time of the role table may be earlier, but the main focus here is on the integrity of the relationships. S330. When it is detected that a data node is carrying an accuracy anomaly flag and its lag time vector exceeds the allowable range, causing its association with an upstream or downstream node to be interrupted, it is determined that a link break has occurred at that location, and the abnormal propagation path is recorded; the abnormal propagation path records the complete transmission process from the accuracy anomaly node to the link break location. Specifically, because R03 carries an accuracy anomaly marker, its association with upstream or downstream nodes is interrupted, and a link break is determined to have occurred at this location; the abnormal propagation path is recorded: from the accuracy anomaly node R03 to the break location; S340. Record the authorized link identifier of the link break, the data node information corresponding to the break location, the accuracy anomaly marker type that caused the break, and the anomaly propagation path. Specifically, the broken link information is recorded as follows: authorized link identifier L001, broken node is role R003, the accuracy anomaly marker type for the broken link is missing foreign key, and the anomaly propagation path is: R03 - missing role. This is only an example and is not a limitation.
[0022] Specifically, step S400 includes: S410. Obtain the timeliness anomaly list, accuracy anomaly markers, link break node records, and anomaly propagation paths, and extract all the anomaly data contained therein and their corresponding anomaly dimension identifiers; the anomaly dimension identifiers are used to distinguish whether the anomaly belongs to the timeliness, accuracy, or consistency dimension. Specifically, obtain the timeliness anomaly list, accuracy anomaly markers, link break node records, and anomaly propagation paths; extract all abnormal data and their dimension identifiers: R03: Timeliness anomaly, accuracy anomaly, and consistency anomaly, with dimension identifiers of timeliness, accuracy, and consistency; R04: Timeliness anomaly only; R05: Timeliness anomaly only. S420. According to the causal association rule base, the extracted abnormal data are sorted in chronological order of occurrence, and the triggering paths between each abnormality are traced based on the dependencies between data nodes and the abnormal transmission paths. The causal association rule base contains causal transmission rules for timeliness abnormalities leading to accuracy abnormalities and accuracy abnormalities leading to consistency abnormalities. The rule weights in the causal association rule base are dynamically adjusted by the historical repair success rate. The rule weights reflect the credibility of each causal rule in root cause localization. The higher the weight, the more reliable the transmission relationship described by the rule. Specifically, the causal relationship rule base is used for time-series sorting and path tracing. The causal rule base contains two rules: Rule 1: Timeliness anomalies cause accuracy anomalies, with an initial weight of 1.0; Rule 2: Accuracy anomalies cause consistency anomalies, with an initial weight of 1.0. The rule weights are dynamically adjusted based on the historical repair success rate, and both are initially set to 1.0. The abnormal data were sorted chronologically: R03 lagged by 124 hours, R04 by 341 hours, and R05 by 354 hours; sorted by lag time from smallest to largest: R03, R04, R05; the triggering path was traced based on data dependencies and the anomaly propagation path; it was found that the accuracy anomaly of R03 directly caused the link break, while the timeliness anomaly of R03 preceded the accuracy anomaly because the data lag led to the discovery of missing foreign keys during foreign key checks, but the actual missing foreign keys may be an independent issue; according to rule 1, timeliness anomalies may trigger accuracy anomalies, but here the accuracy anomaly is not directly caused by timeliness, but rather by the existence of missing foreign keys; however, tracing requires considering the propagation path: the anomaly propagation path shows from R03 to the break, without pointing to other anomalies; at the same time, R03 is earlier than other anomalies in chronological order, but the other anomalies are only timeliness anomalies and did not trigger subsequent anomalies, therefore, the root cause needs to be determined; S430. When the tracing results show that the abnormal data of a certain abnormal dimension precedes other abnormalities in time sequence, and serves as the data dependency basis for subsequent abnormalities and the abnormal propagation path points to this dimension, the abnormal dimension and its corresponding data node are located as the root cause; the root cause is the initial source that triggers the entire abnormal chain reaction. Specifically, the tracing results show that the accuracy anomaly of R03 precedes the consistency anomaly it caused in time, and serves as the data dependency basis for subsequent anomalies. Furthermore, the anomaly propagation path points to this dimension. Therefore, the accuracy anomaly dimension of R03 and its data nodes are identified as the root cause. S440. Based on the anomaly dimension type to which the root cause belongs and the current dynamic timeliness judgment threshold, match the corresponding self-healing operation instruction from the self-healing strategy set and execute it; the self-healing strategy set includes the correspondence between different anomaly dimension types and self-healing operation instructions; the self-healing operation instructions include specific repair actions such as re-pulling data, updating configuration, and rolling back status. Specifically, the root cause type is accuracy anomaly, and the current dynamic timeliness threshold is 72 hours; the initial historical repair success rate is 0, so it has no impact; the corresponding self-healing operation instruction is matched from the self-healing strategy set: the self-healing operation for accuracy anomaly is to repair foreign key references, that is, to retrieve the correct role information from the source system and update it; when the self-healing operation is executed, the system automatically fills in the missing role R003 information from the role table and updates the R03 record; this is only an example and is not a limitation.
[0023] Specifically, step S500 includes: S510. After the self-healing operation is completed, obtain the authorized data after self-healing repair and update the lag time vector synchronously; the update of the lag time vector is to recalculate the lag duration based on the repaired data timestamp. Specifically, after the self-healing operation is completed, the repaired authorization data is obtained; the role ID of R03 has been updated to the correct R003, and the timestamp has been updated to the current time 2025-03-15 13:00:00; the lag time vector is updated synchronously: the lag time of R03 is recalculated, which is now 1 hour; the new lag time vector becomes [1,341,354]; S520. Re-perform timeliness checks, accuracy checks, and consistency checks on the repaired authorized data; Specifically, the timeliness, accuracy, and consistency checks were re-performed on the repaired data; Timeliness check: R03 was 1 hour behind, no longer behind; R04 was 341 hours behind, still behind; R05 was 354 hours behind, still behind; The timeliness anomaly list was updated to [R04, R05], and the lag time vector was updated to [341, 354]; Accuracy check: The accuracy of R04 and R05 was checked and found to be normal, with no accuracy anomaly markers; Consistency check: Based on the new accuracy anomaly markers, no link breaks were found; S530. Obtain the timeliness anomaly list, accuracy anomaly marker, link break node record and anomaly propagation path generated after re-verification; Specifically, obtain the re-verification results: Timeliness anomaly list [R04, R05], accuracy anomaly marker is empty, link break node record is empty, and anomaly propagation path is empty; S540. Determine whether there is any abnormal data in any dimension after re-verification. If there is abnormal data, perform root cause localization and self-healing operations again in combination with the updated lag time vector. Specifically, if the timeliness anomaly is still found, the root cause localization and self-healing need to be performed again; however, at this time there are no accuracy anomalies or consistency anomalies, only timeliness anomalies. S550. Repeat steps S520 to S540 until the test results of the three dimensions of timeliness, accuracy and consistency are qualified after re-verification. Specifically, in the second cycle: Timeliness detection: R04 is lagging by 341 hours and R05 is lagging by 354 hours, still lagging; Accuracy verification: both are normal; Consistency verification: no abnormalities; Root cause location: only timeliness anomaly, no causal transmission, the root cause is the timeliness anomaly itself; Self-healing operation: according to the self-healing strategy set, the self-healing operation for timeliness anomaly is to trigger data synchronization, that is, to pull the latest data from the source system again; after executing the self-healing operation, after pulling again, the timestamps of R04 and R05 are updated to 2025-03-15 14:00:00; Third cycle: Re-test: R04 lagged by 1 hour, R05 lagged by 1 hour, both less than the threshold, the timeliness abnormality is eliminated; accuracy and consistency are both qualified; at this point all dimension test results are qualified, and the cycle terminates.
[0024] Specifically, step S600 includes: S610. After each round of closed-loop process, obtain the result data of root cause localization and self-healing operation, abnormal transmission path and lag time vector distribution of this round, calculate the repair success rate of timeliness dimension, accuracy dimension and consistency dimension and store it; the repair success rate is statistically analyzed by dimension to evaluate the repair effect of each dimension. Specifically, upon completion of this closed-loop process, the following results are obtained: the repair success rate is calculated separately for timeliness, accuracy, and consistency. In this closed loop, two timeliness anomalies were successfully repaired, resulting in a 100% timeliness repair success rate; one accuracy anomaly was successfully repaired, resulting in a 100% accuracy repair success rate; and one consistency anomaly was successfully repaired, resulting in a 100% consistency repair success rate. These repair success rates are stored. This is merely an example and is not intended to impose any limitations. S620. Retrieve historically stored data on repair success rates, anomaly propagation paths, and lag time vectors for each dimension, and comprehensively analyze the repair stability and anomaly triggering patterns for each dimension. The comprehensive analysis includes statistically analyzing the changing trends of repair success rates for each dimension, identifying high-frequency anomaly propagation paths, and analyzing the distribution characteristics of lag time vectors. Specifically, the success rate of repairs across various dimensions was retrieved from historical storage, along with the distribution of anomaly propagation paths and lag time vectors. Comprehensive analysis revealed that the anomaly was triggered by an accuracy anomaly leading to a consistency anomaly, while the timeliness anomaly occurred independently. S630. Dynamically update the timeliness judgment threshold and the time sequence weight of the causal association rule base based on the comprehensive analysis results, and adjust the execution order and execution frequency of each dimension detection in the next round of closed-loop process; the dynamic update enables the judgment threshold and rule weight to adaptively optimize with historical data; the adjustment of the detection order and execution frequency enables the detection resources to focus on dimensions with low repair success rate, thereby improving the overall detection efficiency; the adjusted detection order and execution frequency need to be recorded and filed for traceability in subsequent process optimization. Specifically, the parameters are dynamically updated based on the comprehensive analysis results; Furthermore, the dynamic timeliness determination threshold is updated: The timeliness judgment threshold is used to measure the tolerance for data lag. The size of the timeliness judgment threshold directly affects the judgment result of timeliness anomalies. If the timeliness judgment threshold is too small, a large amount of normal data will be misjudged as lagging, and if the timeliness judgment threshold is too large, real lagging data will be missed. Therefore, the timeliness judgment threshold is adaptively adjusted according to the historical repair success rate to match the judgment standard with the actual data quality. This embodiment adopts a dynamic adjustment algorithm based on the historical average repair success rate, with the formula: T_new=T_base×(1-α×R_avg); where T_new is the updated timeliness judgment threshold; T_base is the base threshold, preset according to business needs, and is initially set to 72 hours in this embodiment; α is the adjustment coefficient, with a value range of (0,1), used to control the sensitivity of the threshold to the repair success rate, and is set to 0.2 in this embodiment; R_avg is the historical average repair success rate, statistically analyzed according to the timeliness dimension; adjustment The logic is as follows: When the historical repair success rate is high, i.e., close to 100%, the current threshold setting is reasonable or too lenient, and the threshold should be appropriately lowered to improve detection sensitivity; when the historical repair success rate is low, the current threshold may be too strict, leading to more false alarms, and the threshold should be appropriately increased to reduce the false alarm rate; the adjustment coefficient α controls the threshold change range to avoid a single fluctuation causing drastic threshold oscillations; in this embodiment, the repair success rate of each dimension in this closed-loop process is 100%, i.e., R_avg=1.0, which is calculated by substituting into the formula: T_new=72×(1-0.2×1)=57.6 hours, rounded down to 57 hours; therefore, the timeliness judgment threshold for the next closed-loop process is adjusted to 57 hours; to avoid frequent threshold fluctuations, a threshold update step size limit is set, such as the single adjustment range not exceeding 10% of the basic threshold; at the same time, upper and lower limits of the threshold are set to ensure that the threshold is within a reasonable range, such as the lower limit not less than 24 hours and the upper limit not exceeding 168 hours; Furthermore, the dynamic updating of the weights of the causal association rules: The rule weights in the causal association rule base reflect the credibility of each causal rule in root cause localization; the higher the weight, the more reliable the causal transmission relationship described by the rule; the dynamic update of the weights should be based on the proportion of the rule that has been verified as effective in the historical repair success rate; this embodiment adopts an iterative weighted algorithm based on the historical success rate, the formula is: W_new=(W_old×N+R) / (N+1); where W_new is the updated rule weight; W_old is the current rule weight, the initial value is set to 1.0; N is the number of times the rule has been applied in history, that is, the number of times the rule has been used for root cause localization; R is the repair success rate corresponding to the rule in this round of closed-loop process; the rule base contains two causal transmission rules: rule 1 is that timeliness anomalies cause accuracy anomalies, and rule 2 is that accuracy anomalies cause consistency anomalies; the adjustment logic is: the weight update of rule 1 is based on the proportion of cases where accuracy anomalies are caused by timeliness anomalies and accuracy anomalies are successfully repaired; the weight update of rule 2 is based on the proportion of cases where consistency anomalies are caused by accuracy anomalies and consistency anomalies are successfully repaired. The proportion of rules that are frequently and successfully repaired; when the repair success rate of a rule is high, the weight gradually increases; conversely, the weight gradually decreases; the historical application count N plays a smoothing role, avoiding drastic fluctuations in weight due to a single abnormal result; in this embodiment, rule 1 timeliness causes accuracy, there are no application cases in this round, that is, there is no timeliness causing an anomaly in accuracy, so N=0, and the weight remains unchanged at 1.0; rule 2 accuracy causes consistency, it was applied once in this round, that is, the accuracy anomaly of R03 caused a consistency anomaly, and the anomaly was successfully repaired, so R=100%, N=0, W_old=1.0, substituting into the formula, we get W_new=(1.0×0+1.0) / (0+1)=1.0, that is, the weight of rule 2 is still 1.0; if the rule is applied multiple times in subsequent rounds, the weight will be dynamically adjusted according to the historical repair success rate; rule weights can be normalized to ensure that the sum of all weights is 1, or stored separately for priority ranking when root cause localization; rules with weights below a certain threshold can be temporarily disabled and re-evaluated when new valid data is available; Furthermore, the detection order and execution frequency are dynamically adjusted: The detection order and execution frequency should be optimized according to the repair success rate of each dimension; a dimension with a lower repair success rate indicates that the abnormality in this dimension is more complex or difficult to repair, and it should be detected first and the detection frequency should be increased to discover problems as early as possible; for dimensions with a higher repair success rate, the detection priority or frequency can be appropriately reduced to save detection resources; in this embodiment, a priority adjustment algorithm based on the repair success rate ranking is adopted, and the steps are as follows: Obtain the repair success rates of the three dimensions of timeliness, accuracy, and consistency from historical storage, and count them separately according to the timeliness dimension, accuracy dimension, and consistency dimension; Generate a dimension priority ranking table: Sort the three dimensions in ascending order of repair success rate, and the dimension with the lowest repair success rate obtains the highest detection priority, and so on; The sorting result is used as the basis for the detection order in the next round of closed-loop process; For dimensions with a repair success rate lower than the preset threshold, increase their detection frequency; For dimensions with a repair success rate higher than the preset threshold, reduce their detection frequency; The specific rules are as follows: The reference detection frequency is 1, that is, each round of closed-loop detection is performed once; The success rate threshold T_success, for example, T_success = 90%; The adjustment step size Δf, for example, Δf = 0.2; For dimension i, if its repair success rate R_i < T_success, the detection frequency is adjusted to F_i = 1 + Δf × (T_success - R_i) / T_success; If R_i ≥ T_success, the detection frequency is adjusted to F_i = 1 - Δf × (R_i - T_success) / (1 - T_success), and ensure that F_i is not lower than the minimum detection frequency; Record and file the adjusted detection order and execution frequency for traceability of subsequent process optimization; At the same time, these parameters are used as the input for the next round of closed-loop process; In this embodiment, the repair success rates of the three dimensions of timeliness, accuracy, and consistency are all 100%, and the repair success rates of each dimension are the same and higher than the threshold T_success = 90%; According to the sorting rule, the priorities of each dimension are the same, so the detection order remains unchanged and is still executed in the order of timeliness detection, accuracy verification, and consistency verification; According to the frequency adjustment rule, substituting R_i = 100% into F_i = 1 - 0.2 × (1.0 - 0.9) / (1 - 0.9) = 1 - 0.2 × 0.1 / 0.1 = 1 - 0.2 = 0.8, that is, the detection frequency of each dimension can be adjusted to 0.8 times / round; However, considering that this is the first round of closed-loop, for the sake of conservatism, the frequency is not adjusted temporarily and still executed by detecting each dimension once per round; If the repair success rate of a certain dimension is significantly lower than other dimensions in the future, the detection priority of this dimension will be advanced and its detection frequency will be appropriately increased.
[0025] The present invention provides a technical solution for a data quality analysis system of an application authorization platform, and the system includes a multi-dimensional detection module, a root cause location module, a self-healing execution module, a closed-loop verification module, and an adaptive optimization module; The multi-dimensional detection module is used to sequentially perform timeliness detection, accuracy verification, and consistency verification on authorized data, and generate a timeliness anomaly list, lag time vector, accuracy anomaly marker, link break node, and anomaly propagation path; The root cause localization module is used to trace the path of an anomaly and locate the root cause based on the causal relationship rules of timeliness, accuracy and consistency, combined with the anomaly propagation path and the lag time vector. The self-healing execution module is used to match and execute the corresponding self-healing operation instructions from the self-healing strategy set based on the anomaly dimension type to which the root cause belongs, the current dynamic timeliness judgment threshold, and the historical repair success rate. The closed-loop verification module is used to re-execute multidimensional detection after self-healing is completed, iteratively performing root cause localization and self-healing until all dimensions of detection are qualified. The adaptive optimization module is used to count and store the repair success rate of each closed loop, dynamically update the timeliness judgment threshold and the weight of causal association rules based on the repair success rate, and adjust the detection order and execution frequency.
[0026] Specifically, the multi-dimensional detection module includes a timeliness detection unit, an accuracy verification unit, and a consistency verification unit. The timeliness detection unit is used to extract the timestamp information of the authorized data, calculate the time difference, identify lagging data based on the dynamic timeliness judgment threshold, and generate a timeliness anomaly list and a lag time vector. The accuracy verification unit is used to perform field-level accuracy verification on the lagging data in conjunction with the lag time vector, generate and associate an accuracy anomaly markers, and store them. The consistency verification unit is used to verify the authorization link in conjunction with the accuracy anomaly markers and the lag time vector, identify link breaks, and record the anomaly propagation path.
[0027] Specifically, the root cause localization module includes a time-series analysis unit and a causal tracing unit. The time-series analysis unit is used to sort the occurrence time of abnormal data in each dimension and determine the sequence of abnormalities by combining the lag time vector. The causal tracing unit is used to trace the dependency relationship between abnormalities based on the causal association rule base and the abnormal transmission path, and locate the initial abnormal dimension that triggered the chain reaction and its corresponding data node as the root cause.
[0028] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A data quality analysis method for application authorization platforms, characterized in that: The method includes: S100. Perform timeliness detection on authorized data, compare the data timestamp with the current system time and calculate the time difference, determine the lagging data according to the dynamic timeliness judgment threshold, and generate and output a timeliness anomaly list and a lagging time vector. S200. The timeliness anomaly list and the lag time vector are used as input data to perform accuracy verification, generate accuracy anomaly markers, and establish a correlation between the accuracy anomaly markers and the corresponding lag data. S300. Based on the accuracy anomaly marker and lag time vector, perform consistency verification, analyze the link breakage caused by data errors, and record the breakage node and abnormal propagation path. S400, based on the causal relationship rules between the three dimensions of timeliness, accuracy and consistency, combined with the abnormal transmission path and the lag time vector, analyzes the sequence of occurrence of detected abnormalities and traces the data dependencies, locates the initial abnormal dimension and corresponding data node that triggers the chain reaction as the root cause, and triggers the corresponding self-healing operation according to the type of the root cause and the historical repair success rate. S500, re-execute timeliness detection, accuracy verification and consistency verification, update the timeliness anomaly list, accuracy anomaly markers and breakpoint records, and iteratively execute root cause determination and self-healing operations based on the update results until all dimension detection results are qualified; S600: Store the repair success rate of each closed-loop process, and dynamically update the timeliness judgment threshold and causal association rule weight based on the repair success rate, and adjust the detection order or execution frequency of timeliness detection, accuracy verification and consistency verification.
2. The data quality analysis method for an application licensing platform according to claim 1, characterized in that: Step S100 includes: S110. Perform timeliness detection on the authorized data, extract the timestamp information carried by the authorized data, compare the timestamp information with the current system time, and calculate the time difference between the timestamp information and the current system time. S120. Based on the time difference and the historical repair success rate, a dynamic timeliness judgment threshold is generated to determine whether the authorized data is lagging data. All authorized data records that are determined to be lagging data are collected to generate a timeliness anomaly list. S130. Construct a lag time vector based on the time difference value corresponding to each lag data, and output the timeliness anomaly list and the lag time vector.
3. The data quality analysis method for an application licensing platform according to claim 1, characterized in that: Step S200 includes: S210. Obtain the timeliness anomaly list and the lag time vector, and extract the lag data recorded therein as the data to be verified. S220. Combining the time series characteristics reflected by the lag time vector, perform field-level accuracy verification on each piece of data to be verified. S230. Generate corresponding accuracy anomaly markers for the data to be verified that contain anomalies based on the verification results; S240. Establish a correspondence between the accuracy anomaly marker and the corresponding data to be verified and the lag time vector, and store them together, so that each piece of lag data carries time sequence characteristics and accuracy anomaly marker.
4. The data quality analysis method for an application licensing platform according to claim 1, characterized in that: Step S300 includes: S310. Obtain all lagged data carrying accuracy anomaly markers and lag time vectors, and determine the authorized link and corresponding data node to which each lagged data belongs. S320. For each authorized link, traverse its data nodes and check the integrity of the association relationship and timing consistency between the data nodes by combining the lag time vector. S330. When it is detected that a data node is carrying an accuracy abnormality marker and its lag time vector exceeds the allowable range, causing its association with upstream or downstream nodes to be interrupted, it is determined that a link break has occurred at that location, and the abnormal propagation path is recorded. S340. Record the authorized link identifier of the link break, the data node information corresponding to the break location, the accuracy anomaly marker type that caused the break, and the anomaly propagation path.
5. The data quality analysis method for an application licensing platform according to claim 1, characterized in that: Step S400 includes: S410. Obtain the timeliness anomaly list, accuracy anomaly markers, link break node records, and anomaly propagation paths, and extract all the anomaly data contained therein and their corresponding anomaly dimension identifiers. S420. According to the causal association rule base, the extracted abnormal data are sorted in chronological order of occurrence, and the triggering paths between each abnormality are traced based on the dependency relationship between data nodes and the abnormal transmission path; the causal association rule base includes causal transmission rules for timeliness abnormalities leading to accuracy abnormalities and accuracy abnormalities leading to consistency abnormalities, and the rule weights in the causal association rule base are dynamically adjusted by the historical repair success rate. S430. When the tracing results show that the abnormal data of a certain abnormal dimension precedes other abnormalities in time sequence, and serves as the data dependency basis for subsequent abnormalities and the abnormal propagation path points to this dimension, the abnormal dimension and its corresponding data node are identified as the root cause. S440. Based on the anomaly dimension type to which the root cause belongs and the current dynamic timeliness judgment threshold, match the corresponding self-healing operation instruction from the self-healing strategy set and execute it; the self-healing strategy set includes the correspondence between different anomaly dimension types and self-healing operation instructions.
6. The data quality analysis method for an application licensing platform according to claim 1, characterized in that: Step S500 includes: S510. After the self-healing operation is completed, obtain the authorized data after self-healing repair and update the lag time vector synchronously. S520. Re-perform timeliness checks, accuracy checks, and consistency checks on the repaired authorized data; S530. Obtain the timeliness anomaly list, accuracy anomaly marker, link break node record and anomaly propagation path generated after re-verification; S540. Determine whether there is any abnormal data in any dimension after re-verification. If there is abnormal data, perform root cause localization and self-healing operations again in combination with the updated lag time vector. S550. Repeat steps S520 to S540 until the results of the timeliness, accuracy and consistency tests are all qualified after re-verification.
7. The data quality analysis method for an application licensing platform according to claim 1, characterized in that: Step S600 includes: S610. After each round of closed-loop process, obtain the result data of root cause localization and self-healing operation, abnormal transmission path and lag time vector distribution, calculate the repair success rate of timeliness dimension, accuracy dimension and consistency dimension and store it. S620: Retrieve historical storage data on repair success rates, anomaly propagation paths, and lag time vectors for each dimension, and comprehensively analyze the repair stability and anomaly triggering patterns for each dimension. S630. Based on the comprehensive analysis results, dynamically update the timeliness judgment threshold and the time sequence weight of the causal association rule base, and adjust the execution order and execution frequency of each dimension detection in the next round of closed-loop process.
8. A data quality analysis system for an application licensing platform, applied to the data quality analysis method for an application licensing platform as described in any one of claims 1-7, characterized in that: The system includes a multi-dimensional detection module, a root cause localization module, a self-healing execution module, a closed-loop verification module, and an adaptive optimization module; The multi-dimensional detection module is used to sequentially perform timeliness detection, accuracy verification, and consistency verification on the authorized data, and generate a timeliness anomaly list, a lag time vector, an accuracy anomaly marker, a link break node, and an anomaly propagation path; The root cause localization module is used to trace the path of an anomaly and locate the root cause based on the causal association rules of timeliness, accuracy and consistency, combined with the anomaly propagation path and the lag time vector. The self-healing execution module is used to match and execute the corresponding self-healing operation instructions from the self-healing strategy set according to the abnormal dimension type to which the root cause belongs, the current dynamic timeliness judgment threshold, and the historical repair success rate. The closed-loop verification module is used to re-execute multidimensional detection after self-healing is completed, iteratively perform root cause localization and self-healing until all dimensions of detection are qualified; The adaptive optimization module is used to count and store the repair success rate of each closed loop, dynamically update the timeliness judgment threshold and the weight of the causal association rule based on the repair success rate, and adjust the detection order and execution frequency.
9. The data quality analysis system for an application licensing platform according to claim 8, characterized in that: The multidimensional detection module includes a timeliness detection unit, an accuracy verification unit, and a consistency verification unit; The timeliness detection unit is used to extract the timestamp information of the authorized data, calculate the time difference, identify the lagging data according to the dynamic timeliness judgment threshold, and generate a timeliness anomaly list and a lag time vector. The accuracy verification unit is used to perform field-level accuracy verification on the lagged data by combining the lag time vector, and to generate and associate and store accuracy anomaly markers. The consistency verification unit is used to combine the accuracy anomaly marker with the lag time vector to verify the authorized link, identify link breakage, and record the abnormal propagation path.
10. The data quality analysis system for an application licensing platform according to claim 8, characterized in that: The root cause localization module includes a time series analysis unit and a causal tracing unit; The time-series analysis unit is used to sort the occurrence time sequence of abnormal data in each dimension and determine the order of abnormalities by combining the lag time vector. The causal tracing unit is used to trace the dependency relationship between anomalies based on the causal association rule base and the anomaly propagation path, and to locate the initial anomaly dimension that triggers the chain reaction and its corresponding data node as the root cause.