Data quality analysis system and method

By introducing metadata management, rule parsing, data collection, and monitoring task scheduling modules into the data quality analysis system, and establishing rule relationships using rule attributes, the problems of repetitive configuration and reading in data quality analysis are solved, system performance is improved, and resource consumption is reduced.

CN114385437BActive Publication Date: 2026-02-13MICRO DREAM TECHTRONIC NETWORK TECH CHINACO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111509003.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2026-02-13
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

Existing technologies require multiple configurations and repeated data readings during data quality analysis, leading to the problem of abnormal data being recorded multiple times.

Method used

By introducing a metadata management module, a rule parsing module, a data acquisition module, a monitoring task scheduling module, and a data quality analysis module into the data quality analysis system, the system establishes connections between monitoring rules using rule attributes, generates rule statements that the rule engine can recognize, automatically parses monitoring rules, and avoids duplicate reading of data and duplicate recording of abnormal data.

Benefits of technology

It improves the performance of the data quality analysis system, reduces resource consumption, and avoids duplicate data reading and duplicate recording of abnormal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114385437B_ABST
    Figure CN114385437B_ABST
Patent Text Reader

Abstract

The application discloses a data quality analysis system and method, which comprises a data warehouse, a metadata management module, a rule analysis module, a data collection module, a monitoring task mobilization module and a data quality analysis module. The metadata management module is used for storing rule metadata and data source metadata provided by the data warehouse. The rule metadata comprises a plurality of monitoring rules. Each monitoring rule is configured with corresponding rule attributes. The rule attributes are used to establish rule connection among the monitoring rules. The rule analysis module is used for generating rule statements recognizable by a rule engine according to the rule metadata. The data collection module is used for collecting data source data from the data warehouse according to the data source metadata. The rule engine is used for registering monitoring tasks according to the rule statements and executing the monitoring tasks on the data source data. The monitoring tasks are used for executing each monitoring rule on the data source data according to the rule connection. The data quality analysis module is used for performing quality analysis on the data source data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of big data, and particularly relates to a data quality analysis system and method. BACKGROUND

[0002] Data quality of a data set refers to a degree to which data in the data set is suitable for use and meets a specific user expectation.

[0003] The data quality analysis framework in the related art mainly provides a function of dynamically formulating and modifying data quality monitoring rules, but lacks management of relationships between rules. For example, when a field needs to be monitored in multiple aspects such as non-empty, length, and size, since there is no connection between the monitoring rules, multiple configurations and repeated reading of data are required through the current data quality analysis framework, thereby causing the problem of abnormal data being recorded multiple times. SUMMARY

[0004] Embodiments of the application provide a data quality analysis system and method, which can solve the problem of abnormal data being recorded multiple times due to multiple configurations and repeated reading of data when data quality analysis is performed in the related art.

[0005] In a first aspect, embodiments of the application provide a data quality analysis system, comprising a data warehouse, a metadata management module, a rule analysis module, a data collection module, a monitoring task mobilization module, and a data quality analysis module. The data warehouse is configured to provide data source metadata for the metadata management module and data source data for the data collection module. The metadata management module is configured to store the data source metadata and rule metadata. The rule metadata is formulated by a user in a front end and comprises multiple monitoring rules. Each monitoring rule is configured with corresponding rule attributes, and the rule attributes establish a rule connection between the monitoring rules. The rule analysis module is configured to generate rule statements that can be recognized by a rule engine in the monitoring task mobilization module according to the rule metadata, and save the rule statements in the metadata management module. The data collection module is configured to collect the data source data according to the data source metadata stored in the metadata management module. The rule engine is configured to register monitoring tasks according to the rule statements, and execute the monitoring tasks on the data source data to obtain corresponding abnormal statistical information. The monitoring tasks execute each monitoring rule on the data source data according to the rule connection. The data quality analysis module is configured to perform quality analysis on the data source data according to the abnormal statistical information.

[0006] In a second aspect, the embodiments of the present application provide a data quality analysis method, comprising: generating rule statements recognizable by a rule engine according to rule metadata, wherein the rule metadata comprises a plurality of monitoring rules, each monitoring rule is configured with a corresponding rule attribute, and each monitoring rule establishes a rule connection through the rule attribute; the rule engine registers a monitoring task according to the rule statements, and executes the monitoring task on data source data to obtain corresponding abnormal statistical information, wherein the monitoring task executes each monitoring rule on the data source data according to the rule connection; and performing quality analysis on the data source data according to the abnormal statistical information.

[0007] In a third aspect, the embodiments of the present application provide an electronic device, which comprises a processor, a memory, and a program or instruction stored on the memory and executable on the processor, and the program or instruction is executed by the processor to implement the steps of the method according to the second aspect.

[0008] In a fourth aspect, the embodiments of the present application provide a readable storage medium, which stores a program or instruction, and the program or instruction is executed by a processor to implement the steps of the method according to the second aspect.

[0009] In a fifth aspect, the embodiments of the present application provide a chip, which comprises a processor and a communication interface, the communication interface is coupled with the processor, and the processor is used to run a program or instruction to implement the method according to the second aspect.

[0010] In the embodiments of the present application, the metadata management module stores data source metadata and rule metadata, the rule parsing module can generate rule statements recognizable by a rule engine according to the rule metadata, then the monitoring task mobilization module registers a monitoring task according to the rule statements, and executes the monitoring task on the data source data collected by the data collection module to obtain corresponding abnormal statistical information, and then the data quality analysis module performs quality analysis on the data source data according to the abnormal statistical information. The rule parsing module can automatically parse the monitoring rules configured with rule attributes to generate rule statements recognizable by the rule engine, so that the rule engine executes the monitoring rules on the data source data through the rule connection between the monitoring rules, thereby avoiding repeated reading of data and repeated recording of abnormal data, and further improving the execution performance of the data quality analysis system and reducing resource consumption. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 is a structural schematic diagram of a data quality analysis system provided by the embodiments of the present application;

[0012] Figure 2 is another structural schematic diagram of a data quality analysis system provided by the embodiments of the present application;

[0013] Figure 3 is a flowchart of a data quality analysis method provided by an embodiment of the present application;

[0014] Figure 4 is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0015] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.

[0016] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a category and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.

[0017] The data quality analysis system and method provided by the embodiments of the present application will be described in detail below with reference to the drawings, specific embodiments and application scenarios.

[0018] Figure 1 is a structural diagram of a data quality analysis system provided by an embodiment of the present application, as shown in Figure 1 The data quality analysis system includes a data warehouse 600, a metadata management module 100, a rule analysis module 200, a data collection module 300, a monitoring task mobilization module 400 and a data quality analysis module 500.

[0019] Specifically, the data warehouse 600 can be connected with the metadata management module 100 and the data collection module 300 respectively, the metadata management module 100 is connected with the rule analysis module 200, the data collection module 300 and the monitoring task mobilization module 400 respectively, the data collection module 300 is connected with the monitoring task mobilization module 400, and the monitoring task mobilization module 400 is connected with the data quality analysis module 500.

[0020] The data warehouse 600 is configured to provide data source metadata for the metadata management module 100 and to provide data source data for the data collection module 300. The metadata management module 100 is configured to store data source metadata and rule metadata. The rule metadata is configured by a user in a front end. The rule metadata includes a plurality of monitoring rules. Each monitoring rule is configured with a corresponding rule attribute. The monitoring rules are connected by the rule attributes. The rule analysis module 200 is configured to generate rule statements that can be recognized by the rule engine 410 in the monitoring task mobilization module 400 according to the rule metadata, and to save the rule statements in the metadata management module. The data collection module 300 is configured to collect data source data according to the data source metadata stored in the metadata management module, and to transmit the data source data to the monitoring task mobilization module 400. The rule engine 410 is configured to register monitoring tasks according to the rule statements, and to execute the monitoring tasks on the data source data collected by the data collection module 300 to obtain corresponding abnormal statistical information. The data quality analysis module 500 is configured to analyze the quality of the data source data according to the abnormal statistical information.

[0021] Specifically, the metadata management module 100 stores data source metadata and rule metadata. The data source metadata refers to structure information of data, such as field name, field type, data collection rate, etc. The rule metadata is configured by a user in a front end. The rule metadata includes a plurality of monitoring rules. Each monitoring rule is configured with a corresponding rule attribute. The monitoring rules are connected by the rule attributes. The data collection module 300 is configured to collect data source data in the data warehouse 600 according to the data source metadata stored in the metadata management module 100, and to transmit the data source data to the monitoring task mobilization module 400. The rule analysis module 200 is configured to analyze the monitoring rules in the rule metadata according to the rule metadata stored in the metadata management module 100, to generate rule statements that can be recognized by the rule engine 410, and to store the rule statements in the same file in the metadata management module 100. Then, the rule engine 410 registers monitoring tasks according to the rule statements, executes the monitoring tasks on the data source data, and obtains corresponding abnormal statistical information. Then, the data quality analysis module 500 analyzes the quality of the data source data according to the abnormal statistical information.

[0022] The rule analysis module can automatically analyze the monitoring rules configured with the rule attributes, generate rule statements that can be recognized by the rule engine, and make the rule engine execute the monitoring rules on the data source data through the rule connection between the monitoring rules. Thus, the data quality analysis system can avoid repeated reading of data and repeated recording of abnormal data, improve the execution performance of the data quality analysis system, and reduce resource consumption.

[0023] In a possible implementation, the rule attribute referred to in the present application can include grouping information, priority, and flow transfer mode. Specifically, the grouping information rule attribute can include an agenda-group rule attribute and an activation-group rule attribute, etc. For monitoring rules in the same group, i.e., monitoring rules that can be activated or closed at the same time, the same agenda-group rule attribute can be specified. For mutually exclusive rules, i.e., there is an intersection between each monitoring rule, the same activation-group rule attribute can be specified. When each monitoring rule corresponding to the agenda-group rule attribute is executed on the data source data, there is no intersection between each monitoring rule, and thus each monitoring rule can be executed at the same time. When each monitoring rule corresponding to the activation-group rule attribute is executed on the data source data, there is an intersection between each monitoring rule, and thus each monitoring rule can be executed on the data source data according to the priority attribute corresponding to each monitoring rule.

[0024] For the priority attribute, specifically, the salience rule attribute can be included, the salience rule attribute determines the priority of the monitoring rule matching, and the higher the salience level, the earlier the execution order. The flow transfer mode rule attribute can include update, insert, and retract, for example, according to the update rule attribute, the data in the working memory can be updated, according to the insert rule attribute, new data can be inserted in the working memory, and according to the retract rule attribute, the data can be removed from the working memory, and the rule with low priority will no longer match.

[0025] By configuring the above rule attribute for each monitoring rule, the rule relationship between each monitoring rule can be established, so that the problem of repeated reading of data and recording of abnormal data multiple times can be avoided, and thus the consumption of resources can be reduced.

[0026] In a possible implementation, as shown in Figure 2 The metadata management module 100 includes a data source metadata database 110 and a rule metadata database 120, the data source metadata database 110 is configured to store the data source metadata, and the rule metadata database 120 is configured to store the rule metadata.

[0027] In further implementations, the monitoring task mobilization module 400 can further include a working memory 420 and a rule library 430. Specifically, the data source meta-database 110 can be connected with the data collection module 300, the data collection module 300 is connected with the working memory 420, the rule meta-database 120 and the rule library 430 are connected, the working memory 420 and the rule library 430 are connected with the rule engine 410 respectively, wherein the rule library 430 is used to call the rule statements stored in the rule meta-database 120; the working memory 420 is used to store the data source data collected by the data collection module 300.

[0028] In specific applications, the user compiles the rule meta-data in the front end and specifies the rule attributes of each monitoring rule in the rule meta-data, and then adds the rule meta-data to the rule meta-database 120. The rule parsing module 200 generates rule statements that can be recognized by the rule engine 410 according to the rule meta-data, and requests the meta-data management module 100 to store in the database, i.e. to store in the rule meta-database 120. After the meta-data management module 100 successfully saves the rule statements, the rule library 430 calls the rule statements stored in the rule meta-database 120, and the rule engine 410 reads the rule statements in the rule library 430 first. If the monitoring rule changes, a new monitoring rule is generated according to the new rule statement, thereby realizing the function of dynamic rules. Then, the rule engine 410 registers the monitoring task according to the rule statement, and executes the monitoring task on the data source data collected by the data collection module 300. Specifically, the data source data collected by the data collection module 300 is inserted into the working memory 420, and the rule engine 410 triggers the data source data for rule matching. In the case of rule matching, the monitoring task is executed on the data source data to obtain corresponding abnormal statistical information, and the abnormal statistical information is transmitted to the data quality analysis module 500. The data quality analysis module 500 analyzes the quality of the data source data according to the abnormal statistical information. Specifically, after the upstream monitoring task is completed, the scoring task calculates the data quality score according to the abnormal statistical information and according to the specified algorithm. The abnormal statistical information can include abnormal details, i.e. which data source data has abnormality, and can also include abnormal statistical data, i.e. how many data source data has abnormality.

[0029] Optionally, the data quality analysis module 500 can include a visualization module 510 and an export module 520. After the quality analysis of the data source data is completed, the visualization module 510 can display the quality score curve, the abnormal rate curve, and the abnormal details on the front end. The user can export the required data through the export module 520. Further, the data quality analysis module 500 can also include a monitoring and alarm module 530. The monitoring and alarm module 530 can configure monitoring and alarm tasks for key data fields. The monitoring and alarm tasks can be integrated into monitoring tasks or scoring tasks according to different granularities, so as to obtain the data quality of the key data fields and further provide convenience for the user.

[0030] Optionally, the rule engine 410 in the present application can adopt the Drools rule engine, and thus the corresponding rule statement is a DRL format rule statement recognizable by the Drools rule engine. It should be noted that, in addition to the Drools rule engine, the present application can also adopt other rule engines, which are not limited here.

[0031] Figure 3 is a flow diagram of a data quality analysis method provided by an embodiment of the present application, as shown in Figure 3 , the data quality analysis method includes the following steps.

[0032] S310, generating a rule statement recognizable by a rule engine according to rule metadata.

[0033] The rule metadata includes a plurality of monitoring rules, each monitoring rule is configured with a corresponding rule attribute, and each monitoring rule establishes a rule connection through the rule attribute.

[0034] In specific applications, step S310 can be executed by the rule parsing module 200 shown in the above Figure 1 and Figure 2 , and the specific implementation manner can be referred to the description in the above data quality analysis system embodiment, which will not be described here again.

[0035] S320, the rule engine registers a monitoring task according to the rule statement, and executes the monitoring task on the data source data to obtain corresponding abnormal statistical information.

[0036] The monitoring task executes each monitoring rule on the data source data according to the rule connection.

[0037] In specific applications, step S320 can be executed by the monitoring task mobilization module 400 shown in the above Figure 1 and Figure 2 , and the specific implementation manner can be referred to the description in the above data quality analysis system embodiment, which will not be described here again.

[0038] S330, performing quality analysis on the data source data according to the abnormal statistical information.

[0039] In a specific application, step S330 can be performed by the data quality analysis module 500 shown in FIG. 5, and the specific implementation manner can refer to the description in the data quality analysis system embodiment described above, which will not be repeated here. Figure 1 and Figure 2 The data quality analysis module 500 shown in FIG. 5, and the specific implementation manner can refer to the description in the data quality analysis system embodiment described above, which will not be repeated here.

[0040] The data quality analysis method provided by the embodiment of the present application can analyze the rule metadata by the rule analysis module 200 to generate rule statements recognizable by the rule engine 410, and then the monitoring task mobilization module 400 can register monitoring tasks according to the rule statements, that is, execute monitoring rules on the data source data according to the rule relationship between the rule attributes. Then, the data quality analysis module 500 performs quality analysis on the data source data according to the abnormal statistical information to obtain the quality of the data source data. Since the rule relationship is established between the monitoring rules, the repeated reading of data can be avoided when the monitoring task is executed, so that the repeated recording of abnormal data is avoided, and the execution performance of the data quality analysis system is improved, and the resource consumption is reduced.

[0041] In a possible implementation manner, the rule attributes include grouping information, priority and flow transfer manner. Specifically, the description of the rule attributes in the data quality analysis system embodiment described above can be referred to, which will not be repeated here.

[0042] In a possible implementation manner, before the rule statements recognizable by the rule engine 410 are generated according to the rule metadata, the method can further include: obtaining the rule metadata compiled by the user in the front end and storing it into the rule metadata database. Specifically, when the user compiles the rule metadata, the user can configure the rule attributes of the rule metadata in the front end, such as rule_group, mutual_exclusion, salience, is_retract, need_gather_execption_datail, need_statistics and the like.

[0043] In a possible implementation manner, the registration of the monitoring tasks according to the rule statements and the execution of the monitoring tasks on the data source data can include: generating the monitoring tasks according to the rule statements; performing rule matching on the data source data according to the rule attributes of the monitoring rules in the rule statements; and executing the monitoring tasks on the data source data in the case of successful rule matching.

[0044] In this possible implementation, the data source data will be triggered to perform rule matching. Since the monitoring task executes each monitoring rule on the data source data according to the rule relationship established between the rule attributes, the monitoring task will be executed on the data source data according to the grouping information, priority and flow method in the rule attributes, thereby avoiding the duplicate recording of abnormal data.

[0045] In one possible implementation, the quality analysis of the data source based on the anomaly statistics may include: calculating a data quality score for the data source according to a preset algorithm based on the anomaly statistics. Specifically, the data quality analysis module 500 schedules a quality scoring task, which depends on the upstream monitoring task. After the upstream monitoring task is completed, the scoring task calculates the data quality score according to a specified algorithm based on the anomaly statistics.

[0046] In one possible implementation, after performing quality analysis on the data source data based on the anomaly statistics, the method may further include: displaying the quality score curve, anomaly rate curve, and anomaly details of the data source data on the front end; and displaying prompt information on the front end when anomalies occur in key fields. Specifically, the data quality analysis module 500 provides a data interface for visualizing data quality on the front end, displaying the quality score curve, anomaly rate curve, and anomaly details of the data source data for user reference. Additionally, alarms can be configured for key fields, and prompt information can be displayed on the front end. Optionally, depending on the granularity of the key fields, it can be integrated into monitoring tasks or scoring tasks to further provide users with information on the quality status of the key fields.

[0047] Optionally, embodiments of this application also provide an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 401, a memory 402, and a program or instructions stored in the memory 402 and executable on the processor 401. When the program or instructions are executed by the processor 401, they implement the various processes of the above-described data quality analysis method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0048] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described data quality analysis method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0049] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0050] The chip provided in the embodiments of the present application includes a processor and a communication interface, the communication interface is coupled with the processor, the processor is used to run programs or instructions, realizes each process of the data quality analysis method embodiments, and can achieve the same technical effects. To avoid repetition, details are not described here.

[0051] It should be understood that the chip mentioned in the embodiments of the present application can also be referred to as a system-level chip, a system chip, a chip system, or a system-on-chip chip, etc.

[0052] It should be noted that in this document, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited, and the functions can be performed in the order shown or discussed, or in a substantially simultaneous manner or in reverse order, for example, the described method can be performed in a different order from that described, and various steps can be added, omitted or combined. In addition, the features described with reference to some examples can be combined in other examples.

[0053] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a plurality of instructions for making a terminal (which can be a mobile phone, computer, server, or network equipment, etc.) execute the method described in each embodiment of the present application.

[0054] The embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the specific embodiments described above, and the specific embodiments described above are merely illustrative, but not restrictive, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope protected by the claims.

Claims

1. A data quality analysis system, characterized in that, include: The system includes a data warehouse, metadata management module, rule parsing module, data acquisition module, monitoring task scheduling module, and data quality analysis module. The data warehouse is used to provide data source metadata for the metadata management module and also to provide data source data for the data acquisition module. The metadata management module is used to store the data source metadata and rule metadata. The rule metadata is compiled by the user on the front end. The rule metadata includes multiple monitoring rules. Each monitoring rule is configured with corresponding rule attributes. The monitoring rules are linked together through the rule attributes. The rule parsing module is used to generate rule statements that the rule engine in the monitoring task dispatching module can recognize based on the rule metadata, and save the rule statements to the metadata management module; The data acquisition module is used to acquire data from the data source from the data warehouse based on the data source metadata stored in the metadata management module; The rule engine is used to register monitoring tasks according to the rule statements and execute the monitoring tasks on the data source data to obtain corresponding anomaly statistics. The monitoring tasks execute each monitoring rule on the data source data according to the rules. The data quality analysis module is used to perform quality analysis on the data source based on the abnormal statistical information; The data source metadata refers to the structural information of the data, including at least one of the following: Field name, field type, data collection rate; The rule attributes include grouping information and priority; The grouping information rule attributes include the agency-group rule attribute and the activation-group rule attribute. For monitoring rules within the same group, they can be activated or deactivated simultaneously and are assigned the same agency-group rule attribute. For mutually exclusive monitoring rules, if the same activation-group rule attribute is specified, when executing the monitoring rules corresponding to the agency-group rule attribute on the data source, there is no overlap between the monitoring rules, and all monitoring rules can be executed simultaneously. If there is overlap between the monitoring rules corresponding to the activation-group rule attribute on the data source, then the monitoring rules on the data source will be executed according to the priority attribute corresponding to each monitoring rule.

2. The system according to claim 1, characterized in that, The rule attributes also include the flow method.

3. The system according to claim 1, characterized in that, The metadata management module includes a data source metadata database and a rule metadata database. The data source metadata database is used to store the data source metadata, and the rule metadata database is used to store the rule metadata.

4. The system according to claim 3, characterized in that, The monitoring task dispatch module includes working memory and a rule base, wherein... The rule base is used to call the rule statements stored in the metadata management module; The working memory is used to store the data source data collected by the data acquisition module.

5. A data quality analysis method, characterized in that, include: Data is collected from the data warehouse based on the data source metadata; The rule statement that the rule engine can recognize is generated based on the rule metadata. The rule metadata includes multiple monitoring rules, each of which is configured with corresponding rule attributes. The rules are linked together through the rule attributes. The rule engine registers monitoring tasks according to the rule statements and executes the monitoring tasks on the data source data to obtain corresponding anomaly statistics. The monitoring tasks execute each monitoring rule on the data source data according to the rules. Based on the aforementioned anomaly statistics, a quality analysis is performed on the data source data. The data source metadata refers to the structural information of the data, including at least one of the following: Field name, field type, data collection rate; The rule attributes include grouping information and priority; The grouping information rule attributes include the agency-group rule attribute and the activation-group rule attribute. For monitoring rules within the same group, they can be activated or deactivated simultaneously and are assigned the same agency-group rule attribute. For mutually exclusive monitoring rules, if the same activation-group rule attribute is specified, when executing the monitoring rules corresponding to the agency-group rule attribute on the data source, there is no overlap between the monitoring rules, and all monitoring rules can be executed simultaneously. If there is overlap between the monitoring rules corresponding to the activation-group rule attribute on the data source, then the monitoring rules on the data source will be executed according to the priority attribute corresponding to each monitoring rule.

6. The method according to claim 5, characterized in that, The rule attributes include grouping information, priority, and flow method.

7. The method according to claim 5, characterized in that, Before generating rule statements that the rule engine can recognize based on rule metadata, the method further includes: Obtain the rule metadata compiled by the user on the front end and store it in the rule metadata database.

8. The method according to claim 5, characterized in that, The step of registering a monitoring task according to the rule statement and executing the monitoring task on the data source includes: The monitoring task is generated based on the rule statement; Based on the rule attributes of each monitoring rule in the rule statement, the data source data is matched according to the rules. If the rule is successfully matched, the monitoring task is executed on the data source data.

9. The method according to claim 5, characterized in that, The step of performing quality analysis on the data source based on the abnormal statistical information includes: Based on the anomaly statistics, the quality score of the data source is calculated according to a preset algorithm.

10. The method according to claim 5, characterized in that, After performing quality analysis on the data source based on the anomaly statistics, the method further includes: The front end displays the quality score curve, anomaly rate curve, and anomaly details of the data source. If an anomaly occurs in a critical field, a prompt message will be displayed on the front end.

Citation Information

Patent Citations

  • Intelligent auditing system oriented to business flow

    CN107644323A