Creating and executing a processing workflow to correct data quality issues in a dataset

The data processing system addresses the challenge of identifying and correcting data quality issues by generating workflows and assigning responsibility, effectively managing data quality across complex networks.

JP7827735B2Active Publication Date: 2026-03-10AB INITIO TECHNOLOGY LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-01
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing systems struggle to efficiently identify and correct data quality issues in datasets without proper context analysis, leading to difficulties in determining appropriate corrective actions.

Method used

A data processing system that identifies data quality problems, generates workflows for automatic correction, and assigns responsibility based on metadata, allowing for customizable and scalable data quality issue management.

Benefits of technology

Enables efficient recognition and resolution of recurring data quality issues, with automatic escalation and customizable workflows, ensuring reliable data quality across complex networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007827735000002
    Figure 0007827735000002
  • Figure 0007827735000003
    Figure 0007827735000003
  • Figure 0007827735000004
    Figure 0007827735000004
Patent Text Reader

Abstract

The system and method are for executing, by a data processing system, a workflow for processing result data indicative of an output of a data quality test on a data record in response to receiving the result data and metadata describing the result data by generating a data quality problem associated with a state of the workflow and one or more processing steps to resolve a data quality error associated with the data quality test. The operations include generating a workflow for processing the result data based on a state specified by the data quality problem. Generating the workflow includes assigning an entity responsible for resolving the data quality error based on the result data and the state of the data quality problem, determining one or more actions to satisfy a data quality condition specified in the data quality test based on the metadata, and updating the state associated with the data quality problem.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Priority Claim) This application claims priority to U.S. Patent Application No. 63 / 155,148, filed March 1, 2021, the entire contents of which are incorporated herein by reference.

[0002] FIELD OF THE INVENTION The present disclosure relates to generating workflows for automatically correcting data quality problems in a dataset. More particularly, the present disclosure relates to generating data processing workflows for identifying data quality problems in a dataset and correcting the data quality problems based on whether the data quality problems are new or reoccurring. [Background technology]

[0003] A computer system may be used to send, receive, and / or process data. For example, a server computer system may be used to receive and store resources (e.g., web content such as web pages) and make the content available to one or more client computer systems. Upon receiving a request for content from a client computer system, the server computer system may retrieve the requested content and send the content to the client computer system to fulfill the request.

[0004] A set of data may include data values ​​stored in data fields. A data field may include one or more of the data values. Many instances of a data field may exist within a data set (e.g., in different fields).

[0005] A data field may have a data quality error in which the value of the data field does not meet one or more criteria for the value of that data field. Once a data quality error is identified, it may be difficult to determine how to correct the data quality error without analyzing the context of the data quality error. Summary of the Invention

[0006] The data processing system described herein is configured to identify data quality problems in datasets and generate a data processing workflow for automatically correcting the data quality problems. The data processing system receives results from a data quality analysis of one or more datasets. The data quality analysis determines whether values ​​in the datasets pass or fail one or more data quality tests for those datasets. The data processing system receives result data from the data quality analysis.

[0007] The data processing system is configured to generate a new data quality problem representing a data quality test or determine whether the results data represent an existing data quality problem based on data included in the results data. A data quality problem includes data associated with a data quality test for a particular dataset. A data quality problem includes data specifying which rules apply to the data quality test, the results for testing each data quality rule, and an identification of the particular dataset being tested. Thus, a data quality problem includes result data of the data quality test and metadata that uniquely identifies the data quality test.

[0008] Data quality problems generated by the data processing system are associated with an identifier that uniquely identifies the data quality problem and a state that represents the location in the data processing workflow that corresponds to the data quality problem. The identifier may include a key value used to associate the data quality problem with additional data. For example, the additional data may include an identifier for the client device associated with the data quality problem, the owner of the dataset that has the data quality problem, the entity responsible for resolving the data quality problem, a workflow for correcting the data quality problem, and the state of the workflow. The data processing system is configured to generate new data quality problem data for new data quality problems identified in the result data or to update existing data quality problem data for reoccurring data quality problems in the result data.

[0009] When a data quality issue is created or updated, the data processing system configures a data processing workflow for resolving the data quality issue. The data processing workflow includes a series of data processing steps that are selected based on the data in or associated with the data quality issue. For example, the data processing workflow may include actions or steps that include assigning responsibility for resolving the data quality issue and generating a timeline of actions to be taken regarding the data quality issue.

[0010] Implementations can include one or more of the following features.

[0011] In one aspect, a process includes executing, by a data processing system, a workflow for processing result data indicating an output of a data quality test on a data record by generating a data quality issue associated with a workflow state and one or more processing steps in response to receiving the result data and metadata describing the result data to resolve a data quality error associated with the data quality test. The process includes receiving the result data specifying the output of the data quality test on the data record. The process includes receiving metadata associated with the result data that specifies a data quality condition of the data quality test and the objects that constituted the data quality test. The process includes generating, based on the result data and the metadata, a data quality issue that specifies a state of a workflow for processing the result data. The process includes generating a workflow for processing the result data based on the state specified by the data quality issue. Generating the workflow includes assigning an entity responsible for resolving the data quality error based on the result data and the state of the data quality issue, determining one or more actions to satisfy the data quality condition specified in the data quality test based on the metadata, and updating the state associated with the data quality issue. The process includes sending data representing the workflow to a workflow execution system.

[0012] In some implementations, the data quality issue is associated with a key value that identifies the data quality issue. The process includes, in response to receiving the metadata, determining that a data quality issue has already been generated for the data source for the data quality test by evaluating the key value. The process includes updating the data quality issue based on result data of the data quality test, where updating the data quality issue includes changing a state associated with the data quality issue. In some implementations, updating the data quality issue includes reassigning the result data to a different entity. In some implementations, updating the data quality issue includes adding a timeline of a data processing action associated with the data quality issue. In some implementations, the process includes, in response to updating the data quality issue, determining that a data quality condition is satisfied. Changing the state associated with the data quality issue includes indicating that the data quality issue is associated with a resolved data quality error.

[0013] In some implementations, a process includes receiving scheduling data to associate with a data quality issue, the scheduling data causing a data quality test to be run at a predetermined time interval. The process includes identifying the time interval associated with the data quality test. The process includes sending data representing the time interval to an assigned entity.

[0014] In some implementations, assigning the entities includes determining, from the data source, a set of eligible entities associated with the data source that are responsible for resolving the data quality errors; receiving selection criteria from the data source for assigning entities from the set of eligible entities; and assigning the entities based on the selection criteria. In some implementations, the selection criteria specify a type of data quality error, and metadata associated with the result data specifies that a data quality issue is associated with the type of data quality error, and the entity is eligible to resolve the type of data quality error. In some implementations, the selection criteria specify a particular state of the data quality issue for each entity in the set of eligible entities, and the state of the data quality issue is the particular state for the assigned entity.

[0015] In some implementations, the entity is a particular device of a data processing system.

[0016] In some implementations, the entity is a user of a data processing system.

[0017] In some implementations, executing the workflow further includes generating a timeline of actions taken to process the result data and sending data representing the timeline of actions taken to the entity.

[0018] In some implementations, each state of a data quality problem is associated with one or more candidate actions for processing the resulting data.

[0019] In some implementations, the process includes determining a category of the data quality error selected from one of a data error, a rule error, or a reference data error based on the results data and the metadata.

[0020] Any of these processes may be implemented as a system including one or more processing devices and a memory that stores instructions that, when executed by the one or more processing devices, are configured to cause the one or more processing devices to perform the operations of the process. In some implementations, one or more non-transitory computer-readable media may be configured to store instructions that, when executed by the one or more processing devices, cause the one or more processing devices to perform the operations of the process.

[0021] There may be a feedback loop for data quality (DQ) issues in the DQ result data, and recurring issues are updated with new result data and new metadata. Different DQ items can be generated / updated based on which DQ tests failed and how they failed (e.g., which threshold conditions were not met).

[0022] Data items (representing DQ problems) can be repeatedly updated with new metadata as new DQ results are received. Corresponding workflows are automatically updated. The system can avoid generating new DQ items for recurring DQ problems; instead, existing DQ items are updated.

[0023] There may be a customizable workflow used for each data quality issue, and the workflow is automatically configured based on the metadata associated with the DQ issue. The metadata can be used to configure the workflow associated with the DQ item, and there may be different behaviors for different DQ issues.

[0024] The workflow may include an assignment of responsibility for handling the DQ issue, processing actions to take to handle the DQ issue (e.g., further analysis of the data record values ​​to determine if the DQ rule definitions are flawed), and the assignment of responsibility may change based on the current state associated with the DQ item (escalations may occur as the workflow state changes).

[0025] The system can generate a timeline of responsibilities associated with a DQ item based on what the workflow execution state of the DQ item is.

[0026] The system generates a timeline of responsibilities associated with the DQ items based on what the workflow execution state of the DQ items is.

[0027] Aspects may include one or more advantages: Recurring data quality issues can be recognized and assigned to different entities (escalation); A timeline can be generated that shows how an issue evolves over time and who is working to resolve it; Metadata associated with each data item identifies the exact data quality threshold that was not met. This can indicate that the issue may not be that the data has low data quality, but that the rules are improperly defined; Each workflow is configurable, as each state can specify a different set of actions available to resolve data quality errors.

[0028] Additionally, there is considerable flexibility in configuring the rules for populating data quality issues and generating workflows. The flexibility to generate different workflows for similar data quality issues allows a wide range of users to tailor the system to meet their specific needs and get the most benefit from the automated issue generation.

[0029] A feedback loop exists for handling data quality (DQ) issues in the DQ result data. Recurring issues are updated with new result data and new metadata. A customizable workflow for each data quality issue is automatically configured based on the metadata associated with the DQ issue. A state is maintained for each DQ issue. A state corresponds to a portion of a workflow. If a data quality issue recurs, different workflow actions can be generated based on the current state of the workflow. This allows for automatic escalation of the assigned responsible entity (e.g., for a recurring data quality issue).

[0030] A data processing system is configured to manage data quality issues on large, complex data networks. At scale, data is generated and transferred in complex ways across many software environments, different data sets, different countries, etc. Because people are constantly changing and moving, the data processing system can reliably delegate responsibility for resolving data quality issues and automatically escalate to another entity as needed. In some implementations, the data processing system can generate workflows configured to isolate data with serious data quality issues to prevent them from contaminating original data in other parts of the networked system. For example, data quality issues are repeatedly updated with new metadata as new data quality test results are received. Corresponding workflows are automatically updated.

[0031] A workflow is automatically generated that includes the assignment of responsibility for handling the data quality issue and / or the processing action to take to handle the data quality issue (e.g., further analysis of the values ​​of the data records to determine whether the data quality rule definitions are flawed). Metadata from the data quality tests is used to configure the workflow associated with the data quality issue using configurable rules. As a result, different workflows are suggested by the data processing system based on which data quality tests failed and how they failed (e.g., which data quality threshold conditions were not met).

[0032] The system avoids creating new data quality issues for recurring data quality issues. Instead, existing data quality issues are updated. The assignment of responsibility can change based on the current state associated with the data quality issue. For example, escalations can occur when the state of a workflow changes. In some implementations, the data processing system can generate a timeline of responsibility associated with a data quality issue based on what the state of the workflow execution for the data quality issue is. The data processing system allows for automatic ownership assignment from metadata associated with the data quality issue, rather than explicitly specified or defined for each data quality issue. Additionally, associated attributes (e.g., metadata) are stored with the generated data quality issue. The attribute values ​​are automatically updated when new data quality results are received from the data quality testing process. Thus, the data processing system is configured to automatically recognize recurring data quality issues and update existing DQ items with new metadata for comprehensive presentation to the user. The data processing system is configured to automatically determine when a data quality issue is resolved and tag the issue as potentially resolved. For example, a workflow can be configured to specify that a follow-up will be provided to the responsible entity. An entity (eg, a user) then decides when to mark a data quality issue as resolved.

[0033] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0034] [Figure 1A] FIG. 1 is a block diagram of a data processing system for automatic data quality issue generation and workflow execution. [Figure 1B] FIG. 1 is a block diagram of a data processing system for automatic data quality issue generation and workflow execution. [Figure 2A] FIG. 1 is a block diagram of a data processing system for automatic data quality issue generation and workflow execution. [Figure 2B] FIG. 1 is a block diagram of a data processing system for automatic data quality issue generation and workflow execution. [Figure 2C] FIG. 1 is a block diagram of a data processing system for automatic data quality issue generation and workflow execution. [Figure 2D] FIG. 1 is a block diagram of a data processing system for automatic data quality issue generation and workflow execution. [Figure 3] 1A-1B for automatic data quality issue generation and workflow execution; [Figure 4] 1A-1B for automatic data quality issue generation and workflow execution; [Figure 5] 1A-1B for automatic data quality issue generation and workflow execution; [Figure 6A] 1A-1B for automatic data quality issue generation and workflow execution; [Figure 6B]1A-1B for automatic data quality issue generation and workflow execution; [Figure 7] 1A-1B for automatic data quality issue generation and workflow execution; [Figure 8] 1 shows a flow chart including an exemplary process. [Figure 9] FIG. 1 is a block diagram of a data processing system. DETAILED DESCRIPTION OF THE INVENTION

[0035] FIG. 1A shows a block diagram of an environment 100a including a data processing system 102. The data processing system 102 is configured to process data quality test result data 105 (also referred to as result data or data quality result data). The result data includes output from data quality testing of data records. The result data 105 specifies whether portions of the data records (e.g., values ​​of fields) satisfy conditions specified in data quality rules applied to the data records during data quality testing applied by a data quality testing system 104. The result data 105 includes reports identifying data quality rules that were not satisfied during testing and which records did not satisfy the data quality rules. The result data is described in more detail below.

[0036] The data processing system 102 is configured, via the data quality issue identification module 112, to identify data quality issues in data records from the results data 105. The data quality issue includes data 130 specifying a particular data quality error, issue, or failure from a data quality test, and possibly other data associated with the particular data quality error, issue, or failure. The other data includes attributes (e.g., metadata) specifying one or more of the name of the data record being tested, the owner of the data record being tested, the data domain, the data type, a rule identifier, data indicating whether the data record has been previously tested, the category or type of data quality issue, the entity responsible for handling the data quality issue, a time period, a timestamp, and other metadata describing the data quality test. The data quality issue data 130 may include portions of the results data 105, such as data quality test report data. The data included in a data quality issue is described in more detail below with reference to FIGS. 6A-6B . The other data may include one or more data objects associated with a workflow 132. As described below, a workflow includes a series of data processing actions or steps to perform to resolve a data quality issue.

[0037] Each data quality test is configured using a control point 136. The control point 136 includes a data structure that allows configuration of one or more attribute values ​​associated with the workflow 132. The control point specifies which data quality test is to be run on which data. For example, the control point 136 can specify a specific data quality rule (or set of rules) to test a specific field or record of a specific dataset. The control point 136 specifies the dataset to test, the rule to apply, the owner of the tested data and test results, the schedule for running the test, etc. If a data quality issue is found in the results of the data quality test, the control point is associated with the data quality issue. The control point includes an identifier used to identify any existing data quality issues previously generated from a previous iteration of the data quality test configured by that control point. The control point is used to populate the associated data quality issue with attributes according to one or more inheritance rules, as described below. The control point associates the data quality issue with the generated workflow 132.

[0038] The data quality problem identification module 112 can distinguish between existing data quality problems and new data quality problems specified in the results data 105, such as by referencing the control point that generated the data quality problem. An existing data quality problem is a data quality problem that references a data quality fault or problem in a previously identified data record from the results data for that data record. Thus, a data quality problem is not a new problem in the data, but rather a continuation or persistence of a data quality problem that has already been handled (e.g., by a system administrator or automatically by the data processing system 102). A new data quality problem represents a data quality problem or fault in a data record that is identified for the first time by a data quality test in the results data 105.

[0039] The data quality problem data 130 is stored in the data store as existing data quality problems 122a-n. The data quality problems 122a-n are referenced by the data quality problem identification module 112 (by their associated control points 136) when additional data quality test results 105 are received from the data quality testing system 104. Thus, the data processing system 102 includes a feedback loop in which recurring data quality problems are mapped to previously identified (e.g., configured) data quality problems in the data and previously generated workflows, and new data quality problems are newly associated with workflows to process those new data quality problems. For example, configured data quality problems 122 are retrieved from the data store 120 by the data quality problem identification module 112 when result data 105 is received from the data quality testing system 104. The existing data quality problems 122 are associated with attributes 124. The attributes 124 have values ​​for populating the workflows. The values ​​of the attributes can be configured by a system administrator or can be automatically populated from one or more data objects. For example, the values ​​of the attributes 124 may be retrieved from the respective control points 136. The existing data quality issues 122 and their attributes 124 are retrieved by the data quality issue identification module 112. The module 112 uses the existing data quality issues 122 and their attributes 124 to determine whether new data quality issues should be created or existing data quality issues should be updated in response to receiving the results data 105.

[0040] The data processing system 102 generates a workflow 132 for resolving a data quality issue via the workflow generation module 118. The workflow 132 includes a series of data processing steps for addressing the data quality issue (e.g., for resolving a data quality issue in a data record). The workflow 132 may be partially automated or fully automated. The workflow 132 is automatically configured by the data processing system 102. The workflow 132 may include an assignment of responsibility for addressing the data quality issue depending on the context of the data quality issue. For example, a data quality issue is repeatedly identified when data quality tests repeatedly indicate that a particular data quality issue persists in the data record. This may be due to a systematic error in generating the data record, updating the data record, an issue with data quality rule definition, or some other issue. A persistent data quality issue may be escalated to a new entity if the data quality error persists in the data record over time (e.g., across several data quality tests). An entity in this context includes a user (e.g., a system administrator, an employee, etc.) or a device (e.g., a client device). As new results data 105 is received by the data processing system 102, the workflow generated by the workflow generation module 118 may be updated or reconfigured to change how data quality issues are being resolved (e.g., if the data quality issues have persisted for a period of time).

[0041] The workflow 132 includes processing actions to perform to address data quality issues. For example, further analysis of the data record values ​​can be scheduled to determine whether a data quality rule definition is flawed. In another example, data records from sources that cause a threshold number of data quality issues can be quarantined from other devices or systems in a computing network to prevent the spread of low-quality data to other systems. For example, in a dynamic networked system, new systems and devices may be constantly added and removed. A new device may be uploading low-quality data to data records stored in a computing system (e.g., a global repository). To prevent the low-quality data from propagating to other repositories or to prevent further uploads of low-quality data, the data processing system 102 can temporarily and automatically restrict the new device from uploading data without performing a data quality check on the data before uploading. This reduces or eliminates data quality errors in the system. At scale, this can greatly assist system administrators, who may not be able to constantly monitor every device on the system for data quality issues.

[0042] A data quality problem is associated with a state. The state indicates the position within an associated workflow 132 that is created to resolve the data quality problem. Generally, if a data quality problem is unresolved, the state is maintained, indicating that further analysis or processing by an entity assigned responsibility for resolving the data quality problem is required. When new results data 105 is received, the data quality problem identification module 112 can identify that a data quality problem has previously been created by referencing the identifier of that data quality problem with a state that indicates that a resolution has not yet been achieved for the data quality problem.

[0043] The generated workflow 132 is sent from the data processing system 102 to the workflow execution system 128. The workflow execution system 128 may include any device configured to execute the logic of the workflow 132. For example, the workflow execution system 128 may include a client device, such as a system administrator device. In some implementations, the workflow execution device 128 includes a server or other back-end device. Generally, the workflow execution device 128 is configured to present user interface data, described below, that indicates data quality issues and data associated with one or more control points associated with the data quality issues in accordance with the workflow 132.

[0044] The data processing system 102 is configured to automatically populate attributes (e.g., responsible entities) in the data quality issue and / or data quality workflow 132 based on a set of attribute inheritance rules, where attributes are part of metadata associated with the data quality issue and / or data quality test result 105. The attribute inheritance rules are configurable (e.g., can be predefined) for generating data quality issues by the module 112. As described below, the attribute inheritance rules determine which attributes of the data quality issue are automatically populated and which data source is used for the values ​​of those attributes. For example, the source attribute is either received from a data quality control point or from another object directly or indirectly associated with the data quality control point, as determined by a source attribute specifier. The source attribute specifier includes a domain name or owner identifier of the data source. The data source may include a device operated by a system administrator, an employer, etc.

[0045] The workflow generation module 118 uses the attributes to construct a workflow 132 associated with a data quality issue. As described above, the actions of the workflow may depend on the values ​​of the attributes (e.g., metadata) associated with the data quality issue. The attribute values ​​may be refreshed or updated when new data quality results 105 are received from the batch testing process.

[0046] FIG. 1B shows a block diagram of the environment 100b of the data processing system 102 of FIG. 1A, including additional details for processing the data quality results 105. As described above, the data processing system 102 is configured to receive the data quality results 105 from the data quality test system 104. The data processing system 102 generates or updates data quality issues 130 that represent data quality tests that have failed due to the tested data. The data quality issues 130 are stored in the data store 120 along with associated attributes 124. The workflow generation module 118 uses the attributes and the configured data quality issues 122 to generate a workflow 132. The workflow 132 includes executable logic that can be sent to a remote device, such as a workflow execution system 128 (e.g., a client device). The workflow execution system 128 executes the logic of the workflow 132 using an execution module 134. The execution module 134 may include an application layer that allows a user to manipulate the workflow 132, the data quality issues 122, or the attributes 124 (e.g., through a user interface).

[0047] The data quality testing system 104 includes a device or system configured to receive datasets (not shown) and apply data quality tests to those datasets. The datasets can include any data having data quality criteria. The data quality testing system 104 generally tests datasets by applying one or more data quality rules to the datasets. Generally, data quality tests are configured based on metadata 108 (e.g., attributes) of control points. Control points are used to configure data quality tests. The attributes of the control points specify how the data quality tests should be performed (e.g., which rules), when the data quality tests are scheduled, what the target data for the data quality are, etc., as described below.

[0048] Data quality rules contain logic that indicates criteria for the values ​​of records. Generally, the specific data quality rules applied to a dataset depend on the desired data quality instructions for the dataset. In some implementations, the instructions can be based on the source of the dataset. For example, the source of a dataset can be associated with data quality metadata 108 that specifies which data quality rules should be applied to data from that source. The metadata 108 can be associated with a data quality control point 136 that scheduled data quality tests. In another example, the data quality rules applied can depend on the particular semantic type of the data being analyzed. A semantic type is the meaning of the data. For example, a semantic type can be an address, a birth date (not just an arbitrary date), a driver's license identifier, or other particular meaning of the data. For example, a particular set of data quality rules may apply to a birth date but not to a credit card expiration date. A semantic type can also refer to the language type of the data (e.g., French, British English, US English, etc.), which may impose different formatting requirements on the data. For example, date formats may vary for British English data (e.g., day / month / year) compared to US English (e.g., month, day, year). These attributes may be used to select which data quality rules are applied to the data. In some implementations, as described below, a data quality issue may indicate that an incorrect rule was applied to a given data set, causing it to fail a data quality test.

[0049] The data quality test metadata 108 can indicate which attributes were used to select one or more data quality rules for a dataset. The data quality test metadata 108 can be obtained from a control point 136 associated with a particular data quality test. The control point 136 includes attribute values ​​that specify how the data quality test should be performed. For example, a particular attribute can indicate one or more rules that can be applied to selected fields or data records of a dataset. In this context, an attribute can indicate a set or range of allowable values, a relationship between the value of the field and the values ​​of one or more other fields (e.g., a function such as always greater than, always less than, or equal to), a data type, the length of an expected value, or other characteristics of the value of the field. For example, as described below, an attribute can indicate what values ​​are allowed in a field or data record and what values ​​are not allowed in the field. For example, a given range indicated in an attribute associated with a field's label can also be a data quality rule for that field. However, an attribute associated with a field and a rule generated for the field do not necessarily match. For example, a data quality rule for a field may require no duplicate values ​​or no empty values, but the attribute need not indicate this rule. Conversely, an attribute may indicate that a selected field conforms to a particular requirement, but the data quality rule being generated does not necessarily enforce this requirement.

[0050] Any number of data quality rules can be selected using any combination of attributes. Rules can enforce requirements for each value of a field, independent of any other fields or values ​​in the dataset. For example, a data quality rule may require that the format of values ​​in a field associated with a social security number field have the format "XXX-XX-XXXX," where X can be any digit between 0 and 9. This requirement is based on each individual value and does not require further analysis to enforce. In another example, a data quality rule may check the type of each character in a value to ensure that all X values ​​are actually digits between 0 and 9, and not other characters. In another example, a data quality rule may require a check against the entire selected field. For example, a data quality rule may require that each value in a selected field be unique. This requires checking values ​​in the selected field, in addition to the value currently being processed, to determine whether each value is unique.

[0051] Other examples of data quality rules are possible. In some examples, a rule used to characterize the quality of a set of data may indicate acceptable or prohibited values ​​for data records in a dataset. An acceptable value can be a single value or a range of values. A rule indicating an acceptable value is satisfied when the profile contains the acceptable value. An example of an acceptable value for a field could be the maximum and minimum acceptable values ​​for that field; the rule is satisfied if the average value of the field falls between the maximum and minimum acceptable values. A rule indicating prohibited values ​​for a dataset is satisfied as long as the dataset does not contain prohibited values. An example of a prohibited attribute for a field could be a list of prohibited values ​​for that field; the rule is not satisfied if the field contains any of the prohibited values.

[0052] A data quality rule may indicate an acceptable deviation between the value(s) of a field and a specified value for the field. A deviation between the specified value and the value of the field that is greater than the acceptable deviation indicated by the corresponding rule may be an indication of a data quality problem in the dataset, and therefore, that the dataset is a possible root cause of existing or potential data quality problems in downstream sets of data. In some examples, the acceptable deviation may be specified as a range of values, such as a maximum and minimum acceptable value. In some examples, the acceptable deviation may be specified as a standard deviation from a single value, which may be a mean value (e.g., the average or median of values ​​in historical datasets).

[0053] In some examples, rules used to characterize the quality of a set of data may indicate allowed or prohibited characteristics of values ​​in each of one or more fields of a data record, such as based on the validity of values ​​in the field. A rule indicating allowed characteristics of a field is satisfied if a value in the field meets the allowed characteristics. A rule indicating prohibited characteristics of a field is satisfied as long as a value in the field does not meet the prohibited characteristics. A value that satisfies a rule may be referred to as a valid value, and a value that does not meet a rule may be referred to as an invalid value. Various characteristics of values ​​in a field may be indicated as allowed or prohibited characteristics by a rule. An example rule may indicate allowed or prohibited characteristics of the contents of a field, such as a range of allowed or prohibited values, a maximum allowable value, a minimum allowable value, or a list of one or more specific values ​​that are allowed or prohibited. For example, a birth year field with a value less than 1900 or greater than the current year may be considered invalid. An example rule may indicate allowed or prohibited characteristics of the data type of a field. An example rule may indicate whether the absence of a value (or the presence of a NULL) in a particular field is allowed or prohibited. For example, a last name field containing a string value (e.g., "Smith") may be considered valid, while a last name field that is blank or contains a numeric value may be considered invalid. Exemplary rules may indicate allowed or prohibited relationships between two or more fields within the same data record. For example, a rule may specify a list of values ​​for the zip code field that correspond to each possible value of the state field, and may specify that any combination of values ​​for the zip code field and the state field that is not supported by the list is invalid.

[0054] In some examples, rules can be selected based on an analysis of historical data. For example, a rule may indicate an acceptable deviation between a profile of a field in a particular dataset and a determined historical profile for that field in the dataset. A historical profile of a dataset can be based on historical data; for example, the historical profile can be a profile of the same dataset from the previous day, an average profile of the same dataset from several days ago (e.g., over the past week or month), or a lifetime average profile of the same dataset. More generally, a profile can hold a wide variety of reference information to utilize various types of statistical analysis. For example, a profile can include information regarding a standard deviation or other indication of the distribution of values. For purposes of the following example, and without limiting the generality of the present application, a profile can include a numerical mean, and possibly also a standard deviation, of a conventional dataset.

[0055] In some examples, machine learning techniques are used to select data quality rules. For example, data can be analyzed over a learning period to determine which characteristics, for a user setting or application, affect the data quality of that application and therefore require data quality rules to enforce those characteristics. For example, if an application often fails because values ​​in a field are in the wrong format, the rule generation module generates a rule to enforce a specific format. If an application fails because numeric values ​​are outside of an acceptable range, the rule generation module 206 generates a data quality rule to enforce a specific value range for the field. Other such examples are possible. The learning period can be a specific duration or the amount of time until average or expected values ​​converge to a stable value. The machine learning logic can be configured by the control point 136.

[0056] The data quality testing system 104 can apply one or more data quality rules based on the expected characteristics of the data records processed by the system. In a particular example, the source data are credit card transaction records for transactions occurring in the United States. The source data is streaming data processed in hourly increments. Based on the attributes identified for the fields and data from the application indicating the operations to be performed when processing the credit card transaction records, a user may identify a transaction identifier field, a card identifier field, a status field, a date field, and a total amount field as important data elements to be profiled.

[0057] In a particular example where the source data are credit card transaction records, the data quality testing system 104 may receive an attribute indicating that there are only 50 allowable values ​​for the status field. The data quality testing system 104 may select a rule that causes a warning flag to be used if a profile of a set of source data identifies more than 50 values ​​in the status field, regardless of the standard deviation of the profile of the set of source data relative to the norm. The data quality testing system 104 may receive an attribute that indicates only credit card transaction records for transactions completed on the same day that processing should be present in the set of source data. The data quality testing system 104 may create a rule that causes a warning message to be sent if any source data records have a date that does not match the date of processing.

[0058] In some examples, the data quality testing system 104 may specify one or more rules being generated through a user interface. Through the user interface, a user can provide feedback on the rules for one or more fields or approve pre-populated default rules for the fields. Further description of the user interface can be found in U.S. patent application Ser. No. 13 / 653,995, filed Oct. 17, 2012, the contents of which are incorporated herein by reference in their entirety. Other implementations of the user interface are possible.

[0059] In some examples, when a possible data quality problem is detected in a dataset, such as in a new version of a set of reference data or in a set of source data, an identifier of the dataset having the possible data quality problem is placed in a list of root cause datasets stored in the database. When a data quality problem is later detected for a set of output data, the database can be queried to identify upstream data lineage elements of that set of output data and to determine which, if any, of those upstream data lineage elements are included in the list of root cause datasets.

[0060] In some examples, if a possible data quality issue is detected in a dataset, such as a new version of a set of reference data or a set of source data, a user notification may be enabled. In some examples, a warning flag may be stored to indicate the data quality issue. For example, if a possible data quality issue is detected in a new version of a set of reference data, a warning flag may be stored in conjunction with profile data for the new version of the reference data. If a possible data quality issue is detected in a set of source data, a warning flag may be stored in conjunction with profile data for that set of source data. In some examples, a warning message may be communicated to a user to indicate the existence of a possible data quality issue. The warning message may be, for example, as a message, an icon, or a pop-up window on a user interface, as an email or short message service (SMS) message, or in another form. These data may be part of the data quality issue attributes 124 stored in the data store 120.

[0061] In some examples, a rule may specify actions for the data processing workflow 132 to be performed in response to how the data quality rule is not satisfied. For example, one or more threshold deviations from the data quality profile may be specified in the data quality test metadata 108. The threshold deviation may indicate the entity to which the data quality issue 122 is assigned, the data quarantine procedure, etc., when a warning flag or warning message is used in the workflow 132. For example, if the deviation between the profile of the current dataset and the profile of that dataset is small, such as between 1 and 2 standard deviations, a warning flag may be stored, and if the deviation exceeds 2, a warning message may be communicated and the data may be quarantined. The threshold deviation may be specific to each set of source data and reference data.

[0062] In some examples, if the deviation is severe, e.g., more than three standard deviations from the profile, further processing by the associated processing system using the data may be stopped until a user intervenes. For example, any further processing affected by source data or reference data with severe deviations is stopped. The transformations to be stopped may be identified by data referencing data lineage elements downstream of the affected source data or reference data.

[0063] In some examples, the profile data is determined automatically. For example, the profile data for a given dataset may be automatically updated, e.g., as a running historical average of past profile data for that dataset, e.g., by recalculating the profile data whenever new profile data for that dataset is determined. In some examples, a user may provide initial profile data, e.g., by profiling a dataset with desired characteristics.

[0064] The data quality testing system 104 applies the selected data quality rules and generates data quality results data 105. The data quality results data 105 indicates whether the values ​​of the analyzed dataset satisfied the applied data quality rules. Data quality test metadata 108 associated with the results data 105 includes data describing the dataset, the data quality tests that were performed. For example, the test metadata 108 may include identification of which rules were applied to which portions of the dataset, when the data quality tests were performed, the owner of the dataset, etc.

[0065] In some implementations, the data quality testing system 104 is configured to perform data quality tests and send the result data 105 to the data processing system 102 in a bulk analysis rather than a data stream. For example, the data quality testing system 104 can generate the result data 105 for an entire data set before sending it to the data processing system 102 for data quality issue generation and workflow generation. Generally, the result data 105 indicates whether the data quality test failed for one or more portions of the data set.

[0066] The data processing system 102 receives the result data 105 and the test metadata 108 from the data quality testing system 104 and determines whether the result data 105 warrants the generation of a data quality issue. As described above, the data processing system 102 can generate a data quality issue identifying a failed test if the failure is severe enough to meet the conditions for the generation of a data quality issue. For example, a data quality issue is generated by the data processing system 102 if a threshold percentage of data in a dataset fails to satisfy one or more applied data quality rules. In some implementations, a dataset may be required to fully satisfy all data quality requirements. In some implementations, one or more portions of a dataset may be ignored so that data quality failures of any magnitude do not result in the generation of a corresponding data quality issue by the data processing system 102.

[0067] The data quality problem identification module 112 of the data processing system 100 is configured to, for a given data quality test, identify potential data quality problems for generation for the data quality test. A data quality problem specifies the results data of the data quality test, the source of the tested data, and details about how the data quality test was not met, such as one or more data quality thresholds that were not met during testing for a particular field or record in a dataset. The data quality problem identification module 112 analyzes the values ​​of test metadata 108 associated with the data quality test results 105. The values ​​of the test metadata 108 are compared with attributes 124 associated with one or more existing data quality problems 122, as described in more detail below with respect to FIG. 2A. When a comparison of the test metadata 108 and the attribute 124 values ​​indicates that the detected data quality problem matches an existing data quality problem 122, the existing data quality problem 122 is updated by the data quality update module 116 using the new results data 105 and test metadata, as described in connection with FIG. 2C. If no existing data quality problems are found, the data quality problem generation module 114 generates new data quality problems using the test metadata 108 and the results data 105, as described in connection with Figure 2B. The generated or updated data quality problem data 130 is stored in the data store 120 as data quality problems 122a-n.

[0068] To generate a workflow, the workflow generation module 118 receives data quality issue data 130 from the data store 120 and any associated data objects, such as data quality issue control points 136, including their respective attributes 124. The workflow generation module 118 applies a set of configured rules to the data quality issues 122. The rules are generally configured prior to processing. The rules specify how to process the data quality issues 122 based on the values ​​of their attributes and their associated data objects (e.g., control points 136). For example, a particular user may be selected to be assigned as the responsible entity for a particular severity of data quality issues or for data quality issues for data from a particular source. In some implementations, the rules can specify actions to be automatically taken in the workflow 132 based on the data quality issue 122. For example, if a data quality issue reoccurs, the responsible entity may be changed or promoted to a different entity than the current or previously assigned entity. In another example, the state associated with a data quality issue is referenced to determine which actions to include in the data processing workflow 132.

[0069] The generated workflow 132 is sent to a workflow execution system 128 configured to execute one or more actions of the workflow. In some implementations, the workflow execution module 134 executes logic for the workflow. The logic may include automatically handling data quality issues, such as redefining data quality rules based on predetermined criteria, isolating the dataset from other devices and systems in the networked system (e.g., devices other than the source of the dataset), deleting portions of the dataset, sending alerts to one or more responsible or interested entities, activating one or more computing applications (e.g., data masking applications), or causing other data processing actions.

[0070] In some implementations, execution of workflow 132 includes one or more actions that require input from a responsible entity (e.g., a user or administrator), such as requesting approval to perform an action on a dataset or change a rule definition, proposing one or more rules to be changed, displaying a report, displaying a timeline related to resolving the data quality issue 122, displaying or updating the status of the data quality issue, displaying an indication of the responsible entity, generating and presenting an alert to the responsible entity, or proposing one or more other actions to resolve the data quality issue.

[0071] When the workflow 132 is executed, a check is performed to determine whether the data quality issue 122 has been resolved. The check may include holding the data quality issue 122 in an unresolved state until subsequent results data 105 indicates that the data quality issue no longer exists for that dataset. In some implementations, a user can affirmatively indicate that the data quality issue 122 has been resolved. In some implementations, completion of the workflow is an indication that the data quality issue 122 has been resolved. If the data quality issue 122 cannot be determined to be resolved (e.g., either automatically or after required human verification), the data quality issue remains active and is stored in the data store 120. A data quality issue 122 may expire (e.g., if it is not updated for a set period of time). When a data quality issue 122 expires, it is deleted from the data quality data store 120.

[0072] 2A-2D illustrate example data processing actions of modules 112, 114, 166, and 118, respectively, of data processing system 100 of Figures 1A-1B. Figure 2A includes a process 200a by which data processing system 100 identifies data quality issues from data quality results data 105 and data quality test metadata 108.

[0073] The data quality problem identification module 112 is configured to determine whether a new data quality problem is represented in the test result data 105 or whether one or more existing data quality problems are represented in the test result data. Generally, a control point 136 associated with an executed data quality test specifies whether a data quality problem exists for that particular data quality test (e.g., one or more rules tested against a particular field in the data quality test). When a new data quality problem is identified, the data quality test metadata 108 and data quality result data are sent to the data quality problem generation module 114 for generation of a new data quality problem with attribute values ​​populated using values ​​from the test metadata. The process of generating a new data quality problem is shown in FIG. 2B. When an existing data quality problem 122 is identified, the data quality problem identification module 112 sends the data quality problem attributes 124 of the existing data quality problem 122, the existing data quality problem data, the data quality test result data 105, and the data quality test result metadata 108 to the data quality problem update module 116. The data quality issue update module updates existing data quality issues using values ​​from the test metadata 108 and test results in the context of attributes 124 already populated for the existing data quality issues 122. The process for updating an existing data quality issue 122 is described in connection with FIG. 2C.

[0074] The data quality problem identification module receives 202 attributes from a data quality control point 136 associated with a data quality test. The data quality control point 136 includes a data object containing the attributes that make up the data quality test, as described above. The data quality problem identification module 112 analyzes 204 the attributes of the data quality control point 136 to determine whether there are any existing data quality problems 122 in the data store 120. The control point 136 includes an identifier (described below). The identifier associates the control point 136 with an existing data quality problem, if one exists. If the data quality problem already exists, the data quality identification module is configured to initiate 210 a process to update the attribute values ​​of the existing data quality problem with metadata 108 from the data quality test. If the data quality problem does not already exist, the data quality problem identification module 112 is configured to initiate 208 a process to create a new data quality problem and populate 208 the attributes of the created data quality problem with values ​​from the metadata 108 of the data quality test.

[0075] Each existing data quality issue 122a-c is associated with a respective attribute 124 having a value populated based on previous results data and processing in the associated workflow. For example, an existing data quality issue 122 may be associated with a status value that indicates how the associated workflow is performing for that data quality issue. For example, the status may indicate whether the existing data quality issue 122 is new, in triage analysis, published (e.g., to the data store 120), being resolved by a responsible entity, resolved, or unresolved. The status generally indicates the status of the workflow for that data quality issue. An existing data quality issue 122 may be associated with other attributes, including the owner of the dataset represented by the data quality issue, the responsible entity for resolving the data quality issue, a timeline of events for resolution of the data quality issue, timestamp data, a data quality score for the data represented by the issue, etc., as described above.

[0076] In one embodiment, data quality issues from a data quality test can be matched against stored data quality issues in the data store 120. In this example, the data quality issue identification module 112 is configured to analyze attribute values ​​of existing data quality issues 122a-n with values ​​in the test metadata 108 and the test result data 105. This comparison includes matching values ​​for one or more of the attribute values, such as key values, field names, data source names, or other similar attributes that identify a data quality issue. In some implementations, the data quality results 105 include an identifier that can be stored with the data quality issue as an attribute value. The same identifier can be referenced when subsequent data quality tests are run. If the value of the attribute 124 matches or is similar within a threshold similarity to the test metadata 108, the data quality issue identification module 112 determines that the result data 105 corresponds to an existing data quality issue. The existing data quality issue 122 is updated based on the attribute value 124 of the existing data quality issue and based on the test metadata 108 and the test results 105. If the value of the attribute 124 does not match the data quality test metadata 108, the data quality identification engine 112 determines, based on the metadata 108 and the test results 105, that a new data quality question should be created and the attribute value populated.

[0077] FIG. 2B illustrates a process 200b for generating a new data quality problem by the data quality problem generation module 114. The data quality problem generation module 114 is configured to populate a data quality problem with attribute values ​​from the metadata 108 and the results of data quality tests 105 if the data quality problem is not identified by the data quality identification module 112. As described above, a data quality problem 122 includes a data object or data structure that stores attributes with populated values. The attribute values ​​are automatically populated using a configurable set of attribute inheritance rules. The attribute inheritance rules are pre-set rules. The attribute inheritance rules specify which attributes of the data quality problem are automatically populated and the source of the values ​​of those attributes. The source attributes can be specified in the data quality control point 136 associated with the data quality problem or on another object directly or indirectly associated with the data quality control point. The source is specified by a source attribute specifier field in the rule set. Examples of source attribute specifiers are shown in Table 1 below.

[0078] [Table 1]

[0079] For example, the first rule in Table 1 indicates that the data quality owner value for the data quality issue being created comes from the business owner attribute in the risk data domain for the fiscal year associated with the controlled asset controlled by the data quality control point 136. This rule references four data objects associated with the data quality issue to find the attribute's value. The attribute's value is used to populate the data quality issue's attribute. In Table 1, precedence is a priority value for the attribute source. Higher priority sources are used to populate the attribute value before lower priority sources. The attribute inheritance rule populates the data quality issue's attributes, including the data quality issue's owner, the identification of the data quality rule being tested, descriptive or comment data, data and time metadata, timeline data, etc., as described below in connection with Figures 6A and 6B.

[0080] 2B , the data quality problem generation module 114 generates a data object that includes a data quality problem 220. The data quality problem 220 includes attributes with values ​​populated from the data quality test result data 105 and the data quality test metadata 108. The data quality problem 220 receives values ​​from the metadata 108 and the data quality test result data 105 that specify one or more of the following values: An identifier 210 may be received that specifies the source of the data being tested by the data quality rule. This identifier may populate the data quality problem 220 with a source name. The generated data quality problem 220 may include at least a portion of the test result data 212. The test result data 212 includes either the actual results for a given field being tested (e.g., a pass or fail indication for each value in a particular field) or a summary of the data quality test results (e.g., the percentage of values ​​in a field that passed or failed).

[0081] The generated data quality problem 220 can include data quality rule data 214. Data quality rule data includes definitions of specific data quality rules or logic to be applied to records or fields selected for data quality testing. The data quality rule data 214 can specify a single rule or a set of rules for a data quality test. These rules together comprise a data quality test. For example, a data quality test can include a requirement that a field value be populated with exactly seven characters, all of which must be numeric. If any rule fails, the data quality test fails. In some implementations, the results data 212 specifies how the data quality test failed (e.g., which data quality condition was not met). In some implementations, the data quality results data 212 specifies that at least one data quality condition was not met.

[0082] A created data quality issue 220 can inherit a list of entities eligible to resolve the data quality error. The entities can include systems, devices, or users assigned responsibility for resolving the error in the tested data that causes the data quality issue. The eligible entities can be based on preconfigured logic. For example, a specific group of entities can be eligible for the newly created data quality issue 220. A specific group of entities can be eligible to resolve data quality issues for data owned by a specific owner or set of owners. The eligible entities can include only automated systems and devices for a first type of data quality rule being tested, but only system users or administrators for manual review of other types of data quality rules being tested. Other similar examples are possible. One or more entities are selected from the eligible entities 216 to handle the data quality issue 220. This selection can be part of the generated workflow 132 and can be based on the logic of that workflow.

[0083] The generated data quality problem 220 can receive attribute values ​​from associated data objects, such as the control point identifier 218. The associated data objects can provide one or more values ​​for the attributes stored by those data objects.

[0084] The generated data quality issue 220 is populated based on data from at least the sources mentioned above. The data quality issue generation module 114 assigns a state value 222 to the data quality issue 220. The state data 222 indicates the state of a workflow execution (e.g., by the associated workflow 132) to process the data quality issue 220. For example, the data quality issue may be assigned a "new issue" state, a "triaging" state, a "published" state, a "completed" state, or one or more other states. In one example, the data quality issue may be escalated if it reoccurs. Because the data quality issue 220 is newly generated, an initial state of "new data quality issue" is assigned to the data quality issue.

[0085] When a new data quality issue 220 is created, the data quality issue data is sent to the workflow execution module 118. In some implementations, the data quality issue data is stored in the data store 120. The workflow execution module 118 generates a workflow 132 for execution based on the attribute values ​​of the data quality issue. The workflow 132 is generated based on the workflow logic 224 of the data quality issue 220. The workflow logic 224 indicates the processing logic or actions for resolving the data quality issue 220. For example, the workflow logic indicates the entity assigned to resolve the data quality issue 220. The workflow logic 224 indicates steps or options for data quarantine. The workflow logic includes guidance for debugging data quality rules or identifying sources of low-quality data.

[0086] 2C shows a process 200c for updating an existing data quality problem 240 based on results data from a data quality test. The process for updating a data quality problem 240 is similar to creating a new data quality problem, except that not all attributes of an existing data quality problem are populated from control points or other data objects. Instead, only attributes that reflect the most recent data quality test are updated based on the results data 105 of that data quality test and its metadata 108.

[0087] When a data quality problem is identified by the data quality problem identification module 112, the existing data quality problem 240 is updated with data from the most recent data quality test configured with the control point that specifies the data quality problem. The status 230 of the existing data quality problem is identified. In some implementations, the status is updated to indicate that a second data quality test identified a recurring data quality problem. The updated status changes what actions are performed by the workflow 132 or whether the workflow 132 is updated in response to the new test.

[0088] Existing data quality issues 240 are updated with timeline data 232. The timeline data 232 includes a series of one or more actions or operations performed in connection with the data quality issue 240. For example, an action may include a data quality test to be scheduled or executed. Each executed data quality test is associated with respective result data in the timeline of the data quality issue 240. This allows a user or entity to quickly review trends in recurring data quality issues to determine whether any changes made to the data quality tests or the tested data improved the results. For example, if the number of records or entries for a field that pass a data quality test increases over time, a user may determine that the data is being improved to eliminate the data quality issue. The timeline indicates when changes are made (e.g., by an assigned entity) to one or more data quality rules or to the tested data. Instructions for the changes or updates may be linked to the respective versions of the rules or tested data. In some implementations, the change indication helps an entity determine whether changes being made to tested data or data quality rules are improving data quality results or worsening data quality issues.

[0089] Existing data quality issues 240 are updated with updated data quality test result data 234. The updated data quality test result data 234 indicates how the tested data performed against the most recent data quality test. The updated test result data 234 is included in the timeline of the updated data quality issue 240 or is otherwise indicated in the data quality issue. The entity responsible for resolving the data quality issue 240 reviews the updated result data 234 to determine whether the recurring data quality issue 240 has improved or worsened relative to the previous data quality test.

[0090] Updated data quality issues 240 include an updated list of entities eligible to resolve the data quality errors that cause the data quality issues. The eligible entities are based on the state of the data quality issue 240, the type of data quality error indicated in the data quality test results data 234, or one or other metrics. For example, the list of eligible entities is based on the owner or source of the tested data. The list of eligible entities is updated for each data quality test to ensure that eligible entities are assigned (or reassigned) responsibility for resolving the data quality issue 240.

[0091] The updated data quality issue 240 receives metadata or attribute values ​​from one or more other objects associated with the data quality test, such as from the control point that configured the data quality test. Thus, the updated data quality issue 240 includes data describing the most recent data quality test that caused the data quality issue to reoccur, as well as data from previous data quality tests that raised the data quality issue. The result is a data quality issue 240 that includes trend data from previous occurrences, which helps the assigned entity debug data quality rules, remove data with low quality from the associated system, or identify sources of low-quality data. Sources of low-quality data include specific data sources (e.g., applications, databases, etc. that generate low-quality data), but can also include poorly defined data quality tests. Data from each occurrence of a reoccurring data quality issue is correlated to allow the entity to determine how the source of low-quality data is causing the low-quality data. Previous and current data regarding data quality issues 240 can assist an entity in correcting the issues, such as by restricting data from low-quality data sources, modifying data quality rule definitions, isolating low-quality data, adjusting application or data processing workflow logic, or taking some other corrective action.

[0092] A data quality issue 240 is associated with state data 242. The state data 242 indicates the state of the workflow that processes the data quality issue 240. The state of the data quality issue 240 is similar to the state 222 of the data quality issue 220 described above.

[0093] The data quality issue 240 is associated with updated workflow logic 244. The workflow logic 244 indicates processing logic or actions for resolving the data quality issue 240. For example, the workflow logic indicates an entity assigned to resolve the data quality issue 240. The workflow logic 244 indicates steps or options for data isolation. The workflow logic includes guidance for debugging data quality rules or identifying sources of low-quality data. In some implementations, some or all of the workflow actions are performed automatically. In some implementations, the workflow actions are suggestions that, once assigned, are presented to a responsible entity. In some implementations, the workflow logic specifies the configuration of a user interface for presenting the data quality issue 240 to a user, such as an assigned entity.

[0094] The data including the updated data quality issues 240 is sent to the workflow execution module 118 for generation of the workflow 132. In some implementations, the updated data quality issues 240 data is sent to the data store 120.

[0095] 2D illustrates a process 200d for generating a workflow 132 (e.g., as workflow data 132a) from a new data quality problem 220 or an updated data quality problem 240. Process 200d is an example and may include additional or different actions depending on the values ​​of attributes of the data quality problems 220, 240. To execute process 200d, the workflow generation module 118 receives data quality problem data from either the data quality problem generation module 114 (if a new data quality problem is being generated) or the data quality problem update module (if an existing data quality problem is being updated). The workflow generation module 118 is configured to generate a workflow 132 based on the state of the received data quality problems 220, 240 and the values ​​of the attributes of each data quality problem. As described above, the workflow 132 includes logic for resolving the data quality problem. The logic specifies one or more actions to be performed by the workflow execution system 128. The logic of the workflow may include displaying one or more different user interfaces, generating alerts, quarantining data, assigning an entity responsible for resolving data quality issues, invoking one or more third-party services, or other similar actions. The logic of the workflow may be based on pre-configured logic created by a user (e.g., a system administrator). In this manner, the workflow 132 may be application-specific.

[0096] The workflow generation module 118 is configured to check 260 the status of the data quality issue 220, 240. If the status specifies a new data quality issue, the workflow generation module 118 generates a new workflow for processing the data quality issue and resolving the data quality issue. If the status specifies an existing data quality issue, the workflow generation module 118 determines whether the workflow associated with the data quality issue 240 should be updated depending on the attribute values ​​of the data quality issue.

[0097] The workflow generation module 118 is configured to derive 262 workflow logic based on the identified state of the data quality issue. The workflow logic is configured based on the requirements of the system. In one example, the workflow logic is configured via control points. In this example, attribute values ​​of the control points populate the data quality issue. The attribute values ​​of the data quality issue are analyzed to determine the workflow logic for resolving the data quality issue. For example, an existing workflow can be updated with new logic for assigning an entity responsible for resolving a data quality issue if the data quality issue recurs or worsens. In another example, the workflow 132 logic can be changed if the target execution system 128 executing the workflow (or a portion thereof) changes, if the owner changes, if the data source changes, etc. For example, if the target execution system 128 is a client device, the specific user interface generated by the assigned entity, such as for viewing attribute values ​​of the data quality issue, is tailored based on the device type.

[0098] The workflow generation module 118 is configured to assign 264 an entity to resolve the data quality issue. The entity is assigned based on the attribute values ​​of the data quality issue 220, 240. In some implementations, the entity is not actually assigned, but logic for assigning the entity is extracted and included in the workflow. In some implementations, the entity is assigned such that a target workflow execution system 128 is identified to send the workflow data 132a. For example, the entity can be the workflow execution system 128 itself or a user thereof.

[0099] The workflow generation module 118 is configured to update 266 the data quality issue timeline associated with the data quality issue 220, 240. The data quality issue timeline represents what actions are being taken with respect to the data quality issue and when those actions are being taken.

[0100] The workflow generation module 118 is configured to update 268 the status of the data quality issue 220, 240 to indicate that the data quality issue is being processed. In some implementations, the status of the data quality issue is updated after execution of the workflow 132. In this example, the workflow generation module 118 generates a workflow, sends workflow data 132a to the execution system 128, and receives a response from the execution system indicating that one or more actions of the generated workflow 132 have been executed. In response to this data from the execution system 128, the workflow generation module 118 updates the data quality status.

[0101] The generated workflow 132 is sent as workflow data 132a to the execution system 128. The workflow is executed by the execution system 128 either automatically or in response to user input. The workflow 128 executes selected logic based on attributes of the data quality problem. As mentioned above, the actual logic is specific to a given application and is pre-configured by a system administrator.

[0102] 3 shows an exemplary user interface screen 300 representing a user interface for interacting with data quality control points. A list of control points 302a-d is shown. The exemplary representative control point 302a includes an identifier 304, a type 306 (here, all are data quality control objects), a description 308, and an indication of any controlled assets 310. The identifier 304 is a unique code that identifies the data quality control point. The object type 306 distinguishes the control point from one or more other data objects associated with the data being tested in a data quality test. The controlled asset 308 specifies the application being tested.

[0103] The screen 300 includes a search parameters window 312. The parameters in window 312 are used to refine the search through data quality control points. In some implementations, the parameters include filter criteria.

[0104] 4 shows a screen 400 including a user interface representing a list of data quality issues 402a-h in the system. Each data quality issue 402a-h includes an identifier 406, a state value 408, an associated object 410 (e.g., a control point) specified by its object identifier, a list of one or more assigned entities 414, and workflow data 416, such as creation time or other timeline data. A control 418 for manually creating a new data quality issue is shown.

[0105] The screen 400 includes a search parameters window 404. The parameters in the window 404 are used to refine the search through data quality issues. In some implementations, the parameters include filter criteria.

[0106] 5 shows a screen 500 that displays an overview of a data quality issue. Screen 500 shows a data quality issue identifier 502. Screen 500 shows summary data 504 about the data quality issue. Screen 500 shows the entity 506 that reported the data quality issue (if applicable). Screen 500 shows the assigned entity 508 for resolving the data quality issue. Screen 500 shows a portion of timeline data 510 about the data quality issue. Screen 500 is typically inserted into a larger user interface, but provides a quick summary of the data quality issue by showing some of its attribute values.

[0107] Figure 6A shows a screen 600a representing data for data quality issue #2 (referred to as data quality issue DQI-2). Screen 600a shows a more complete view of a data quality issue, such as the data quality issue of Figure 5. Screen 600a includes a navigation window 602 that allows a user to access data including attribute values ​​for the data quality issue, previous versions of the data quality issue (if applicable), and permissions associated with the data quality issue.

[0108] Screen 600a includes a general information window 604 that shows more basic attribute values ​​about the data quality issue. Window 604 shows a summary 604a that describes the name of the object, in this case, an automated data quality issue. Window 604 shows the type of data quality issue 604b. Window 604 shows the category of the data quality issue 604c. Window 604 shows the control point identifier 604d associated with the data quality issue (e.g., the control that configured the data quality issue). In this example, the control point is DQCP-2010, i.e., Data Quality Control Point #2010. Window 604 shows the reporter 604e of the data quality issue (if the data quality issue was manually generated or reported). Window 604 shows the assigned entity 604f for resolving the data quality issue. In this case, the entity is a person called Abbey A. Williams. In other examples, the entity may be a device name or a system name. Window 604 shows a list of one or more listeners 604g who can observe the data quality issue or receive alerts or other messages based on a workflow associated with the data quality issue. Here, the listener is a person called Melodie M. Pickett. In some implementations, the listener may be a device or a system.

[0109] Window 606 shows a list of recommended actions for the data quality issue. In one example, the workflow for the data quality issue can dictate the actions shown in window 606. Here, the "Full Triage" action item is shown because the state 616 of the data quality issue is "In Triage."

[0110] Window 608 shows data describing data quality measurements related to the data quality issue. Timeline data 610 is included in window 608. Data quality measurement data specifies when data quality tests (also called measurements) are performed. Window 608 indicates when the result data was received. Window 608 shows a portion of the result data from one or more previous data quality tests, if applicable. If a data quality issue is reoccurring, the data in window 608 allows a user to quickly see trends in the data quality issue over time.

[0111] Window 612 provides space for additional comments provided (e.g., by the assigned entity). The comments can supplement the attribute values ​​inherited by the data quality issue.

[0112] The data quality issue shown in screen 600a is associated with state 616. In this example, the data quality issue is in an early stage called "in triage," where a responsible entity is required to triage the data quality issue. If the data quality issue is determined to be critical, the data quality issue can be escalated, or one or more other actions can be taken to resolve the data quality issue.

[0113] Figure 6B shows a screen 600b depicting control point #2010 (called DQCP-2010) associated with DCI-2 in Figure 6A. This control point constitutes the data quality test that causes the data quality problem in Figure 6A.

[0114] Screen 600b includes a general information window 624 similar to window 604 of FIG. 6A. Window 624 includes a control point name 624a. In this example, the control point name 624a is DQCP-2010, or Data Quality Control Point #2010. Window 624 shows the type of represented object 624b, which is the data quality control point. Window 624 shows a data quality control point description 624c. Here, the description describes the data quality test configured by the data quality control point. The goal of the data quality test for this control point is "Product Type should be a valid value in Auto Loan." This indicates that the data quality test is configured to check for a valid value in a field called Auto Loan to check for Product Type (e.g., in another field). Description 624d provides further information, specifying that "Product Type cannot be NULL or empty in Auto Loan." This description 624d indicates which data quality rule is being applied. For example, values ​​in the Auto Loan field cannot be NULL or empty. If any rule condition is not met, that value of the field fails the data quality test. In some implementations, the result data specifies exactly which rule or condition for the data quality test was not met. In some implementations, the result data indicates that at least one condition of a data quality test was not met for a particular field without specifying which condition was not met. Window 624 specifies which data quality metric 624e is being tested. In the illustrated example, metric 624f is the completeness of the values ​​contained in the field of each record. Window 624 specifies test unit 624f, which indicates how much data is being tested.

[0115] Screen 600b includes a window 626 for specifying the source of the data being tested. For example, window 626 specifies a business data element 626a that contains the product being tested. The path 626b of data is from the Data Lake > Retail Banking > Auto Loans dataset.

[0116] The Data Management window 628 specifies aspects of data quality testing. Window 628 specifies data quality business requirements 628a, which are the goals of a particular data quality test. Here, the goal is completeness of the data set, and more specifically, that each record in the data set be populated with a valid value. The data quality test category 628b is monitoring. The frequency 628c of data quality testing is shown in window 628. Here, the frequency 628c is weekly data quality testing.

[0117] The control ownership window 630 specifies one or more responsible entities for resolving the data quality issue. The owner of data quality control point #2010 is the same as its associated data quality issue DCI-2. The assigned entity is Abbey A. Williams. The data quality analyst assigned to the hearing is Melodie M. Pickett.

[0118] The specification window 632 includes a specification (e.g., workflow) identifier. In this example, a workflow is not identified. The specification can be a graph or other logic.

[0119] The implementation window 634 includes data specifying when the data quality tests are being run, where an initial data quality test timestamp 634a and a most recent data quality test timestamp 634b are shown.

[0120] Window 620 shows data quality test result data 638. Result data 638a includes a representation of the data quality test results, such as the percentage of all records that passed, or the absolute number of records that failed or passed. Result data 638 includes a representation of trend data 638b. Trend data 638b indicates how the results of the data quality test have changed from one or more previous tests to the most recent test, either as an improvement or a deterioration in data quality.

[0121] The data quality control point is associated with a status value 636, which is different from the state of the data quality issue 638 associated with the data quality control point. Here, the status 636 is "Published."

[0122] The data quality control point screen 600b shows each of the data quality issues 638 associated with the control point, where data quality issue DQI-2 is shown inserted into the data quality control data (e.g., as the summary data shown in FIG. 5).

[0123] FIG. 7 shows a portion of a user interface screen 700 depicting a timeline 702 of a data quality issue. Here, the data quality issue state 704 is "In Progress." The timeline 702 shows one or more entities and when any actions were taken on each data quality issue. Here, user Wade L. Register commented, edited, and accepted an assignee change to the data quality issue. The managing entity changed the state of the data quality issue by completing triage. In one example, this may include an automated process.

[0124] FIG. 8 illustrates a process 800 for automated data quality issue generation and workflow generation. Process 800 may be implemented by a data processing system such as those described above. Process 800 includes receiving result data specifying the output of a data quality test on a data record (802); receiving metadata associated with the result data specifying data quality conditions of the data quality test and the objects that comprised the data quality test (804); generating a data quality issue specifying a workflow state for processing the result data based on the result data and the metadata (806); and generating a workflow for processing the result data based on the state specified by the data quality issue (808). Generating the workflow includes assigning an entity responsible for resolving the data quality error based on the result data and the state of the data quality issue (810); determining one or more actions to satisfy the data quality conditions specified in the data quality test based on the metadata (812); and updating the state associated with the data quality issue (814). The process also includes sending data representing the workflow to a workflow execution system (816).

[0125] Some implementations of the subject matter and operations described herein can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, or in one or more combinations thereof, including the structures disclosed herein and their structural equivalents. For example, in some implementations, the modules of data processing device 102 can be implemented using digital electronic circuitry, or in computer software, firmware, or hardware, or in one or more combinations thereof. In another example, processes 900, 1000, 1100, and 1200 can be implemented using digital electronic circuitry, or in computer software, firmware, or hardware, or in one or more combinations thereof.

[0126] Some implementations described herein can be implemented as one or more groups or modules of digital electronic circuitry, computer software, firmware, or hardware, or as a combination of one or more of them. Although different modules can be used, each module need not be separate; multiple modules can be implemented in the same digital electronic circuitry, computer software, firmware, or hardware, or as a combination of them.

[0127] Some implementations described herein may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by or controlling the operation of a data processing apparatus. A computer storage medium may be or be included in a computer-readable storage device, a computer-readable storage substrate, a random-access or serial-access memory array or device, or one or more combinations thereof. Further, a computer storage medium is not a propagating signal, although a computer storage medium may be the source or target of computer program instructions encoded in an artificially generated propagating signal. A computer storage medium may also be or be included in one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).

[0128] The term "data processing device" encompasses all kinds of devices, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, a system-on-chip, or a plurality or combination of the above. In some implementations, data processing system 100 comprises a data processing device as described herein. The device may include special-purpose logic circuitry, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). In addition to hardware, the device may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or one or more combinations thereof. The device and execution environment may implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0129] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted, declarative or procedural languages. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program can be deployed to run on one computer, or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.

[0130] Some of the processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform actions by manipulating input data and generating output. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, such as, for example, an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0131] Processors suitable for executing a computer program include, for example, both general-purpose and special-purpose microprocessors, as well as processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. A computer includes a processor for performing actions in accordance with the instructions and one or more memory devices for storing instructions and data. A computer may also include, or be operatively coupled to receive data from, transfer data to, or do both, one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data. However, a computer need not have such devices. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices (e.g., EPROM, EEPROM, flash memory devices, etc.), magnetic disks (e.g., internal hard disks, removable disks, etc.), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.

[0132] To provide for user interaction, actions can be performed on a computer that has a display device (e.g., a monitor or another type of display device) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse, trackball, tablet, touch-sensitive screen, or another type of pointing device) by which the user can provide input to the computer. Other types of devices can also be used to provide for user interaction. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or haptic feedback, and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, a computer can interact with a user by sending documents to and receiving documents from a device used by the user, e.g., by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0133] A computer system may include a single computing device or multiple computers operating within close proximity or generally remotely from each other and typically interacting through a communications network. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), internetworks (e.g., the Internet), networks including satellite links, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks). The relationship of client and server may arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0134] An exemplary computer system includes a processor, memory, storage devices, and input / output devices. Each of the components may be interconnected, for example, by a system bus. The processor may process instructions for execution within the system. In some implementations, the processor is a single-threaded processor, a multi-threaded processor, or another type of processor. The processor may process instructions stored in memory or storage devices. The memory and storage devices may store information within the system.

[0135] The input / output devices provide input and output operations to the system. In some implementations, the input / output devices may include one or more of a network interface device, such as an Ethernet card, a serial communication device, such as an RS-232 port, and / or a wireless interface device, such as an 802.11 card, a 3G wireless modem, a 4G wireless modem, a 5G wireless modem, etc. In some implementations, the input / output devices may include a driver device configured to receive input data and send output data to other input / output devices, such as keyboards, printers, and display devices. In some implementations, mobile computing devices, mobile communication devices, and other devices may be used.

[0136] While this specification contains many details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features specific to particular examples. Certain features described herein in the context of separate implementations may also be combined. Conversely, various features that are described in the context of a single implementation may also be implemented in multiple embodiments separately or in any suitable subcombination.

[0137] Although several embodiments have been described, it will be understood that various modifications can be made without departing from the scope of the data processing system described herein. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. 1. A method for executing, by a data processing system, a workflow for resolving data quality errors associated with the data quality tests, by processing result data indicative of an output of a data quality test on a data record in response to receiving the result data and metadata describing the result data to generate data quality issues associated with a state of the workflow, the method comprising: The method comprises: receiving the results data specifying the output of the data quality test on the data record; receiving the metadata associated with the result data specifying the data quality test; generating a data quality problem based on the results data and the metadata, the data quality problem including a state of the workflow, the state indicating one or more processing actions in the workflow for processing the results data to resolve the data quality problem; executing the workflow to process the results data based on the state specified by the data quality issue; Executing the workflow, assigning an entity responsible for resolving the data quality error based on the results data and the status of the data quality problem; determining one or more processing actions for the workflow to resolve the data quality problem based on the metadata; and updating the status of the data quality problem when the one or more processing actions are completed; storing data representing an updated state of the data quality problem, the updated state representing one or more subsequent processing actions for resolving the data quality problem by the workflow; and and A method comprising:

2. The data quality problem is associated with a key value that identifies the data quality problem, and the process comprises: determining, in response to receiving the metadata, by evaluating the key value, that the data quality problem has already been generated for the data source of the data record for the data quality test; updating the already generated data quality questions based on the received result data of the data quality test and the received metadata; The method of claim 1 , wherein updating the data quality issue comprises changing the state specified by the data quality issue.

3. The method of claim 2 , wherein updating the data quality problem comprises reallocating the resulting data to a different entity based on the updated state specified by the data quality problem.

4. 3. The method of claim 2, wherein updating the data quality issue includes adding a timeline for a processing action associated with the data quality issue based on the updated state specified by the data quality issue.

5. the metadata specifies data quality conditions for the data quality test, and the method comprises:

3. The method of claim 2, further comprising: determining that the data quality condition is satisfied in response to updating the data quality problem; and wherein changing the state associated with the data quality problem comprises indicating that the data quality problem is associated with a resolved data quality error.

6. receiving scheduling data to associate with the data quality problem, the scheduling data causing the data quality test to be run at predetermined time intervals; identifying the time interval associated with the data quality test; The method of claim 1 , further comprising: transmitting data representing the time interval to the assigned entity.

7. allocating the entity, determining, from a data source of the data records, a set of eligible entities associated with the data source that are responsible for resolving the data quality errors; receiving selection criteria from the data source for allocating the entity from the set of eligible entities; and assigning the entity based on the selection criteria.

8. 8. The method of claim 7, wherein the selection criteria specify a type of data quality error, the metadata associated with the results data specifies that the data quality problem is associated with the type of data quality error, and the entity is eligible to resolve the type of data quality error.

9. 8. The method of claim 7, wherein the selection criteria specify a particular state of the data quality problem for each entity in the set of eligible entities, the state of the data quality problem being the particular state for the assigned entity.

10. 2. The method of claim 1, wherein the entity is a particular device of the data processing system.

11. 2. The method of claim 1, wherein the entity is a user of the data processing system.

12. Executing the workflow, generating a timeline of actions taken to process the results data; The method of claim 1 , further comprising: sending data to the entity representing a timeline of the actions taken.

13. The method of claim 1 , wherein each state of the data quality problem is associated with one or more candidate actions for processing the resulting data.

14. 10. The method of claim 1, further comprising determining a category of the data quality error selected from one of a data error, a rule error, or a reference data error based on the results data and the metadata.

15. 2. The method of claim 1, wherein the data quality problem includes an identification of the data record being tested by the data quality test to obtain the received results data, data specifying which one or more data quality rules were applied to the data record in the data quality test, and a result for the application of each of the data quality rules to the data record.

16. 3. The method of claim 2, wherein the key value associates the data quality issue with additional data, the additional data including an identifier for a client device associated with the data quality issue, an owner of the data record associated with the data quality issue, an entity responsible for resolving the data quality issue, the workflow for correcting the data quality issue, and the state of the workflow.

17. The method of claim 1 , wherein the data quality problem specifies one or more processing actions for the workflow to process the resulting data to resolve the data quality problem.

18. 2. The method of claim 1, wherein the results data specifies, for a portion of the data record, whether the portion satisfies conditions specified in one or more data quality rules applied to the data record during the data quality test, and the results data identifies one or more conditions of the data quality rules that were not met during the test and which portions of the data record did not satisfy the conditions.

19. The method of claim 1 , wherein the result data specifies how the data quality test failed, including by failing to meet a data quality condition of a data quality rule of the test.

20. The method of claim 1 , wherein the one or more processing actions include analyzing values ​​of the data records to determine whether a data quality rule definition of the data quality test is flawed.

21. 2. The method of claim 1 , wherein the metadata specifies data quality conditions of the data quality test and objects that comprise the data quality test, and the actions are for satisfying the data quality conditions specified in the data quality test.

22. 10. The method of claim 1, further comprising configuring the workflow using configurable data processing rules for the workflow based on the received metadata such that different workflows are generated based on which data quality tests failed and how the tests failed.

23. The method of claim 2 , wherein the updating of the data quality problem is performed without creating a new data quality problem.

24. 1. A system for executing, by a data processing system, a workflow for processing result data indicative of an output of a data quality test on a data record by generating data quality issues associated with a state and one or more processing steps of the workflow in response to receiving the result data and metadata describing the result data to resolve data quality errors associated with the data quality test, the system comprising: at least one processor; a memory for storing instructions; The instructions, when executed by the at least one processor, cause the at least one processor to: receiving result data specifying an output of a data quality test on the data record; receiving metadata associated with the results data that specifies data quality conditions of the data quality test and objects that constituted the data quality test; generating a data quality problem based on the results data and the metadata, the data quality problem including a state in the workflow for the data quality problem, the state indicating one or more processing actions in the workflow for processing the results data to resolve the data quality problem; executing the workflow for processing the results data based on the status of the data quality problem; Executing the workflow, assigning an entity responsible for resolving the data quality error based on the results data and the status of the data quality problem; determining, based on the metadata, one or more actions to satisfy the data quality conditions specified in the data quality test; updating the status associated with the data quality issue when the one or more processing actions are completed; storing data representing an updated state of the data quality problem, the updated state representing one or more subsequent processing actions for resolving the data quality problem by the workflow; and and A system that causes an operation including

25. The data quality problem is associated with a key value that identifies the data quality problem, and the process comprises: responsive to receiving the metadata, determining that the data quality issue has already been generated for the data source of the data record for the data quality test by evaluating the key value; and updating the data quality issue based on the result data of the data quality test, wherein updating the data quality issue includes changing the state associated with the data quality issue.

26. 26. The system of claim 25, wherein updating the data quality issue comprises reassigning the results data to a different entity.

27. 26. The system of claim 25, wherein updating the data quality issue includes adding a timeline of a processing action associated with the data quality issue.

28. one or more non-transitory computer-readable media storing instructions for executing, by a data processing system, a workflow for processing result data indicative of an output of a data quality test on a data record in response to receiving the result data and metadata describing the result data by generating data quality issues associated with a state and one or more processing steps of the workflow to resolve data quality errors associated with the data quality test, wherein execution of the instructions by at least one processor causes the at least one processor to: receiving result data specifying an output of a data quality test on the data record; receiving metadata associated with the results data that specifies data quality conditions of the data quality test and objects that constituted the data quality test; generating a data quality issue based on the results data and the metadata, the data quality issue including a state for the data quality issue in the workflow, the state indicating one or more processing actions in the workflow for processing the results data; executing the workflow for processing the results data based on the status of the data quality problem; Executing the workflow, assigning an entity responsible for resolving the data quality error based on the results data and the status of the data quality problem; determining, based on the metadata, one or more actions to satisfy the data quality conditions specified in the data quality test; updating the status associated with the data quality issue when the one or more processing actions are completed; storing data representing an updated state of the data quality problem, the updated state representing one or more subsequent processing actions for resolving the data quality problem by the workflow; and and One or more non-transitory computer-readable media for causing operations to be performed, including:

29. The data quality problem is associated with a key value that identifies the data quality problem, and the process comprises: responsive to receiving the metadata, determining that the data quality issue has already been generated for the data source of the data record for the data quality test by evaluating the key value; and updating the data quality problem based on the result data of the data quality test, wherein updating the data quality problem comprises changing the state associated with the data quality problem.

Citation Information

Patent Citations

  • Inspection management system and inspection management device and inspection management method

    JP2007179363A

  • Management device, management method, and program

    JP2014048699A

  • Proactive automated data validation

    US20200210401A1