Methods, systems, and media for predicting data quality in dataset

By analyzing the data set using an associative machine learning algorithm, identifying the relationship between data columns and identifying exceptions, the problems of low efficiency and poor accuracy of data exception confirmation in the prior art are solved, and automated data quality prediction and exception repair are achieved.

CN120197086APending Publication Date: 2025-06-24COLLIBRA NV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510256218.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2021-04-21
Filing Date
2022-04-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively confirm data abnormalities in the data set, and manually confirming the relationship between data points is time-consuming and error-prone, resulting in inaccuracy and inefficiency in predicting data quality.

Method used

The data set is analyzed using associative machine learning algorithms (such as random forests and frequent pattern growth), identify the relationships between data columns, and identify exceptions and make repair suggestions by generating project sets and applying a second machine learning algorithm.

Benefits of technology

Automatic data exception identification and repair is realized, which improves the accuracy and efficiency of data quality prediction and reduces the possibility of manual intervention and errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197086A_ABST
    Figure CN120197086A_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods, systems, and media for predicting data quality in a dataset. The method comprises the following steps: receiving and analyzing a first data set; applying a first machine learning algorithm to the first data set, the first machine learning algorithm identifying at least one relationship between a first data column and a second data column in the first data set; generating a second data set based on the identification of at least one relationship between the first data column and the second data column, the second data set being a subset of the first data set; connecting a plurality of column headers in the second data set; generating an item set based on the connection of the plurality of column headers in the second data set; applying a second machine learning algorithm to the item set, wherein the second machine learning algorithm identifies at least one frequency value related to at least one relationship between a first data column and a second data column in the first data set; identifying at least one anomaly in the item set; and generating at least one suggestion for repairing the at least one anomaly.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application is a divisional application of the following invention patent application:

[0002] Application No.: 202280042703.3

[0003] Filing Date: April 5, 2022 Invention Title: Systems and Methods for Predicting Correct or Missing Data and Data Anomalies Technical Field

[0004] This disclosure relates to systems and methods for predicting correct or missing data, data anomalies, and overall data quality using associative machine learning algorithms such as random forests and frequent pattern growth. Background Art

[0005] Entities maintain large amounts of data that may be related (or associated) with certain data classifications. Compared to column-level validation, such entities may have difficulty understanding the attribute relationships within a dataset and validating record-level information. For example, a zip code is typically related to a street address and a state. An entity may have a valid zip code and a valid state code, but the zip code and state code together do not form a valid record (e.g., because the zip code is not within the state). In these datasets, there may be data anomalies such as the one just described. Currently, entities must manually confirm these data anomalies by determining whether certain data points are actually anomalous. To make this determination, entities typically must confirm certain relationships between data points (e.g., a zip code is related to a street address and has a specific numerical format consisting of nine digits). Confirming these relationships in a dataset is both cumbersome and time-consuming. The current process of attempting to confirm relationships also produces errors because the manual confirmation of data relationships often overlooks or simply misses the relationships between certain data points. Thus, the current methods for confirming data anomalies are inconsistent and inefficient, partly because confirming data anomalies typically requires evaluating the relationships (or lack thereof) between data points. If the specific data received appears unusual compared to the data received previously, this may indicate an anomaly, but it may also be a false positive. Similarly, the specific data received is not frequently received from a particular entity can also indicate an anomaly, or alternatively, it may not be an anomaly but just a less frequently received data point. Unraveling large datasets to confirm data anomalies is necessary for predicting the data quality of a dataset, and currently, due to the error-prone and inefficient methods for confirming data anomalies today, accurately predicting the correct or missing data for a given record and the overall data quality of a dataset is difficult and unreliable.

[0006] Thus, there is an increasing need for systems and methods that can address the challenges of confirming data anomalies, confirming data relationships, and accurately and effectively predicting correct datasets.

[0007] In view of these and other general considerations, various aspects are disclosed herein. Additionally, although relatively specific problems may be discussed, it should be understood that the examples should not be limited to solving specific problems identified in the background of the present disclosure or elsewhere. Summary of the Invention

[0008] In some aspects, a method for predicting data quality in a dataset is provided, including: receiving a first dataset; profiling the first dataset; applying a first machine learning algorithm to the first dataset, wherein the first machine learning algorithm identifies at least one relationship between a first data column and a second data column in the first dataset; generating a second dataset based on the identification of the at least one relationship between the first data column and the second data column in the first dataset, wherein the second dataset is a subset of the first dataset; concatenating a plurality of column headers in the second dataset; generating an item set based on the concatenation of the plurality of column headers in the second dataset; applying a second machine learning algorithm to the item set, wherein the second machine learning algorithm identifies at least one frequency value associated with the at least one relationship between the first data column and the second data column in the first dataset; identifying at least one anomaly in the item set; and generating at least one recommendation for fixing the at least one anomaly.

[0009] In some aspects, a system for correcting data anomalies in a dataset is provided, including: a memory configured to store non-transitory computer-readable instructions; and a processor communicatively coupled to the memory, wherein the processor, when executing the non-transitory computer-readable instructions, is configured to: receive a dataset; profile the dataset; apply a first machine learning algorithm to the dataset, wherein the first machine learning algorithm identifies at least one relationship between a first data column and a second data column in the dataset; generate an item set based on the identification of the at least one relationship between the first data column and the second data column in the dataset; apply a second machine learning algorithm to the item set, wherein the second machine learning algorithm identifies at least one frequency value associated with the at least one relationship between the first data column and the second data column in the dataset; identify at least one anomaly in at least one data record in the item set; and automatically fix the at least one anomaly in at least one data record in the item set.

[0010] In some aspects, a computer-readable medium is provided that stores non-transitory computer-executable instructions that, when executed, cause a computing system to perform a method for predicting and correcting data anomalies. The method includes: receiving a first data set; parsing the first data set; comparing a first data column in the first data set with multiple other data columns in the first data set; based on a comparison result of the first data column and the multiple other data columns, identifying a first relationship between the first data column and a second data column in the first data set; applying a first machine learning algorithm to the first data column and the second data column, where the first machine learning algorithm identifies a second relationship between the first data column and the second data column in the first data set; based on the identification of the second relationship between the first data column and the second data column in the first data set, generating a second data set, where the second data set is a subset of the first data set; concatenating multiple column headers in the second data set; based on the concatenation of the multiple column headers in the second data set, generating an item set; applying a second machine learning algorithm to the item set, where the second machine learning algorithm identifies at least one frequency value related to the second relationship between the first data column and the second data column in the first data set; based on the at least one frequency value, identifying at least one outlier in the item set, where the at least one outlier is related to a low frequency value; and fixing the at least one outlier by substituting the at least one outlier with a substitute value related to a high frequency value. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 Shows an example of a distributed system for predicting correct or missing data and data quality as described herein.

[0012] Figure 2 Shows an exemplary input processor for predicting correct or missing data and data quality as described herein.

[0013] Figure 3A and 3B Shows an exemplary method for predicting correct or missing data and data quality of a data set as described herein.

[0014] Figures 4A - 4C Shows an exemplary architecture for implementing systems and methods for predicting correct or missing data and data quality.

[0015] Figure 5 Shows an example of a suitable operating environment in which one or more embodiments of the present embodiment can be implemented. DETAILED DESCRIPTION

[0016] Aspects of the present disclosure are described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show specific exemplary aspects. However, the different aspects of the present disclosure may be implemented in many different forms and should not be construed as limited to the aspects set forth herein; rather, these aspects are provided so that this disclosure will be thorough and complete, and will fully convey the scope of these aspects to those skilled in the art. The various aspects may be practiced as a method, system, or apparatus. Accordingly, the various aspects may take the form of a hardware implementation, an entirely software implementation, or an implementation combining software and hardware aspects. Thus, the following detailed description should not be construed as limiting.

[0017] Figure 1 An example of a distributed system for predicting correct or missing data and data quality, as described herein, is shown. The proposed exemplary system 100 is a combination of interdependent components that interact to form an integrated whole for integrating and enriching data in a data marketplace. The components of the system may be hardware components or software implemented on and / or executed by the hardware components of the system. For example, system 100 includes client devices 102, 104, and 106, local databases 110, 112, and 114, network 108, and server devices 116, 118, and / or 120.

[0018] Client devices 102, 104, and 106 can be configured to receive and transmit data. For example, client devices 102, 104, and 106 can include client-specific data. The client devices can also be configured to receive data that can be ingested and analyzed at client devices 102, 104, and / or 106 from various data sources (e.g., data lakes, databases, flat files, data streams, etc.) via one or more networks 108. The received data can be stored in local databases 110, 112, and 114. The received data can be analyzed by applying at least one machine learning (ML) algorithm to the data. The received data can be transmitted to one or more servers 116, 118, and / or 120 via one or more networks 108 and / or satellite 122. One or more servers 116, 118, and / or 120 can be third-party servers and / or client-owned servers. Depending on, for example, the size of the data, the complexity of the analysis to be performed on the data, etc., the data can be transmitted to one or more servers 116, 118, and / or 120 for further processing, analysis, and / or storage. In other examples, the received data can be stored in a server (in addition to or instead of the local client devices and local databases) and can be analyzed to determine correct or missing data and the data quality of the data set. The received data along with a data quality predictor (e.g., a score, a table, etc.) can be transmitted from the client server to a third-party server via one or more networks 108 and / or satellite 122.

[0019] In various aspects, client devices (e.g., client devices 102, 104, and 106) can have access to one or more data sets or data sources and / or databases that include client-specific data. In other aspects, client devices 102, 104, and 106 can be equipped to receive broadband and / or satellite signals carrying customer data. The data can include raw data sets, data sets that have been analyzed (e.g., that have had at least one ML algorithm applied), and / or a mixture of both. The signals and information that client devices 102, 104, and 106 can receive can be transmitted from satellite 122. Satellite 122 can also be configured to communicate with one or more networks 108 and, in addition, be capable of communicating directly with client devices 102, 104, and 106. In some examples, client devices can be mobile phones, laptops, tablets, smart home devices, landline phones, and wearable devices (e.g., smartwatches) and other devices.

[0020] To further elaborate on the network topology, client devices 102, 104, and / or 106 (along with their corresponding local databases 110, 112, and 114) can receive data from the client. On client devices 102, 104, and / or 106, a profiler and / or a rule engine can be installed and applied to the received data. The comparison and analysis of the received data can occur locally on client devices 102, 104, and / or 106. In an alternative embodiment, the received data can be compared and analyzed at one or more of servers 116, 118, and / or 120. Once the comparison and analysis are performed, the received data can be manipulated to identify which data columns have the strongest relationships. In some examples, a predictive relationship / association representor can be generated. The predictive representor can be generated at one or more client devices 102, 104, and / or 106, one or more servers 116, 118, and / or 120, or a combination of both. The predictive representor output can be generated locally or externally and shared via one or more networks 108.

[0021] Once the associations of the data set are identified, transaction item sets can be generated. The item sets can be generated at one or more client devices 102, 104, and / or 106 or externally at one or more of servers 116, 118, and / or 120. A second ML algorithm (e.g., the frequent pattern growth (FPG) algorithm) can then be used to analyze the transaction item sets to confirm the frequency of the identified data associations. A data quality representor can be generated, which represents which data associations are likely to be less frequent and present in the data set. The less frequent data associations can indicate data anomalies. The output can be generated locally at one or more client devices 102, 104, and / or 106 and / or one or more of servers 116, 118, and / or 120. Regardless of where the DQ output is generated, the DQ output can be shared with various devices via one or more networks 108.

[0022] Figure 2An exemplary input processor for predicting correct or missing data and data quality as described herein is shown. The input processor 200 can be embedded in client devices (e.g., client devices 102, 104, and / or 106), remote network server devices (e.g., devices 116, 118, and / or 120), and other devices capable of implementing systems and methods for predicting correct or missing data and data quality. The input processing system includes one or more data processors and is capable of executing algorithms, software routines, and / or instructions based on processed data provided by at least one data source (e.g., data lake, database, flat file, data stream, etc.). The input processing system can be a factory-installed system or an add-on unit for a particular device. Additionally, the input processing system can be a general-purpose computer or a specialized dedicated computer. There is no limitation on the location of the input processing system relative to client or remote network server devices, etc.

[0023] According to Figure 2 In the embodiment shown, the disclosed system can include a memory 205, one or more processors 210, a communication module 215, a data collection module 220, an ML algorithm module 225, and a data or DQ recommendation module 230. Other embodiments of the present technology can include some, all, or none of these modules and components, as well as other modules, applications, data, and / or components. However, some embodiments can combine two or more of these modules and components into a single module and / or associate a portion of the functionality of one or more of these modules with a different module.

[0024] Memory 205 may store instructions for running one or more applications or modules on one or more processors 210. For example, memory 205 may be used in one or more embodiments to hold all or part of the instructions required to perform the functions of the data collection module 220, the ML algorithm module 225, and / or the data / DQ recommendation module 230, as well as the communication module 215. Generally, memory 205 may include any device, mechanism, or populated data structure for storing information. According to some embodiments of the present disclosure, memory 205 may encompass, but is not limited to, any type of volatile memory, non-volatile memory, and dynamic memory. For example, memory 205 may be random access memory, a memory storage device, an optical memory device, a magnetic medium, a floppy disk, a magnetic tape, a hard disk drive, a SIMM, SDRAM, RDRAM, DDR, RAM, a SODIMM, an EPROM, an EEPROM, an optical disc, a DVD, and / or the like. According to some embodiments, memory 205 may include one or more disk drives, flash drives, one or more databases, one or more tables, one or more files, local cache memory, processor cache memory, relational databases, flat databases, and / or the like. Additionally, those of ordinary skill in the art will understand that many additional devices and techniques for storing information may be used as memory 205.

[0025] The communication module 215 is related to sending / receiving information (e.g., data collected by the data collection module 220, data analyzed / processed according to the ML algorithm module 225, and / or predicting correct or missing data, as well as data quality information from the DQ recommendation module 230), and commands received via a client device or a server device, other client devices, a remote network server, etc. These communications may employ any suitable type of technology, such as Bluetooth, WiFi, WiMax, cellular (e.g., 5G), single-hop communication, multi-hop communication, dedicated short-range communication (DSRC), or proprietary communication protocols. In some embodiments, the communication module 215 sends information (e.g., data) collected by the data collection module 220 (e.g., data in batches, large quantities, and / or continuous feeds), processed by the ML algorithm module 225 (e.g., a dataset can be profiled and analyzed using a random forest algorithm to predict associations, and the frequency of associations can be further analyzed using a frequent pattern growth algorithm), and generated by the DQ recommendation module 230 (e.g.), and / or sends it to the client devices 102, 104, and / or 106, as well as the memory 205 for storage for future use. In some examples, the communication module may be built on the HTTP protocol using one or more secure REST servers using RESTful services.

[0026] The data collection module 220 is configured to receive and profile data. For example, the data collection module 220 can receive data related to a data set. The data collection module 220 can then profile the data, including capturing statistical data, including but not limited to one or more minimum values, one or more maximum values, cardinality, confidence frequency, and categorical value encoding, as well as other statistical data. In some examples, cardinality can refer to the number of distinct values captured in a data set.

[0027] The data collection module 220 can also be configured to identify certain data values with inherently low confidence, low support, and / or high frequency. Such data values can be excluded by the data collection module 220 before being sent to the ML algorithm module 225.

[0028] One or more ML algorithm modules 225 are configured to receive a data set and apply at least one machine learning (ML) algorithm to the data set. For example, after the data set is received and profiled by the data collection module 220, the data set can be transmitted to the ML algorithm module 225. The ML algorithm module 225 can be configured to analyze the data set, and specifically, analyze the columns of the data set. Each column label can be compared with other columns to identify the relationships between the columns in the data set.

[0029] To identify the relationships between columns, the ML algorithm module 225 can apply a random forest (RF) algorithm to the data set. The features and attributes required to initiate / build the ML algorithm can be obtained during profiling or other analysis steps. The data set can be profiled first, and the descriptive statistics obtained from the profiling step can be used as input to the RF algorithm. The random forest algorithm can build multiple decision trees and combine them together to obtain a more accurate and stable prediction of the relationships between the data sets. Other algorithms that can be applied to the data set include but are not limited to classifier models, XGBoost (and other tree-based ensemble machine learning algorithms that utilize gradient boosting), and other ML models from libraries such as MLlib (the machine learning library of Apache Spark).

[0030] After applying the RF algorithm (or other ML algorithms), an output can be generated that represents the strongest relationships between the data columns. For example, a relationship representing the relationship between the column label "zip code" and the column label "street address" can be represented, as zip codes are typically associated with street addresses. The label "SSN" can be related to the column labels "first name" and "last name" as social security numbers are typically associated with an individual's identity. These data columns can be grouped together based on feature importance, interpretation coefficients, and other model accuracy / relationship determination.

[0031] Once the relevant data columns are identified, the headers of the data can be joined for each row. The ML algorithm module 225 can be configured to create a new transaction dataset called an "itemset" based on the joining of the column headers for each data row. The ML algorithm module 225 can then apply a second ML algorithm, such as the frequent pattern growth algorithm, to further predict the frequency of the relationships. In other words, the application of the RF algorithm can indicate whether certain data items are related, and the additional application of the frequent pattern growth algorithm can indicate how frequently these relationships occur in the dataset.

[0032] The ML algorithm module 225 can be configured to calculate a probability matrix based on the frequency of the data relationships. The frequency of the relationship baskets can indicate whether there are certain data anomalies in the dataset. Less frequent relationships can indicate a higher likelihood of data anomalies, while more frequent relationships can indicate a lower likelihood of data anomalies.

[0033] After the dataset is processed by the ML algorithm module 225, the DQ recommendation module 230 can be configured to receive processing information from the ML algorithm module 225 representing the frequency of certain data relationships (i.e., certain data columns are related to other data columns). In some instances, the DQ recommendation module 230 can output intelligent recommendations suggesting certain actions to be taken on the data to fix any identified data anomalies. In other examples, a table representing the associations between certain data records and the frequency of these data relationships can be displayed (e.g., the system can highlight potential relationships in an output table showing two records with the same SSN, name, and gender but two different email addresses). Less frequent data relationships can be highlighted to indicate outlier data points or suspicious data that may be incorrect.

[0034] Figure 3A and 3BIllustrates an exemplary method for predicting data quality of a data set as described herein. Method 300 is divided into two separate sub-methods - sub-method 300A and sub-method 300B. Sub-method 300A is a relationship finder sub-method that identifies relationships (if any) between data sets. Sub-method 300A begins with the step of receiving data 302. Data can be received in the form of a data lake, database, flat file, data stream, or other ingestible data format. The data can then be analyzed at step 304 to ensure that the data is in a format readable by the systems and methods described herein (e.g., step 304 can include a user interface button that can be pressed to initiate profiling). The data can be profiled at step 306, where certain data statistics can be captured, such as minimum, maximum, average, cardinality indicators, categorical columns, etc. After profiling the data values, a brute-force comparison can then be applied to the data set at step 308. For example, exemplary attribute A can be compared to all other attributes in the data set, and exemplary attribute B can be compared to all other attributes in the data set. The brute-force comparison output can display a table representing which column headers may be related to the data values. For example, the following table can be displayed for an exemplary comparison of the 1st column ("col1") in the data set with other columns:

[0035] Comparison column Relationship percentage (brute force) Col5 98% Col2 72% Col3 10% Col6 1%

[0036] At step 310, a first machine learning algorithm can be applied to the data set. The first ML algorithm can be a random forest algorithm, which can be applied to understand which data columns have potential relationships with each other. At step 312, those column relationships can be confirmed based on the output of the RF algorithm. For example, a zip code data column can be related to a street address. A bank account data column can be related to a bank name data column. An SSN data column can be related to a first name data column and a last name data column. An IP address data column can be related to a machine identifier data column.

[0037] Once the relationships between the data columns are confirmed, at step 314, based on the automatic selection of the data column relationships and / or the user-selected relationships of the data columns, the results of these relationships are output and added to a basket. The data column relationship basket is then analyzed at step 316, where the data column headers are joined to create transaction item sets at step 318. Then at step 320, the transaction item sets from step 318 are provided as input to a second ML algorithm, such as a frequent pattern growth model. At step 322, the FPG model can calculate a probability matrix of item set combinations to determine specific patterns in the item sets. The patterns confirmed at step 322 can represent the frequency (or lack of relationships) between certain data columns.

[0038] Based on the frequency of data relationships confirmed by the FPG model and the information derived from the RF model, a data prediction can be generated at step 324, which can predict correct or missing data and the data quality of the data set. For example, less frequent relationships in the data may indicate that certain data values are abnormal or incorrect. More frequent relationships in the data may indicate that the data values include fewer anomalies and errors.

[0039] In one specific example, a bank may be analyzing loan data. The data may be limited to a specific geographical area. Some relationships in the data may be confirmed by the RF algorithm at step 310, such as the relationship between certain bank customer account numbers and SSN numbers. Additionally, the frequency of these relationships may be confirmed by applying the FPG model at step 320. However, although the DQ prediction system described herein can confirm a certain relationship, the relationship may be suspect, for example, because the claimed account number may not comply with the standard format guidelines for account numbers (e.g., an account number with 6 digits instead of 8 digits). Thus, the system can flag the data value as an anomaly because the data value has a low confidence score. Another example of anomaly detection can be based on the transaction volume in a specific geographical area. Although there may be a relationship between data columns, the frequency of a specific transaction volume associated with a specific type of bank transaction (e.g., deposit, withdrawal, loan, etc.) may be low. Since the frequency of a specific relationship in the data set is low, the system can also flag this transaction as a potential anomaly.

[0040] In yet another example, the data set may include columns with zip codes and state identifier codes. One data entry may include a zip code "21087" with a state identifier code of "MD" (Maryland), while another data entry may include the same zip code "21087" but with a state code of "NY" (New York). The system described herein can confirm the data anomaly, apply at least one machine learning algorithm to the data set, and suggest at least one corrective action (i.e., confirm which data entry is correct and which data entry is incorrect). Although the zip code itself is in the zip code column and is in the correct format, one of these data records is incorrect (21087 is a zip code associated with Maryland).

[0041] In an additional example, the data set may include asset symbols, such as publicly traded stocks or cryptocurrencies. An asset may be traded on multiple exchanges under the same symbol. However, if a specific asset appears on an exchange where the asset has never been traded before, the system can confirm the anomaly and suggest one or more corrective actions. From the perspective of data quality, the asset symbol may be correct and in the correct format, but the potential quality of the data record (asset symbol plus exchange) may be incorrect because the exchange does not actually support trading of the asset.

[0042] Figures 4A - 4C An exemplary architecture for implementing systems and methods for predicting correct or missing data and data quality is shown. Environment 400 includes data sources, which may include, but are not limited to, data lake 402, database 404, flat file 406, and data stream 408. Once data is received by the system, data extraction service 410 can be applied to the raw input data. The extraction service 410 can reformat the data into a format suitable for processing. The data can be fed continuously / batch / large batch into profiler 412. The profiler 412 can be configured to implement rule engine 414, which may include analyzing statistical data of the data. Such statistical data may include the maximum value, minimum value, average value, and cardinality of the data values. Certain data columns can be classified, and confidence values can be initially assigned to certain data. The frequency of certain data types can also be profiled by the profiler 412.

[0043] The profiled data can be compared and analyzed at 416. In some examples, the profiled data can be further analyzed by applying certain rules to prepare a dataset for the application of ML algorithms. For example, certain missing values in the dataset can be filled with dummy values (e.g., n / a for categorical values, 0 for numerical values, or the replacement average). In another example, the profiled data can be decomposed to provide more relevant information to the ML algorithms. For example, if sales data varies according to a day of the week, that day of the week (e.g., Friday) may be separated from the actual date (April 15, 2021). Certain min-max normalization can also be applied to the dataset, such as applying confidence intervals or filtering based on the maximum and / or minimum extreme values. For example, in some instances, the top 2.5% and bottom 2.5% of numerical values may be discarded before applying a certain ML algorithm.

[0044] Multiple artificial intelligence / machine learning tools 418 can be applied to classifier module 420. The classifier module 420 can classify the profiled data into certain data categories, such as strings, integers, and other discrete objects.

[0045] The features of the dataset can be compared at 422, which can indicate which ML algorithm to apply to the dataset. Once the ML algorithm is applied to the dataset, at 426, the output can represent the predicted relationships and / or associations between the datasets.

[0046] Once the relationships of the data sets are confirmed, the problem confirmation process can begin. First, a data set can be selected at 428. The selected data set can be a subset of the initial data set input into the system. The selected data set can be the data set of the confirmed related data columns and / or possible anomalies. The selected data set can be transformed into a transaction item set 430, where the column headers with data values can be concatenated together. Once the item set is created, additional AI / ML tool sets can be applied to the item set, such as an ML algorithm like frequent pattern growth, which can confirm the frequency of the relationships in the data set (i.e., how frequently certain data columns appear in relationships with other data columns). Based on the output of the FPG model, a prediction pattern can be established at 434, where data anomalies can be confirmed. In addition, recommended repair actions for cleaning the data can be presented to the user of the system. For example, both data anomalies and / or potential errors can be presented to the user. The most likely results and / or predictions for a specific data set can also be presented to the user. Specifically, the user can accept a specific predicted value to clean the data set (e.g., the system can predict the correct value for a specific data entry that may be in error, and the user can accept the proposed change).

[0047] Figure 5 An example of a suitable operating environment in which one or more embodiments of the present embodiment can be implemented is shown. This is only one example of a suitable operating environment and is not intended to impose any limitation on the scope of use or functionality. Other well-known computing systems, environments, and / or configurations that may be suitable for use include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, programmable consumer electronics (e.g., smartphones), network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.

[0048] In its most basic configuration, the operating environment 500 generally includes at least one processing unit 502 and memory 504. Depending on the exact configuration and type of the computing device, the memory 504 (which stores information related to the detected device, association information, personal gateway settings, and instructions for performing the methods disclosed herein, etc.) can be volatile (e.g., RAM), non-volatile (e.g., ROM, flash memory, etc.), or some combination of both. This most basic configuration is in Figure 5as shown by the dashed line 506. Additionally, the environment 500 may also include storage devices (removable 508 and / or non-removable 510), including but not limited to magnetic disks, optical disks, or magnetic tapes. Similarly, the environment 500 may also have one or more input devices 514, such as keyboards, mice, pens, voice inputs, etc., and / or one or more output devices 516, such as monitors, speakers, printers, etc. The environment may also include one or more communication connections 512, such as LAN, WAN, point-to-point, etc.

[0049] The operating environment 500 generally includes at least some form of computer-readable medium. Computer-readable media can be any available media that can be accessed by the processing unit 502 or other devices including the operating environment. By way of example, and not limitation, computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other tangible medium that can be used to store the desired information. Computer storage media does not include communication media.

[0050] Communication media embodies non-transitory computer-readable instructions, data structures, program modules, or other data. Computer-readable instructions can be transmitted in a modulated data signal such as a carrier wave or other transmission mechanism and includes any information delivery medium. The term "modulated data signal" refers to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.

[0051] The operating environment 500 can be a single computer operating in a networked environment using a logical connection to one or more remote computers. The remote computers can be personal computers, servers, routers, network PCs, peer devices, or other common network nodes, and typically include many or all of the above elements as well as other elements not so mentioned. The logical connection can include any method supported by the available communication media. Such networked environments are often found in offices, enterprise-wide computer networks, intranets, and the Internet.

[0052] For example, aspects of the present disclosure have been described above with reference to block diagrams and / or operational illustrations of methods, systems, and computer program products according to various aspects of the present disclosure. The functions / actions noted in the blocks may occur out of the order shown in any flowchart. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality / action involved.

[0053] The description and illustration of one or more aspects provided in this application are not intended to limit or restrict the scope of the claimed disclosure in any way. The aspects, examples, and details provided in this application are considered sufficient to convey possession thereof and enable others to make and use the best mode of the claimed disclosure. The claimed disclosure should not be construed as limited to any aspect, example, or detail provided in this application. Various features (structures and methods), whether shown and described in combination or separately, are intended to be selectively included or omitted to yield embodiments having a particular set of features. After the description and illustration of this application have been provided, those skilled in the art may envision variations, modifications, and alternative aspects that fall within the spirit of the broader aspects of the general inventive concept embodied in this application, without departing from the broader scope of the claimed disclosure.

[0054] In summary, it should be understood that, for purposes of illustration, specific embodiments of the present invention have been described herein, but various modifications may be made without departing from the scope of the present invention. Accordingly, the present invention is not limited except as by the appended claims.

Claims

1. A method for predicting data quality in a dataset, comprising: Receiving a first dataset; Parsing the first dataset; Applying a first machine learning algorithm to the first dataset, wherein the first machine learning algorithm identifies at least one relationship between a first data column and a second data column in the first dataset; Generating a second dataset based on the identification of the at least one relationship between the first data column and the second data column in the first dataset, wherein the second dataset is a subset of the first dataset; Concatenating multiple column headers in the second dataset; Generating an item set based on the concatenation of the multiple column headers in the second dataset; Applying a second machine learning algorithm to the item set, wherein the second machine learning algorithm identifies at least one frequency value related to the at least one relationship between the first data column and the second data column in the first dataset; Identifying at least one anomaly in the item set; and Generating at least one suggestion for fixing the at least one anomaly.

2. The method according to claim 1, wherein Identifying at least one anomaly in the item set includes: identifying at least one missing value in the item set.

3. The method according to claim 2, wherein The at least one suggestion for fixing the at least one anomaly includes: at least one suggested value for filling the at least one missing value in the item set.

4. The method according to claim 3, wherein The at least one suggested value is at least one of the following: zero, average value, minimum value, and maximum value.

5. The method according to claim 1, wherein, The first machine learning algorithm is a random forest algorithm.

6. The method according to claim 1, wherein, The second machine learning algorithm is a frequent pattern growth algorithm.

7. The method according to claim 1, wherein The at least one anomaly in the item set is a relationship anomaly in at least one data record.

8. The method according to claim 7, wherein, The at least one suggestion is a suggestion for correcting the relationship anomaly in the at least one data record.

9. The method according to claim 8, wherein, The suggestion for correcting the relationship anomaly includes at least one alternative value.

10. The method according to claim 1, further comprising: Receiving at least one user input that accepts the at least one suggestion; And Applying the at least one suggestion to the item set based on the at least one user input.

11. The method according to claim 1, wherein, The first dataset and the second dataset are at least one of the following: a data file, a data lake, a flat file, and a data stream.

12. A system for correcting data anomalies in a dataset, comprising: A memory configured to store non-transitory computer-readable instructions; And A processor communicatively coupled to the memory, wherein the processor, when executing the non-transitory computer-readable instructions, is configured to: Receive a dataset; Parse the dataset; Apply a first machine learning algorithm to the dataset, wherein the first machine learning algorithm identifies at least one relationship between a first data column and a second data column in the dataset; Generating an item set based on the identification of the at least one relationship between the first data column and the second data column in the dataset; Apply a second machine learning algorithm to the item set, where the second machine learning algorithm identifies at least one frequency value related to the at least one relationship between the first data column and the second data column in the first data set; Identify at least one anomaly in at least one data record in the item set; and Automatically repair the at least one anomaly in the at least one data record in the item set.

13. The system according to claim 12, wherein, The at least one anomaly in the at least one data record is a relationship anomaly between at least one data value in the first data column and at least one data value in the second data column.

14. The system according to claim 13, wherein, The relationship anomaly is related to a geographical location.

15. The system according to claim 13, wherein, The relationship anomaly is related to an asset symbol, where the asset symbol is at least one of the following: a stock code and a cryptocurrency code.

16. The system according to claim 12, wherein, The first machine learning algorithm is a random forest algorithm.

17. The system according to claim 12, wherein The second machine learning algorithm is a frequent pattern growth algorithm.

18. The system according to claim 12, wherein, The at least one anomaly is at least one of the following: a missing value, an incorrect value, and a formatting error.

19. The system according to claim 12, further configured to generate at least one output table that displays the at least one anomaly in the at least one data record in the item set and a corrected version of the at least one data record without the at least one anomaly.

20. A computer-readable medium storing non-transitory computer-executable instructions that, when executed, cause a computing system to perform a method for predicting and correcting data anomalies, the method comprising: Receive a first data set; Parse the first data set; Compare a first data column in the first data set with multiple other data columns in the first data set; Based on the comparison result of the first data column and the multiple other data columns, identify a first relationship between the first data column and a second data column in the first data set; Apply a first machine learning algorithm to the first data column and the second data column, where the first machine learning algorithm identifies a second relationship between the first data column and the second data column in the first data set; Based on the identification of the second relationship between the first data column and the second data column in the first data set, generate a second data set, where the second data set is a subset of the first data set; Concatenate multiple column headers in the second data set; Generate an item set based on the concatenation of the multiple column headers in the second data set; Apply a second machine learning algorithm to the item set, where the second machine learning algorithm identifies at least one frequency value related to the second relationship between the first data column and the second data column in the first data set; Based on the at least one frequency value, identify at least one outlier value in the item set, where The at least one outlier value is related to a low frequency value; and Repair the at least one outlier value by substituting the at least one outlier value with a substitute value related to a high frequency value.

Citation Information

Patent Citations

  • Machine learning for isolated data sets

    CN111566640A

  • Explanations for artificial intelligence based recommendations

    CN112189204A

  • Data quality detection and compensation for machine learning

    US20170372232A1

  • Suggestion of views based on correlation of data

    US20180357292A1