Method, system and medium for continuous data profiling
By implementing a continuous data analysis manager in the client database, directly using the computing resources in the database for data analysis, the problems of low efficiency and poor security in the existing technology are solved, and more efficient resource utilization and data security are achieved.
Patent Information
- Application Number
- CN202411908931.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2021-06-30
- Filing Date
- 2022-07-07
- Publication Date
- 2025-05-13
AI Technical Summary
When performing data analysis, the prior art needs to export data from the local database to a third-party platform for analysis, resulting in inefficiency, unpredictable security and excessive use of computing resources.
By implementing the continuous data analysis manager (CDP manager) in the client database, it directly communicates with the database management system, uses the computing resources in the database to perform data analysis, and avoids the data export and import process.
It improves the efficiency of computer resources, reduces the transmission of data between different platforms, and enhances the security of data, because the data is processed only in a single secure location throughout the process.
Smart Images

Figure CN119988356A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of July 7, 2022, application number 202280056985.2, and invention name “System and method for continuous data analysis”.
[0002] Cross-reference to related applications
[0003] This application is related to U.S. patent application No. 16 / 844,927, entitled “CONTEXTDRIVEN DATA PROFILING”; and U.S. patent application No. 17 / 236,823, entitled “SYSTEMS AND METHODS FOR PREDICTING CORRECT OR MISSING DATA AND DATA ANOMALIES”, which are incorporated herein by reference in their entirety. Technical Field
[0004] The present disclosure relates to continuous data profiling, and in particular to performing continuous data profiling to derive insights into data while conserving computing power. Background Art
[0005] An entity may maintain large amounts of data digitally across multiple computing devices. For example, an organization may maintain multiple columns of data across a series of interconnected servers. It may often be necessary to inspect and evaluate these large amounts of data to determine multiple insights into various characteristics of the data. However, acquiring and processing large amounts of data may require significant computing resources. Furthermore, given the large amount of information contained in the data, it may often be difficult to derive high-quality data.
[0006] As previously described in patent application Ser. No. 16 / 844,927, which is incorporated herein by reference in its entirety, a solution to this problem of gaining insight into large amounts of data is data profiling, which is a process that can include validating attributes in client data, normalizing those attributes in a standardized format, and then processing the normalized attributes to derive insights from the data.
[0007] However, as data continues to grow, it becomes cumbersome to profile in an efficient manner. Currently, entities that want to profile their data sets typically use specialized third-party tools that require the client data to be exported from their local platform to a separate third-party platform for profiling. There are many problems with this process, including inefficiencies in the export and import of large amounts of data, unpredictable security measures of third-party platforms, and excessive use of computer resources. In practice, an entity first exports its data from its local database (usually by creating a copy), then imports the copy of the data into a third-party profiling runtime environment, then exports the profiled data from the third-party runtime environment, and finally imports the profiled copy of the data back into the local database environment from which the initial data set originated. In addition, because copies of data sets are often used in data profiling, clients typically need to reconcile the profiled data set imported back into the database with the unprofiled data left in the database. This is another additional step that requires time and intensive computing power.
[0008] Therefore, there is an increasing need for systems and methods that can address the challenges of external and one-time data profiling, including profiling data in a computationally efficient manner using fewer resources and requiring fewer import-export operations, which will further improve the security of the data because the data is less mobile.
[0009] It is with respect to these and other general considerations that the various aspects disclosed herein have been made.In addition, although relatively specific problems may be discussed, it should be understood that the examples should not be limited to solving specific problems identified in the background of this disclosure or elsewhere. Summary of the invention
[0010] It is with respect to these and other general considerations that the various aspects disclosed herein have been made.
[0011] According to some embodiments of the present invention, a method for continuously profiling data is disclosed, the method comprising: receiving at least one input data stream; profiling the at least one input data stream in the following manner, confirming at least one attribute in the at least one input data stream, wherein the at least one attribute is related to a series of features, determining a profiling score of the at least one attribute based on the source of the at least one input data stream and an aggregation of the series of features, determining that the at least one attribute represents an address, and in response to determining that the at least one attribute represents an address, processing the at least one attribute through an address library engine, the address library engine adding the at least one attribute to an address library; generating a profiled set of data based on the profiling of the at least one input data stream, wherein the profiled set of data includes the profile score of the at least one attribute; and storing the profiled set of data in at least one client database.
[0012] According to some embodiments of the present invention, a system is disclosed, comprising: one or more processors; and one or more memories storing instructions, wherein when the instructions are executed by the one or more processors, the system executes a process for continuously profiling data, the process comprising: receiving at least one input data stream; profiling the at least one input data stream in the following manner, confirming at least one attribute in the at least one input data stream, wherein the at least one attribute is related to a series of features, determining a profiling score of the at least one attribute based on the source of the at least one input data stream and an aggregation of the series of features, determining an address representing the at least one attribute, and in response to determining the address representing the at least one attribute, processing the at least one attribute through an address library engine, wherein the address library engine adds the at least one attribute to an address library; generating a profiled set of data based on the profiling of the at least one input data stream, wherein the profiled set of data includes the profile score of the at least one attribute; and storing the profiled set of data in at least one client database.
[0013] According to some embodiments of the present invention, a non-transitory computer-readable medium storing instructions is disclosed, which, when executed by a computing system, causes the computing system to perform operations for continuously profiling data, the operations comprising: receiving at least one input data stream; profiling the at least one input data stream in the following manner, confirming at least one attribute in the at least one input data stream, wherein the at least one attribute is related to a series of features, determining a profiling score of the at least one attribute based on a source of the at least one input data stream and an aggregation of the series of features, determining an address representing the at least one attribute, and in response to determining the at least one attribute representing the address, processing the at least one attribute through an address library engine, the address library engine adding the at least one attribute to an address library; generating a profiled set of data based on the profiling of the at least one input data stream, wherein the profiled set of data includes a profiling score of the at least one attribute; and storing the profiled set of data in at least one client database.
[0014] Furthermore, although relatively specific problems may be discussed, it should be understood that the examples should not be limited to solving specific problems identified in the background of this disclosure or elsewhere. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Non-limiting and non-exhaustive examples are described with reference to the following figures.
[0016] Figure 1 An example of a distributed system for continuous data profiling as described herein is shown.
[0017] Figure 2 An exemplary input processor for continuous data profiling as described herein is shown.
[0018] Figure 3 An exemplary architecture for continuous data profiling is shown.
[0019] Figure 4 An exemplary method for continuous data profiling as described herein is shown.
[0020] Figure 5 An exemplary architecture of a continuous data profiling manager and a database management system is shown.
[0021] Figure 6 An exemplary environment for continuous data profiling is shown.
[0022] Figure 7 An example of a suitable operating environment is shown in which one or more of the present embodiments may be implemented. DETAILED DESCRIPTION
[0023] Aspects of the present disclosure are described more fully below with reference to the accompanying drawings, which form a part of this document and illustrate specific exemplary aspects. However, different aspects of the present disclosure can be implemented in many different forms and should not be construed as being limited to the aspects described herein; rather, these aspects are provided so that the present disclosure will be thorough and complete and the scope of these aspects will be fully conveyed to those skilled in the art. Various aspects can be practiced as methods, systems, or devices. Therefore, various aspects can take the form of hardware implementations, fully software implementations, or implementations combining software and hardware aspects. Therefore, the following detailed description should not be taken as limiting.
[0024] Embodiments of the present application relate to systems and methods for continuous data profiling. Many entities (e.g., companies, organizations) maintain large amounts of data. The data may be stored in a variety of registers or databases on a computing device. In many cases, these entities may need to confirm and match records in different data sets and gain insights into the data sets. For example, given multiple similar data sets, an organization may attempt to confirm and select a high-quality and accurate data set in the similar data sets.
[0025] This embodiment relates to continuously processing and analyzing data and generating insights into the ingested data. The continuous data analysis process may include verifying attributes of client data, normalizing the attributes into a standardized format, and processing the data via one or more rule engines. Other information may be generated based on the input information obtained, such as usage rankings or value scores.
[0026] The data profiling process can allow insights into the data to be generated, thereby improving data quality. Examples of insights can include duplication or multiple instances of data attributes within and across domains, including percentage overlap. As further examples, insights can include data quality reports from normalization and standardization (what percentage are standard versus non-standard) or trends based on label processing (e.g., records with the same home address).
[0027] As previously described, current systems and methods for data profiling typically require entities to export their data sets from their local runtime environment into a specialized third-party profiling runtime environment. This process is both unsafe and inefficient from a computing resource perspective. To remedy these issues, the present system and method disclose an efficient continuous data profiling process in which an entity's data set can be locally profiled within the database where it is stored. This is facilitated via a continuous data profiling (CDP) manager, which is a lightweight front-end application that communicates directly with a database management system (e.g., a software application that is locally coupled to a database that stores an entity's data set). The CDP manager can take the form of an application programming interface (API), wherein the CDP installs certain profiling logic directly into the database management system, allowing the database management system to handle all profiling (e.g., tracking, scheduling, calculating, and storing profiled data). As a further example, the CDP manager can allow the database management system to generate and store statistical tables, change data capture (CDC) tables, profiling programs, and profiling triggers.
[0028] Thus, the present disclosure provides multiple technical benefits, including, but not limited to, enabling more efficient use of computer resources because entities no longer need to export and import their data from their local database systems into third-party profiling systems. Instead, the systems and methods disclosed herein enable entities to simply call CDP APIs that communicate directly with the entity's local database management system, utilizing the entity's database computing resources for the profiling process. Another technical benefit is the increased security of the entity's data. By avoiding the continuous export-import process into an unknown and unpredictable third-party runtime environment, the risk of security breaches or exposure of personally identifiable information is significantly reduced because the entity's data is not transferred outside of its local runtime environment before, during, and after profiling. The data remains in a single secure location. In short, the continuous data profiling process provides a more efficient use of computer resources and processing power, and also provides improved security and protection of sensitive data.
[0029] Figure 1An example of a distributed system for continuous data profiling as described herein is shown. The proposed exemplary system 100 is a combination of interdependent components that interact to form an integrated whole for consolidating and augmenting data on a data marketplace. The components of the system can be hardware components or software implemented on and / or executed by hardware components of the system. For example, the system 100 includes client devices 102, 104, and 106, local databases 110, 112, and 114, one or more networks 108, and server devices 116, 118, and / or 120.
[0030] The client devices 102, 104, and 106 may be configured to receive and transmit data. For example, the client devices 102, 104, and 106 may include client-specific data having client-specific data terms and tags. The client devices may download a CDP manager program via one or more networks 108 that may be communicatively coupled to one or more databases 110, 112, and / or 114 where the client data resides. In other embodiments, instead of directly downloading the CDP manager, one or more client devices 102, 104, and / or 106 may simply call a CDP manager API via one or more networks 108, wherein activation of the API allows the CDP manager (which may be remotely operated on one or more servers 116, 118, and / or 120) to communicate directly with and parse data stored on the one or more databases 110, 112, and / or 114. Because the profiling of the data occurs at the local location of the client's data set, the client data stored on one or more databases 110, 112, and / or 114 is not transmitted via one or more networks 108 for remote profiling, such as on one or more third-party servers 116, 118, and / or 120. Client-specific data can be stored in local databases 110, 112, and 114. The original un-profiled data is stored on the local databases 110, 112, and 114, and the profiled data (after the CDP process is run on the data) is also stored on the one or more local databases 110, 112, and / or 114. The one or more servers 116, 118, and / or 120 can be third-party servers owned by an administrator of the CDP manager and / or CDP API. In other examples, once the data is profiled, the profiled client-specific data can be stored in a remote server (in addition to or in lieu of the local client device and local database) and can be transmitted from the client server to the third-party server via one or more networks 108 and / or satellites 122.
[0031] In other examples, one or more servers 116, 118, and / or 120 may be owned by the client. The one or more servers 116, 118, and / or 120 may be client-owned cloud servers where client data resides. In this example, client data may be transferred from a client-owned local database 110, 112, and / or 114 to a client-owned database 116, 118, and / or 120. The CDP manager may be communicatively coupled to a local or remote database owned by the client. The communication channel between the CDP manager and the client-owned database may be facilitated via one or more networks 108 and / or satellites 122. This example is applicable to scenarios where the remote database / server is owned by the client rather than a third party managing the CDP manager and / or API.
[0032] In various aspects, client devices (e.g., client devices 102, 104, and 106) may have access to one or more data sets or data sources and / or databases that include client-specific data. In other aspects, client devices 102, 104, and 106 may be equipped to receive broadband and / or satellite signals carrying CDP management software and / or CDP API files that are necessary to be installed on client-owned databases for profiling to be performed. The signals and information that client devices 102, 104, and 106 may receive may be transmitted from satellite 122. Satellite 122 may also be configured to communicate with one or more networks 108, in addition to being able to communicate directly with client devices 102, 104, and 106. In some examples, client devices may be mobile phones, laptops, tablets, smart home devices, landlines, and wearable devices (e.g., smart watches), among other devices.
[0033] To further illustrate the network topology, once the CDP manager is communicatively coupled to the local databases 110, 112, and / or 114, the client devices 102, 104, and / or 106 (and their corresponding local databases 110, 112, and 114) can receive CDP management files and information. Note that this also applies to scenarios where one or more remote databases 116, 118, and / or 120 are client-owned. CDP management files can include, but are not limited to, statistics tables, CDC tables, profiling procedures, and profiling triggers. Once profiling of a data set is completed, the profiled data can be stored in the original database where the original, unprofiled data was stored.
[0034] Figure 2An exemplary input processor for continuous data profiling as described herein is shown. Input processor 200 can be embedded within a client device (e.g., client devices 102, 104, and / or 106), a remote network server device (e.g., devices 116, 118, and / or 120), and other devices capable of implementing systems and methods for continuous data profiling. The input processing system includes one or more data processors and is capable of executing algorithms, software routines, and / or instructions based on processed data provided by at least one client data source. The input processing system can be a factory-installed system or an add-on unit to a particular device. In addition, the input processing system can be a general purpose computer or a specialized, special-purpose computer. There is no limitation on the location of the input processing system relative to a client or remote network server device, etc. According to Figure 2 In the illustrated embodiment, the disclosed system may include a memory 205, one or more processors 210, a communication module 215, a continuous data profiling (CDP) module 220, and a database management system (DMS) module 225. Other embodiments of the present technology may include some, all, or none of these modules and components, as well as other modules, applications, data, and / or components. However, some embodiments may combine two or more of these modules and components into a single module and / or associate a portion of the functionality of one or more of these modules with different modules.
[0035] The memory 205 may store instructions for running one or more applications or modules on one or more processors 210. For example, the memory 205 may be used in one or more embodiments to hold all or part of the instructions required to perform the functions of the CDP module 220 and / or the DMS module 225 and the communication module 215. In general, the memory 205 may include any device, mechanism, or populated data structure for storing information. According to some embodiments of the present disclosure, the memory 205 may cover but is not limited to any type of volatile memory, non-volatile memory, and dynamic memory. For example, the memory 205 may be a random access memory, a memory storage device, an optical memory device, a magnetic medium, a floppy disk, a tape, a hard disk drive, a SIMM, a SDRAM, a RDRAM, a DDR, a RAM, a SODIMM, an EPROM, an EEPROM, an optical disk, a DVD, and / or the like. According to some embodiments, the memory 205 may include one or more disk drives, a flash drive, one or more databases, one or more tables, one or more files, a local cache memory, a processor cache memory, a relational database, a flat database, and / or the like. In addition, a person of ordinary skill in the art will understand that many additional devices and techniques for storing information may be used as the memory 205.
[0036] In some exemplary aspects, the memory 205 may store certain files from the CDP module 220 that may originate from the CDP manager, such as a software application that enables one or more client databases to generate, display, and store statistics, CDC tables, profiling procedures, and profiling triggers. The CDP manager may also enable a user to configure any of the CDP files, which may allow for customization of statistics and CDC tables as well as profiling procedures and profiling triggers. In a further example, the memory 205 may store certain profiling statistics and profiled data that may be used to facilitate profiling of data on the client databases and data flows between the CDP manager and the DMS.
[0037] The communication module 215 is associated with sending / receiving information (e.g., CDP applications from the CDP module 220 and data (unparsed and parsed) from the DMS module 225), commands received via a client device or server device, other client devices, remote network servers, etc. These communications can use any suitable type of technology, such as Bluetooth, WiFi, WiMax, cellular (e.g., 5G), single-hop communication, multi-hop communication, dedicated short-range communication (DSRC), or proprietary communication protocols. In some embodiments, the communication module 215 sends information output by the CDP module 220 (e.g., software applications and / or logic to be installed on the DMS) and / or information output by the DMS module 225 (e.g., parsed data, such as tracking, scheduling, calculating, and storing parsed data statistics for each data table), and / or sends to the client devices 102, 104, and / or 106, and the memory 205 to be stored for future use. In some examples, the communication module can be built on the HTTP protocol through one or more secure REST servers using RESTful services. In yet further examples, the CDP module 220 can communicate with the DMS module 225 via a CDP API. In other examples, an external application can request profiled data statistics, and the communication module 215 can facilitate the transfer of profiled data from the DMS module 225 to a third-party external service.
[0038] The CDP module 220 is configured to install certain logic and software functions on the database, specifically configuring the database management system that manages the client database. The logic and / or software that can be provided by the CDP module 220 may include functions for facilitating the construction and storage of statistical tables, CDC tables, analysis programs, and analysis triggers. For example, the CDP module 220 can enable change data capture methods to be run on the client database via the DMS. These methods may include: initialization timestamps or version numbers, table triggers (e.g., so that when data changes, the administrator of the database or data table receives a push notification), snapshots or table comparisons, and log crawling. Each of these methods allows real-time reporting capabilities of database status.
[0039] The CDP module 220 may also be configured with an API that allows a DMS (e.g., DMS module 225) to communicate with the CDP module 220 and receive downloads and functionality designed and supported by the CDP manager. Once the CDP module 220 is communicatively coupled to the local database in which profiling should be performed, profiling may be performed continuously based on different factors. For example, a profiling trigger may be established via the CDP module 220 that triggers profiling of new data that has been added to a data set every 24 hours. In another example, the profiling trigger may be based on the amount of new data added to a certain data set or data table. Once the amount of new data reaches or exceeds, for example, 10 gigabytes, the profiling process is triggered and the new data is automatically profiled.
[0040] The DMS module 225 is configured to manage at least one local database storing client-specific data. The DMS module 225 is configured to operate change tracking, scheduling, calculation, and storage of profiling statistics for each data table. Most of the computing resources are managed by the DMS module 225 because the CDP system and method described herein uses local database resources to profile and store data. The DMS module 225 is also configured to generate and store certain timeline statistical tables that allow the DMS module 225 to capture the entire history of the profiled data. The statistical tables can be displayed via the CDP module 220 based on queries received by the CDP module 220.
[0041] Figure 3 An exemplary architecture for continuous data profiling is shown. A context-driven data profiling process can help determine the data quality of source data. Data profiling can include several processing steps that modify input information to generate insights into the data that help optimize applications such as matching accuracy. For example, data profiling can standardize and validate the data before tokenizing the profiled data.
[0042] Figure 3is an exemplary architecture for continuous data profiling, showing an exemplary profiling process 300. The continuous data profiler may include flexible data processes. Data may be accessed and / or processed from a data source in multiple batches, continuous streams, or bulk loads. As previously described, the present application relates to continuous data profiling flows. One or more data sources 302 may include nodes (e.g., database devices 304a-d) configured to store / maintain data (e.g., data lake 306a, database 306b, flat file 306c, data stream 306d). For example, a data source 302 may include a single column of data, a series of relational databases with multiple data tables, or a data lake with a large number of data assets.
[0043] Data quality can be resolved in the data profiler by use case or client. For example, the context can be based on a column of data, a combination of multiple columns of data, or a data source. During the data profiling process, multiple data can be derived and a data summary can be generated. For example, a summary of a column of data can be confirmed in the form of a data sketch. The data sketch may include numerical data and / or string data. Examples of numerical data included in the data sketch may include any of a number of missing values, the mean / variance / maximum / minimum value of the numerical data, an approximate quantile estimate of the numerical data that can be used to generate a distribution or histogram, and the like. Examples of string data may include a number of missing values, a maximum character length, a minimum character length, an average character length, a label frequency table, a set of frequency items, a different value estimate, and the like.
[0044] Once any of a series of metrics is calculated in the summary of the data, a data profiling score can be calculated. The data profiling score can be used to determine data quality and / or confirm optimal data, data composition, and targeted data quality enhancement activities. At user-set intervals, data profiling can be re-executed to recalculate metrics. These user-set intervals can be time-based (e.g., profile new data received by the data lake 306a every 24 hours) or size-based (e.g., profile every 1 GB of data added to the flat file 306c). In addition to efficiently using computer resources to continuously profile data streams rather than manually batch processing, this can also be used to track data score history throughout the data life cycle and enable flagging of data quality issues.
[0045] In some embodiments, the data summary may include a proportion of values (eg, reference data) that follow a particular regular expression. For example, for a phone number that follows a particular format, the data summary may indicate that there are multiple formats.
[0046] In some embodiments, the data summary may include multiple anonymous values. For example, known anonymous names (eg, John Doe) may be identified in the source data to determine the proportion of the data that includes anonymous values.
[0047] In other embodiments, the data summary may include a set of data quality indicators based on a data quality rule base. The data summary may be used to learn data quality rules based on reference data related to the attribute. The data summary may also be used to learn data quality rules directly from the source data (e.g., between which values the source data should be included, what the minimum character length should be).
[0048] As a first example, the source data may be reviewed to derive a data quality score. The data quality score may include a score calculated at a column level or a record level of the source data. The data quality score may be derived by calculating any metric included in the data summary.
[0049] As another example, the source data can be reviewed to confirm the quality data. For the data profiling score of each column of data in each data source, the most likely set of data can be matched to a specific client. For example, a table can be prepared that shows a set of columns / attributes (e.g., name, address, phone, date of birth, email address), and the data profiling scores of different sources (CRM, ERP, order management, network) where the columns / attributes exist. Using the data included in such a table, a set of data with the highest data quality can be selected for a specific client. In some instances, multiple sources can be matched to receive data of the highest possible quality. This can be performed without over-processing the source data.
[0050] As another example, the source data can be reviewed to derive historical data profiling scores and perform what-if analysis. What-if analysis can include analysis of what would happen if other (certain) rules were invoked on the data. To facilitate the calculation of these, this process can be done on sample data collected from the data summary created during the calculate metric phase. If the results of the what-if analysis are sufficient, a new full calculation of the metric can be performed using the new rules selected in the what-if analysis.
[0051] Data extracted from a data source (e.g., data lake 306a, database 306b, flat file 306c, data stream 306d) can be fed into a parser (e.g., parser 310a-n) via a data feed 308. Data feed 308 can include a continuous data feed to the parser. Parser 310a-n can be installed on a local database via a CDP manager, which can be communicatively coupled to one or more databases 304a-d via CDP module 220, such as Figure 2 The data fed into the parser may include attributes (eg, attributes 312a-n). Attributes may be part of the data in a table, source, or part of the same record.
[0052] In such Figure 3 In the illustrated embodiment, a first parser 310a may process attribute 1 312a, and a second parser 310b may process attribute 2 312b. Any suitable number of parsers (e.g., parser N 310n) may process any number of attributes (e.g., attribute N 312n). Each parser 310a-n may include a set of standardized rules 314a-n and a set of rule engines 316a-n. The standardized rules 314a-n and the rule engines 316a-n may be installed on one or more databases 304a-n via a CDP manager communicatively coupled to the database, providing continuous parsing of data stored on the repository and provided to the parsers 310a-n via a data feed 308. The standardized rules 314a-n and / or the rule engines 316a-n may be modular, wherein each set of rules may be processed for an attribute. Each parser may process a corresponding set of standardized rules and a set of rule engines to process a corresponding attribute. In some embodiments, each profiler may implement a number of machine learning and / or artificial intelligence techniques and statistical tools to improve data quality when processing attributes. Resulting data from each profiler 310a-n may include insights 318 representing a number of characteristics of the attribute.
[0053] In some embodiments, data quality rules may be adjusted, which may result in different determinations being made when performing data quality improvement tasks. For example, a data set may have a good score, but it was not previously known that the name "John Doe" was an anonymous (forged or false) value. By updating the rules to confirm that "John Doe" is an anonymous value, the change in the data profiling score and the history of the score may be modified. This change in the data profiling score may enable confirmation of multiple data included in the data set.
[0054] As another example, source data may be reviewed to derive automatic data quality improvement requests. A trigger may be associated with a data profiling score for a particular attribute or set of attributes. A trigger may specify that if the data profiling score is below a threshold, source data associated with the attribute may be reviewed. If the source data has a validation value that indicates how the data is used in multiple contexts, the source data may potentially be improved.
[0055] As another example, source data can be reviewed to derive data insights. Processing the data profiling scores of the source data can generate data distribution and other insights that can be used to understand the characteristics of the data before initiating another data analysis.
[0056] As another example, source data can be reviewed to make intelligent data quality-based data selection decisions. Based on mapping the source data to a model (e.g., a canonical model), when the data quality score is better than another data set with similar attributes, highly relevant profiling / sampling outputs, relevant definitions, and / or similar endpoint consumption relationship patterns can provide suggestions for alternatives worth reviewing. Side-by-side comparisons can be run based on user-initiated requests to help users confirm overlapping measurements and express relative preferences. This can be stored / recorded with users and communities to provide recommendations calibrated with user-specific needs over the long term. For example, statistical tables can be stored and generated via a database management system that manages the data source 302. The statistical tables can be provided to the CDP manager for display after the CDP manager receives a query to display the statistical tables.
[0057] Figure 4 An exemplary method for continuous data profiling as described herein is shown. Method 400 begins by receiving a first input data stream 402; the data stream may come from any number of client-owned data sources, such as Figure 3 The data sources described in . The data stream corresponding to the client may include one or more columns of client data.
[0058] Once the first input data stream is received at step 402, the first input data stream may be parsed at step 404, where at least one attribute from the data stream may be identified. Further steps in the data parsing process may include obtaining a set of validation rules and a set of normalization rules corresponding to the attributes. The set of validation rules may provide rules indicating whether the attribute corresponds to the attribute. The set of normalization rules may provide rules for modifying the attribute into a standardized format.
[0059] The data profiling process step 404 may include comparing the attribute to the set of validation rules to determine whether the attribute corresponds to a rule. If it is determined that the attribute corresponds to a rule, the attribute may be modified, as described herein. In some embodiments, verifying the attribute may include determining whether the attribute includes an invalid value identified in the set of validation rules. The attribute may be verified in response to determining that the attribute does not include an invalid value.
[0060] The data profiling process may include modifying the attribute into a standardized format according to the set of standardized rules. This may be performed in response to determining that the attribute is validated via the validation rules.
[0061] The data profiling process step 404 may include processing the attribute through a plurality of rule engines. The rule engines may include a name engine that, in response to determining that the attribute represents a name, associates the attribute with a commonly related name included in a list of related names. The rule engines may also include an address library engine that, in response to determining that the attribute represents an address, adds the attribute to an address library associated with the client.
[0062] In some embodiments, processing the modified attribute by the set of rule engines at step 404 may include, in response to determining that the attribute represents a name, processing the modified attribute by a name engine associating the attribute with a related name included in a related name list. Processing the modified attribute by the set of rule engines may also include, in response to determining that the attribute represents an address, processing the modified attribute by an address library engine adding the attribute to an address library associated with the data object.
[0063] In some embodiments, method 400 may include comparing multiple instances of the attribute relative to other attributes in the data stream at data profiling step 404. A usage ranking may be generated for the attribute. The usage ranking may be based on the number of instances of the attribute in the data stream, and the usage ranking may represent multiple insights that can be derived from the attribute.
[0064] In some embodiments, a series of features related to the attribute and identified relative to other attributes in the data stream can be identified. Exemplary features in the series of features can include quality features, availability features, cardinality features, etc. A value score for the attribute can be derived based on the aggregation of the series of features.
[0065] In some embodiments, at step 404, deriving a value score for an attribute based on the aggregation of a series of features may include: processing the attribute to derive a quality feature of the attribute, the quality feature confirming a plurality of differences between the attribute confirmed in the data stream and the modified attribute modified according to a set of standardized rules. Deriving a value score for an attribute based on the aggregation of a series of features may also include: processing the attribute to derive an availability feature of the attribute, the availability feature representing a plurality of invalid entries in a portion of data in the data stream corresponding to the attribute. Deriving a value score for an attribute based on the aggregation of a series of features may also include: processing the attribute to derive a cardinality feature of the attribute, the cardinality feature representing a difference of the attribute relative to other attributes in the data stream. Deriving a value score for an attribute based on the aggregation of a series of features may also include: aggregating the derived quality feature, availability feature, and cardinality feature of the attribute to generate a value score for the attribute.
[0066] Once the first input data stream is profiled at step 404, a first set of profiled data may be generated at step 406. The profiled data may be structured into statistical tables and displayed via the CDP manager at step 406. The system described herein may also maintain profiled insights / rankings / scores over a series of processed and profiled attributes, which allows data quality insights to be derived from the original input data stream.
[0067] Once the first set of profiled data is generated at step 406, the system may receive a second input data stream at step 408. In some examples, the second input data stream may trigger the profiling process at step 410. The trigger may be based on a timing factor (e.g., profile a new input data stream every 24 hours) or a size factor (e.g., process a new input data stream once it reaches 1 GB in size). In other examples, the second input data stream at step 408 may be stored in the client database until the profiling process is triggered in step 410. Thus, new data received by the client data store between the generation of the first set of profiled data and the triggering of a subsequent profiling process may be defined as a "second input data stream."
[0068] Once the profiling process is triggered again at step 410 , the second input data stream is parsed at step 412 according to the parsing steps and processing described above with respect to the parsing step 404 .
[0069] Similarly, once the second input data stream is profiled at step 412 , a second set of profiled data is generated at step 414 , wherein new statistics and data quality insights may be derived from the input data.
[0070] This process can continue to repeat as long as the profiling process step is triggered when a new input data stream is received by the client data store, which is connected to the CDP manager. The CDP manager can monitor the data flow into one or more client data stores, and once the profiling trigger is triggered, the new data flow can be profiled in the client database.
[0071] Figure 5An exemplary architecture of a continuous data profiling manager and a database management system is shown. The exemplary architecture 500 includes a CDP manager 502, which is a lightweight user interface software application that provides communication between the underlying client database and the CDP tool. In some examples, the CDP manager 502 can manage the CDP API and provide access (or call access) to the CDP API. The CDP manager 502 can be communicatively coupled to a database management system 506. The CDP manager 502 can also install certain profiling tools from the CDP toolset on the database management system 506, such as the ability for the DMS 506 to generate and store statistics tables, CDC tables, profiling programs, and profiling triggers. The CDP manager 502 can also provide the DMS 506 with tools for configuring certain stored procedures and profiling triggers. For example, the CDP manager 502 can allow a user to configure which profiling triggers are set for automatic data profiling, such as time or size-based triggers, as previously described.
[0072] In some cases, where the data warehouse is a public cloud hosted or managed (such as Snowflake, BigQuery, Redshift, etc.), the manager plays a limited role. Schedules and triggers can be provided by cloud services that are local to the service provider but external to the database itself. In another example, the Amazon Web Services (AWS) event bridge handles the scheduling and triggering of analytics execution within Redshift (e.g., Redshift is AWS's database).
[0073] The architecture 500 also includes an external process 508 that may be involved if the DMS 506 is configured to use the external process 508. For example, once the data is profiled and stored in a client database, the DMS 506 may transmit the stored profiled data via the API 508 to an external process that may further analyze the profiled data. In other examples, the external process 508 may include a data marketplace where clients may wish to enhance and / or buy / sell certain data assets related to the profiled data sets stored on the client database.
[0074] Figure 6 An exemplary environment for connecting continuous data profiling to external applications via an API for analysis / insight is shown. Environment 600 includes client feeds 602 that include data from various data sources, such as Figure 3Each of the data sources has its own CDP environment where profiling statistics are continuously stored. CDP feeds are readable via an API gateway to quickly analyze and provide insights without excessive processing time delays. The API gateway can be provided by any third party with data profiling or data quality capabilities.
[0075] The API gateway 610 is a continuous data profiling (CDP) gateway managed by the CDP manager. The CDP manager can be a top-level lightweight software interface that is communicatively coupled to the client environment 604. The CDP manager can derive its functionality from the CDP environment where certain data profiling and data quality analysis tools reside. Certain CDP tool sets can be used on client data sets via the CDP API 610. The client CDP data feed and the API gateway work as a lock-and-key mechanism that the client can use to benefit from profiling insights into its data from third parties. Once the connection is established, the CDP API can install tools within the client environment 604 and / or provide access to certain CDP tools via the CDP API, which can be utilized (e.g., via a cloud server) to analyze data stored within the client environment 604. It is important to note that client data (e.g., CDP feeds) are not transmitted from outside the client environment 604 to, for example, the CDP environment 606.
[0076] Figure 7 An example of a suitable operating environment in which one or more embodiments of the present embodiment can be implemented is shown. This is only an example of a suitable operating environment and is not intended to limit the scope of use or functionality. Other well-known computing systems, environments, and / or configurations that may be suitable for use include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, programmable consumer electronics (e.g., smartphones), network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.
[0077] In its most basic configuration, the operating environment 700 typically includes at least one processing unit 702 and memory 704. Depending on the exact configuration and type of computing device, the memory 704 (which stores information related to detected devices, related information, personal gateway settings, and instructions for executing the methods disclosed herein, etc.) can be volatile (e.g., RAM), non-volatile (e.g., ROM, flash memory, etc.), or some combination of the two. Figure 7In the embodiment shown by dashed line 706. In addition, environment 700 may also include storage devices (removable 708 and / or non-removable 710), including but not limited to disks or optical disks or tapes. Similarly, environment 700 may also have one or more input devices 714, such as keyboard, mouse, pen, voice input, etc., and / or one or more output devices 716, such as display, speaker, printer, etc. The environment may also include one or more communication connections 712, such as LAN, WAN, point-to-point, etc.
[0078] The operating environment 700 typically includes at least some form of computer-readable media. Computer-readable media can be any available media that can be accessed by the processing unit 702 or other devices that comprise the operating environment. By way of example, and not limitation, computer-readable media can include computer storage media and communication media. Computer storage media include volatile and nonvolatile removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical storage, cassettes, magnetic tape, disk storage, or other magnetic storage devices, or any other tangible media that can be used to store the desired information. Computer storage media do not include communication media.
[0079] Communication media embodies non-transitory computer-readable instructions, data structures, program modules, or other data. Computer-readable instructions may be transmitted in a modulated data signal such as a carrier wave or other transport mechanism, and include any information delivery media. The term "modulated data signal" refers to a signal that has one or more of its characteristics set or changed in a manner that encodes the information in the signal. By way of example, and not limitation, communication media include wired media, such as a wired network or a direct wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.
[0080] Operating environment 700 can be a single computer that operates in a networked environment using a logical connection to one or more remote computers. The remote computer can be a personal computer, a server, a router, a network PC, a peer device, or other public network node, and typically includes many or all of the above elements and other elements that are not mentioned in this way. The logical connection can include any method supported by an available communication medium. Such a networked environment is often located in an office, an enterprise-wide computer network, an intranet, and the Internet.
[0081] For example, the above text describes various aspects of the present disclosure with reference to the block diagrams and / or operational diagrams of the methods, systems, and computer program products according to various aspects of the present disclosure. The functions / actions marked in the blocks may not occur in the order shown in any flow chart. For example, two blocks shown in succession may in fact be substantially executed concurrently, or the blocks may sometimes be executed in reverse order, depending on the functions / actions involved.
[0082] The description and illustration of one or more aspects provided in this application are not intended to limit or restrict the scope of the claimed disclosure in any way. The multiple aspects, examples, and details provided in this application are considered sufficient to convey the possession thereof and enable others to make and use the best mode of the claimed disclosure. The claimed disclosure should not be interpreted as being limited to any aspect, example, or detail provided in this application. Whether shown and described in combination or individually, a variety of features (both structures and methods) are intended to be selectively included or omitted to produce an embodiment with a specific feature set. After providing the description and illustration of this application, those skilled in the art can envision changes, modifications, and alternative aspects that fall within the spirit of the broader aspects of the general inventive concept embodied in this application without departing from the broader scope of the claimed disclosure.
[0083] In summary, it should be understood that specific embodiments of the present invention have been described herein for illustrative purposes, but that various modifications may be made without departing from the scope of the present invention. Therefore, the present invention is not to be limited except as set forth in the appended claims.
Claims
1. A method for continuously analyzing data, the method comprising: receiving at least one input data stream; parsing the at least one input data stream by: identifying at least one attribute in the at least one input data stream, wherein the at least one attribute is associated with a set of features, determining a profile score for the at least one attribute based on a source of the at least one input data stream and an aggregation of the set of features, determining that the at least one attribute represents an address, and In response to determining that the at least one attribute represents an address, processing the at least one attribute by an address repository engine, the address repository engine adding the at least one attribute to an address repository; generating a profiled set of data based on profiling the at least one input data stream, wherein the profiled set of data includes a profile score for the at least one attribute; as well as The parsed set of data is stored in at least one client database.
2. The method according to claim 1, further comprising: Obtaining at least one set of analysis rules and at least one set of processing rules corresponding to the at least one attribute; comparing the at least one attribute to at least one set of profiling rules and the at least one set of processing rules to verify information included in the at least one attribute; In response to determining that the information included in the at least one attribute is parsed according to the at least one set of processing rules, storing the information included in the at least one attribute in at least one parsed format according to the at least one set of processing rules, wherein the information is provided via at least one API gateway; as well as The at least one attribute is processed by at least one set of rule engines.
3. The method according to claim 1, further comprising: In response to determining that the at least one attribute represents a name, the at least one attribute is processed by a name engine that correlates the at least one attribute with related names included in a list of related names.
4. The method according to claim 1, further comprising: At least one set of continuous data profiling tools is received from a continuous data profiling manager, wherein the continuous data profiling manager is a front-end software application communicatively coupled to the at least one client database.
5. The method according to claim 1, further comprising: connecting at least one continuous data profiling (CDP) manager application to the at least one client database; and At least one instruction or at least one function is received via the at least one continuous data profiling (CDP) manager application.
6. The method according to claim 5, wherein the at least one function is a first function for generating a statistical table of the profiled data, wherein the at least one function is a second function for generating a change data storage table for the parsed data, wherein the at least one function is a third function for managing at least one profiling process, and Wherein the at least one function is a fourth function for managing at least one profile triggered by comparison with the at least one input data stream.
7. A system comprising: one or more processors; and one or more memories storing instructions that, when executed by the one or more processors, cause the system to perform a process for continuously profiling data, the process comprising: receiving at least one input data stream; parsing the at least one input data stream by: identifying at least one attribute in the at least one input data stream, wherein the at least one attribute is associated with a set of features, determining a profile score for the at least one attribute based on a source of the at least one input data stream and an aggregation of the set of features, determining that the at least one attribute represents an address, and In response to determining that the at least one attribute represents an address, processing the at least one attribute by an address repository engine, the address repository engine adding the at least one attribute to an address repository; generating a profiled set of data based on profiling the at least one input data stream, wherein the profiled set of data includes a profile score for the at least one attribute; and The parsed set of data is stored in at least one client database.
8. The system according to claim 7, wherein: The process also includes: Obtaining at least one set of analysis rules and at least one set of processing rules corresponding to the at least one attribute; comparing the at least one attribute to at least one set of profiling rules and the at least one set of processing rules to verify information included in the at least one attribute; In response to determining that the information included in the at least one attribute is parsed according to the at least one set of processing rules, storing the information included in the at least one attribute in at least one parsed format according to the at least one set of processing rules, wherein the information is provided via at least one API gateway; and The at least one attribute is processed by at least one set of rule engines.
9. The system according to claim 7, wherein: The process also includes: In response to determining that the at least one attribute represents a name, the at least one attribute is processed by a name engine that correlates the at least one attribute with related names included in a list of related names.
10. The system according to claim 7, wherein: The process also includes: At least one set of continuous data profiling tools is received from a continuous data profiling manager, wherein the continuous data profiling manager is a front-end software application communicatively coupled to the at least one client database.
11. The system according to claim 7, wherein: The process also includes: connecting at least one continuous data profiling (CDP) manager application to the at least one client database; and At least one instruction or at least one function is received via the at least one continuous data profiling (CDP) manager application.
12. The system according to claim 11, wherein the at least one function is a first function for generating a statistical table of the profiled data, wherein the at least one function is a second function for generating a change data storage table for the parsed data, wherein the at least one function is a third function for managing at least one profiling process, and Wherein the at least one function is a fourth function for managing at least one profile triggered by comparison with the at least one input data stream.
13. A non-transitory computer-readable medium storing instructions which, when executed by a computing system, cause the computing system to perform operations for continuously profiling data, the operations comprising: receiving at least one input data stream; parsing the at least one input data stream by: identifying at least one attribute in the at least one input data stream, wherein the at least one attribute is associated with a set of features, determining a profile score for the at least one attribute based on a source of the at least one input data stream and an aggregation of the set of features, determining that the at least one attribute represents an address, and In response to determining that the at least one attribute represents an address, processing the at least one attribute by an address repository engine, the address repository engine adding the at least one attribute to an address repository; generating a profiled set of data based on profiling the at least one input data stream, wherein the profiled set of data includes a profile score for the at least one attribute; as well as The parsed set of data is stored in at least one client database.
14. The non-transitory computer readable medium of claim 13, wherein the operations further comprise: Obtaining at least one set of analysis rules and at least one set of processing rules corresponding to the at least one attribute; comparing the at least one attribute to at least one set of profiling rules and the at least one set of processing rules to verify information included in the at least one attribute; In response to determining that the information included in the at least one attribute is parsed according to the at least one set of processing rules, storing the information included in the at least one attribute in at least one parsed format according to the at least one set of processing rules, wherein the information is provided via at least one API gateway; as well as The at least one attribute is processed by at least one set of rule engines.
15. The non-transitory computer readable medium of claim 13, wherein the operations further comprise: In response to determining that the at least one attribute represents a name, the at least one attribute is processed by a name engine that correlates the at least one attribute with related names included in a list of related names.
16. The non-transitory computer readable medium of claim 13, wherein the operations further comprise: At least one set of continuous data profiling tools is received from a continuous data profiling manager, wherein the continuous data profiling manager is a front-end software application communicatively coupled to the at least one client database.
17. The non-transitory computer readable medium of claim 13, wherein the operations further comprise: connecting at least one continuous data profiling (CDP) manager application to the at least one client database; and receiving at least one instruction or at least one function via the at least one continuous data profiling (CDP) manager application, wherein the at least one function is a first function for generating a statistical table of the profiled data, wherein the at least one function is a second function for generating a change data storage table for the parsed data, wherein the at least one function is a third function for managing at least one profiling process, and Wherein the at least one function is a fourth function for managing at least one profile triggered by comparison with the at least one input data stream.
Citation Information
Patent Citations
Context driven data profiling
US20210319027A1