Context-driven data profiling
By employing verification, standardization, and rule engine processing during the data analysis process, the problems of high computational resource requirements and low insight quality when processing large amounts of data are solved, achieving efficient data quality improvement and insight generation.
Patent Information
- Application Number
- CN202180041734.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-09
- Filing Date
- 2021-04-09
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2041-04-09
AI Technical Summary
When processing and analyzing large amounts of data, existing technologies require significant computing resources and struggle to yield high-quality data insights.
Through data profiling processes, including validation, standardization, and rule engine processing, insights are generated from the data, and machine learning and artificial intelligence technologies are used to improve data quality.
It improves data quality, generates valuable insights from the data, and optimizes data matching accuracy and processing efficiency.
Smart Images

Figure CN115698977B_ABST
Abstract
Description
[0001] Cross-referencing related applications
[0002] This application claims the benefit and priority of U.S. Patent Application No. 16 / 844,927, filed April 9, 2020, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates to data profiling, and more particularly to performing data profiling to gain insights into the data. Background Technology
[0004] Multiple entities may maintain large amounts of data digitally across multiple computing devices. For example, an organization may maintain multiple columns of data on a series of interconnected servers. It may often be necessary to inspect and evaluate this large amount of data to determine multiple insights into its various characteristics. However, acquiring and processing large amounts of data can require significant computing resources. Furthermore, given the vast amount of information contained within the data volume, deriving high-quality data can often be challenging. Attached Figure Description
[0005] By studying the specific embodiments in conjunction with the accompanying drawings, those skilled in the art will gain a clearer understanding of the various features and characteristics of this technology. Embodiments of this technology are shown in the drawings by way of example rather than limitation, wherein the same reference numerals may denote the same elements.
[0006] Figure 1 This is an example network architecture that can be implemented in this embodiment.
[0007] Figure 2 This is a block diagram illustrating the example data profiling process.
[0008] Figure 3 This is a block diagram illustrating the example validation and standardization process.
[0009] Figure 4 It is a diagram used to generate example value scores for multiple attributes.
[0010] Figure 5 This is a block diagram of an example method for implementing the data profiling process.
[0011] Figure 6 This is a flowchart illustrating an example method for securely classifying and tokenizing data during the data registration process.
[0012] Figure 7 This is a block diagram illustrating an example of a processing system that can at least implement some of the operations described herein.
[0013] The various embodiments depicted in the accompanying drawings are for illustrative purposes only. Those skilled in the art will recognize that alternative embodiments can be employed without departing from the technical principles. Therefore, although specific embodiments are shown in the drawings, the technology is suitable for various modifications. Detailed Implementation
[0014] Many entities (e.g., companies, organizations) maintain large amounts of data. This data may be stored in various registers or databases on computing devices. In many cases, these entities may need to identify and match records in different datasets and gain insights into those datasets. For example, based on multiple similar datasets, an organization might attempt to identify and select a high-quality and accurate dataset from among those similar datasets.
[0015] This embodiment relates to processing data and generating insights from ingested data. The data profiling process may include validating the attributes of client data, standardizing the attributes into a standardized format, and processing the data via one or more rule engines. Additional information, such as rankings or value scores, may be generated based on the acquired input information.
[0016] Data profiling processes can generate insights into data, thereby improving data quality. Examples of insights can include duplicates or multiple instances of data attributes within and across domains, including percentage overlap. As further examples, insights can include trends from normalized and standardized data quality reports (what percentage of standard data is relative to non-standard data) or label-based processing (e.g., records with the same home address).
[0017] The embodiments described below represent the necessary information to enable those skilled in the art to implement these embodiments and illustrate the best mode of practice. Upon reading the following description in conjunction with the accompanying drawings, those skilled in the art will understand the concept of the art and will recognize the application of these concepts not specifically set forth herein. These concepts and applications fall within the scope of this disclosure and the appended claims.
[0018] Implementations may be described with reference to specific computer programs, system configurations, networks, etc. However, those skilled in the art will recognize that these features are equally applicable to other types of computer programs, system configurations, network types, etc. For example, although the term "Wi-Fi network" may be used to describe a network, the relevant implementations may be deployed in another type of network.
[0019] Furthermore, the techniques disclosed herein can be implemented using dedicated hardware (e.g., circuitry), programmable circuitry suitably programmed with software and / or firmware, or a combination of dedicated hardware and programmable circuitry. Therefore, embodiments may include machine-readable media having instructions for programming computing devices (e.g., computing devices or network-accessible server systems) to examine and process data as described herein.
[0020] the term
[0021] The terminology used herein is for the purpose of describing embodiments only and is not intended to limit the scope of this disclosure. Where the context permits, words used in singular or plural forms may also include either the singular or plural forms, respectively.
[0022] As used herein, unless otherwise expressly stated, terms such as “processing,” “computing,” “calculating,” “determining,” “displaying,” “generating,” or similar terms refer to the actions and processes of a computer or similar electronic computing device that manipulates and converts data represented as physical (electronic) quantities in computer memory or registers into other data similarly represented as physical quantities in the computer’s memory, registers, or other such storage medium, transmission, or display device.
[0023] As used herein, terms such as “connection,” “coupled,” or similar terms may refer to any direct or indirect connection or coupling between two or more elements. The coupling or connection between elements may be physical, logical, or a combination of both.
[0024] A reference to "an embodiment" or "one embodiment" means that a particular feature, function, structure, or characteristic described is included in at least one embodiment. The appearance of such phrases does not necessarily refer to the same embodiment, nor does it necessarily refer to mutually exclusive alternative embodiments.
[0025] Unless the context explicitly requires otherwise, “comprise” and “comprising” should be understood as inclusive rather than exclusive or exhaustive (i.e., “including but not limited to”).
[0026] The term "based on" should also be understood in an inclusive sense, rather than an exclusive or exhaustive one. Therefore, unless otherwise stated, the term "based on" is intended to mean "at least partially based on".
[0027] The term "module" broadly refers to software components, hardware components, and / or firmware components. A module is typically a functional component that can generate useful data or other outputs based on specified inputs. A module can be self-contained. A computer program may include one or more modules. Therefore, a computer program may include multiple modules responsible for performing different tasks or a single module responsible for performing multiple tasks.
[0028] When used to refer to a list of multiple items, the word "or" is intended to cover all of the following interpretations: any item in the list, all items in the list, and any combination of items in the list.
[0029] The sequence of steps performed in any process described herein is exemplary. However, unless contrary to physical possibilities, these steps may be performed in a variety of orders and combinations. For example, steps may be added to or removed from the process described herein. Similarly, steps may be substituted or reordered. Therefore, any description of a process is intended to be open-ended.
[0030] Data profiling overview
[0031] Context-driven data profiling processes can help determine the quality of source data. Data profiling can include several processing steps that modify input information to generate insights into the data that help optimize applications such as matching accuracy. For example, data profiling can standardize and validate data before tokenizing the data being profiled.
[0032] Figure 1 This is a block diagram of an example profiling flow 100. A data profiler can incorporate flexible data flows. Data can be accessed and / or processed from a data source in multiple batches, continuous streams, or large-volume loads. Data source 102 may include nodes (e.g., devices 104a-d) configured to store / maintain data (e.g., data lake 106a, database 106b, flat file 106c, data stream 106d). For example, data source 102 may include a single column of data, a series of relational databases with multiple data tables, or a data lake with a large amount of data assets.
[0033] Data quality can be resolved by use case or client within the data profiler. For example, the context can be based on a single data column, a combination of multiple data columns, or a data source. During data profiling, various data types can be derived, and data summaries can be generated. For example, a summary of a data column can be represented as a data sketch. A data sketch can include numerical data and / or string data. Examples of numerical data included in a data sketch can include any of the following: missing values, mean / variance / maximum / minimum values of the numerical data, approximate quantile estimates of the numerical data that can be used to generate distributions or histograms, etc. Examples of string data can include missing values, maximum character length, minimum character length, average character length, label frequency tables, sets of frequency items, estimates of distinct values, etc.
[0034] Once any of a set of metrics has been calculated in the data summary, a data profiling score can be generated. The data profiling score can be used to determine data quality and / or identify best-in-class data, data composition, and target data quality enhancement activities. Data profiling can be re-executed at user-defined time intervals to recalculate the metrics. This can be used to track the history of data scores throughout the data lifecycle and to flag data quality issues.
[0035] In some embodiments, the data summary may include a proportion of data that follows a specific regular expression (e.g., reference data). For example, for telephone numbers that follow a specific format, the data summary may represent several existing formats.
[0036] In some embodiments, the data summary may include several anonymous values. For example, known anonymous names (e.g., John Doe) may be identified in the source data to determine the proportion of data that includes anonymous values.
[0037] In other embodiments, the data digest may include a set of data quality metrics based on a data quality rule base. Data digests can be used to learn data quality rules based on reference data associated with attributes. Data digests can also be used to learn data quality rules directly from source data (e.g., between which values the source data should contain, and what the minimum character length should be).
[0038] As a first example, the source data can be examined to derive a data quality score. A data quality score can include scores calculated at the column level or record level of the source data. A data quality score can be derived by calculating any metrics included in the data summary.
[0039] As another example, source data can be examined to identify data quality. For each column of data in each data source, a data profile score can be used to match the most likely set of data with a specific client. For instance, a table can be prepared displaying a set of columns / attributes (e.g., name, address, phone number, date of birth, email address) and their data profile scores from different sources (CRM, ERP, order management, web). Using the data contained in such a table, the set of data with the highest quality can be selected for a specific client. In some cases, multiple sources can be matched to receive the highest possible data quality. This can be done without over-processing the source data.
[0040] As another example, source data can be examined to derive historical data profile scores and perform what-if analysis. What-if analysis can include analyses of what would happen if other (certain) rules were applied to the data. To facilitate these calculations, this process can be performed on sample data collected from the data summary created during the indicator calculation phase. If the results of the what-if analysis are sufficient, a new, complete calculation of the indicator can be performed using the new rules selected during the what-if analysis.
[0041] Data extracted from a data source (e.g., data lake 106a, database 106b, flat file 106c, data stream 106d) can be fed to a profiler (e.g., profilers 110a-n) via a data feed 108. The data feed 108 can include batch, bulk, or continuous data feeds to the profiler. Data fed into the profiler can include attributes (e.g., attributes 112a-n). Attributes can be portions of data in a table, a source, or a part of the same record.
[0042] In such Figure 1 In the illustrated embodiment, a first profiler 110a can process attribute 1 112a, and a second profiler 110b can process attribute 2 112b. Any suitable number of profilers (e.g., profiler N 110n) can process any number of attributes (e.g., attribute N 112n). Each profiler 110a-n may include a set of normalization rules 114a-n and a set of rule engines 116a-n. The normalization rules 114a-n and / or rule engines 116a-n may be modular, wherein each set of rules can be used to process attributes. Each profiler can use a corresponding set of normalization rules and a set of rule engines to process a corresponding attribute. In some embodiments, each profiler may implement various machine learning and / or artificial intelligence techniques and statistical tools to improve data quality when processing attributes. The resulting data from each profiler 110a-n may include insights 118 representing various characteristics of the attribute.
[0043] In some embodiments, data quality rules can be adjusted, which can lead to different determinations when performing data quality improvement tasks. For example, a dataset may have good scores, but it was previously unaware that the name "John Doe" was an anonymous (forged or fake) value. By updating the rules to identify "John Doe" as an anonymous value, changes in data profile scores and score history may be modified. Such changes in data profile scores can enable the identification of multiple data types contained in the dataset.
[0044] As another example, source data can be reviewed to generate automatic data quality improvement requests. Triggers can be associated with the data profile score of a specific attribute or a range of attributes. A trigger can specify that if the data profile score falls below a threshold, the source data associated with that attribute can be reviewed. If the source data has identifying values that indicate how the data is used in various contexts, it can potentially be improved.
[0045] As another example, source data can be examined to derive data insights. Processing the data profile score of the source data can generate data distributions and other insights that can be used to understand the characteristics of the data before initiating another data profile.
[0046] As another example, source data can be reviewed to arrive at data selection decisions based on intelligent data quality. Based on mapping source data to a model (e.g., a canonical model), highly relevant profiling / sampling outputs, relevant definitions, and / or similar endpoint consumption relationship patterns can provide recommendations for worthwhile alternatives when the data quality score is better than another dataset with similar attributes. Side-by-side comparisons can be run upon user-initiated requests to help users identify overlapping measurements and express relative preferences. This can be stored / documented with users and the community to provide recommendations calibrated to user-specific needs over the long term.
[0047] Figure 2 This is a block diagram 200 illustrating an example data profiling process. (For example...) Figure 2 As shown, data profiling 200 may include acquiring input information. Example input information may include generated context / classification information (or "labels") 202 and / or ingested data 204. Ingested data 204 may include client-side data.
[0048] The data profiling process 200 may include defining attributes 206. Attribute 206 may represent characteristics or features of the client data. For example, attribute 206 may include date of birth (e.g., January 1, 1990). This may include month, day, year, and / or full date of birth (DOB). Other example attributes 206 may include address, name, email address, gender, phone number, Social Security number, etc. Attribute 206 may also include tags / categories representing the client data.
[0049] Data profiling 200 may include the normalization 208 of attribute 206. Normalization 208 may include validating the data corresponding to attribute 206 contained in attribute 206 and normalizing the format of attribute 206 to a uniform format. Data profiling 200 may include multiple normalization processes that can normalize various types of attributes. In many examples, normalization may be horizontally and / or vertically modular. Figure 3 The standardization of attributes is discussed in more detail.
[0050] The normalized attributes can be processed via one or more rule engines 210. The rule engines can further process the normalized attributes, deriving more insights from them. Example rule engines 210 may include nickname engine 212a, address database engine 512b, or any other number of rule engines (e.g., rule engine N 212n).
[0051] Address database engine 512b may include an identifier attribute to determine whether an address is included and add that address to a repository / list containing multiple addresses. Address database engine 512b may associate addresses with clients / entities. During processing via rule engine 210, data profiling may output profiling data 514.
[0052] like Figure 2 As shown, the profiling process can output any of the following: a ranking of 216 and / or a value score of 518. A ranking of 216 can represent the rank of an attribute type relative to other attribute types. For example, the attribute "First Name" can have a higher ranking than the attribute "Gender". A ranking of 216 can represent the information quality of an attribute type and / or several insights associated with that attribute type. For example, an attribute type with a higher ranking of 216 can indicate that more insights can be derived from that attribute type.
[0053] As an example, a ranking of 216 can be used to represent the data quality of each attribute. For instance, in a healthcare context, data can be linked based on patient availability. In this example, the data value in identifiers (such as Social Security Numbers (SSNs)) would typically be high, but in a healthcare context, the data value in patient identifiers would be greater than that in SSNs. In this example, the ranking could be a series of scores representing the most unique identifiers used to identify patients, as the resulting data would have greater value. Therefore, in this example, the patient identifier attribute might have a higher ranking than the SSN.
[0054] As another example, in a corporate context, given that SSNs are used to identify employees, such as for payroll purposes, employers can assign the highest usage ranking to the SSN that represents an employee.
[0055] A value score of 218 can represent the value of multiple characteristics of an attribute type. For example, a value score of 218 can represent the aggregated value of multiple characteristics of an attribute type relative to other attribute types. Value scores can provide additional insights into the attributes of ingested data. Figure 4 A more detailed discussion was conducted in the middle.
[0056] Rankings and value scores can be provided to a server system with network access. In some embodiments, profiling 200 may include performing a series of steps to process and normalize the input information. For example, processing the input information may include removing outliers (e.g., foreign characters) or false values from the raw data. This can normalize the input information and provide different levels of insight into data quality.
[0057] Figure 3 This is a block diagram 300 illustrating an example validation and standardization process. The process may include acquiring an attribute and processing that attribute to validate and standardize information including that attribute.
[0058] As mentioned above, example attributes can include name, date, address, etc. Figure 3 In the example shown, attribute 302 can include date of birth. Date of birth can include multiple features, such as month 304a, day 304b, year 304c, and full date of birth (DOB) 304d. For example, date of birth could be set to January 1, 1990.
[0059] Attributes can be validated via validation process 306. A set of validation rules 308 can be compared with the attributes' characteristics (e.g., 304a-d) to determine if the attribute is correctly identified as such. For example, validation rules can determine whether the characteristic of birth date actually represents a birth date. For instance, if the attribute is a credit card number instead of a birth date, validation rules can identify that the attribute is incorrectly identified as a birth date. In this example, the attribute can be processed through a separate validation and normalization process associated with the credit card number. If the attribute fails the validation rules, the attribute can be null or empty 310.
[0060] Validation rule 308 can include a set of characteristics of attribute 302 that identify whether the attribute contains information representing the attribute. For example, validation rule 308 can examine an attribute to determine if it is an invalid value. For instance, if an attribute contains an invalid value, then the attribute does not identify a date of birth and should be identified as an invalid value 310. Other example validation rules 308 can include determining whether an attribute contains at least one character of a name, whether a US phone number is no more than 10 digits, and whether it contains no punctuation marks other than dashes, forward slashes, and periods. A set of validation rules 308 can be provided for each type of attribute. In some examples, validation rules can be updated to modify / add / delete any validation rule.
[0061] You can view the processed attributes to generate attribute value scores. Value scores can include values related to various features of the attributes aggregated and ingested from the data.
[0062] Figure 4 This is a diagram 400 used to generate example value scores for multiple attributes. For example... Figure 4 As shown, you can view various attributes 402 of the ingested data to derive a value score for each attribute 402. Example attributes may include address 404a, name 404b, phone number 404c, and any number of other attribute types (e.g., attribute 1 404d, attribute N 404n).
[0063] Value scores can be generated using multiple features for each attribute. For example, each attribute can be examined to derive its quality feature 406. Quality feature 406 can represent the relative difference between an attribute and its standardized version. Generally, if an attribute closely corresponds to its standardized version, the overall quality of the attribute is likely to be higher. Therefore, quality feature 406 can represent several modifications made to the attribute to provide a standardized format. The number of modifications made to the attribute to provide a standardized format can be converted into a value for quality feature 406.
[0064] Another example feature could be usability feature 408. Usability feature 408 could represent a number of invalid / empty entries for an attribute in a subset of ingested data. For example, as the number of invalid / empty entries in a column of data increases, the overall quality of the attributes in that column of data may also decrease. Therefore, a value for usability feature 408 could be derived based on a number of invalid / empty entries for that attribute type relative to other attribute types.
[0065] Value scores can be based on any suitable number of features (e.g., feature 1410, feature N412). Any feature used to derive an attribute type can include looking at a subset of the ingested data (e.g., a column of data) and comparing the characteristics of the ingested data with other attribute types to derive the feature for the attribute type. As an example, feature 412 could include the cardinality of the attribute, which can represent the uniqueness of the attribute relative to other attributes.
[0066] Value scores can be based on the weight of each attribute relative to other attributes. Each attribute type can be weighted based on other data in a reference dataset, which can adjust the values of other features for that attribute type.
[0067] The attributes' features (e.g., features 406, 408, 410, 412) and the determined weights 414 can be used to derive a default score 416. The default score 416 can be an initial value / score that aggregates values associated with the attributes' features and can be adjusted based on the attribute's weights 414. In some embodiments, various techniques (e.g., machine learning, neural networks) can be used to increase the accuracy of the default score for the attribute. For example, the default score can be dynamically adjusted using training data that can improve the accuracy of the default score 416.
[0068] A value score of 418 can be derived from the default score of 416. As mentioned above, the value score 418 can include an aggregation of various characteristics of the attribute type. In some cases, the value score 418 can be encrypted and maintained by a server system with access to a network.
[0069] Example methods for implementing the data profiling process
[0070] Figure 5 This is a block diagram 500 illustrating an example method for implementing a data profiling process. The method may include ingesting a data stream corresponding to a client (block 502). The data stream corresponding to the client may include one or more columns of client data.
[0071] The method may include identifying attributes from a data stream (box 504). The method may include processing attributes via a data profiling process (box 506). The data profiling process may include obtaining a set of validation rules and a set of normalization rules corresponding to the attributes (box 508). The set of validation rules can provide rules indicating whether an attribute corresponds to a given attribute. The set of normalization rules can provide rules for modifying attributes to a normalized format.
[0072] The data profiling process may include comparing an attribute to a set of validation rules to determine whether the attribute corresponds to the specified attribute (box 510). If the attribute is determined to correspond to the specified attribute, it can be modified as described herein. In some embodiments, validating an attribute may include determining whether the attribute includes invalid values, which are identified in a set of validation rules. An attribute may be validated to determine that it does not contain invalid values.
[0073] The data profiling process may include modifying attributes to a standardized format according to a set of standardization rules (Box 512). This can be performed in response to determining that an attribute has been validated by the validation rules.
[0074] The data profiling process may include processing attributes through multiple rule engines (Box 514). A rule engine may include a name engine, which, in response to determining that the attribute represents a name, associates the attribute with common associated names contained in a list of associated names. A rule engine may also include an address database engine, which, in response to determining that the attribute represents an address, adds the attribute to an address database associated with the client.
[0075] In some embodiments, processing a modified attribute through a set of rule engines may include, in response to determining that the attribute represents a name, processing the modified attribute through a name engine, which associates the attribute with an associated name included in a list of associated names. Processing a modified attribute through a set of rule engines may also include, in response to determining that the attribute represents an address, processing the modified attribute through an address database engine, which adds the attribute to an address database associated with the client.
[0076] In some embodiments, the method may include comparing several instances of the attribute relative to other attributes in the data stream. A ranking can be generated for the attribute. The ranking can be based on the number of instances of the attribute in the data stream, and the ranking can represent several insights that can be derived from the attribute.
[0077] In some embodiments, a set of features may be identified, which are associated with an attribute and identified relative to other attributes in the data stream. Example features of a set of features may include quality features, availability features, cardinality features, etc. The value score of an attribute can be derived based on the aggregation of a set of features.
[0078] In some embodiments, deriving a value score for an attribute based on the aggregation of a set of features may include processing the attribute to derive a quality feature that identifies several differences between the attribute identified in the data stream and a modified attribute modified according to a set of standardization rules. Determining a value score for an attribute based on the aggregation of a set of features may also include processing the attribute to derive an availability feature that represents several invalid entries in a subset of data corresponding to the attribute in the data stream. Determining a value score for an attribute based on the aggregation of a set of features may also include processing the attribute to derive a cardinality feature that represents the difference between the attribute and other attributes in the data stream. Determining a value score for an attribute based on the aggregation of a set of features may also include aggregating the derived quality feature, availability feature, and cardinality feature of the attribute to generate a value score for the attribute.
[0079] This method may include outputting processed insights / profiles / rankings / scores of attributes to a network-accessible server system (Box 516). The network-accessible server system may maintain insights / profiles / rankings / scores of a range of processed attributes and generate data quality insights into client data.
[0080] Example methods for implementing the data registration process
[0081] In some embodiments, the data profiling process described herein can be implemented as part of an example data registration process. Figure 6 This is a block diagram illustrating an example method 600 for securely classifying and tokenizing data during the data registration process. (See diagram for example.) Figure 6 As shown, the method may include ingesting a dataset corresponding to the client (box 602). The dataset may include a series of data columns associated with the client. This information may be maintained at the client node. In some examples, at least a portion of the data included in the dataset includes personally identifiable information (PII).
[0082] The method may include examining the dataset to identify classifiers, which represent characteristics of attributes included in the dataset (box 604). In some embodiments, the classifiers include any one of domain classifiers, subdomain classifiers, attribute classifiers, and entity classifiers. In some examples, each classifier may be determined based on examining the dataset.
[0083] The method may include obtaining client-specific encrypted information and client-specific configuration information, the information including a list of anonymous tags representing the types of information included in the dataset (box 606). In some embodiments, the client-specific encrypted information may be obtained from a security server, the client-specific encrypted information may be encrypted using the Hash Message Authentication Code (HMAC) protocol, and the hash code may include a computer-generated SHA2 512 / 256 token.
[0084] The method may include identifying a first tag (box 608) in an anonymous tag list corresponding to an information type in the attribute, based on an identified classifier. A tag can provide an anonymous identifier for an information type represented by the attribute. The tag can be generated based on any of the attribute and the classifier. For example, if the attribute is related to a name, the corresponding tag could be "La1". In these embodiments, only entities authorized to access the information list corresponding to the tag can identify the information type identified by each tag, thereby anonymizing the data.
[0085] The method may include processing the attributes of the dataset to generate modified attributes that have been adapted to a standardized format (box 610). This may include the profiling process as described herein.
[0086] In some embodiments, processing the attributes of the dataset to generate the modified attributes further includes obtaining a set of validation rules and a set of normalization rules corresponding to the first label. The set of validation rules can provide rules indicating whether the attribute corresponds to the first label. The set of normalization rules can provide rules for modifying the attribute into a normalized format. The attribute can be compared with the set of validation rules to determine whether the attribute corresponds to the first label. The attribute can be modified into a normalized format according to the set of normalization rules in response to determining that the attribute corresponds to the first label.
[0087] In some embodiments, processing the attributes of the dataset to generate the modified attributes further includes processing the attributes using a series of rule engines. The rule engines may include a name engine that, in response to determining that the attribute represents a name, associates the attribute with commonly associated names included in a list of associated names. The rule engines may also include an address database engine that adds the attribute to an address database associated with the client in response to determining that the attribute represents an address.
[0088] The method may include generating a tokenized version of the modified attribute (box 612). Generating a tokenized version of the modified attribute may include hashing the modified attribute using a hash code included in client-specific cryptographic information to generate a hashed modified attribute (box 614). The hashed modified attribute may be compressed from a 64-character token to a 44-character string using an encoding scheme.
[0089] Generating a tokenized version of the modified attribute may further include comparing a first tag with a token store, the token store including a series of client-specific tokens to identify a first token corresponding to the first tag (box 616). Generating a tokenized version of the modified attribute may further include generating a context token for the modified attribute that includes the first token (box 618).
[0090] In some embodiments, a tokenized version of the modified attributes can be sent from a remote node to a network-accessible server system.
[0091] In some embodiments, in response to identifying the first label, the method may include generating a first set of insights for the dataset based on the first label and attributes. In response to generating modified attributes, the method may further include generating a second set of insights for the dataset based on the modified attributes. The first and second sets of insights may be stored on a network-accessible server system.
[0092] Example processing system
[0093] Figure 7 This is a block diagram illustrating an example of a processing system 700, which can implement at least some of the operations described herein. Figure 7 As shown, the processing system 700 may include one or more central processing units (“processors”) 702, main memory 706, non-volatile memory 710, network adapter 712 (e.g., network interface), video display 718, input / output devices 720, control devices 722 (e.g., keyboard and pointing devices), a drive unit 724 including storage medium 726, and signal generation devices 730 communicatively connected to bus 716. Bus 716 is shown as an abstraction, representing any one or more individual physical buses, point-to-point connections, or both connected via appropriate bridges, adapters, or controllers. Therefore, bus 716 may include, for example, a system bus, a Peripheral Component Interconnect (PCI) bus or PCI-Express bus, HyperTransport or Industry Standard Architecture (ISA) bus, Small Computer System Interface (SCSI) bus, Universal Serial Bus (USB), IIC (I2C) bus, or the Institute of Electrical and Electronics Engineers (IEEE) standard 1394 bus, also known as “FireWire”.
[0094] In various embodiments, the processing system 700 operates as part of a user equipment, although the processing system 700 may also be connected (e.g., wired or wireless) to the user equipment. In a networked deployment, the processing system 700 may operate as a server or client in a client-server network environment, or as a peer in a peer-to-peer (or distributed) network environment.
[0095] The processing system 700 may be a server computer, client computer, personal computer, tablet computer, laptop computer, personal digital assistant (PDA), cellular phone, processor, network device, network router, switch or bridge, console, handheld console, gaming device, music player, networked (“smart”) TV, TV connection device, or any portable device or machine capable of executing a set of instructions (sequential or otherwise) that specifies the actions to be taken by the processing system 700.
[0096] Although main memory 706, non-volatile memory 710, and storage medium 726 (also referred to as "machine-readable medium") are shown as a single medium, the terms "machine-readable medium" and "storage medium" should be understood to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) storing one or more sets of instructions 728. The terms "machine-readable medium" and "storage medium" should also be understood to include any medium capable of storing, encoding, or carrying a set of instructions for execution by a computing system and causing the computing system to perform any one or more methods of the currently disclosed embodiments.
[0097] Generally, routines executed to implement embodiments of the present disclosure may be implemented as part of an operating system or a specific application, component, program, object, module, or sequence of instructions referred to as a "computer program." A computer program typically includes one or more instructions (e.g., instructions 704, 708, 728) located in various memories and storage devices within a computer at multiple times, and when read and executed by one or more processing units or processors 702, causes the processing system 700 to operate to perform elements relating to various aspects of the present disclosure.
[0098] Furthermore, although embodiments have been described in the context of fully functional computers and computer systems, those skilled in the art will understand that the various embodiments described herein can be distributed as program products in various forms, and this disclosure applies equally to any particular type of machine or computer-readable medium on which the distribution is actually implemented. For example, the techniques described herein can be implemented using virtual machines or cloud computing services.
[0099] Further examples of machine-readable storage media, or computer-readable (storage) media, include, but are not limited to, recordable type media such as volatile and non-volatile storage devices 810, floppy disks and other removable disks, hard disk drives, optical disks (e.g., optical disk read-only memories (CD-ROMs), digital universal optical disks (DVDs)), and transmission type media such as digital and analog communication links.
[0100] Network adapter 712 enables processing system 700 to process data formed in network 714 by entities outside system 700 via any known and / or convenient communication protocols supported by processing system 700 and external entities. Network adapter 712 may include one or more of the following: network adapter card, wireless network interface card, router, access point, wireless router, switch, multilayer switch, protocol converter, gateway, bridge, bridging router, hub, digital media receiver, and / or repeater.
[0101] Network adapter 712 may include a firewall, which in some embodiments can control and / or manage permissions to access / proximize data in a computer network and track different trust levels between different machines and / or applications. The firewall can be any number of modules with any combination of hardware and / or software components capable of enforcing a predetermined set of access permissions between a specific set of machines and applications, between machines, and / or between applications, for example, to regulate traffic and resource sharing between these different entities. The firewall may additionally manage and / or access access control lists that detail permissions, including, for example, access and manipulation permissions for objects by individuals, machines, and / or applications, and under what circumstances the rights to those permissions are granted.
[0102] As described above, the techniques described herein are implemented through, for example, programmable circuits (e.g., one or more microprocessors), programmed with software and / or firmware, entirely as dedicated hardwired (i.e., non-programmable) circuits, or a combination of these forms. Dedicated circuits can take the form of, for example, one or more application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), etc.
[0103] As can be understood from the foregoing, specific embodiments of the present invention have been described herein for illustrative purposes, but various modifications can be made without departing from the scope of the invention. Therefore, the present invention is not limited except for the appended claims.
Claims
1. A computer-implemented method, comprising: Ingest the data stream corresponding to the client; Identify the attributes contained in the data stream; The attribute is processed during the data profiling process, which includes: Obtain a set of validation rules and a set of standardization rules corresponding to the attribute; The attribute is compared with the set of validation rules to validate the information contained in the attribute; In response to determining that the information contained in the attribute has been validated according to the set of validation rules, the attribute is modified to a standardized format according to the set of standardization rules; Modified attributes are processed through a set of rules engines; The processed attributes are output to a network-accessible server system; and The quality feature value of the processed attribute is obtained, wherein the quality feature value is calculated based on the number of modifications made to transform the processed attribute into a standardized version of the processed attribute.
2. The computer-implemented method of claim 1, wherein the attribute includes an impression of a portion of data contained in the data stream that prevents the transmission of information contained in the data stream from a client node maintaining the data stream.
3. The computer-implemented method according to claim 1, wherein processing the modified attributes through the set of rule engines further includes: In response to determining that the attribute represents a name, the modified attribute is processed by a name engine that associates the attribute with associated names contained in a list of associated names; as well as In response to determining that the attribute represents an address, the modified attribute is processed by an address library engine, which adds the attribute to the address library associated with the client.
4. The computer-implemented method according to claim 1 further includes: Compare the attribute with respect to several instances of other attributes in the data stream; as well as Generate a usage ranking for the attribute, the usage ranking being based on the number of instances of the attribute in the data stream, wherein the usage ranking represents several insights that can be derived from the attribute.
5. The computer-implemented method according to claim 1, further comprising: A set of features that identify the attribute as associated with other attributes in the data stream; as well as The value score of the attribute is derived from the aggregation of the aforementioned series of features.
6. The computer-implemented method of claim 5, wherein deriving the value score of the attribute based on the aggregation of the series of features further comprises: The attribute is processed to derive a quality feature of the attribute, the quality feature identifying several differences between the attribute identified in the data stream and the modified attribute modified according to the set of standardization rules; The attribute is processed to obtain an availability feature of the attribute, the availability feature representing a number of invalid entries in a portion of the data corresponding to the attribute in the data stream; The attribute is processed to derive a cardinality feature of the attribute, the cardinality feature representing the difference of the attribute relative to other attributes in the data stream; The derived quality features, usability features, and cardinality features of the attribute are aggregated to generate the value score of the attribute.
7. The computer-implemented method according to claim 5, wherein comparing the attribute with the set of verification rules to verify the information contained in the attribute further comprises: Determine whether the attribute contains an invalid value identified in the set of validation rules, wherein the attribute is validated in response to determining that the attribute does not contain the invalid value.
8. The computer-implemented method according to claim 1, further comprising: Obtain client-specific configuration information containing a list of tags, wherein each tag in the list provides a client-specific representation of the type of information contained in the data stream; as well as The identifier represents a first tag included in the list of tags that indicates the information contained in the attribute, wherein the set of validation rules and the set of standardization rules correspond to the first tag.
9. A method for generating modified attributes of a dataset, performed by a computing node, the method comprising: Retrieve the dataset corresponding to the client from the client node; Identify attributes from the dataset, the attributes comprising an impression of a portion of the data in the dataset; Compare several instances of the attribute with other attributes in the dataset; A usage ranking of the attribute is generated based on the number of instances of the attribute in the dataset; Identify a set of features associated with the attribute, which is identified relative to other attributes in the dataset; The value score of the attribute is derived from the aggregation of the aforementioned series of features; Obtain a set of validation rules and a set of standardization rules corresponding to the attribute; The attribute is compared with the set of validation rules to validate the information contained in the attribute; In response to determining that the information contained in the attribute has been validated according to the set of validation rules, the attribute is modified to a standardized format according to the set of standardization rules; Modified attributes are processed through a set of rules engines; The processed attributes are output to a network-accessible server system; as well as The quality feature value of the processed attribute is obtained, wherein the quality feature value is calculated based on the number of modifications made to transform the processed attribute into a standardized version of the processed attribute.
10. The method of claim 9, wherein processing the modified attribute through the set of rule engines further comprises: In response to determining that the attribute represents a name, the modified attribute is processed by a name engine that associates the attribute with associated names contained in a list of associated names; as well as In response to determining that the attribute represents an address, the modified attribute is processed by an address library engine, which adds the attribute to the address library associated with the client.
11. The method of claim 9, wherein deriving the value score of the attribute based on the aggregation of the series of features further comprises: The attribute is processed to derive a quality feature of the attribute, the quality feature identifying several differences between the attribute identified in the data stream and the modified attribute modified according to the set of standardization rules; The attribute is processed to obtain an availability feature of the attribute, the availability feature representing a number of invalid entries in a portion of the data corresponding to the attribute in the data stream; The attribute is processed to derive a cardinality feature of the attribute, the cardinality feature representing the difference of the attribute relative to other attributes in the data stream; The resulting quality features, usability features, and cardinality features of the attribute are aggregated to generate the value score of the attribute.
12. The method of claim 9, wherein comparing the attribute with the set of validation rules to validate the information contained in the attribute further comprises: Determine whether the attribute contains an invalid value identified in the set of validation rules, wherein the attribute is validated in response to determining that the attribute does not contain the invalid value.
13. The method of claim 9, further comprising: Obtain client-specific configuration information containing a list of tags, wherein each tag in the list provides a client-specific representation of the type of information contained in the data stream; as well as The identifier represents a first tag included in the list of tags that indicates the information contained in the attribute, wherein the set of validation rules and the set of standardization rules correspond to the first tag.
14. A tangible, non-transitory computer-readable medium having instructions stored thereon that, when executed by a processor, cause the processor to: Ingest the data stream corresponding to the client; Identify the attributes contained in the data stream; The attribute is processed during the data profiling process, which includes: Obtain a set of validation rules and a set of standardization rules corresponding to the attribute; The attribute is compared with the set of validation rules to validate the information contained in the attribute; In response to determining that the information contained in the attribute has been validated according to the set of validation rules, the attribute is modified to a standardized format according to the set of standardization rules; and Modified attributes are processed through a set of rules engines; The processed attributes are output to a network-accessible server system; and The quality feature value of the processed attribute is obtained, wherein the quality feature value is calculated based on the number of modifications made to transform the processed attribute into a standardized version of the processed attribute.
15. The computer-readable medium of claim 14, wherein the attribute includes an impression of a portion of data contained in the data stream that prevents the transmission of information contained in the data stream from a client node maintaining the data stream.
16. The computer-readable medium of claim 14, wherein processing the modified attributes via the set of rule engines further comprises: In response to determining that the attribute represents a name, the modified attribute is processed by a name engine that associates the attribute with associated names contained in a list of associated names; as well as In response to determining that the attribute represents an address, the modified attribute is processed by an address library engine, which adds the attribute to the address library associated with the client.
17. The computer-readable medium of claim 14, further comprising the processor: Compare the attribute with respect to several instances of other attributes in the data stream; Generate a usage ranking for the attribute, the usage ranking being based on the number of instances of the attribute in the data stream, wherein the usage ranking represents several insights that can be derived from the attribute; A set of features is used to identify the value score of the attribute relative to other attributes in the data stream.
18. The computer-readable medium of claim 17, further comprising the processor: The attribute is processed to derive a quality feature of the attribute, the quality feature identifying several differences between the attribute identified in the data stream and the modified attribute modified according to the set of standardization rules; The attribute is processed to obtain an availability feature of the attribute, the availability feature representing a number of invalid entries in a portion of the data corresponding to the attribute in the data stream; The attribute is processed to derive a cardinality feature of the attribute, the cardinality feature representing the difference of the attribute relative to other attributes in the data stream; The resulting quality characteristics, usability characteristics, and cardinality characteristics of the attribute are aggregated to derive the value score of the attribute.
19. The computer-readable medium of claim 14, wherein comparing the attribute with the set of verification rules to verify the information contained in the attribute further comprises: Determine whether the attribute contains an invalid value identified in the set of validation rules, wherein the attribute is validated in response to determining that the attribute does not contain the invalid value.
20. The computer-readable medium of claim 14, further comprising: Obtain client-specific configuration information containing a list of tags, wherein each tag in the list provides a client-specific representation of the type of information contained in the data stream; as well as The identifier represents a first tag included in the list of tags that indicates the information contained in the attribute, wherein the set of validation rules and the set of standardization rules correspond to the first tag.
Citation Information
Patent Citations
Architecture for a data cleansing application
US20040107203A1
Selecting an anomaly for presentation at a user interface based on a context
US20180082449A1