Standardization in the context of data integration
By using machine learning data standardization models in a cloud computing environment to automatically classify and process data points, generate and apply data standardization rules, the problem of manual standardization in data integration in a cloud computing environment is solved, and the automation of data integration and data quality are improved.
Patent Information
- Application Number
- CN202280016349.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-24
- Filing Date
- 2022-02-18
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-02-18
AI Technical Summary
Existing technologies require extensive manual standardization during data integration in cloud computing environments, which consumes time and labor, and lacks automated data standardization methods.
The dataset is automatically classified and standardized using a machine learning data standardization model. The model classifies data points and handles outliers, generates data standardization rules, and applies these rules for automatic data standardization after client review.
It reduces reliance on subject matter experts and data administrators, improves the automation and quality of data integration, reduces the need for manual correction, and enables continuous improvement in data integration.
Smart Images

Figure CN116888584B_ABST
Abstract
Description
Background Technology
[0001] The various embodiments described herein generally relate to standardization in the context of data integration. More specifically, the embodiments describe techniques for creating a set of data standardization rules to facilitate data integration in the managed services domain of a cloud computing environment. Summary of the Invention
[0002] The various embodiments described herein provide techniques for automatic data normalization in the context of data integration. An associated computer-implemented method includes receiving a dataset during a data onboarding process. The method further includes classifying data points within the dataset. The method further includes applying a machine learning data normalization model to the data points of each class within the dataset. The method further includes obtaining a proposed set of data normalization rules for the dataset based on any normalization modifications determined as the machine learning data normalization model is applied. In one embodiment, optionally, the method includes presenting the proposed set of data normalization rules for client review, and, in response to acceptance of the proposed set of data normalization rules, applying the proposed set of data normalization rules to the dataset. In a further embodiment, the method includes updating the machine learning data normalization model based on the proposed set of data normalization rules in response to acceptance of the proposed set of data normalization rules.
[0003] One or more further embodiments relate to a computer program product including a computer-readable storage medium having program instructions. According to such embodiments, the program instructions are executable by a computing device to cause the computing device to perform one or more steps associated with the computer-implemented method described above and / or implement one or more embodiments associated with the computer-implemented method described above. One or more further embodiments relate to a system having at least one processor and a memory storing an application program, which, when executed on the at least one processor, performs one or more steps associated with the computer-implemented method described above and / or implements one or more embodiments associated with the computer-implemented method described above. Attached Figure Description
[0004] To obtain and understand in detail the above aspects, a more specific description of the embodiments briefly outlined above can be obtained with reference to the accompanying drawings.
[0005] However, it should be noted that the accompanying drawings only illustrate typical embodiments of the invention and should not be considered as limiting the scope of the invention, as the invention can allow for other equally effective embodiments.
[0006] Figure 1 A cloud computing environment according to one or more embodiments is shown.
[0007] Figure 2 An abstract model layer provided by a cloud computing environment is shown according to one or more embodiments.
[0008] Figure 3 A management service domain in a cloud computing environment according to one or more embodiments is illustrated.
[0009] Figure 4 A method for creating data standardization rules to facilitate data integration in a management service domain, according to one or more embodiments, is illustrated.
[0010] Figure 5 A method for configuring a machine learning data normalization model according to one or more embodiments is shown.
[0011] Figure 6 A method for applying a machine learning data normalization model to each data point in a dataset, according to one or more embodiments, is illustrated.
[0012] Figure 7 A method for determining at least one standardized modification to process data points within a dataset, according to one or more embodiments, is illustrated.
[0013] Figure 8 A method is shown that uses an application data scraping algorithm according to one or more embodiments to evaluate infrequently occurring anomalous string values within a dataset.
[0014] Figure 9 A method for determining remediation of anomalous non-string values within a dataset is illustrated according to one or more embodiments.
[0015] Figure 10 A method for determining corrections for anomalous non-string values within a dataset is illustrated according to one or more other embodiments.
[0016] Figure 11 This illustrates a method for determining corrections for anomalous non-string values within a dataset, according to one or more other embodiments. Detailed Implementation
[0017] The various embodiments described herein are techniques for automatic data standardization for the purpose of data integration in the management service domain of a cloud computing environment. In the context of these various embodiments, data integration encompasses data governance, and further encompasses both data governance and integration. An example data integration solution is... Unified governance and integration platform. A cloud computing environment is a virtualized environment in which one or more computing capabilities can be used as services. A cloud server system configured to implement the automatic data standardization techniques associated with the embodiments described herein can leverage the artificial intelligence capabilities of machine learning knowledge models (specifically, machine learning data standardization models) and information from a knowledge base associated with such models.
[0018] Various embodiments may offer advantages over conventional techniques. Traditional data integration requires data loading, which necessitates manual standardization, such as the correction of outliers or inconsistencies within the dataset by subject matter experts or data administrators. While manual correction is useful in improving data quality, it can be resource-intensive in terms of time and labor. The various embodiments described herein focus on providing automated standardization in the context of data integration, thereby reducing the need for manual correction. Specifically, different embodiments facilitate automated data standardization of datasets by reducing the need for manual data standardization by subject matter experts and / or data administrators, while still allowing for review and approval when necessary. Furthermore, various embodiments improve computing technology by promoting continuous improvement in data integration through machine learning based on the continuous application of data standardization models. Additionally, different embodiments improve computing technology by applying data scraping and / or crowdsourcing techniques to handle and / or investigate outliers in datasets. Some of the different embodiments may not include all of these advantages, and such advantages are not necessarily required in all embodiments.
[0019] In the following text, reference is made to various embodiments of the invention. However, it should be understood that the invention is not limited to the embodiments specifically described. Rather, any combination of the following features and elements (whether or not related to different embodiments) is contemplated to implement and practice the invention. Furthermore, while embodiments may achieve advantages over other possible solutions and / or over the prior art, they are not limiting regardless of whether a given embodiment achieves a particular advantage. Therefore, unless expressly stated in the claims, the following aspects, features, embodiments, and advantages are merely illustrative and should not be considered elements or limitations of the appended claims. Similarly, unless expressly stated in one or more claims, references to “the invention” should not be construed as a generalization of any inventive subject matter disclosed herein and should not be considered elements or limitations of the appended claims.
[0020] This invention can be a system, method, and / or computer program product with any possible level of technical detail integration. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to execute aspects of the invention.
[0021] Computer-readable storage media can be tangible means for retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital universal disk (DVD), memory sticks, floppy disks, mechanical encoding devices such as punched cards or protrusions in slots having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.
[0022] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or external storage device. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.
[0023] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may be personalized to execute computer-readable program instructions by utilizing state information from the computer-readable program instructions in order to perform aspects of this invention.
[0024] The present invention will now be described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0025] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium storing the instructions includes an article of manufacture containing instructions that implement aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0026] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0027] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the figures. For example, two blocks shown consecutively may actually be completed as a single step, executed simultaneously, substantially simultaneously, or with partial or complete temporal overlap, or the blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0028] Specific embodiments are described that relate to techniques for automatic data standardization for data integration purposes in a management services domain. However, it should be understood that the techniques described herein are applicable to various purposes in addition to those specifically described herein. Therefore, references to specific embodiments are included in this document for illustrative purposes and not for limitation.
[0029] The various embodiments described herein can be provided to end users via cloud computing infrastructure. It should be understood that while this disclosure includes a detailed description of cloud computing, implementations of the teachings cited herein are not limited to cloud computing environments. Rather, the various embodiments described herein can be implemented in conjunction with any other type of computing environment now known or developed hereafter.
[0030] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. Therefore, cloud computing allows users to access virtual computing resources (e.g., storage, data, applications, and even complete virtualized computing systems) in the cloud, regardless of the underlying physical systems used to provide those computing resources (or the location of those systems). This cloud model may include at least five features, at least three service models, and at least four deployment models.
[0031] The features are as follows:
[0032] On-demand self-service: Cloud consumers can unilaterally and automatically provide computing power, such as server time and network storage, as needed, without having to interact with the service provider.
[0033] Extensive network access: Capabilities are available through the network and accessed via standard mechanisms that facilitate the use of heterogeneous thin client platforms or thick client platforms (e.g., mobile phones, laptops, and personal digital assistants (PDAs)).
[0034] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned as needed. There is a sense of location independence because consumers typically do not have control or knowledge of the exact location of the resources provided, but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center).
[0035] Rapid flexibility: The ability to provide capacity quickly and flexibly, automatically scaling down and up rapidly in some situations to scale up rapidly. For consumers, the available supply capacity often appears unlimited and can be purchased in any quantity at any time.
[0036] Measuring services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.
[0037] The service model is as follows:
[0038] Software as a Service (SaaS): This provides consumers with the ability to use the provider's applications running on cloud infrastructure. The applications can be accessed from different client devices through a thin client interface such as a web browser (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even individual application capabilities, with possible exceptions such as limited user-specific application configuration settings.
[0039] Platform as a Service (PaaS): This provides consumers with the ability to deploy applications created by the consumer or acquired using programming languages and tools supported by the provider onto cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the deployed applications and the configuration of any application hosting environment.
[0040] Infrastructure as a Service (IaaS): "The capabilities provided to consumers are processing, storage, networking, and other basic computing resources that enable consumers to deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but rather have control over the operating system, storage, deployed applications, and potentially limited control over selected networking components (e.g., host firewalls)."
[0041] The deployment model is as follows:
[0042] Private cloud: A cloud infrastructure designed solely for organizational operations. It can be managed by the organization or a third party and can exist either on-site or off-site.
[0043] Community cloud: A cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on-site or off-site.
[0044] Public cloud: Cloud infrastructure available to the general public or large industry groups and owned by organizations that sell cloud services.
[0045] Hybrid cloud: Cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that enable data and applications to be ported (e.g., cloud bursting for load balancing between clouds).
[0046] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure comprising a network of interconnected nodes.
[0047] Figure 1 A cloud computing environment 50 according to one or more embodiments is illustrated. As shown, the cloud computing environment 50 may include one or more cloud computing nodes 10, with local computing devices used by cloud consumers (e.g., personal digital assistants or cellular phones 54A, desktop computers 54B, laptop computers 54C, and / or automotive computer systems 54N) communicating with the cloud computing nodes 10. The nodes 10 may communicate with each other. They may be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. Therefore, the cloud computing environment 50 can provide infrastructure, platforms, and / or software as services that cloud consumers do not need to maintain resources on their local computing devices. It should be understood that... Figure 1The types of computing devices 54A-N shown are intended to be illustrative only, and computing node 10 and cloud computing environment 50 can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).
[0048] Figure 2 This illustrates a set of functional abstraction layers provided by a cloud computing environment 50 according to one or more implementations. It should be understood in advance that... Figure 2 The components, layers, and functions shown are intended to be illustrative only; the different embodiments described herein are not limited thereto. As described, different layers and corresponding functions are provided. Specifically, hardware and software layer 60 includes hardware and software components. Examples of hardware components may include host 61, RISC (Reduced Instruction Set Computer) based server 62, server 63, blade server 64, storage device 65, and network and network components 66. In some embodiments, software components may include network application server software 67 and database software 68. Virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 71; virtual memory 72; virtual network 73, including virtual private network; virtual application and operating system 74; and virtual client 75.
[0049] In one example, management layer 80 can provide the following functionalities: Resource provisioning 81 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment 50. Metering and pricing 82 provides cost tracking as resources are utilized within the cloud computing environment 50 and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides cloud resource allocation and management to ensure that required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-scheduling and procurement of cloud resources, anticipating future requirements for those resources according to the SLA.
[0050] Workload layer 90 provides examples of functionalities that can be utilized from cloud computing environment 50. Examples of workloads and functionalities that can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom education provision 93; data analytics and processing 94; transaction processing 95; and data standardization 96. Data standardization 96 enables automatic data standardization for the purpose of data integration via machine learning knowledge models according to the different embodiments described herein.
[0051] Figure 3A management service domain 300 within a cloud computing environment 50 is illustrated. Functions related to data standardization 96 and other workloads / functions can be performed within the management service domain 300. The management service domain 300 includes a cloud server system 310. In an embodiment, the cloud server system 310 includes a data loading repository 320, a database management system (DBMS) 330, and a cloud server application 340 including a machine learning knowledge model 350, which incorporates the capabilities of at least a machine learning data standardization model. The cloud server application 340 represents a single application or multiple applications. The machine learning knowledge model 350 is configured to implement and / or facilitate automatic data standardization according to the various embodiments described herein. Furthermore, the management service domain 300 includes a data loading interface 360, one or more external database systems 370, and multiple application server clusters 3801 to 380. n The data loading interface 360 enables communication between the cloud server system 310 and the data client / system, which interacts with the management service domain 300 to facilitate data loading in a data integration context. In one embodiment, the cloud server system 310 is configured to interact with one or more external database systems 370 and multiple application server clusters 3801 to 3802. n Communication. Additionally, application server clusters 3801 to 380... n The corresponding servers within can be configured to communicate with each other and / or with server clusters in other domains.
[0052] In one embodiment, the loading data repository 320, representing a single data repository or a collection of data repositories, includes unnormalized and / or pre-normalized datasets received from the data loading interface 360 during the data loading process, as well as categorized and normalized datasets processed according to the various embodiments described herein. Alternatively, one or more data repositories include pre-normalized and / or unnormalized datasets received during the data loading process, while one or more additional data repositories include categorized and normalized datasets. According to another alternative embodiment, one or more data repositories include pre-normalized and / or unnormalized datasets received during the data loading process, one or more additional data repositories include categorized datasets, and one or more other additional data repositories include normalized datasets. In an embodiment, the DBMS 330 coordinates and manages the knowledge base of the machine learning knowledge model 350. The DBMS 330 may include one or more database servers that can coordinate and / or manage various aspects of the knowledge base. In another embodiment, the DBMS 330 manages or otherwise interacts with one or more external database systems 370. One or more external database systems 370 include one or more relational databases and / or one or more database management systems configured to interface with DBMS 330. In a further embodiment, DBMS 330 stores multiple application server clusters 3801 to 380. n The relationship with the knowledge base. In a further embodiment, DBMS 330 includes or is operatively coupled to one or more databases, some or all of which may be relational databases. In a further embodiment, DBMS 330 includes one or more ontology trees or other ontology structures. Application server clusters 3801 to 380 n Hosting and / or storing various aspects of different applications, and also providing management server services to one or more client systems and / or data systems.
[0053] Figure 4A method 400 for creating data standardization rules to facilitate data integration is illustrated. The creation of data standardization rules according to method 400 facilitates automatic data standardization within the context of the various embodiments described herein. In one embodiment, one or more steps associated with method 400 are performed in an environment where computing power is provided as a service (e.g., cloud computing environment 50). According to such an embodiment, one or more steps associated with method 400 are performed in a management service domain (e.g., management service domain 300) within that environment. This environment may be a hybrid cloud environment. In another embodiment, one or more steps associated with method 400 are performed in one or more other environments (such as a client-server network environment or a peer-to-peer network environment). A centralized cloud server system in the management service domain (e.g., cloud server system 310 in management service domain 300) can facilitate processing according to method 400 and other methods further described herein. More specifically, a cloud server application in the cloud server system (e.g., cloud server application 340) can perform or otherwise facilitate one or more steps of method 400 and other methods described herein. Automatic data standardization techniques facilitated or otherwise implemented by cloud server systems in the management service domain can be associated with data standardization within workload layers in functional abstraction layers provided by the environment (e.g., data standardization 96 within workload layer 90 of cloud computing infrastructure 50).
[0054] Method 400 begins at step 405, where a cloud server application receives a dataset during the data loading process. The cloud server application receives the dataset via a data loading interface in the management service domain (e.g., via data loading interface 360) (according to step 405). The loaded data is stored wholly or partially in at least one loaded data repository (e.g., loaded data repository 320), which may or may otherwise be associated with a data lake. In the context of the various embodiments described herein, a data lake is a storage repository that holds data in its natural or raw format from a number of sources. The data in a data lake may be structured, semi-structured, or unstructured. In embodiments, the data loading process includes loading a dataset into a data lake. According to such embodiments, the dataset may optionally be loaded in Hadoop Distributed File System (HDFS) format. Alternatively, according to such embodiments, the dataset is bound as normalized data, for example, by manually or automatically eliminating data redundancy. In further embodiments, during the data loading process, only metadata associated with the data located in the data lake is loaded, rather than the data itself. According to this further embodiment, based on the execution of the steps of method 400, the cloud server application may optionally create a standardized version of the data located in the data lake based on loaded metadata. Alternatively, according to this further embodiment, based on the execution of the steps of method 400, the cloud server application may directly apply data standardization to the data located in the data lake based on loaded metadata. In summary, in the context of method 400, the cloud server application may optionally load a dataset into the data lake, or alternatively load metadata associated with a dataset already in the data lake. In a further embodiment, the data loading process includes bringing the dataset online from an online transaction processing (OLTP) system to an online analytical processing (OLAP) system.
[0055] In step 410, the cloud server application classifies the data points within the dataset received in step 405. In one embodiment, the cloud server application detects data categories by applying at least one data classifier algorithm, and in response to detecting data categories, classifies the data points by labeling (e.g., tagging) associated data fields with terms or categories (e.g., business-related terms or categories). Examples of corresponding data categories include municipalities and postal codes. According to such an embodiment, at least one data classifier algorithm incorporates regular expressions to identify data within data fields or based on data field names. Additionally or alternatively, at least one data classifier algorithm includes one or more lookup operations to reference tables containing values that make up the data class. Additionally or alternatively, at least one data classifier algorithm includes custom logic written in a programming language (e.g., Java). In a further embodiment, the cloud server application classifies data points based on the evaluation of individual values. Additionally or alternatively, the cloud server application classifies data points based on the evaluation of multiple values within a data column (e.g., in the context of structured or semi-structured data). Additionally or alternatively, the cloud server application classifies data points based on the evaluation of multiple values within multiple related data columns. In another embodiment, the cloud server application associates confidence level values with data point classifications made via at least one data classifier algorithm. According to such a further embodiment, the cloud server application prioritizes and / or labels data point classifications based on confidence level classifications made via at least one data classifier. Optionally, a client, such as a subject matter expert or data manager (i.e., a data editor or data engineer), overrides one or more data point classifications made via at least one data classifier algorithm and manually assigns one or more data point classifications.
[0056] In one embodiment, within the context of step 410, the cloud server application randomly samples a subset of categorized data points within the dataset. According to this embodiment, the cloud server application determines the value of the randomly sampled data points and then determines a specific data class and / or data column associated with each randomly sampled data point. According to this embodiment, the cloud server application determines the data class and / or data column associated with other data points within the subset based on the association with the randomly sampled data points. Thus, random sampling according to this embodiment allows the cloud server application to avoid iterating through the entire dataset to classify the data points. Additionally or alternatively, the cloud server application classifies one or more data points in the dataset directly by assigning data categories. Therefore, the cloud server application can classify each data point within the dataset based on random sampling and / or based on direct manual classification.
[0057] In step 415, the cloud server application applies a machine learning data normalization model (e.g., machine learning knowledge model 350) to the data points for each category within the dataset. In an embodiment, the cloud server application applies the data normalization model based on the data category. According to such embodiments, the type of detected anomaly and the corresponding normalization relative to such anomaly type are determined at least in part based on the data category type. For example, data points with a telephone number data category only require validation regarding format (i.e., only format anomalies are relevant), while data points with a city data category require validation regarding at least spelling and possible grammar (i.e., spelling and grammar anomalies are potentially relevant). As further described herein, the data category is also relevant relative to the distinction between string values and non-string values in the dataset.
[0058] In an embodiment, the cloud server application applies multiple data quality rules within the context of a data normalization model to determine whether data points within the dataset are valid or outliers, and further determines how to handle any values determined to be outliers. The multiple data quality rules include multiple baseline data integration rules, which are normalization baselines used to determine whether data points are valid or outliers. In the context of the various embodiments described herein, the multiple baseline data integration rules are hard-coded or fixed rules applicable to data categories or groups of data categories and are integrated into the data normalization model during model initialization. Optionally, the multiple baseline data integration rules are customized based on the dataset size and / or dataset source. The multiple baseline data integration rules include pre-established rules for identifying outliers associated with data points within the dataset. In a further embodiment, the multiple baseline data integration rules are specific to a particular data class or group of data classes. According to such a further embodiment, the cloud server application applies multiple baseline data integration rules based on data categories. For example, according to a baseline integration rule regarding birthday data category values, such values cannot be future dates. According to such a further embodiment, the cloud server application may optionally analyze null or duplicate values based on the data category associated with such null or duplicate values. In various embodiments, valid values are values that conform to multiple baseline data integration rules, while outliers are values that do not conform to multiple baseline data integration rules. The cloud server application may optionally determine that a value does not conform to multiple baseline data integration rules, and therefore is an outlier if the value does not satisfy one or more conditions established for the associated data category and / or for the associated data column (i.e., if the value includes at least one outlier). Additionally or alternatively, the cloud server application determines that a value does not conform to multiple baseline data integration rules and is therefore an outlier if the value is inconsistent with other values in a defined group (e.g., a data column or a group of associated data columns) within the dataset. In some instances, such inconsistencies are apparent in scenarios where a first data column is associated with another data column or other data columns, such that outliers in the first data column are inconsistent with corresponding values in the other data column or other data columns associated with the first data column. In the context of different embodiments, the corresponding column is associated if the value in one column is predictable based on the values in other columns. In other instances, such inconsistencies are evident where data values within a column are correlated based on its column position, leading to discrepancies between outliers within the column and their respective column positions. In other instances, such inconsistencies are manifested through outliers that are inconsistent both with respect to their corresponding values in the relevant columns and with respect to their column positions.
[0059] In addition to multiple baseline data integration rules, the multiple data quality rules include any set of previously obtained and accepted data standardization rules for handling values identified as outliers due to the application of the multiple baseline data integration rules. The multiple data quality rules are fully bound together to avoid any uncertain range. Furthermore, the multiple data quality rules include variables bound to data columns in the dataset. In an embodiment, the cloud server application calculates the number of outliers identified within the dataset as a result of the application of the data standardization model. According to such an embodiment, the cloud server application calculates a quantitative data quality score for the dataset based on compliance with the multiple baseline data integration rules. The cloud server application may optionally calculate the data quality score based on the percentage of non-compliant / outliers among all categorized data points within the dataset. By applying the data standardization model according to step 415, the cloud server application can apply previously obtained and accepted standardization rules to handle outliers, and furthermore, as described below, the cloud server application can determine additional standardization modifications, based on which a new set of proposed standardization rules can be obtained. (Refer to...) Figure 6 A method is described for applying the data normalization model to each data point within the dataset according to step 415.
[0060] In step 420, the cloud server application obtains a proposed set of data normalization rules for the dataset based on any normalization modifications determined as the machine learning data normalization model is applied. The proposed set of data normalization rules includes one or more rules for automatically facilitating the normalization of the dynamic dataset during subsequent model application. In an embodiment, the proposed set of data normalization rules includes multiple mappings, which may optionally be stored in a normalization lookup table. According to such an embodiment, the multiple mappings include corresponding mappings or other relational aspects from outliers to valid values. Specifically, the multiple mappings may include mappings between outlier string values containing one or more misspelled letters and valid string values with correct spelling; for example, the outlier string value "Dehli" may be mapped to the valid string value "Delhi". Embodiments regarding the determination of normalization modifications are further described herein.
[0061] Optionally, in step 425, the cloud server application presents the proposed set of data standardization rules obtained in step 420 for client review. In one embodiment, the cloud server application presents the set by publishing the proposed set of data standardization rules in a forum accessible to the client. In a further embodiment, the cloud server application presents the set by transmitting the proposed set of data standardization rules to a client interface. According to this further implementation, the client interface is a user interface in the form of a graphical user interface (GUI) and / or a command-line interface (CLI) presented by at least one client application installed on or accessible to the client system or device. Optionally, the cloud server application facilitates client acceptance of the proposed set of data standardization rules by generating form-response interface elements for display to the client (e.g., in a forum or client interface accessible to the client). Clients include subject matter experts in a specific domain, data administrators, and / or any other entity associated with the dataset or cloud server system.
[0062] In one embodiment, the cloud server application automatically accepts the proposed set of data standardization rules based on a quantitative confidence level value belonging to the proposed set of data standardization rules and / or the corresponding standardization rules therein, without client review. According to this embodiment, in response to determining that the proposed set of data standardization rules corrects data inaccuracies at a confidence level exceeding a predetermined confidence level threshold, the cloud server application automatically accepts the proposed set of data standardization rules and thus abandons client review. The predetermined confidence level threshold quantitatively measures the confidence in the automatic data correction capability (i.e., the correction capability without subject matter experts / data administrators). According to this embodiment, the quantitative confidence level value belonging to the proposed set of data standardization rules and / or the corresponding standardization rules therein can optionally be set by the client based on relevant usage. Alternatively or additionally, the confidence level value is determined based on data relationships identified via machine learning analytics in the context of previous model applications and / or previous model iterations (i.e., previous iterations within the current model application). For example, suppose the state capital data class value is "Memphis," the corresponding state data class value will always be "Tennessee," and therefore the confidence level value for one or more applicable data normalization rules will be relatively high. In another example, suppose the state capital data class value is "Hyderabad," depending on the context, the corresponding state data class value could be "Telangana" or "Andhra Pradesh" (because Hyderabad is the capital of both states), and therefore the confidence level value for one or more applicable data normalization rules will be relatively low. Therefore, assuming the predetermined confidence level thresholds are the same in the contexts of these examples, if one or more rules within the proposed set resolve state data class values within the dataset corresponding to "Memphis" (as opposed to "Hyderabad"), then in the context of determining the normalized state data class value, it is relatively more likely that the cloud server application will accept the proposed set of data normalization rules without client review.
[0063] In step 430, the cloud server application determines whether the proposed set of data standardization rules has been accepted. Optionally, acceptance of the proposed set of data standardization rules is determined based on client review, or alternatively, acceptance is automatic in response to determining that the proposed set of data standardization rules remedies data inaccuracies at a confidence level exceeding a predetermined confidence level threshold. In response to determining that the proposed set of data standardization rules has not yet been accepted, the cloud server application optionally adjusts one or more of the proposed set of data standardization rules and then repeats step 430. Optionally, such rule adjustment includes collecting client feedback on the proposed set of data standardization rules, in which case the cloud server application may apply one or more modifications based on the client feedback. In an embodiment, such client feedback includes one or more rule coverage requests with rule modifications imposed by one or more clients. The cloud server application may optionally adjust the proposed set of data standardization rules based on client feedback.
[0064] In response to determining that the proposed set of data standardization rules has been accepted, in step 435, the cloud server application applies the proposed set of data standardization rules to the dataset. In one embodiment, the cloud server application applies the proposed set of data standardization rules to each data point in the dataset. In a further embodiment, the cloud server application confirms that the dataset data quality has improved with the application of the proposed set of data standardization rules. According to such a further embodiment, the cloud server application confirms that the dataset data quality has improved by conducting a client review of the dataset data quality before and after applying the proposed set of data standardization rules. Additionally or alternatively, the cloud server application confirms that the dataset data quality has improved by comparing the dataset data quality before and after applying the proposed set of data standardization rules.
[0065] Furthermore, in response to the acceptance of the proposed set of data standardization rules, in step 440, the cloud server application updates the machine learning data standardization model based on the proposed set of data standardization rules. In an embodiment, the cloud server application integrates (e.g., adds or otherwise associates) the proposed set of data standardization rules into multiple data quality rules associated with the data standardization model, such that the cloud server application considers the proposed set of data standardization rules along with multiple baseline data integration rules and any previously accepted set of data standardization rules during subsequent model applications. According to such an embodiment, the cloud server application adds the proposed set of data standardization rules to a knowledge base associated with the data standardization model. The knowledge base associated with the model optionally includes all repositories, ontologies, files, and / or documents associated with the multiple data quality rules and any proposed set of data standardization rules. The cloud server application optionally accesses the knowledge base via a DBMS (e.g., DBMS 330) associated with the cloud server system.
[0066] Based on the integration of the proposed set of data standardization rules into multiple data quality rules, the cloud server application optionally trains a data standardization model to identify patterns of data points that will allow for automatic standardization during subsequent model applications. Upon acceptance, the cloud server application facilitates adaptation to the proposed set of data standardization rules to ensure that outliers previously unexpected but automatically handled by the proposed set are standardized when multiple data quality rules associated with the machine learning data standardization model are subsequently applied. Thus, based on the integration of the proposed set of data standardization rules, during subsequent model applications, if the standardization of a dataset data point identified as an outlier is handled by one or more of the proposed set of data standardization rules, the cloud server application automatically standardizes that data point. When updating the data standardization model according to step 440, the cloud server application optionally returns to step 415 to reapply the data standardization model to the dataset to refine the dataset standardization. Such a further embodiment is consistent with the application of multi-pass algorithms as envisioned in the context of different embodiments. Alternatively, when updating the data standardization model according to step 440, the cloud server application may proceed to the end of method 400.
[0067] Figure 5A method 500 for configuring a machine learning data normalization model is illustrated. Method 500 begins at step 505, where a cloud server application samples multiple datasets during multiple corresponding data loading scenarios. In one embodiment, the cloud server application samples the datasets based on one or more associated data classes. According to an alternative embodiment, to focus on one or more data classes, the cloud server application samples datasets with a number of values greater than a threshold number for classification according to one or more associated data classes. According to another alternative embodiment, to ensure data diversity, the cloud server application samples datasets with values associated with a number of data classes greater than a threshold number. In step 510, the cloud server application identifies corresponding data anomaly scenarios based on the multiple sampled datasets. In an embodiment, the cloud server application identifies outliers in the multiple sampled datasets by applying multiple data quality rules associated with the data normalization model. Optionally, multiple baseline data integration rules among the multiple data quality rules include at least one correlation rule that identifies expected relationships between or among corresponding values in related data columns. For example, a correlation rule might require that the value of a row in a first data column is twice the value of a row in a second data column. Additionally or alternatively, the multiple baseline data integration rules include at least one data column correlation rule that identifies the expected relationship between or among relevant values within a data column. Optionally, the multiple data quality rules also include any adaptable rules obtained from external sources such as ontology or repositories. For example, a cloud server application may obtain a list of valid values for one or more data classes from a data repository and obtain rules for interpreting the valid values included in that list. After the initial construction of the data normalization model, the cloud server application may optionally update multiple data quality rules beyond the multiple baseline data integration rules based on the acceptance of the corresponding proposed set of data normalization rules (e.g., in the context of method 400) and / or based on the adoption of additional rules from external sources.
[0068] In step 515, the cloud server application identifies existing data anomaly correction techniques to handle the corresponding data anomaly scenario. In one embodiment, the cloud server application identifies existing data anomaly correction techniques by consulting any previously accepted set of data normalization rules and / or any previously proposed set of data normalization rules. Such a set of rules may be stored in a knowledge base associated with the data normalization model, or may otherwise be recorded in relation to the model, such as being retrieved from external sources. According to such an embodiment, the cloud server application identifies any previously accepted and / or previously proposed set of data normalization rules for datasets in multiple sampled datasets. In step 520, the cloud server application trains a machine learning data normalization model based on the applicability of existing data anomaly correction techniques to the corresponding data anomaly scenario. According to step 520, the cloud server application trains the model by creating and analyzing association data that correlates the handling of the corresponding data anomaly scenario with existing data anomaly correction techniques. The association data may contain confidence level data regarding the effectiveness of one or more of the existing anomaly correction techniques in addressing one or more of the corresponding data anomaly scenarios. Alternatively, the associated data may include comparisons of multiple existing anomaly correction techniques in handling a given data anomaly scenario. By analyzing the associated data, the cloud server application can determine which of the existing anomaly correction techniques will be relatively more effective in handling the corresponding data anomaly scenario or similar scenarios during future model applications. Therefore, the cloud server application can calibrate the model based on such analysis. In an embodiment, the cloud server application records the associated data about the model in, for example, a knowledge base associated with the model.
[0069] In one embodiment, the cloud server application trains a data normalization model based on analysis of any previously accepted set of data normalization rules and / or any previously proposed set of data normalization rules consulted, in order to enhance or otherwise adapt multiple data quality rules beyond multiple baseline data integration rules. Optionally, after training at step 520, the cloud server application integrates rules from the previously accepted set of data normalization rules into multiple data quality rules. Such training can enhance the data normalization model as a complement or alternative to updating the data normalization model, according to step 440 in the context of method 400, as the proposed set of data normalization rules is accepted. In another embodiment, the cloud server application initially performs one or more model configuration steps, particularly external source negotiation, to build the data normalization model. In a further embodiment, the cloud server application performs one or more model configuration steps, particularly training, at periodic time intervals and / or each time the cloud server application processes a threshold number of datasets. The cloud server application can continue until the end of method 500 after training the machine learning data normalization model.
[0070] In summary, configuring a machine learning data normalization model according to Method 500 includes: sampling multiple datasets during multiple corresponding data loading scenarios; identifying corresponding data anomaly scenarios based on the multiple sampled datasets; identifying existing data anomaly correction techniques to address the corresponding data anomaly scenarios; and training the machine learning data normalization model based on the applicability of existing data anomaly correction techniques to the corresponding data anomaly scenarios.
[0071] Figure 6 Method 600, which applies a machine learning data normalization model to data points for each category within a dataset, is illustrated in the context of step 415 of method 400. Method 600 begins at step 605, where the cloud server application selects data points within the dataset for evaluation. At step 610, the cloud server application determines whether the data point selected in step 605 is a valid value. In an embodiment, the cloud server application determines whether a data point is a valid value by determining whether it conforms to multiple baseline data integration rules. In response to determining that the data point selected in step 605 is a valid value, no further normalization processing is required for that data point, and the cloud server application therefore proceeds to step 630. In response to determining that the data point is not a valid value, i.e., in response to determining that the data point does not conform to multiple baseline data integration rules, at step 615, the cloud server application determines whether the data point is an outlier processed by pre-existing normalization rules incorporated into the data normalization model (e.g., rules from a set of previously accepted data normalization rules integrated into multiple data quality rules, and / or rules adapted from external sources as the model is trained). In response to the determination that the data point is an outlier handled by pre-existing standardization rules incorporated into the data standardization model, in step 620, the cloud server application dynamically modifies the outlier by applying the pre-existing standardization rules. By dynamically modifying outliers based on the application of pre-existing standardization rules, the cloud server application can leverage the model to automatically resolve standardization issues without further analysis and processing. After step 620, the cloud server application proceeds to step 630.
[0072] In response to determining that a data point is an outlier that has not been processed by pre-existing standardization rules incorporated into the data standardization model, in step 625, the cloud server application determines at least one standardization modification to correct the data point. (Refer to...) Figure 7 A method is described for determining at least one standardized modification to correct the data point according to step 625. After step 625, the cloud server application proceeds to step 630. In step 630, the cloud server application determines whether there is another data point in the dataset to be evaluated. In response to determining that there is another data point in the dataset to be evaluated, the cloud server application returns to step 605. In response to determining that there are no other data points in the dataset to be evaluated, the cloud server application may proceed to the end of method 600.
[0073] In summary, applying the machine learning data normalization model to data points of each category within the dataset according to method 600 includes: dynamically modifying the outlier by applying the pre-existing normalization rules in response to determining that a data point is an outlier processed by pre-existing normalization rules incorporated in the machine learning data normalization model. Furthermore, applying the machine learning data normalization model to data points of each category within the dataset according to method 600 includes: determining at least one normalization modification to correct the data point in response to determining that a data point is an outlier not processed by any pre-existing normalization rules.
[0074] Figure 7 A method 700 is illustrated in the context of step 625 of method 600 to determine at least one standardized modification to correct data points within a dataset. Method 700 begins at step 705, where a cloud server application determines whether a data point is null. In response to determining that a data point is not null, the cloud server application proceeds to step 715. In response to determining that a data point is null, in step 710, the cloud server application determines a replacement for the null value based on the relationship between a first data column containing the null value and at least one related data column (i.e., at least one data column associated with the first data column) or based on the relationship between related values within the first data column. In an embodiment, if the null value is in a first data column associated with one or more other data columns, the cloud server application may optionally replace the null value with a value consistent with the correlation of the data column. According to such an embodiment, the cloud server application may optionally analyze corresponding values within one or more other data columns (e.g., corresponding values in data columns adjacent to or otherwise associated with the first data column containing the null value) to determine the correlation between the first data column containing the null value and (one or more) other data columns, and replace the null value with a valid value identified based on that correlation. Additionally or alternatively, the cloud server application may optionally analyze corresponding values within one or more other data column rows (e.g., corresponding values in rows adjacent to or otherwise associated with the row containing the null value) to determine the correlation between the first data column containing the null value and (one or more) other data columns, and replace the null value with a valid value identified based on that correlation. The cloud server application may optionally apply at least one pattern matching algorithm and / or at least one clustering algorithm to determine any such correlation(s). In another embodiment, if the null value is located in a data column with correlated values based on data column position (e.g., data column row), the cloud server application may optionally replace the null value with a value consistent with the data column position of the null value.
[0075] In one embodiment, according to step 710, the cloud server application determines the replacement of null values based on relevant and / or cross-correlation value information obtained from multiple baseline data integration rules, relevant and / or cross-correlation value information obtained from adaptable rules and relational data obtained from a repository or ontology, and / or relevant and / or cross-correlation value information obtained through machine learning via the application of a data normalization model. In a further embodiment, the cloud server application records, marks, or otherwise records the replacement of null values as a normalization modification in the context of the application model. According to such a further embodiment, the cloud server application obtains at least one rule in the proposed set of data normalization rules based on the normalization modification. In the context of the data normalization model, more generally, any such proposed data normalization rule may optionally supplement or otherwise modify the data column relevant and / or cross-correlation rules within multiple baseline data integration rules or multiple data quality rules. In response to performing step 710, the cloud server application may proceed to the end of method 700.
[0076] In step 715, the cloud server application determines whether a data point is a frequent anomalous string value. The cloud server application identifies frequent anomalous string values as those with an occurrence rate greater than or equal to a predetermined dataset frequency threshold. In the context of different embodiments, the cloud server application may optionally measure the occurrence rate of anomalous string values based on their occurrence within a dataset processed according to the method described herein, or alternatively, it may measure the occurrence rate based on their occurrence within a specified plurality of datasets, which may or may not include datasets within the processed dataset and / or sampled datasets. In various embodiments, string values are data types used to represent text. Such string values may include sequences of letters, numbers, and / or symbols. In embodiments, such string values may be or otherwise include variable character fields (varchar), which in the context of different embodiments are sets of characters of indeterminate length. A data category may be defined to include string values or may be otherwise associated with string values, for example, as opposed to non-string values. For example, a postal code data category value may be defined as a string value comprising a specific number of characters (numbers and / or letters). Therefore, in some embodiments, string values and non-string values are distinguishable based on data category type and / or data column type. In response to determining that a data point is not a frequent anomalous string value, the cloud server application proceeds to step 725. In response to determining that a data point is a frequent anomalous string value, in step 720, the cloud server application classifies the frequent anomalous string value as a valid value. In embodiments, the cloud server application stores, marks, or otherwise marks the frequent anomalous string value as a valid value in the context of a data normalization model. According to such an embodiment, the cloud server application may optionally store the frequent anomalous string value in a knowledge base associated with the model. In a further embodiment, the cloud server application records, marks, or otherwise records the validation of the frequent anomalous string value as a normalization modification in the context of applying the model. According to such a further embodiment, the cloud server application derives at least one rule from the proposed set of data normalization rules based on the normalization modification. By validating the frequent anomalous string value, the cloud server application facilitates automatic normalization of the frequent anomalous string value during subsequent model application. In response to performing step 720, the cloud server application may proceed to the end of method 700.
[0077] In step 725, the cloud server application determines whether the data point is an infrequent anomalous string value. The cloud server application identifies infrequent anomalous string values as those with an occurrence rate less than a predetermined dataset frequency threshold. In response to determining that the data point is not an infrequent anomalous string value, the cloud server application proceeds to step 735. In response to determining that the data point is an infrequent anomalous string value, in step 730, the cloud server application applies a data crawling algorithm to evaluate infrequent anomalous string values via an automated application. In related embodiments, the input to the data crawling algorithm includes infrequent anomalous string values and associated data class names. In further related embodiments, the data crawling algorithm is a web crawling algorithm designed to systematically access web pages. In the context of different embodiments, the automated application may be or otherwise combined with a robot. Additionally or alternatively, in the context of step 730, the cloud server application applies a data crawling algorithm to retrieve data from sources other than web pages. In response to performing step 730, the cloud server application may proceed to the end of method 700. About Figure 8 The method described in step 730 for applying a data scraping algorithm to evaluate infrequent anomalous string values is described.
[0078] In step 735, the cloud server application determines whether a data point is an anomalous non-string value. In various embodiments, a non-string value represents a data type that combines one or more aspects beyond text, such as a regular expression. A regular expression comprises a sequence of characters defining a search pattern. While string values are typically defined by a finite list of values, non-string values are evaluated based on pattern analysis (e.g., regular expression analysis). In some embodiments, non-string values are distinguishable from string values based on data category type and / or data column type. A data category can be defined to include non-string values or can be associated with non-string values, for example, the opposite of string values.
[0079] In one embodiment, a cloud server application determines that a non-string value is an outlier, and more specifically, identifies one or more outliers based on the non-string value being an outlier, according to regular expression analysis (i.e., search pattern analysis). The cloud server application identifies one or more outliers as format violations during regular expression analysis. For example, the cloud server application may determine that a non-string Social Security Number data class value is an outlier based on a missing hyphen, because regular expression validation of non-string values in the Social Security Number data class requires a defined hyphen pattern. In another example, the cloud server application may determine that a non-string Email Address data class value is an outlier based on a missing period, because a period within a domain name is required for regular expression validation of non-string values in the Email Address data class. According to such embodiments, the cloud server application identifies format violations associated with outlier non-string values based at least in part on the application of multiple baseline data integration rules included within multiple data quality rules. In response to the application of the multiple baseline data integration rules, one or more alerts associated with (one or more) non-string value format violations may optionally be automatically triggered. In another embodiment, the cloud server application determines that a data point is an outlier non-string value based on one or more outliers in one or more string portions within the data point. In another embodiment, the cloud server application determines that a data point is an anomalous non-string value based on the fact that the data point is a duplicate value within a data column associated with a unique value. For example, the cloud server application can identify a data point with a non-string Social Security Number data class value as a non-string anomalous value, which is a duplicate value of at least one other data point within the Social Security Number data column.
[0080] In response to determining that a data point is an anomalous non-string value, in step 740, the cloud server application determines a correction for the anomalous non-string value. In response to executing step 740, the cloud server application may continue to the end of method 700. About Figure 9-11 Various methods for correcting the determination of anomalous non-string values according to step 740 are described. In response to determining that a data point is not an anomalous non-string value, optionally at step 745, the cloud server application investigates the identification of the data point via crowdsourcing and / or external repository negotiation, and may then proceed to the end of method 700. In embodiments, the cloud server application facilitates crowdsourcing by consulting public forums of data experts or subject matter experts. The inability to identify a data point before step 745 may indicate a data quality anomaly, for example, caused by a poorly defined value. According to one or more alternative embodiments, the cloud server application performs the steps of method 700 in an alternative configuration. For example, the cloud server application may perform steps 705-710, 715-720, 725-730, and 735-740 in one or more alternative sequences.
[0081] In summary, determining at least one standardized modification for correcting a data point according to method 700 includes: in response to determining that a data point is a null value, determining a replacement for the null value based on the relationship between a first data column including null values and at least one related data column, or based on the relationship between related values within the first data column. Furthermore, determining at least one standardized modification for correcting a data point according to method 700 includes: in response to determining that a data point is a frequent anomalous string value with an occurrence rate greater than or equal to a predetermined dataset frequency threshold, classifying the frequent anomalous string value as a valid value. Furthermore, determining at least one standardized modification for correcting a data point according to method 700 includes: in response to determining that a data point is an infrequent anomalous string value with an occurrence rate less than a predetermined dataset frequency threshold, applying a data scraping algorithm to evaluate the infrequent anomalous string value via automated application, wherein the input to the data scraping algorithm includes the infrequent anomalous string value and an associated data class name. Additionally, determining at least one standardized modification for correcting a data point according to method 700 includes: in response to determining that a data point is an anomalous non-string value, determining a correction for the anomalous non-string value.
[0082] Figure 8 A method 800 is illustrated that applies a data scraping algorithm to evaluate infrequent anomalous string values in the context of step 730 of method 700. Method 800 begins at step 805, where the cloud server application determines whether a threshold number of data points associated with both infrequent anomalous string values and associated data class names has been identified via data scraping. In response to failure to identify a threshold number of data points associated with both infrequent anomalous string values and associated data class names, the cloud server application proceeds to step 815. In response to identifying a threshold number of data points associated with both infrequent anomalous string values and associated data class names, in step 810, the cloud server application classifies the infrequent anomalous string values as valid values. In an embodiment, the cloud server application stores, tags, or otherwise marks infrequent anomalous string values as valid values in the context of a data normalization model. According to such an embodiment, the cloud server application may optionally store infrequent anomalous string values as valid values in a knowledge base associated with the model. In a further embodiment, within the context of the application model, the cloud server application records, marks, or otherwise documents the verification of infrequent anomalous string values as standardized modifications. According to such a further embodiment, the cloud server application derives at least one rule from the proposed set of data standardization rules based on these standardized modifications. By verifying infrequent anomalous string values, the cloud server application facilitates automatic standardization relative to these infrequent anomalous string values during subsequent model application. In response to performing step 810, the cloud server application may proceed to the end of method 800.
[0083] In step 815, the cloud server application determines whether at least one valid string value exists within the dataset, which has a predetermined degree of string similarity to an infrequent anomalous string value. In this embodiment, the cloud server application defines the predetermined degree of string similarity between infrequent anomalous string values and valid string values based on infrequent anomalous string values and valid string values with the largest number of character differences. Alternatively or additionally, the cloud server application defines the predetermined degree of string similarity between infrequent anomalous string values and valid string values based on infrequent anomalous string values and valid string values with the smallest number of common characters. Alternatively or additionally, the cloud server application defines the predetermined degree of string similarity between infrequent anomalous string values and valid string values based on infrequent anomalous string values and valid string values with corresponding string lengths differing by less than a predetermined threshold. In response to failure to identify at least one valid string value within the dataset having a predetermined degree of string similarity to an infrequent anomalous string, the cloud server application proceeds to step 825. In response to identifying at least one valid string value within the dataset that has a predetermined degree of string similarity to an infrequent anomalous string value, at step 820, the cloud server application determines a correction for the infrequent anomalous string value based on the selection of valid string values. In one embodiment, the selection of valid string values includes selecting a valid string value from at least one valid string value, and determining at least one modification to the infrequent anomalous string value based on any identified character differences between the infrequent anomalous string value and the selected valid string value. In another embodiment, within the context of the application model, the cloud server application records, marks, or otherwise records at least one modification to the infrequent anomalous string value as a standardized modification. According to such an embodiment, the cloud server application derives at least one rule from a set of proposed data standardization rules based on the standardized modification.
[0084] Selecting a valid string value from at least one valid string value according to step 820 may optionally include evaluating a quantitative similarity between an infrequent anomalous string value and each of the at least one valid string value. In an embodiment, the cloud server application evaluates quantitative similarity by calculating a quantitative similarity score for each of the at least one valid string value based on the evaluated similarity between the valid string value and the infrequent anomalous string value. According to such an embodiment, the cloud server application selects a valid string value from at least one valid string value based on the highest calculated quantitative similarity score. The cloud server application may optionally calculate the quantitative similarity score based on one or more similarity evaluation factors. The similarity evaluation factors may optionally include the edit distance between the infrequent anomalous string value and each of the at least one valid string value. In the context of different embodiments, the edit distance is the amount of modification required to correct an anomalous value to conform to a particular valid value. Additionally or alternatively, the similarity evaluation factors may include a data source of infrequent anomalous string values compared to each of the at least one valid string value. In such a context, the data source may optionally include the author's starting location and / or identity. Additionally or alternatively, the similarity assessment factor includes a data classification of infrequently occurring outlier string values compared to each of at least one valid string value. In this context, data classification may optionally include data category, data type, subject, (one or more) expected statistics, data age, data modification history, and / or values of associated data fields.
[0085] In relevant embodiments, one or more similarity assessment factors are weighted such that similarity assessment factors with relatively higher weights have a greater impact on the calculated quantitative similarity score than similarity assessment factors with relatively lower weights. To distinguish between two valid values with the same quantitative similarity score (e.g., two city values from a common source and within the same data class with a single character difference), the cloud server application may optionally select one of the two valid values based on the evaluation of the corresponding associated field values in the relevant data class for each of the two valid values (e.g., the postal code data class value for each of the two city data class values). The cloud server application may select one of the two valid values based on which of the two valid values has associated field values that satisfy or more fully comply with one or more conditions established for the associated data category and / or for the associated data column as determined by the set of data standardization rules associated with the data standardization model (e.g., a previously accepted set of data standardization rules). For example, given an infrequent anomalous string city data category value “Middletonw” and valid city data category values “Middleton” and “Middletown”, a cloud server application can select one of two valid city data category values based on which valid value has a postal code that satisfies or more fully conforms to one or more conditions established for the associated postal code data category and / or for the associated postal code data column, as determined by the set of data standardization rules associated with the model.
[0086] In a further related embodiment, the cloud server application evaluates the quantitative similarity between infrequent anomalous string values and each of at least one valid string value by applying an algorithm based on a weighted decision tree. The cloud server application may optionally determine the corresponding weight value of the weighted decision tree based on the corresponding edit distance between the infrequent anomalous string value and each of the at least one valid string value. Additionally or alternatively, the cloud server application may optionally determine the corresponding weight value of the weighted decision tree based on the corresponding value in a data column that is related to (e.g., relevant or otherwise associated with) the data column of infrequent anomalous string values.
[0087] Additionally or alternatively, selecting a valid string value among at least one valid string value according to step 820 may optionally include applying one or more heuristics based at least in part on a string value selection history. In the context of this embodiment and other embodiments described herein, the heuristic is a rule based on machine logic that provides a determination based on one or more predetermined types of input. The predetermined types of input to the one or more heuristics in the context of selecting a valid string value among at least one valid string value may include a valid string selection history in the context of data classes and / or data columns associated with infrequent anomalous string values, such as the most recently selected valid string or the most frequently selected valid string. In response to performing step 820, the cloud server application may proceed to the end of method 800.
[0088] In step 825, the cloud server application marks infrequent anomalous string values as data quality anomalies and applies at least one crowdsourcing technique to determine corrections for these infrequent anomalous string values. In one embodiment, the cloud server application facilitates crowdsourcing by querying public forums of data experts or subject matter experts to obtain one or more correction techniques to identify valid string values, and implements corrections based on these valid string values. In a further embodiment, the cloud server application records, marks, or otherwise documents corrections to infrequent anomalous string values as standardization modifications within the context of the application model. According to such an embodiment, the cloud server application derives at least one rule from a set of proposed data standardization rules based on the standardization modifications. If crowdsourcing fails to identify a valid string value, the cloud server application may optionally replace the infrequent anomalous string value with a null value. In response to performing step 825, the cloud server application may proceed to the end of method 800.
[0089] In summary, applying a data scraping algorithm according to method 800 to evaluate infrequent outlier string values includes: classifying infrequent outlier string values as valid values in response to identifying a threshold number of data points associated with both the infrequent outlier string value and the associated data class name. Additionally, applying a data scraping algorithm according to method 800 to evaluate infrequent outlier string values includes: determining whether there exists at least one valid string value within the dataset that has a predetermined degree of string similarity to the infrequent outlier string value in response to failing to identify a threshold number of data points associated with both the infrequent outlier string value and the associated data class name. In such a case, applying a data scraping algorithm according to method 800 to evaluate infrequent outlier string values further includes: determining a correction to the infrequent outlier string value by selecting a valid string value from among the at least one valid string value in response to identifying at least one valid string value with a predetermined degree of string similarity within the dataset; and determining at least one modification to the infrequent outlier string value based on any identified character differences between the infrequent outlier string value and the selected valid string value. In related embodiments, selecting a valid string value from at least one valid string value includes evaluating a quantitative similarity between an infrequent anomalous string value and each of the at least one valid string value. In further related embodiments, selecting a valid string value from at least one valid string value includes applying one or more heuristics methods based at least in part on a string value selection history. Furthermore, applying a data scraping algorithm to evaluate infrequent anomalous string values according to method 800 also includes: in response to the inability to identify at least one valid string value with a predetermined level of string similarity within the dataset, applying at least one crowdsourcing technique to determine corrections for the infrequent anomalous string values.
[0090] Figure 9A method 900 is illustrated in the context of step 740 of method 700 to determine corrections for anomalous non-string values. Method 900 begins at step 905, where a cloud server application identifies a regular expression format associated with the anomalous non-string value. In embodiments, the cloud server application identifies the regular expression format and / or any predefined regular expression rules based on data categories assigned to the anomalous non-string value. Additionally or alternatively, the cloud server application identifies the regular expression format and / or any predefined regular expression rules based on one or more character, syntactic, and / or symbol patterns identified in the anomalous non-string value. In step 910, the cloud server application determines at least one modification to the anomalous non-string value to conform to the regular expression format. In embodiments, the at least one modification includes corrections to any character, syntactic, and / or symbol differences indicated by the regular expression format. In the context of different embodiments, syntactic differences of the anomalous non-string value relative to the regular expression format may include irregular placement of characters (letters and / or numbers) relative to symbols, irregular character order, irregular punctuation, and / or irregular patterns. For example, anomalous non-string values for date data classes may have irregular patterns in the representation of month, day, and / or year. The symbolic differences of anomalous non-string values relative to the regular expression format may include symbol omission and / or irregular symbol usage, such as using symbols in anomalous non-string values that are similar to but not identical to their corresponding symbols in valid non-string values. For example, annomalous non-string values may use slashes between aspects of a date, while the relevant regular expression format for date data classification values may require hyphens instead of slashes. While an anomalous non-string value in such an instance may represent a date, the value may violate the relevant regular expression format required for each data normalization model, and therefore the cloud server application may modify the value to conform to such a format according to step 910. In another embodiment, the cloud server application records, marks, or otherwise records at least one modification of the anomalous non-string value as a normalization modification in the context of the application model. According to such a further embodiment, the cloud server application derives at least one rule from the proposed set of data normalization rules based on the normalization modification.
[0091] In summary, determining the correction for anomalous non-string values according to method 900 includes: identifying the regular expression format associated with the anomalous non-string value; and determining at least one modification to the anomalous non-string value to conform to the regular expression format.
[0092] Figure 10Method 1000 is shown in the context of step 740 of method 700 to determine the correction for anomalous non-string values. As a supplement to or alternative to the steps of method 900, the cloud server application optionally performs the steps of method 1000. Method 1000 begins at step 1005, where the cloud server application parses the anomalous non-string value into at least one string portion. The cloud server application parses the anomalous non-string value to identify any string portions within the non-string value, and then analyzes each identified string portion to determine any format violations within it. Upon determining any format violation within any such string portion, the cloud server application identifies that string portion as an anomalous string portion. In step 1010, for each anomalous string portion in the at least one string portion, i.e., for each string portion comprising at least one anomalous element, the cloud server application performs two sub-steps. In sub-step 1010a, the cloud server application selects a valid string value from the at least one valid string value by evaluating a quantitative similarity between the anomalous string portion and each of the at least one valid string value. In an embodiment, to assess quantitative similarity, the cloud server application applies at least one technique similar to that described previously with respect to step 820, such as the calculation of a quantitative similarity score or the application of a weighted decision tree-based algorithm. In sub-step 1010b, the cloud server application determines at least one modification to the anomalous string portion based on any identified character differences between the anomalous string portion and the valid string value selected in sub-step 1010a. In an embodiment, the cloud server application records, marks, or otherwise records at least one modification to the anomalous string portion as a standardized modification within the context of the application model. According to such a further embodiment, the cloud server application derives at least one rule from a proposed set of data standardization rules based on the standardized modification.
[0093] In summary, determining the correction for an anomalous non-string value according to method 1000 includes parsing the anomalous non-string value into at least one string portion, and, for each anomalous string portion in the at least one string portion, selecting a valid string value from the at least one valid string value by evaluating a quantitative similarity between the anomalous string portion and each of the at least one valid string value, and further by determining at least one modification to the anomalous string portion based on any identified character differences between the anomalous string portion and the selected valid string value.
[0094] Figure 11Method 1100 for determining correction for abnormal non-string values is illustrated in the context of step 740 of method 700. In addition to, or as an alternative to, the steps of method 900 and / or 1000, the cloud server application may optionally perform the steps of method 1100. Method 1100 begins at step 1105, where the cloud server application determines whether the abnormal non-string value is a duplicate value. In response to determining that the abnormal non-string value is not a duplicate value, the cloud server application may proceed to the end of method 1100. In response to determining that the abnormal non-string value is a duplicate value, in step 1110, the cloud server application determines a replacement for the duplicate value based on the relationship between a first data column including the duplicate value and at least one related data column (i.e., at least one data column associated with the first data column). In an embodiment, if the duplicate value is in a first data column associated with one or more other data columns, the cloud server application may optionally replace the duplicate value with a value consistent with the data column. According to such embodiments, the cloud server application may optionally analyze corresponding values within one or more other data columns (e.g., corresponding values in data columns adjacent to or otherwise associated with the first data column containing duplicate values) to determine a correlation between the first data column containing duplicate values and other data columns, and replace the duplicate values with valid values identified based on that correlation. For example, in response to determining duplicate Social Safety Number data class values within a Social Safety Number data column, the cloud server application may analyze one or more Name data class values in the associated data column to determine appropriate Social Safety Number data class value replacements based on the one or more Name data class values. Additionally or alternatively, the cloud server application may optionally analyze corresponding values within one or more other data column rows (e.g., corresponding values in rows adjacent to or otherwise associated with rows containing duplicate values) to determine a correlation between the first data column containing duplicate values and other data columns, and replace the duplicate values with valid values identified based on that correlation. The cloud server application may optionally apply at least one pattern matching algorithm and / or at least one clustering algorithm to determine any such correlation(s). In a further embodiment, the cloud server application records, marks, or otherwise records the replacement of duplicate values as a standardization modification within the context of the application model. According to such a further embodiment, the cloud server application derives at least one rule from a proposed set of data standardization rules based on the standardization modification. More generally, within the context of the data standardization model, any such proposed data standardization rule may optionally supplement or otherwise modify data column-related rules within multiple baseline data integration rules or multiple data quality rules.
[0095] In summary, determining the correction for an anomalous non-string value according to method 1100 includes: in response to determining that the anomalous non-string value is a duplicate value, determining a replacement for the duplicate value based on the relationship between a first data column including the duplicate value and at least one related data column.
[0096] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. All kinds of modifications to the described embodiments and equivalent arrangements should fall within the scope of this invention. Therefore, the scope of the invention should be interpreted most broadly according to the following claims relating to the specific embodiments, and should cover all possible equivalent variations and arrangements. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or technical improvements superior to those found in the market, or to enable those skilled in the art to understand the embodiments described herein.
Claims
1. A computer implementation method, comprising: Receive the dataset during the data loading process; Classify the data points within the dataset; Apply the machine learning data normalization model to data points of each category within the dataset; A set of proposed data standardization rules for the dataset is obtained based on any standardization modifications determined as the machine learning data standardization model is applied to outliers. The set of proposed data standardization rules includes multiple mappings for automatically promoting the standardization of the dataset when the machine learning data standardization model is subsequently applied. as well as The proposed set of data standardization rules is applied to the dataset to improve its quality.
2. The computer-implemented method according to claim 1, further comprising: The proposed set of data standardization rules is presented for client review. as well as In response to the acceptance of the proposed set of data standardization rules, the proposed set of data standardization rules is applied to the dataset.
3. The computer-implemented method according to claim 2, further comprising: In response to the acceptance of the proposed set of data standardization rules, the machine learning data standardization model is updated based on the proposed set of data standardization rules.
4. The computer-implemented method according to claim 1, wherein, Configuring the machine learning data normalization model includes: Multiple datasets are sampled during multiple corresponding data loading scenarios; Identify corresponding data anomaly scenarios based on the multiple sampled datasets; Identify existing data anomaly correction techniques to address the corresponding data anomaly scenarios; and The machine learning data standardization model is trained based on the applicability of the existing data anomaly correction techniques to the corresponding data anomaly scenarios.
5. The computer-implemented method according to claim 1, wherein, Applying the machine learning data normalization model to data points of each category within the dataset includes: In response to determining that the data point is an outlier processed by pre-existing standardization rules incorporated into the machine learning data standardization model, the outlier is dynamically modified by applying the pre-existing standardization rules.
6. The computer-implemented method according to claim 1, wherein, Applying the machine learning data normalization model to data points of each category within the dataset includes: In response to determining that the data point is an outlier that has not been corrected by any pre-existing standardization rules, at least one standardization modification is determined to correct the data point.
7. The computer-implemented method according to claim 6, wherein, Determining at least one standardized modification to correct the data points includes: In response to determining that the data point is a null value, the replacement of the null value is determined based on the relationship between a first data column including the null value and at least one related data column, or based on the relationship between related values within the first data column.
8. The computer-implemented method according to claim 6, wherein, Determining at least one standardized modification to correct the data points includes: In response to determining that the data point is a frequent anomalous string value with an occurrence rate greater than or equal to a predetermined dataset frequency threshold, the frequent anomalous string value is classified as a valid value.
9. The computer-implemented method according to claim 6, wherein, Determining at least one standardized modification to correct the data points includes: In response to determining that the data point is an infrequent anomalous string value with an occurrence rate less than a predetermined dataset frequency threshold, a data scraping algorithm is applied to evaluate the infrequent anomalous string value via automatic application, wherein the input to the data scraping algorithm includes the infrequent anomalous string value and the associated data class name.
10. The computer-implemented method according to claim 9, wherein, The application of the data scraping algorithm to evaluate the infrequent abnormal string values includes: In response to identifying a threshold number of data points associated with both the infrequent anomalous string value and the associated data class name, the infrequent anomalous string value is classified as a valid value.
11. The computer-implemented method according to claim 9, wherein, The application of the data scraping algorithm to evaluate the infrequent abnormal string values includes: In response to the failure to identify a threshold number of data points associated with both the infrequent anomalous string value and the associated data class name, it is determined whether at least one valid string value exists within the dataset that has a predetermined degree of string similarity to the infrequent anomalous string value; and In response to identifying at least one valid string value within the dataset having the predetermined degree of string similarity, a correction for the infrequent anomalous string value is determined by selecting a valid string value from among the at least one valid string value, and at least one modification to the infrequent anomalous string value is determined based on any identified character differences between the infrequent anomalous string value and the selected valid string value.
12. The computer-implemented method according to claim 11, wherein, Selecting a valid string value from the at least one valid string value includes evaluating the quantitative similarity between the infrequent anomalous string value and each of the at least one valid string value.
13. The computer implementation method according to claim 11, wherein, Selecting a valid string value from the at least one valid string value includes at least partially applying one or more heuristics based on the string value selection history.
14. The computer implementation method according to claim 11, wherein, Applying the data scraping algorithm to evaluate the infrequent abnormal string values further includes: In response to the inability to identify at least one valid string value with a predetermined degree of string similarity within the dataset, at least one crowdsourcing technique is applied to determine the correction for the infrequent anomalous string values.
15. The computer-implemented method according to claim 6, wherein, Determining at least one standardized modification to correct the data points includes: In response to determining that the data point is an anomalous non-string value, a correction for the anomalous non-string value is determined.
16. The computer-implemented method according to claim 15, wherein, Determining the correction for the aforementioned abnormal non-string value includes: Identify the regular expression format associated with the anomalous non-string values; and Determine at least one modification to the abnormal non-string value to conform to the regular expression format.
17. The computer-implemented method according to claim 15, wherein, Determining the correction for the aforementioned abnormal non-string value includes: Parse the abnormal non-string value into at least one string component; and For each abnormal string portion in the at least one string portion: Valid string values are selected from at least one valid string value by evaluating the quantitative similarity between outlier string segments and each of at least one valid string value; and At least one modification to the abnormal string portion is determined based on any identified character differences between the abnormal string portion and the selected valid string value.
18. The computer-implemented method according to claim 15, wherein, Determining the correction for the aforementioned abnormal non-string value includes: In response to determining that the abnormal non-string value is a duplicate value, a replacement for the duplicate value is determined based on the relationship between a first data column including the duplicate value and at least one related data column.
19. A computer program product comprising a computer-readable storage medium having program instructions executable by a computing device to cause the computing device to perform the method steps of any one of claims 1 to 18.
20. A computer system, comprising: At least one processor; as well as A memory that stores an application program that performs operations when executed on the at least one processor, the operations including the steps of the method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Artificial intelligence (AI) based automatic data remediation
US20200349169A1
Active learning for data matching
US20210042330A1