Standardization in the Context of Data Integration

Automated data standardization in cloud environments using machine learning models addresses the inefficiencies of manual methods, enhancing data quality and integration efficiency by classifying and standardizing data points with client review and model updates.

JP7838904B2Active Publication Date: 2026-04-01INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-18
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Existing data integration methods in cloud computing environments rely heavily on manual standardization, which is time-consuming and resource-intensive, requiring expert intervention for data remediation.

Method used

Automated data standardization techniques using machine learning models to classify and standardize data points, allowing for the derivation and application of data standardization rules, with optional client review and model updates for improved accuracy.

Benefits of technology

Reduces the need for manual data standardization efforts, enhances data quality through automated processes, and improves data integration efficiency by leveraging machine learning and data crawling techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007838904000001
    Figure 0007838904000001
  • Figure 0007838904000002
    Figure 0007838904000002
  • Figure 0007838904000003
    Figure 0007838904000003
Patent Text Reader

Abstract

Techniques are described for automated data standardization in a managed services domain of a cloud computing environment. A related computer-implemented method includes receiving a dataset during a data onboarding procedure and classifying data points in the dataset. The method further includes applying a machine learning data standardization model to each classified data point in the dataset and deriving a set of proposed data standardization rules for the dataset based on any standardization modifications determined by application of the model. Optionally, the method includes presenting the set of proposed data standardization rules for a client's review and applying the set of proposed data standardization rules to the dataset in response to acceptance of the set of proposed data standardization rules. The method further includes updating the machine learning data standardization model in response to acceptance of the set of proposed data standardization rules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The various embodiments described herein generally relate to standardization in the context of data integration. More specifically, the various embodiments describe techniques for creating a set of data standardization rules to facilitate data integration in the managed service domain of a cloud computing environment.

Summary of the Invention

[0002] The various embodiments described herein provide techniques for automatic data standardization in the context of data integration. The related computer-implemented method includes receiving a data set during a data onboarding procedure. The method further includes classifying data points within the data set. The method further includes applying a machine learning data standardization model to each classified data point within the data set. The method further includes deriving a set of proposed data standardization rules for the data set based on any standardization corrections determined by the application of the machine learning data standardization model. In one embodiment, optionally, the method includes presenting the set of proposed data standardization rules for review by a client and, in response to acceptance of the set of proposed data standardization rules, applying the set of proposed data standardization rules to the data set. In a further embodiment, the method includes updating the machine learning data standardization model based on the set of proposed data standardization rules in response to acceptance of the set of proposed data standardization rules.

[0003] One or more additional embodiments relate to a computer program product including a computer-readable storage medium in which program instructions are implemented. According to such embodiments, the program instructions may be executable by a computing device and cause the computing device to perform one or more steps of the computer implementation method described above, or implement one or more embodiments related thereto, or both. One or more further embodiments relate to a system having at least one processor and memory for storing an application program, the memory which, when executed on at least one processor, performs one or more steps of the computer implementation method described above, or implements one or more embodiments related thereto, or both.

[0004] Therefore, a more detailed explanation of the embodiments briefly summarized above can be obtained by referring to the attached drawings, so that the embodiments described above can be achieved and understood in detail.

[0005] However, it should be noted that the attached drawings only illustrate typical embodiments of the present invention and should therefore not be considered to limit the scope of the invention, for the present invention may permit other equally effective embodiments. [Brief explanation of the drawing]

[0006] [Figure 1] This document describes a cloud computing environment according to one or more embodiments. [Figure 2] This document describes one or more embodiments of an abstraction model layer provided by a cloud computing environment. [Figure 3] This describes one or more embodiments of a managed services domain in a cloud computing environment. [Figure 4]This document describes how to create data standardization rules to facilitate data integration in a managed services domain, according to one or more embodiments. [Figure 5] This document describes how to construct a machine learning data standardization model according to one or more embodiments. [Figure 6] This document describes how to apply a machine learning data standardization model to each data point in a dataset, according to one or more embodiments. [Figure 7] One or more embodiments describe a method for determining at least one standardization modification to address data points in a dataset. [Figure 8] This document describes how to apply a data crawling algorithm to evaluate non-frequent outlier string values ​​in a dataset, according to one or more embodiments. [Figure 9] This document describes a method for determining remediation measures for outlier non-string values ​​in a dataset, according to one or more embodiments. [Figure 10] One or more further embodiments describe a method for determining a fix for outlier non-string values ​​in a dataset. [Figure 11] One or more further embodiments describe a method for determining a fix for outlier non-string values ​​in a dataset. [Modes for carrying out the invention]

[0007] The various embodiments described herein concern automated data standardization techniques for data integration in the managed services domain of a cloud computing environment. In the context of the various embodiments, data integration encompasses data governance, and further encompasses data governance and integration. An exemplary data integration solution is IBM® Unified Governance and Integration Platform. A cloud computing environment is a virtualized environment in which one or more computing capacities are available as a service. A cloud server system configured to implement automated data standardization techniques relating to the various embodiments described herein may utilize machine learning knowledge models, specifically the artificial intelligence capabilities of machine learning data standardization models, and information in a knowledge base associated with such models.

[0008] Various embodiments may offer advantages over the prior art. Traditional data integration involves manual standardization, such as data onboarding which requires manual remediation of anomalous or other non-conforming data points in a dataset by subject experts or data stewards. While manual remediation is useful for improving data quality, it can consume significant resources in terms of time and effort. The various embodiments described herein focus on providing automated standardization in the context of data integration and reducing the need for manual remediation. Specifically, the various embodiments facilitate automated data standardization of datasets by reducing the need for manual data standardization by subject experts or data stewards or both, while still allowing review and approval where desired. Furthermore, the various embodiments improve computing techniques by facilitating machine learning that increasingly improves data integration based on the successive application of data standardization models. Furthermore, the various embodiments improve computing techniques through the application of data crawling or crowdsourcing or both techniques to address, investigate, or both anomalies in a dataset. Some of the various embodiments may not include all of such advantages, and such advantages are not necessarily required for all embodiments.

[0009] Various embodiments of the present invention will be referred to below. However, it should be understood that the present invention is not limited to any specific embodiment described herein. Rather, any combination of the following features and elements, whether related to a different embodiment or not, is intended to implement and practice the present invention. Furthermore, embodiments may achieve advantages that are to other possible solutions, to the prior art, or both, but it is not limited to whether a particular advantage is achieved by a given embodiment. Accordingly, the following aspects, features, embodiments, and advantages are merely illustrative and shall not be considered elements or limitations of any appended claims unless expressly stated in a claim. Similarly, references to “the present invention” shall not be construed as a generalization of any inventive subject matter disclosed herein and shall not be considered elements or limitations of any appended claims unless expressly stated in one or more claims.

[0010] The present invention may be a system, method, or computer program product or combination thereof, integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0011] A computer-readable storage medium can be a tangible device capable of holding and storing instructions used by an instruction execution device. A computer-readable storage medium may, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific examples of computer-readable storage media include portable computer diskettes, hard disks, RAM, ROM, EPROM (or flash memory), SRAM, CD-ROM, DVD, memory stick, floppy disk, punch cards, or grooved raised structures, and mechanically encoded devices on which instructions are recorded, and suitable combinations thereof. The computer-readable storage medium as used herein should not be interpreted as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted through wires.

[0012] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). The network consists of copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. The network adapter card or network interface of each computing / processing device receives computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on the computer-readable storage medium within each computing / processing device.

[0013] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++ and procedural programming languages ​​such as the C programming language or similar programming languages. The computer-readable program instructions are executable as a standalone software package, either entirely on the user's computer or partially on the user's computer. Alternatively, they may be executable partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by personalizing them using state information of computer-readable program instructions in order to perform aspects of the present invention.

[0014] Aspects of the present invention are described herein with reference to flowcharts or block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block in a flowchart or block diagram, or both, and any combination of blocks in a flowchart or block diagram, or both, can be implemented by computer-readable program instructions.

[0015] These computer-readable program instructions can be provided to a general-purpose computer, a dedicated computer processor, or other programmable data processing device to generate a machine, such that instructions executed via the processor of a computer or other programmable data processing device generate means for implementing functions / operations specified in one or more blocks of a flowchart or block diagram or both. These computer-readable program instructions can also be stored in a computer-readable storage medium that can be connected to a computer, a programmable data processing device, or other device or combination of devices that function in a particular way, such that the computer-readable storage medium on which the instructions are stored constitutes one of the outputs containing instructions that implement the modes of function / operations specified in one or more blocks of a flowchart or block diagram or both.

[0016] Computer-readable program instructions, like instructions that perform a function / action specified in one or more blocks of a flowchart or block diagram or both on a computer, other programmable device, or other device, can also be loaded into a computer, other programmable data processing device, or other device and perform a series of operational steps on the computer, other programmable device, or other device to produce a computer-implemented process.

[0017] The flowcharts and block diagrams in the figures illustrate the structure, functionality, and operation of implementations that can be executed by systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which constitutes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions shown in the blocks may occur in a different order than shown in the figures. For example, two blocks shown in succession may actually be accomplished as one step, and at the same time, may be executed in a partially or wholly temporally overlapping manner, or the blocks may be executed in the reverse order depending on the relevant functions. It should also be noted that each block of the block diagram or flowchart diagram, or both, and combinations of blocks of the block diagram or flowchart diagram, or both, can be implemented by a special-purpose hardware-based system that performs the specified functions or operations, or combinations of special-purpose hardware and computer instructions.

[0018] In certain embodiments, techniques related to automatic data normalization for the purpose of data integration in a managed service domain will be described. However, it should be understood that the techniques described herein can be adapted to various purposes in addition to those specifically described herein. Thus, references to specific embodiments are included by way of illustration and not limitation.

[0019] The various embodiments described herein can be provided to an end user through a cloud computing infrastructure. Although this disclosure includes a detailed description of cloud computing, the implementation forms of the teachings described herein are not limited to a cloud computing environment. Rather, the various embodiments described herein can be implemented with any other type of computer environment currently known or developed in the future.

[0020] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or service provider interaction. Thus, cloud computing allows users to access virtual computing resources (storage, data, applications, and even complete virtualized computing systems) within the cloud regardless of the underlying physical systems (or their locations) used to provide the computing resources. This cloud model may include at least five characteristics, at least three service models, and at least four implementation models.

[0021] The characteristics are as follows.

[0022] On-demand self-service: Cloud consumers can unilaterally provision computing capabilities such as server time and network storage automatically as needed, without the need for human interaction with the service provider.

[0023] Broad network access: Computing capabilities are available over the network and can be accessed via standard mechanisms, thereby facilitating use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, personal digital assistants (PDAs)).

[0024] Resource pooling: A provider's computing resources are pooled and delivered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated as needed. Generally, consumers have a sense of location independence because they do not manage or know the exact location of the resources provided. However, consumers may be able to identify the location at a higher level of abstraction (e.g., country, state, data center).

[0025] Rapid Elasticity: Computing power can be prepared quickly and flexibly, allowing it to scale out automatically and immediately, and to be quickly released and scale in immediately. To consumers, the computing power available for preparation often appears unlimited and can be purchased in any quantity at any time.

[0026] Measured Services: Cloud systems leverage metric capabilities at a certain level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, active user count) to automatically control and optimize resource usage. Resource usage can be monitored, controlled, and reported to provide transparency to both service providers and consumers.

[0027] The service model is as follows:

[0028] Software as a Service (SaaS): The functionality offered to consumers is the ability to use the provider's applications running on a cloud infrastructure. These applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., webmail). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, except for configuring a limited number of user-specific applications.

[0029] Platform as a Service (PaaS): The functionality offered to consumers is the ability to deploy applications they have created or acquired to cloud infrastructure using programming languages ​​and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, and storage, but they can control the deployed applications and, in some cases, the configuration of their hosting environment.

[0030] Infrastructure as a Service (IaaS): The functionality provided to consumers is the provision of processors, storage, networking, and other basic computing resources that enable consumers to deploy and run any software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they can control the operating system, storage, and deployed applications, and in some cases, partially control certain network components (e.g., host firewalls).

[0031] The deployment model is as follows:

[0032] Private Cloud: This cloud infrastructure is operated exclusively for a specific organization. This cloud infrastructure can be managed by that organization or a third party and can reside on-premises or off-premises.

[0033] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with common interests (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by the organization or a third party and can reside on-premises or off-premises.

[0034] Public Cloud: This cloud infrastructure is provided to a large number of people or large industry groups and is owned by organizations that sell cloud services.

[0035] Hybrid Cloud: This cloud infrastructure combines two or more cloud models (private, community, or public). While maintaining the unique entities of each model, they are bound together by standards or individual technologies to achieve data and application portability (e.g., cloud bursting for load balancing across clouds).

[0036] Cloud computing environments are service-oriented environments that emphasize statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is the infrastructure, which includes a network of interconnected nodes.

[0037] Figure 1 shows a cloud computing environment 50 in one or more embodiments. The cloud computing environment 50 includes one or more cloud computing nodes 10. Local computer devices used by cloud consumers (e.g., personal digital assistants or mobile phones 54A, desktop computers 54B, laptop computers 54C, or automotive computer systems 54N, or a combination thereof) can communicate with these nodes. The nodes 10 can communicate with each other. The nodes 10 can be grouped physically or virtually (not shown) in one or more networks, such as the private, community, public, or hybrid clouds or a combination thereof. This allows the cloud computing environment 50 to provide infrastructure, platforms, or software as a service, or a combination thereof, without requiring cloud consumers to maintain resources on their local computer devices. Note that the types of computer devices 54A-N shown in Figure 1 are merely examples, and it should be understood that the computing nodes 10 and the cloud computing environment 50 can communicate with any type of electronic device via any type of network or network addressable connection (e.g., using a web browser) or both.

[0038] Figure 2 shows a set of functional abstraction model layers provided by the cloud computing environment 50 in one or more embodiments. It should be understood that the components, layers, and functions shown in Figure 2 are illustrative only, and the various embodiments described herein are not limited to these. Various layers and corresponding functions are provided as illustrated. Specifically, the hardware and software layer 60 includes hardware and software components. Examples of hardware components include a mainframe 61, a reduced instruction set computer (RISC) architecture-based server 62, server 63, blade server 64, storage 65, and a network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68. The virtualization layer 70 provides an abstraction layer. From this layer, for example, the following virtual entities can be provided: a virtual server 71, virtual storage 72, a virtual network 73 including a virtual private network, a virtual application and operating system 74, and a virtual client 75.

[0039] As an example, the management layer 80 can provide the following functions: Resource preparation 81 enables the dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment 50. Metering and pricing 82 enables cost tracking as resources are used within the cloud computing environment 50 and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only protection of data and other resources but also identification and verification of cloud consumers and tasks. The user portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 enables the allocation and management of cloud computing resources to ensure that requested service levels are met. Service Level Agreement (SLA) planning and execution 85 enables the pre-arrangement and procurement of cloud computing resources that are expected to be needed in the future in accordance with the SLA.

[0040] The workload layer 90 provides examples of the capabilities available to the cloud computing environment 50. Examples of workloads and capabilities available from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom education delivery 93, data analysis processing 94, transaction processing 95, and data standardization 96. Data standardization 96 can enable automated data standardization for data integration via machine learning knowledge models according to various embodiments described herein.

[0041] Figure 3 shows a managed services domain 300 within a cloud computing environment 50. Functions related to data standardization 96 and other workloads / functions may be implemented within the managed services domain 300. The managed services domain 300 includes a cloud server system 310. In one embodiment, the cloud server system 310 includes an onboard data repository 320, a database management system (DBMS) 330, and a cloud server application 340 that includes a machine learning knowledge model 350 incorporating at least the functionality of a machine learning data standardization model. The cloud server application 340 is representative of a single application or multiple applications. The machine learning knowledge model 350 is configured to enable, facilitate, or both enable automated data standardization through various embodiments described herein. Furthermore, the managed services domain 300 includes a data onboard interface 360, one or more external database systems 370, and multiple application server clusters 3801 to 380. n This includes: The data onboard interface 360 ​​enables communication between the cloud server system 310 and data clients / systems interacting with the managed service domain 300 to facilitate data onboarding in the context of data integration. In one embodiment, the cloud server system 310 has one or more external database systems 370 and multiple application server clusters 3801 to 380 n It is configured to communicate with. Furthermore, application server cluster 3801 to 380 n Each server within can be configured to communicate with each other, with server clusters in other domains, or both.

[0042] In one embodiment, an onboard data repository 320, representing either a single data repository or a collection of data repositories, includes unstandardized or pre-standardized datasets, or both, received from the data onboard interface 360 ​​during the data onboard procedure, as well as classified and standardized datasets processed by the various embodiments described herein. Alternatively, one or more data repositories include pre-standardized or unstandardized datasets, or both, received during the data onboard procedure, and one or more additional data repositories include classified and standardized datasets. According to a further alternative, one or more data repositories include pre-standardized or unstandardized datasets, or both, received during the data onboard procedure, one or more additional data repositories include classified datasets, and one or more further additional data repositories include standardized datasets. In one embodiment, a DBMS 330 coordinates and manages the knowledge base of a machine learning knowledge model 350. The DBMS 330 may include one or more database servers, which can coordinate or manage various aspects of the knowledge base, or both. In an additional embodiment, the DBMS 330 manages or interacts with one or more external database systems 370. One or more external database systems 370 include one or more relational databases, or one or more database management systems, or both, configured to interface with DBMS 330. In a further embodiment, DBMS 330 includes multiple application server clusters 3801 to 380 n and stores the relationships between knowledge bases. In a further embodiment, DBMS330 includes or operably coupled to one or more databases, some or all of which may be relational databases. In a further embodiment, DBMS330 includes one or more ontology trees or other ontology structures. Application server clusters 3801 to 380 nIt hosts, stores, or both hosts and stores various application configurations, and provides managed server services to one or more client systems or data systems or both.

[0043] Figure 4 shows a method 400 for creating data standardization rules to facilitate data integration. Creating data standardization rules according to method 400 facilitates automated data standardization in the context of the various embodiments described herein. In one embodiment, one or more steps related to method 400 are performed in an environment where computing power is provided as a service (e.g., a cloud computing environment 50). According to such an embodiment, one or more steps related to method 400 are performed in a managed service domain within the environment (e.g., a managed service domain 300). The environment may be a hybrid cloud environment. In a further embodiment, one or more steps related to method 400 are performed in one or more other environments, such as a client-server network environment or a peer-to-peer network environment. A centralized cloud server system within a managed service domain (e.g., a cloud server system 310 within managed service domain 300) can facilitate processing according to method 400 and other methods further described herein. More specifically, a cloud server application within a cloud server system (e.g., cloud server application 340) can perform or facilitate one or more steps of Method 400 and other methods described herein. Automated data standardization techniques facilitated or performed via a cloud server system within a managed service domain may also be associated with data standardization within the workload layer of a functional abstraction layer provided by the environment (e.g., data standardization within the workload layer 90 of the cloud computing infrastructure 50 96).

[0044] Method 400 begins with step 405, in which a cloud server application receives a dataset during a data onboarding procedure. The cloud server application receives the dataset in accordance with step 405 via a data onboarding interface within a managed service domain (e.g., via data onboarding interface 360). The onboarded data is stored, fully or partially, in at least one onboarded data repository (e.g., onboarded data repository 320), which may be a data lake or associated therewith. In the context of the various embodiments described herein, a data lake is a storage repository that holds data from many sources in its natural or raw form. The data in the data lake may be structured, semi-structured, or unstructured. In one embodiment, the data onboarding procedure includes onboarding a dataset to a data lake. According to such an embodiment, the dataset is optionally onboarded in Hadoop Distributed File System (HDFS) format. Alternatively, according to such an embodiment, the dataset is onboarded as normalized data, for example, through manual or automated removal of data redundancy. In a further embodiment, only metadata related to data located in the data lake is onboarded during the data onboarding procedure, rather than the data itself. According to such a further embodiment, based on the execution of the steps of Method 400, the cloud server application optionally creates a standardized version of the data located in the data lake based on the onboarded metadata. Alternatively, according to such a further embodiment, based on the execution of the steps of Method 400, the cloud server application directly applies data standardization to the data located in the data lake based on the onboarded metadata. In summary, in the context of Method 400, the cloud server application optionally onboards a dataset to the data lake, or alternatively, onboards metadata related to a dataset that is already in the data lake.In further embodiments, the data onboarding procedure includes onboarding a dataset from an online transaction processing (OLTP) system to an online analytical processing (OLAP) system.

[0045] In step 410, the cloud server application classifies the data points in the dataset received in step 405. In one embodiment, the cloud server application discovers a data class based on the application of at least one data classifier algorithm and, in response to the discovery of a data class, classifies the data points by labeling (e.g., tagging) the relevant data fields with terms or classifications (e.g., business-related terms or classifications). Examples of each data class include local government or postal code. According to such an embodiment, at least one data classifier algorithm incorporates regular expressions to identify data within a data field or based on the data field name. In addition or alternatively, at least one data classifier algorithm incorporates lookup operations to reference a table containing constituent data class values. In addition or alternatively, at least one data classifier algorithm includes custom logic written in a programming language, e.g., Java. In a further embodiment, the cloud server application classifies the data points based on the evaluation of each individual value. In addition or alternatively, the cloud server application classifies the data points based on the evaluation of multiple values ​​in a data column, for example, in the context of structured or semi-structured data. In addition or alternatively, the cloud server application classifies data points based on evaluations of multiple values ​​in multiple related data columns. In a further embodiment, the cloud server application associates confidence values ​​with the classification of data points made through at least one data classifier algorithm. According to such a further embodiment, the cloud server application prioritizes or labels or both based on the confidence data point classification made through at least one data classifier. Optionally, a client, such as a subject expert or data steward (i.e., a data editor or data engineer), overrides one or more data point classifications made through at least one data classifier algorithm and manually assigns data point classifications.

[0046] In one embodiment, in the context of step 410, the cloud server application randomly samples a subset of classified data points in the dataset. According to such an embodiment, the cloud server application determines the value of the randomly sampled data point and then determines a specific data class or data column, or both, associated with the randomly sampled data point. According to such an embodiment, the cloud server application determines a data class or data column, or both, associated with the other data points in the subset, based on the association with the randomly sampled data point. Thus, random sampling according to such an embodiment can allow the cloud server application to avoid iterating through the entire dataset to classify the data points. In addition or alternatively, the cloud server application directly classifies one or more data points in the dataset by assigning a data class. Thus, the cloud server application can classify each data point in the dataset based on random sampling, or based on direct manual classification, or both.

[0047] In step 415, the cloud server application applies a machine learning data standardization model (e.g., machine learning knowledge model 350) to each classified data point in the dataset. In one embodiment, the cloud server application applies the data standardization model based on the data class. According to such an embodiment, the type of anomaly detected and the corresponding standardization for such anomaly type are determined at least partially based on the type of data class. For example, a data point with the telephone number data class requires validation only with respect to format (i.e., only format anomalies are relevant). On the other hand, a data point with the municipal data class requires validation at least with respect to spelling, and possibly syntax (i.e., spelling and syntax anomalies are potentially relevant). As will be further discussed herein, data classes are also relevant in distinguishing between string values ​​and non-string values ​​in the dataset.

[0048] In one embodiment, a cloud server application applies a plurality of data quality rules in the context of a data standardization model to determine whether a data point in a dataset is a valid value or an outlier, and further to determine how to handle any values ​​determined to be outliers. The plurality of data quality rules include a plurality of baseline data integration rules, which are standardized baselines for determining whether a data point is a valid value or an outlier. In the context of the various embodiments described herein, the plurality of baseline data integration rules are hardcoded or fixed rules applied to a data class or group of data classes and are integrated into the data standardization model during model initialization. Optionally, the plurality of baseline data integration rules are customized based on the dataset size, the dataset source, or both. The plurality of baseline data integration rules include pre-established rules for identifying anomalies associated with data points in a dataset. In a further embodiment, the plurality of baseline data integration rules are specific to a particular data class or a particular group of data classes. According to such a further embodiment, the cloud server application applies the plurality of baseline data integration rules based on the data class. For example, according to a baseline integration rule for values ​​in the date of birth data class, such values ​​cannot be future dates. According to such further embodiments, the cloud server application optionally analyzes null or duplicate values ​​based on the data class associated with such null or duplicate values. In the context of various embodiments, a valid value is a value that conforms to a set of baseline data integration rules, while an outlier is a value that does not conform to a set of baseline data integration rules. The cloud server application optionally determines that a value does not conform to the set of baseline data integration rules and is therefore an outlier if the value does not satisfy one or more conditions established for the associated data class or the associated data column or both (i.e., the value contains at least one anomaly).In addition or alternatively, a cloud server application determines that a value is an outlier if it contradicts other values ​​within a defined group in the dataset (e.g., a data column or a group of related data columns), as the value does not conform to multiple baseline data integration rules. In a particular example, such contradictions are evident in a scenario where a first data column correlates with other data columns such that an outlier in the first data column contradicts a corresponding value in another data column or other data columns that correlates with the first data column. In the context of various embodiments, columns are correlated if the values ​​in one column are predictable based on the values ​​in other columns. In other examples, such contradictions are evident in a scenario where data values ​​in a data column are related based on the location in the data column, such that an outlier in the data column contradicts the location in that data column. In other examples, such contradictions are evident in an outlier that contradicts both the corresponding value in the correlated data columns and the location in that data column.

[0049] In addition to multiple baseline data integration rules, the multiple data quality rules include a set of arbitrary data standardization rules previously derived and accepted for the purpose of addressing values ​​determined to be outliers by the application of the multiple baseline data integration rules. The multiple data quality rules are fully coupled so as not to have an indeterminate range. Furthermore, the multiple data quality rules include variables to be coupled to data columns in the dataset. In one embodiment, the cloud server application calculates the amount of outliers identified in the dataset by the application of the data standardization model. According to such an embodiment, the cloud server application calculates a quantitative data quality score for the dataset based on its conformance to the multiple baseline data integration rules. The cloud server application optionally calculates a data quality score based on the proportion of non-conformities / outliers among the classified data points in the dataset. By applying the data standardization model according to step 415, the cloud server application can apply previously derived and accepted standardization rules to address outliers, and further, as described below, the cloud server application can determine additional standardization modifications based on the fact that a new set of proposed standardization rules may be derived. The method for applying the data standardization model to each data point in the dataset according to step 415 is illustrated with reference to Figure 6.

[0050] In step 420, the cloud server application derives a proposed set of data standardization rules for the dataset based on the standardization modifications determined by the application of the machine learning data standardization model. The proposed set of data standardization rules includes one or more rules to facilitate automatic and dynamic dataset standardization when the model is subsequently applied. In one embodiment, the proposed set of data standardization rules includes a plurality of mappings, which are optionally stored in a standardization lookup table. According to such an embodiment, the plurality of mappings include each mapping from an outlier to a valid value or other relational aspects. Specifically, the plurality of mappings may include mappings between outlier string values ​​containing one or more misspelled characters and valid string values ​​with the correct spelling, for example, the outlier string value "Dehli" may be mapped to the valid string value "Delhi". Embodiments relating to the determination of standardization modifications are further described herein.

[0051] Optionally, in step 425, the cloud server application presents the proposed set of data standardization rules derived in step 420 for client review. In one embodiment, the cloud server application presents the proposed set of data standardization rules by publishing the set in a forum accessible to the client. In a further embodiment, the cloud server application presents the proposed set of data standardization rules by sending the set to a client interface. According to such a further embodiment, the client interface is a user interface in the form of a graphical user interface (GUI), a command-line interface (CLI), or both, presented via at least one client application installed in or accessible by the client system or device. Optionally, the cloud server application facilitates client acceptance of the proposed set of data standardization rules by generating form-response interface elements for display to the client, for example, in a forum accessible to the client or in the client interface. The client includes subject matter experts in a particular domain, data stewards, or any other entity related to the dataset or the cloud server system, or a combination thereof.

[0052] In one embodiment, the cloud server application automatically accepts a proposed set of data standardization rules without client review, based on quantitative confidence values ​​attributable to the proposed set of data standardization rules, each standardization rule within them, or both. According to such an embodiment, the cloud server application automatically accepts the proposed set of data standardization rules and therefore omits client review in response to determining that the proposed set of data standardization rules improves data inaccuracies with confidence exceeding a predetermined confidence threshold. The predetermined confidence threshold quantitatively measures the reliability of the automated data improvement capability (i.e., improvement capability without subject matter experts / data stewards). According to such an embodiment, the quantitative confidence values ​​attributable to the proposed set of data standardization rules, each standardization rule within them, or both are optionally set by the client based on the relevant use case. In addition or alternatively, the confidence values ​​are determined based on data relationships identified through machine learning analysis in the context of previous model applications or previous model iterations (i.e., previous iterations within the current model application) or both. For example, assuming the capital state data class value is "Memphis," the corresponding state data class value will always be "Tennessee," and therefore the confidence value for one or more applicable data standardization rules will be relatively high. In another example, assuming the capital state data class value is "Hyderabad," the corresponding state data class value will be either "Telangana" or "Andhra Pradesh" depending on the context (since Hyderabad is the capital of both states), and therefore the confidence value for one or more applicable data standardization rules will be relatively low. Thus, assuming the same predetermined confidence threshold in the context of these examples, if one or more rules in the proposed set deal with a state data class value in the dataset corresponding to "Memphis" rather than "Hyderabad," the cloud server application is relatively likely to accept the proposed set of data standardization rules without client review in the context of determining standardized state data class values.

[0053] In step 430, the cloud server application determines whether the proposed set of data standardization rules has been accepted. Acceptance of the proposed set of data standardization rules is optionally determined based on a client review, or alternatively, automatically determined in response to a determination that the proposed set of data standardization rules improves data inaccuracies with a confidence level exceeding a predetermined confidence threshold. In response to a determination that the proposed set of data standardization rules has not been accepted, the cloud server application optionally adjusts one or more of the proposed set of data standardization rules, and then repeats step 430. Optionally, such rule adjustments include collecting client feedback on the proposed set of data standardization rules, in which case the cloud server application may apply one or more modifications based on the client feedback. In one embodiment, such client feedback includes one or more rule override requests with rule modifications imposed by one or more clients. The cloud server application optionally adjusts the proposed set of data standardization rules based on the client feedback.

[0054] In response to the decision that the proposed set of data standardization rules has been accepted, in step 435, the cloud server application applies the proposed set of data standardization rules to the dataset. In one embodiment, the cloud server application applies the proposed set of data standardization rules to each data point in the dataset. In a further embodiment, the cloud server application verifies that the data quality of the dataset has improved as a result of applying the proposed set of data standardization rules. According to such a further embodiment, the cloud server application verifies that the data quality of the dataset has improved through a client review of the data quality of the dataset before and after the application of the proposed set of data standardization rules. In addition or alternatively, the cloud server application verifies that the data quality of the dataset has improved through a comparison of the data quality of the dataset before and after the application of the proposed set of data standardization rules.

[0055] Furthermore, in response to the acceptance of the proposed set of data standardization rules, in step 440, the cloud server application updates the machine learning data standardization model based on the proposed set of data standardization rules. In one embodiment, the cloud server application integrates (e.g., adds or associates) the proposed set of data standardization rules with a plurality of data quality rules associated with the data standardization model, so that the cloud server application considers the proposed set of data standardization rules along with any previously accepted sets of a plurality of baseline data integration rules and data standardization rules during subsequent model application. According to such an embodiment, the cloud server application adds the proposed set of data standardization rules to a knowledge base associated with the data standardization model. The knowledge base associated with the model optionally includes all repositories, ontologities, files, or documents, or a combination thereof, associated with a plurality of data quality rules and any set of proposed data standardization rules. The cloud server application optionally accesses the knowledge base via a DBMS associated with the cloud server system (e.g., DBMS330).

[0056] Based on integrating the proposed set of data standardization rules into multiple data quality rules, the cloud server application optionally trains a data standardization model to identify patterns of data points that enable automatic standardization during subsequent model application. The cloud server application facilitates the adaptation of the proposed set of data standardization rules upon acceptance so that outliers previously unexpected but automatically addressed by the proposed set are standardized during subsequent application of multiple data quality rules associated with the machine learning data standardization model. Thus, based on the integration of the proposed set of data standardization rules, during subsequent model application, the cloud server application automatically standardizes dataset data points determined to be outliers if they are addressed by one or more of the proposed set of data standardization rules. After updating the data standardization model according to step 440, the cloud server application optionally returns to step 415 and reapplies the data standardization model to the dataset to improve dataset standardization. Such further embodiments correspond to the application of a multipath algorithm, as contemplated in the context of various embodiments. Alternatively, after updating the data standardization model according to step 440, the cloud server application can proceed to the termination of method 400.

[0057] Figure 5 illustrates a method 500 for configuring a machine learning data standardization model. Method 500 begins in step 505, in which a cloud server application samples multiple datasets between multiple respective data onboarding scenarios. In one embodiment, the cloud server application samples datasets based on one or more relevant data classes. According to one alternative, in order to focus on one or more specific data classes, the cloud server application samples datasets that have a quantity greater than a threshold of values ​​classified according to one or more relevant data classes. According to another alternative, in order to ensure diversity in the sampled data, the cloud server application samples datasets that have a quantity of values ​​associated with a data class that exceeds a threshold. In step 510, the cloud server application identifies each data anomaly scenario based on the multiple sampled datasets. In one embodiment, the cloud server application identifies outliers from the multiple sampled datasets by applying multiple data quality rules related to the data standardization model. Optionally, multiple baseline data integration rules within the multiple data quality rules include at least one correlation rule that identifies the expected relationship between or among the respective values ​​in the correlated data columns. For example, a correlation rule might require that the numerical value of a particular row in a first data column is twice the numerical value of a particular row in a second data column. In addition or alternatively, multiple baseline data integration rules include at least one data column correlation rule that identifies the expected relationship between or among related values ​​in the data columns. Optionally, multiple data quality rules further include any adaptable rules derived from external sources such as ontologities or repositories. For example, a cloud server application might retrieve a list of valid values ​​from a data repository for one or more data classes and derive a rule that considers the valid values ​​included in the list.Following the initial construction of the data standardization model, the cloud server application optionally updates multiple data quality rules beyond multiple baseline data integration rules, for example, based on acceptance of each set of data standardization rules proposed in the context of Method 400, or based on the adoption of additional rules from external sources, or both.

[0058] In step 515, the cloud server application identifies existing data anomaly correction techniques to address each data anomaly scenario. In one embodiment, the cloud server application identifies existing data anomaly correction techniques by referencing a previously accepted set of data standardization rules, a previously proposed set of data standardization rules, or both. Such rule sets may be stored in a knowledge base associated with the data standardization model, or they may be documented in relation to the model, for example, so that they are available from an external source. According to such an embodiment, the cloud server application identifies any set of previously accepted or previously proposed data standardization rules for a dataset from among a plurality of sampled datasets. In step 520, the cloud server application trains a machine learning data standardization model based on the applicability of existing data anomaly correction techniques to each data anomaly scenario. According to step 520, the cloud server application trains the model by creating and analyzing relevant data that addresses the relationship between each data anomaly scenario and the existing data anomaly correction techniques. The relevant data may include confidence data regarding the effectiveness of one or more existing anomaly correction techniques in addressing one or more of each data anomaly scenario. In addition, or alternatively, the relevant data may include a comparison of multiple existing anomaly correction techniques in terms of addressing each specific data anomaly scenario. By analyzing the relevant data, the cloud server application can determine which of the existing anomaly correction techniques may be relatively more effective in addressing each data anomaly scenario or a similar scenario during future model application. Thus, the cloud server application can calibrate the model based on such analysis. In one embodiment, the cloud server application records the relevant data about the model, for example, in a knowledge base associated with the model.

[0059] In one embodiment, the cloud server application trains a data standardization model based on an arbitrary referenced and previously accepted set of data standardization rules, or an arbitrary referenced and previously proposed set of data standardization rules, or both, to enhance or adapt multiple data quality rules beyond multiple baseline data integration rules. Optionally, as a result of the training in step 520, the cloud server application integrates rules from the previously accepted set of data standardization rules into multiple data quality rules. Such training can enhance the data standardization model in addition to, or as an alternative to, updating the data standardization model resulting from the acceptance of a proposed set of data standardization rules in accordance with step 440 in the context of method 400. In additional embodiments, the cloud server application first performs one or more model configuration steps, in particular external source consultation, to build the data standardization model. In further embodiments, the cloud server application performs one or more model configuration steps, in particular training, at regular time intervals, or whenever the cloud server application processes a threshold amount of the dataset, or both. Following the training of the machine learning data standardization model, the cloud server application can proceed to the end of method 500.

[0060] In summary, constructing a machine learning data standardization model according to Method 500 involves sampling multiple datasets between multiple data onboarding scenarios, identifying each data anomaly scenario based on the multiple sampled datasets, identifying existing data anomaly correction techniques to address each data anomaly scenario, and training a machine learning data standardization model based on the applicability of existing data anomaly correction techniques to each data anomaly scenario.

[0061] Figure 6 illustrates Method 600, in the context of step 415 of Method 400, which applies a machine learning data standardization model to each classified data point in a dataset. Method 600 begins in step 605, when the cloud server application selects data points in the dataset for evaluation. In step 610, the cloud server application determines whether the data points selected in step 605 are valid values. In one embodiment, the cloud server application determines whether a data point is a valid value by determining whether the data point conforms to multiple baseline data integration rules. In response to determining that the data points selected in step 605 are valid values, no further standardization processing is required for the data points, and therefore the cloud server application proceeds to step 630. In response to determining that a data point is not a valid value, i.e., in response to determining that a data point does not conform to multiple baseline data integration rules, in step 615, the cloud server application determines whether the data point is an outlier that can be addressed by existing standardization rules incorporated into the data standardization model, e.g., rules from a previously accepted set of data standardization rules incorporated into multiple data quality rules, or rules adapted from external sources as a result of model training, or both. In response to determining that the data point is an outlier that can be addressed by existing standardization rules incorporated into the data standardization model, in step 620, the cloud server application dynamically corrects the outlier by applying the existing standardization rules. By dynamically correcting the outlier based on the application of existing standardization rules, the cloud server application can leverage the model to automatically resolve standardization issues without further analysis and processing. Following step 620, the cloud server application proceeds to step 630.

[0062] In response to determining that a data point is an outlier not addressed by existing standardization rules incorporated into the data standardization model, in step 625, the cloud server application determines at least one standardization modification to improve the data point. The method relating to determining at least one standardization modification to improve the data point according to step 625 is illustrated with reference to Figure 7. Following step 625, the cloud server application proceeds to step 630. In step 630, the cloud server application determines whether there are any other data points in the dataset to be evaluated. In response to determining that there are any other data points in the dataset to be evaluated, the cloud server application returns to step 605. In response to determining that there are no other data points in the dataset to be evaluated, the cloud server application may proceed to the end of method 600.

[0063] In summary, applying a machine learning data standardization model to each classified data point in a dataset according to Method 600 includes dynamically correcting outliers by applying existing standardization rules in response to determining that a data point is an outlier that is addressed by existing standardization rules incorporated into the machine learning data standardization model. Furthermore, applying a machine learning data standardization model to each classified data point in a dataset according to Method 600 includes determining at least one standardization correction to improve a data point in response to determining that a data point is an outlier that is not addressed by any existing standardization rules.

[0064] Figure 7 illustrates Method 700 in the context of step 625 of Method 600, which determines at least one standardization modification to improve a data point in a dataset. Method 700 begins in step 705, where the cloud server application determines whether a data point is a null value. In response to determining that the data point is not a null value, the cloud server application proceeds to step 715. In response to determining that the data point is a null value, in step 710, the cloud server application determines a replacement for the null value based on the relationship between the first data column containing the null value and at least one correlated data column (i.e., at least one data column correlated with the first data column), or based on the relationship between interrelated values ​​within the first data column. In one embodiment, if the null value is in the first data column correlated with one or more other data columns, the cloud server application optionally replaces the null value with a value that matches the data column correlation. According to such embodiments, the cloud server application optionally analyzes each value in one or more other data columns (e.g., each value in a data column adjacent to or related to the first data column containing the null value) to determine the correlation between a first data column containing a null value and other data columns, and to replace the null value with a valid value identified based on the correlation. In addition or alternatively, the cloud server application optionally analyzes each value in one or more other data column rows (e.g., each value in a row adjacent to or related to the row containing the null value) to determine the correlation between a first data column containing a null value and other data columns, and to replace the null value with a valid value identified based on the correlation. The cloud server application optionally applies at least one pattern matching algorithm or at least one clustering algorithm or both to determine such correlation. In further embodiments, if the null value is in a data column with mutually related values ​​based on the data column location (e.g., data column row), the cloud server application optionally replaces the null value with a value that matches the data column location of the null value.

[0065] In one embodiment, the cloud server application determines the null value substitution in step 710 based on correlation information or interrelated value information or both derived from a plurality of baseline data integration rules and applicable rules and correlation information or interrelated value information or both derived from relational data obtained from a repository or ontology, or based on correlation information or interrelated value information or both obtained by machine learning through the application of a data standardization model, or a combination thereof. In a further embodiment, the cloud server application logs, marks, or records the null value substitution as a standardization modification in the context in which the model is applied. According to such a further embodiment, the cloud server application derives at least one rule from the proposed set of data standardization rules based on the standardization modification. Any such proposed data standardization rule optionally supplements or modifies data column correlation rules or interrelated rules or both in the plurality of baseline data integration rules or more generally in the plurality of data quality rules in the context of the data standardization model. Upon execution of step 710, the cloud server application can proceed to the termination of method 700.

[0066] In step 715, the cloud server application determines whether a data point is a frequent outlier string value. The cloud server application determines a frequent outlier string value as an outlier string value having an occurrence rate greater than or equal to a predetermined dataset frequency threshold. In the context of various embodiments, the cloud server application may optionally measure the occurrence rate of such outlier string values ​​based on their occurrence in a dataset being processed according to the method herein, or alternatively, the cloud server application measures the occurrence rate based on their occurrence in a specified set of datasets, which may or may not include the dataset being processed, or in a sampled set of datasets, or both. In the context of various embodiments, a string value is a data type used to represent text. Such a string value may include a sequence of characters, numbers, or symbols, or combinations thereof. In one embodiment, such a string value may be a variable character field (varchar), which is a set of characters of indefinite length in the context of various embodiments, or may include one. A particular data class may be defined to include string values, or otherwise may be associated with string values, for example, as opposed to non-string values. For example, a postal code data class value may be defined to contain a string value of a certain amount of characters (digits or characters or both). Thus, in certain embodiments, string values ​​are distinguished from non-string values ​​based on the data class type or data column type or both. In response to determining that a data point is not a frequent outlier string value, the cloud server application proceeds to step 725. In response to determining that a data point is a frequent outlier string value, in step 720, the cloud server application classifies the frequent outlier string value as a valid value. In one embodiment, the cloud server application saves, tags, or marks the frequent outlier string value as a valid value in the context of a data standardization model.According to such embodiments, the cloud server application optionally stores the frequent outlier string values ​​in a knowledge base associated with the model. In a further embodiment, the cloud server application logs, marks, or records the validation of the frequent outlier string values ​​as a standardization correction in the context in which the model is applied. According to such further embodiments, the cloud server application derives at least one rule from a proposed set of data standardization rules based on the standardization correction. By validating the frequent outlier string values, the cloud server application facilitates automatic standardization of the frequent outlier string values ​​when the model is subsequently applied. Upon execution of step 720, the cloud server application can proceed to the termination of method 700.

[0067] In step 725, the cloud server application determines whether the data point is a non-frequent outlier string value. The cloud server application identifies a non-frequent outlier string value as an outlier string value having an occurrence rate below a predetermined dataset frequency threshold. In response to determining that the data point is not an outlier string value, the cloud server application proceeds to step 735. In response to determining that the data point is a non-frequent outlier string value, in step 730, the cloud server application applies a data crawling algorithm to evaluate the non-frequent outlier string value via an automated application. In relevant embodiments, the input to the data crawling algorithm includes the non-frequent outlier string value and the associated data class name. In further relevant embodiments, the data crawling algorithm is a web crawling algorithm designed to systematically access web pages. In the context of various embodiments, the automated application may be a bot, or may incorporate a bot. In addition or alternatively, the cloud server application applies a data scraping algorithm in the context of step 730 to retrieve data from other sources beyond web pages. Upon execution of step 730, the cloud server application can proceed to the termination of method 700. A method relating to applying a data crawling algorithm to evaluate non-frequent outlier string values ​​according to step 730 is illustrated with reference to Figure 8.

[0068] In step 735, the cloud server application determines whether a data point is an outlier non-string value. In the context of various embodiments, a non-string value represents a data type that incorporates one or more aspects beyond text, such as a regular expression. A regular expression contains a sequence of characters that define a search pattern. While string values ​​are generally defined through a finite list of values, non-string values ​​are evaluated based on pattern analysis, such as regular expression analysis. In certain embodiments, non-string values ​​are distinguished from string values ​​based on the data class type, data column type, or both. A particular data class may be defined to contain non-string values, or may be associated with non-string values, for example, in contrast to string values.

[0069] In one embodiment, the cloud server application determines, as a result of regular expression analysis, i.e., search pattern analysis, that a non-string value is an outlier, and more specifically, determines one or more anomalies based on the fact that a non-string value is an outlier. During regular expression analysis, the cloud server application identifies one or more anomalies as formal violations. For example, the cloud server application may determine that a non-string social security number data class value is an outlier based on the absence of a hyphen, since a defined pattern of hyphens is required for regular expression validation of non-string values ​​in the social security number data class. In another example, the cloud server application may determine that a non-string email address data class value is an outlier based on the absence of a period, since a period in the domain name is required for regular expression validation of non-string values ​​in the email address data class. According to such an embodiment, the cloud server application identifies formal violations related to outlier non-string values ​​at least in part based on the application of multiple baseline data integration rules contained within multiple data quality rules. In response to the application of multiple baseline data integration rules, alerts related to formal violations of non-string values ​​are optionally and automatically triggered. In a further embodiment, the cloud server application determines that a data point is an outlier non-string value based on one or more anomalies in one or more string portions within such data point. In a further embodiment, the cloud server application determines that a data point is an outlier non-string value based on the fact that the data point is a duplicate value in a data column related to a unique value. For example, the cloud server application may identify a data point as a non-string outlier if it has a non-string social security number data class value that is a duplicate value of at least one other data point in a social security number data column.

[0070] In response to determining that a data point is an outlier non-string value, in step 740, the cloud server application determines a remedial action for the outlier non-string value. Upon execution of step 740, the cloud server application may proceed to the termination of method 700. Several methods for determining a remedial action for an outlier non-string value according to step 740 are described with reference to Figures 9 to 11. In response to determining that a data point is not an outlier non-string value, in step 745, optionally, the cloud server application may investigate the identity of the data point through crowdsourcing, external repository consultation, or both, and then proceed to the termination of method 700. In one embodiment, the cloud server application facilitates crowdsourcing by referring to a public forum of data experts or subject experts. Failure to identify the data point before step 745 may indicate an exception in data quality, e.g., a poorly defined value. According to one or more alternative embodiments, the cloud server application performs the steps of method 700 in an alternative configuration. For example, the cloud server application can execute steps 705-710, 715-720, 725-730, and 735-740 in one or more alternative sequences.

[0071] In summary, determining at least one standardization correction to improve a data point according to Method 700 includes determining a substitute for the null value based on the relationship between a first data column containing the null value and at least one correlated data column, or based on the relationship between interrelated values ​​within the first data column, in response to determining that the data point is a null value. Furthermore, determining at least one standardization correction to improve a data point according to Method 700 includes classifying a frequent outlier string value as a valid value in response to determining that the data point is a frequent outlier string value having an occurrence rate greater than or equal to a predetermined dataset frequency threshold. Furthermore, determining at least one standardization correction to improve a data point according to Method 700 includes applying a data crawling algorithm to evaluate the non-frequent outlier string value via an automated application in response to determining that the data point is a non-frequent outlier string value having an occurrence rate less than a predetermined dataset frequency threshold, with the input to the data crawling algorithm including the non-frequent outlier string value and the associated data class name. Furthermore, determining at least one standardization modification to improve a data point according to Method 700 includes determining a modification for an outlier non-string value in response to determining that the data point is an outlier non-string value.

[0072] Figure 8 illustrates Method 800, which applies a data crawling algorithm to evaluate a non-frequent outlier string value in the context of step 730 of Method 700. Method 800 begins in step 805, where the cloud server application determines whether a threshold amount of data points associated with both the non-frequent outlier string value and the associated data class name has been identified through data crawling. In response to failing to identify a threshold amount of data points associated with both the non-frequent outlier string value and the associated data class name, the cloud server application proceeds to step 815. In response to identifying a threshold amount of data points associated with both the non-frequent outlier string value and the associated data class name, in step 810, the cloud server application classifies the non-frequent outlier string value as a valid value. In one embodiment, the cloud server application stores, tags, or marks the non-frequent outlier string value as a valid value in the context of a data standardization model. According to such an embodiment, the cloud server application optionally stores the non-frequent outlier string value as a valid value in a knowledge base associated with the model. In a further embodiment, the cloud server application logs, marks, or records the validation of non-frequent outlier string values ​​as a standardization modification in the context in which the model is applied. According to such a further embodiment, the cloud server application derives at least one rule from the proposed set of data standardization rules based on the standardization modification. By validating non-frequent outlier string values, the cloud server application facilitates automatic standardization of non-frequent outlier string values ​​during subsequent model application. Upon execution of step 810, the cloud server application can proceed to the termination of method 800.

[0073] In step 815, the cloud server application determines whether there is at least one valid string value in the dataset that has a predefined string similarity to a non-frequent outlier string value. In one embodiment, the cloud server application defines a predetermined degree of string similarity between a non-frequent outlier string value and a valid string value based on the non-frequent outlier string value and the valid string value having the maximum amount of character difference. In addition or alternatively, the cloud server application defines a predetermined string similarity between a non-frequent outlier string value and a valid string value based on the non-frequent outlier string value and a valid string value having the minimum amount of common characters. In addition or alternatively, the cloud server application defines a predetermined string similarity between a non-frequent outlier string value and a valid string value based on the non-frequent outlier string value and a valid string value whose respective string length difference is less than a predetermined threshold. In response to failing to identify at least one valid string value in the dataset that has a predefined string similarity to a non-frequent outlier string value, the cloud server application proceeds to step 825. In response to identifying at least one valid string value in a dataset having a predefined string similarity to a non-frequent outlier string value, in step 820, the cloud server application determines a corrective action for the non-frequent outlier string value based on the selection of the valid string value. In one embodiment, the valid string value selection includes selecting a valid string value from at least one valid string value and determining at least one correction for the non-frequent outlier string value based on any identified character difference between the non-frequent outlier string value and the selected valid string value. In a further embodiment, the cloud server application logs, marks, or records at least one correction of the non-frequent outlier string value as a standardization correction in the context in which the model is applied. According to such an embodiment, the cloud server application derives at least one rule of a proposed set of data standardization rules based on the standardization correction.

[0074] Selecting a valid string value from at least one valid string value according to step 820 optionally includes evaluating the quantitative similarity between each of the non-frequent outlier string values ​​and the at least one valid string value. In one embodiment, the cloud server application evaluates quantitative similarity by calculating a quantitative similarity score for each of the at least one valid string value based on the evaluated similarity between the valid string value and the non-frequent outlier string value. According to such an embodiment, the cloud server application selects a valid string value from at least one valid string value based on the highest calculated quantitative similarity score. The cloud server application optionally calculates a quantitative similarity score based on one or more similarity factors. The similarity factors optionally include the edit distance between each of the non-frequent outlier string values ​​and the at least one valid string value. In the context of various embodiments, the edit distance is the amount of correction required to correct an outlier to fit a particular valid value. In addition or alternatively, the similarity factors include a data source of the non-frequent outlier string values ​​compared to each of the at least one valid string value. In such context, the data source may optionally include the location of origin, the author's identity, or both. In addition or alternatively, the similarity metric may include the data classification of non-frequent outlier string values ​​compared to each of at least one valid string value. In such context, the data classification may optionally include the data class, data type, subject, target stratum, data age, data modification history, or the values ​​of related data fields, or a combination thereof.

[0075] In the relevant embodiments, one or more similarity factors are weighted such that similarity factors with relatively higher weights have a greater impact on the calculated quantitative similarity score than similarity factors with relatively lower weights. To distinguish between two valid values ​​having the same quantitative similarity score (e.g., two municipality values ​​in the same data class that differ by one character from a common source), the cloud server application may optionally select one of the two valid values ​​based on an evaluation of the respective relevant field values ​​of the relevant data class for each of the two valid values ​​(e.g., a postal code data class value for each of the two municipality data class values). The cloud server application may also select one of the two valid values ​​based on which of the two valid values ​​has relevant field values ​​that satisfy, or more completely comply with, one or more conditions established for the relevant data class or the relevant data column or both, as determined by a data standardization rule set associated with the data standardization model, e.g., a previously accepted data standardization rule set. For example, given the non-frequently occurring outlier string municipal data class value "Middletonw" and the valid municipal data class values ​​"Middleton" and "Middletown", the cloud server application should select one of the two valid municipal data class values ​​based on which valid value has a postal code that satisfies, or more completely adheres to, one or more conditions established for the relevant postal code data class or the relevant postal code data column or both, as determined by the data standardization rule set associated with the model.

[0076] In a further relevant embodiment, the cloud server application evaluates the quantitative similarity between each of the non-frequent outlier string values ​​and at least one valid string values ​​by applying an algorithm based on a weighted decision tree. The cloud server application optionally determines the respective weight values ​​of the weighted decision tree based on the respective edit distances between each of the non-frequent outlier string values ​​and at least one valid string values. In addition or alternatively, the cloud server application optionally determines the respective weight values ​​of the weighted decision tree based on the respective values ​​of data columns related to (e.g., correlated or related to) the data column of the non-frequent outlier string values.

[0077] In addition or alternatively, selecting a valid string value from at least one valid string value according to step 820 optionally includes applying one or more heuristics based at least partially on a string value selection history. In the context of this embodiment and other embodiments described herein, a heuristic is a machine logic-based rule that provides a decision based on one or more predetermined types of inputs. The predetermined types of inputs to one or more heuristics in the context of selecting a valid string value from at least one valid string value may include a valid string selection history in the context of a data class or data column or both related to a non-frequently occurring outlier string value, e.g., the most recently selected valid string or the most frequently selected valid string. Upon execution of step 820, the cloud server application may proceed to the termination of method 800.

[0078] In step 825, the cloud server application marks a non-frequent outlier string value as a data quality exception and applies at least one crowdsourcing technique to determine an improvement for the non-frequent outlier string value. In one embodiment, the cloud server application facilitates crowdsourcing to identify a valid string value for which improvement is to be implemented by querying a public forum of data experts or subject experts for one or more improvement techniques. In a further embodiment, the cloud server application logs, marks, or records the improvement of the non-frequent outlier string value as a standardization correction in the context in which the model is applied. According to such an embodiment, the cloud server application derives at least one rule from a proposed set of data standardization rules based on the standardization correction. If crowdsourcing cannot identify a valid string value, the cloud server application optionally replaces the non-frequent outlier string value with a null value. Upon execution of step 825, the cloud server application can proceed to the termination of method 800.

[0079] In summary, applying a data crawling algorithm to evaluate a non-frequent outlier string value according to Method 800 includes classifying the non-frequent outlier string value as a valid value in response to identifying a threshold amount of data points associated with both the non-frequent outlier string value and the associated data class name. Furthermore, applying a data crawling algorithm to evaluate an outlier string value according to Method 800 includes determining whether there is at least one valid string value in the dataset that has a predefined string similarity to the non-frequent outlier string value in response to failing to identify a threshold amount of data points associated with both the outlier string value and the associated data class name. If so, applying a data crawling algorithm to evaluate a non-frequent outlier string value according to Method 800 further includes determining a remedy for the non-frequent outlier string value by selecting a valid string value from at least one valid string value in the dataset that has a predefined string similarity, in response to identifying at least one valid string value in the dataset that has a predefined string similarity, and determining at least one modification to the non-frequent outlier string value based on any identified character differences between the non-frequent outlier string value and the selected valid string value. In a related embodiment, selecting a valid string value from at least one valid string value includes evaluating the quantitative similarity between each of the non-frequent outlier string values ​​and at least one valid string value. In a further related embodiment, selecting a valid string value from at least one valid string value includes applying one or more heuristics based at least partially on the string value selection history. Furthermore, applying a data crawling algorithm to evaluate a non-frequent outlier string value according to Method 800 further includes applying at least one crowdsourcing technique to determine improvements for the non-frequent outlier string value in response to failure to identify at least one valid string value in a dataset having a predefined string similarity.

[0080] Figure 9 illustrates Method 900 for determining a corrective action for an outlier non-string value in the context of step 740 of Method 700. Method 900 begins in step 905, in which a cloud server application identifies a regular expression format associated with the outlier non-string value. In one embodiment, the cloud server application identifies a regular expression format or an arbitrary predefined regular expression rule, or both, based on the data class assigned to the outlier non-string value. In addition or alternatively, the cloud server application identifies a regular expression format or an arbitrary predefined regular expression rule, or both, based on one or more characters, syntax, or symbol patterns, or combinations thereof, identified within the outlier non-string value. In step 910, the cloud server application determines at least one correction to the outlier non-string value to conform to the regular expression format. In one embodiment, at least one correction includes correcting any character, syntax, or symbol mismatch, or combination thereof, indicated by the regular expression format. In the context of various embodiments, syntactic mismatches in outlier nonstring values ​​with respect to regular expression formats may include incompatible arrangement of characters (letters, numbers, or both) for symbols, incompatible character order, incompatible punctuation, or incompatible patterns, or combinations thereof. For example, an outlier nonstring value of a date data class may have an incompatible pattern with respect to the representation of month, day, year, or a combination thereof. Symbolic mismatches in outlier nonstring values ​​with respect to regular expression formats may include the omission of symbols, the use of incompatible symbols, or both (for example, the use of a symbol in an outlier nonstring value that is similar to, but not identical to, the corresponding symbol in a valid nonstring value). For example, an outlier nonstring value may use slashes between aspects of a date, while the relevant regular expression format for a date data class value may require the use of hyphens instead of slashes.In such an example, the outlier non-string value might represent a date, but nevertheless, the value might violate the relevant regular expression format required for each data standardization model, so the cloud server application can modify the value according to step 910 to conform to such format. In a further embodiment, the cloud server application logs, marks, or records at least one modification of the outlier non-string value as a standardization modification in the context in which the model is applied. According to such a further embodiment, the cloud server application derives at least one rule of the proposed set of data standardization rules based on the standardization modification.

[0081] In summary, determining a remediation strategy for an outlier nonstring value according to Method 900 involves identifying the regular expression format associated with the outlier nonstring value and determining at least one modification to the outlier nonstring value in order to conform to the regular expression format.

[0082] Figure 10 shows Method 1000 for determining a corrective action for an outlier non-string value in the context of step 740 of Method 700. The cloud server application optionally performs the steps of Method 1000 in addition to or as an alternative to the steps of Method 900. Method 1000 begins in step 1005, in which the cloud server application parses the outlier non-string value into at least one string portion. The cloud server application parses the outlier non-string value to identify any string portions within the non-string value, and then parses each identified string portion to determine any format violations within it. Having determined any format violations within such string portions, the cloud server application marks the string portion as an outlier string portion. In step 1010, for each outlier string portion of at least one string portion, i.e., for each string portion containing at least one anomaly, the cloud server application performs two substeps. In substep 1010a, the cloud server application selects a valid string value from at least one valid string value by evaluating the quantitative similarity between each of the outlier string portion and at least one valid string value. In one embodiment, to evaluate the quantitative similarity, the cloud server application performs at least one technique similar to the techniques previously described with respect to step 820, e.g., calculation of a quantitative similarity score or application of an algorithm based on a weighted decision tree. In substep 1010b, the cloud server application determines at least one modification to the outlier string portion based on any identified character difference between the outlier string portion and the valid string value selected in substep 1010a. In one embodiment, the cloud server application logs, marks, or records the at least one modification to the outlier string portion as a standardization modification in the context in which the model is applied. According to such a further embodiment, the cloud server application derives at least one rule of the proposed set of data standardization rules based on the standardization modification.

[0083] In summary, determining a corrective action for an outlier non-string value according to Method 1000 includes parsing the outlier non-string value into at least one string part, selecting a valid string value from at least one valid string value by evaluating the quantitative similarity between each outlier string part and at least one valid string value, and further determining at least one correction for the outlier string part based on any identified character differences between the outlier string part and the selected valid string value.

[0084] Figure 11 shows Method 1100 for determining a remedy for an outlier non-string value in the context of step 740 of Method 700. The cloud server application optionally performs the steps of Method 1100 in addition to or as an alternative to the steps of Method 900 or 1000 or both. Method 1100 begins in step 1105, in which the cloud server application determines whether the outlier non-string value is a duplicate value. In response to determining that the outlier non-string value is not a duplicate value, the cloud server application may proceed to the end of Method 1100. In response to determining that the outlier non-string value is a duplicate value, in step 1110, the cloud server application determines a substitute for the duplicate value based on the relationship between the first data column containing the duplicate value and at least one correlated data column (i.e., at least one data column correlated with the first data column). In one embodiment, if the duplicate value is in the first data column correlated with one or more other data columns, the cloud server application optionally replaces the duplicate value with a value that matches the data column correlation. According to such embodiments, the cloud server application optionally determines the correlation between a first data column containing duplicate values ​​and other data columns, and analyzes each value in one or more other data columns (e.g., each value in a data column adjacent to or related to the first data column containing duplicate values) to replace the duplicate values ​​with valid values ​​identified based on the correlation. For example, if duplicates of social security number data class values ​​are determined in a social security number data column, the cloud server application may analyze one or more name data class values ​​in the correlated data columns to determine an appropriate alternative for the social security number data class value based on one or more name data class values. In addition or alternatively, the cloud server application optionally analyzes each value in one or more rows of other data columns (e.g., each value in a row adjacent to or related to the row with duplicate values) to determine the correlation between a first data column containing duplicate values ​​and other data columns, and replace the duplicate values ​​with valid values ​​identified based on the correlation.The cloud server application optionally applies at least one pattern matching algorithm or at least one clustering algorithm or both to determine such correlations. In a further embodiment, the cloud server application logs, marks, or records substitutions of duplicate values ​​as standardization modifications in the context in which the model is applied. According to such a further embodiment, the cloud server application derives at least one rule from the proposed set of data standardization rules based on the standardization modifications. Such proposed data standardization rule optionally supplements or modifies data column correlation rules within a plurality of baseline data integration rules or more generally within a plurality of data quality rules in the context of the data standardization model.

[0085] In summary, determining a remedy for an outlier non-string value according to Method 1100 involves determining a replacement for the duplicate value based on the relationship between the first data column containing the duplicate value and at least one correlated data column, in response to determining that the outlier non-string value is a duplicate value.

[0086] The descriptions of various embodiments of the present invention are presented for illustrative purposes only and are not intended to be exhaustive, nor are they intended to limit the disclosed embodiments. Any modifications of any kind made to the described embodiments and equivalent configurations shall be within the scope of protection of the present invention. Therefore, the scope of the present invention should be most broadly described in accordance with the claims that follow in connection with the detailed description, and should cover all possible equivalent variations and equivalent arrangements. It will be apparent to those skilled in the art that many modifications and changes are possible without departing from the scope of the described embodiments. The terms used herein have been selected to best describe the principles of the embodiments, their actual application to the technology found in the market, or technical improvements, or to enable those skilled in the art to understand the embodiments described herein.

Claims

1. The computer receives a dataset during the data onboard procedure, The computer classifies the data points within the dataset, The computer applies a machine learning data standardization model to each classified data point in the dataset, The computer derives a proposed set of data standardization rules for the dataset based on any standardization modifications determined by applying the machine learning data standardization model, Computer implementation methods, including those mentioned above.

2. The computer presents the proposed set of data standardization rules to the client for review, The computer, in response to acceptance of the proposed set of data standardization rules, applies the proposed set of data standardization rules to the dataset, The computer implementation method according to claim 1, further comprising:

3. The computer updates the machine learning data standardization model based on the proposed set of data standardization rules in response to acceptance of the proposed set of data standardization rules. The computer implementation method according to claim 1, further comprising:

4. Constructing the aforementioned machine learning data standardization model is, Sampling multiple datasets between each of the multiple data onboarding scenarios, Based on the aforementioned multiple sampled datasets, identify each data anomaly scenario, To identify existing data anomaly correction techniques to address each of the aforementioned data anomaly scenarios, The machine learning data standardization model is trained based on the applicability of the existing data anomaly correction techniques to each of the aforementioned data anomaly scenarios. The computer implementation method according to claim 1, including the method described in claim 1.

5. Applying the machine learning data standardization model to each classified data point in the dataset is: In response to determining that the data point is an outlier addressed by an existing standardization rule incorporated into the machine learning data standardization model, the outlier is dynamically corrected by applying the existing standardization rule. The computer implementation method according to claim 1, including the method described in claim 1.

6. Applying the machine learning data standardization model to each classified data point in the dataset is: In response to determining that the data point is an outlier not addressed by any existing standardization rule, determine at least one standardization modification to improve the data point. The computer implementation method according to claim 1, including the method described in claim 1.

7. Determining at least one standardization modification to improve the aforementioned data points is: In response to determining that the data point is a null value, an alternative for the null value is determined based on the relationship between the first data column containing the null value and at least one correlated data column, or based on the relationship between related values ​​within the first data column. The computer implementation method according to claim 6, including the method described in claim 6.

8. Determining at least one standardization modification to improve the aforementioned data points is: In response to determining that the data point is a frequently occurring outlier string value having an occurrence rate equal to or greater than a predetermined dataset frequency threshold, the frequently occurring outlier string value is classified as a valid value. The computer implementation method according to claim 6, including the method described in claim 6.

9. Determining at least one standardization modification to improve the aforementioned data points is: In response to determining that the data point is a non-frequently occurring outlier string value having an occurrence rate below a predetermined dataset frequency threshold, a data crawling algorithm is applied via an automated application to evaluate the non-frequently occurring outlier string value, wherein the input to the data crawling algorithm includes the non-frequently occurring outlier string value and the associated data class name. The computer implementation method according to claim 6, including the method described in claim 6.

10. Applying the data crawling algorithm to evaluate the aforementioned non-frequent outlier string values ​​is, In response to identifying the threshold amount of data points associated with both the non-frequently occurring outlier string value and the associated data class name, classify the non-frequently occurring outlier string value as a valid value. The computer implementation method according to claim 9, including the method described in claim 9.

11. Applying the data crawling algorithm to evaluate the aforementioned non-frequent outlier string values ​​is, In response to the failure to determine the threshold amount of data points associated with both the non-frequently occurring outlier string value and the associated data class name, it is determined whether there is at least one valid string value in the dataset that has a predefined string similarity to the non-frequently occurring outlier string value. In response to identifying at least one valid string value in the dataset having the predefined string similarity, the improvement measures for the outlier string value are determined by selecting a valid string value from among the at least one valid string value and determining at least one modification to the outlier string value based on any identified character difference between the outlier string value and the selected valid string value. The computer implementation method according to claim 9, including the method described in claim 9.

12. The computer implementation method according to claim 11, wherein selecting a valid string value from the at least one valid string value includes evaluating the quantitative similarity between each of the non-frequently occurring outlier string values ​​and the at least one valid string value.

13. The computer implementation method according to claim 11, wherein selecting a valid string value from the at least one valid string value comprises applying one or more heuristics based at least partially on the string value selection history.

14. Applying the data crawling algorithm to evaluate the aforementioned non-frequent outlier string values ​​is, In response to the failure to identify at least one valid string value in the dataset having the aforementioned predefined string similarity, apply at least one crowdsourcing technique to determine a solution for the infrequent outlier string value. The computer implementation method according to claim 11, further comprising:

15. Determining at least one standardization modification to improve the aforementioned data points is: In response to determining that the aforementioned data point is an outlier non-string value, a corrective action for the said outlier non-string value is determined. The computer implementation method according to claim 6, including the method described in claim 6.

16. Determining countermeasures for the aforementioned outlier non-string values ​​is, Identifying the regular expression format associated with the aforementioned outlier non-string value, To conform to the aforementioned regular expression format, determine at least one modification for the outlier non-string value, The computer implementation method according to claim 15, including the method described in claim 15.

17. Determining countermeasures for the aforementioned outlier non-string values ​​is, The outlier non-string value is parsed into at least one string portion, For each outlier string portion among the at least one string portion, By evaluating the quantitative similarity between each of the outlier string portion and at least one valid string value, a valid string value is selected from the at least one valid string value. Determine at least one modification to the outlier string portion based on any identified character difference between the outlier string portion and the selected valid string value, The computer implementation method according to claim 15, including the method described in claim 15.

18. Determining countermeasures for the aforementioned outlier non-string values ​​is, In response to determining that the outlier non-string value is a duplicate value, a substitute for the duplicate value is determined based on the relationship between the first data column containing the duplicate value and at least one correlated data column. The computer implementation method according to claim 15, including the method described in claim 15.

19. A computer program that causes a computer to perform the method described in any one of claims 1 to 18.

20. At least one processor, The processor includes memory for storing an application program, and the at least one processor executes an operation by executing the application program, and the operation is Receiving the dataset during the data onboarding procedure, Classifying the data points within the aforementioned dataset, Applying a machine learning data standardization model to each classified data point in the aforementioned dataset, Based on any standardization modifications determined by applying the aforementioned machine learning data standardization model, a proposed set of data standardization rules for the dataset is derived. A system that includes this.

Citation Information

Patent Citations

  • Device and method for detection of vocabulary error

    JP2010134709A

  • Map data error correction device

    JP2010231560A

  • Techniques for similarity analysis and data enrichment using knowledge sources

    JP2017536601A

  • Data management system, data management method, and data management program

    JP2020129224A

  • Cleansing and standardizing data

    US20140279972A1