Method, system, and computer-readable program
The data AI system addresses the challenge of complex data integration in distributed applications by using machine learning for automated mapping and governance, enhancing efficiency and enabling rapid application development.
Patent Information
- Application Number
- JP2022031862
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-08-22
- Filing Date
- 2022-03-02
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2037-08-22
AI Technical Summary
Existing software design tools for distributed applications are resource-intensive and require domain expertise for managing data integration across different execution environments, leading to significant effort in designing and configuring complex and scalable applications.
A data AI system leveraging machine learning for data flow management and integration, providing automatic mapping, data governance, and lifecycle management, with features like provenance, lineage, security, classification, retention, and validity tracking, and supporting rapid application development in cloud environments.
Enables efficient and automated data integration and governance, reducing the complexity and resource intensity of managing diverse data environments, facilitating rapid application development and discovery of unknown data insights.
Smart Images

Figure 0007715657000022 
Figure 0007715657000023 
Figure 0007715657000024
Abstract
Description
Technical Field
[0001] Copyright Notice Part of the disclosure of this patent document may contain material subject to copyright protection. The copyright owner does not object to the reproduction by anyone of the patent document or the patent disclosure, as long as it appears in the patent file or records of the Patent and Trademark Office, but reserves all copyrights in any case otherwise.
[0002] Claim of Priority This application claims priority to U.S. Provisional Patent Application No. 62 / 378,143, filed on August 22, 2016, entitled "SYSTEM AND METHOD FOR AUTOMATED MAPPING OF DATA TYPES BETWEEN CLOUD AND DATABASE SERVICES", U.S. Provisional Patent Application No. 62 / 378,146, filed on August 22, 2016, entitled "SYSTEM AND METHOD FOR DYNAMIC, INCREMENTAL RECOMMENDATIONS WITHIN REAL-TIME VISUAL SIMULATION", U.S. Provisional Patent Application No. 62 / 378,147, filed on August 22, 2016, entitled "SYSTEM AND METHOD FOR INFERENCING OF DATA TRANSFORMATIONS THROUGH PATTERN DECOMPOSITION", filed on August 22, 2016, entitled "SYSTEM AND METHOD FOR ONTOLOGY INDUCTION THROUGH STATISTICAL PROFILING AND U.S. Provisional Patent Application No. 62 / 378,150, filed on August 22, 2016, entitled "REFERENCE SCHEMA MATCHING", U.S. Provisional Patent Application No. 62 / 378,151, filed on August 22, 2016, entitled "SYSTEM AND METHOD FOR METADATA-DRIVEN EXTERNAL INTERFACE GENERATION OF APPLICATION PROGRAMMING INTERFACES", and U.S. Provisional Patent Application No. 62 / 378,152, filed on August 22, 2016, entitled "SYSTEM AND METHOD FOR DYNAMIC LINEAGE TRACKING AND RECONSTRUCTION OF COMPLEX BUSINESS ENTITIES WITH HIGH-LEVEL POLICIES" are claimed as priority, and each of the above applications is incorporated herein by reference. U.S. Provisional Patent Application No. 62 / 378,151, and U.S. Provisional Patent Application No. 62 / 378,152, filed on August 22, 2016, entitled "SYSTEM AND METHOD FOR DYNAMIC LINEAGE TRACKING AND RECONSTRUCTION OF COMPLEX BUSINESS ENTITIES WITH HIGH-LEVEL POLICIES" is claimed as priority, and each of the above applications is incorporated herein by reference.
[0003] Field of the Invention Embodiments of the present invention generally relate to methods for integrating data obtained from various sources, and more particularly to support for dynamic lineage tracking, reconstruction, and lifecycle management. BACKGROUND ART
[0004] Background Many modern computing environments require the ability to share large amounts of data among various types of software applications. However, distributed applications may have significantly different configurations, for example, due to differences in the supported data types or the respective execution environments. The configuration of an application can vary, for example, depending on its application programming interface, runtime environment, deployment method, lifecycle management, or security management.
[0005] Software design tools intended to be used in the development of such distributed applications tend to be resource-intensive and often require the services of domain model experts to manage application and data integration. As a result, application developers who tackle the challenge of building complex and scalable distributed applications that integrate various types of data across various types of execution environments generally have to expend a great deal of effort in designing, building, and configuring these applications.
SUMMARY OF THE INVENTION
[0006] Summary According to various embodiments, described herein is a system (data artificial intelligence (AI) system, data AI system) used in data integration or other computing environments that utilizes machine learning (ML, dataflow machine learning, DFML) for managing data flow (dataflow, DF) and constructing composite data flow software applications (data flow applications, pipelines). According to one embodiment, the system can provide a data governance function. This can include, for example, provenance (where specific data came from), lineage (how this data was acquired / processed), security (who was responsible for this data), classification (what this data is related to), impact (how much impact this data has on the business), retention (how long this data should persist), and validity (whether this data should be excluded / included for analysis / processing) for each slice of time-related data for a particular snapshot. These can be used in life cycle decisions and data flow recommendations. According to one embodiment, the system can provide a data governance function. This can include, for example, provenance (where specific data came from), lineage (how this data was acquired / processed), security (who was responsible for this data), classification (what this data is related to), impact (how much impact this data has on the business), retention (how long this data should persist), and validity (whether this data should be excluded / included for analysis / processing) for each slice of time-related data for a particular snapshot. These can be used in life cycle decisions and data flow recommendations. These can be used in life cycle decisions and data flow recommendations. These can be used in life cycle decisions and data flow recommendations.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
【Figure A diagram showing a system for displaying one or more semantic actions enabled for accessed data according to an embodiment. A diagram showing a graphical user interface for displaying one or more semantic actions enabled for accessed data according to an embodiment. A diagram further showing a graphical user interface for displaying one or more semantic actions enabled for accessed data according to an embodiment. A diagram showing a process for displaying one or more semantic actions enabled for accessed data according to an embodiment. A diagram showing means for identifying a pattern of transformation within a data flow for one or more functional expressions generated for each of one or more applications according to an embodiment. A diagram showing an example of identifying a pattern of transformation within a data flow for one or more functional expressions according to an embodiment. A diagram showing an object diagram used when identifying a pattern of transformation within a data flow for one or more functional expressions generated for each of one or more applications according to an embodiment. A diagram showing a process of identifying a pattern of transformation within a data flow for one or more functional expressions generated for each of one or more applications according to an embodiment. A diagram showing a system for generating functional rules according to an embodiment. A further diagram showing a system for generating functional rules according to an embodiment. A diagram showing an object diagram used for generating functional rules according to an embodiment. A diagram showing a process of generating a functional system based on one or more generated rules according to an embodiment. A diagram showing a system for identifying a pattern used in providing recommendations regarding a data flow based on information provided via an interface for functions in another language according to an embodiment. A diagram showing identifying a pattern used in providing recommendations regarding a data flow based on information provided via an interface for functions in another language according to an embodiment. A further diagram showing identifying a pattern used in providing recommendations regarding a data flow based on information provided via an interface for functions in another language according to an embodiment. A diagram showing a process of identifying a pattern used in providing recommendations regarding a data flow based on information provided via an interface for functions in another language according to an embodiment. A diagram showing the management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. A further diagram showing the management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. A further diagram showing the management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. A further diagram showing the management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. A further diagram showing the management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. A further diagram showing the management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. A further diagram showing the management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. A diagram showing the process of managing sampled or accessed data for lineage tracking across one or more layers, according to an embodiment.
DETAILED DESCRIPTION OF THE INVENTION
[0008] <id= The foregoing description will be apparent from the following description including the specification and claims, along with the accompanying drawings, in conjunction with other embodiments and their features. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of the various embodiments of the invention. It will be apparent, however, that the various embodiments may be practiced without these specific details. The following description including the specification and claims, along with the accompanying drawings, is not intended to be limiting.
[0009] First In accordance with various embodiments, described herein is a system (data artificial intelligence system, data AI system) used in a data integration or other computing environment that leverages machine learning (ML, data flow machine learning, DFML) for the management of data flow (DF) and the construction of composite data flow software applications (data flow applications, pipelines).
[0010] According to one embodiment, the system can provide support for the automatic mapping of complex data structures, data sets, or entities between one or more data sources or data targets, which are referred to as HUBs in some embodiments herein. The automatic mapping can be driven by metadata, schema, and statistical profiling of the data sets, and by using the automatic mapping, the source data sets or entities associated with the input HUB can be mapped to the target data sets or entities, or vice versa, to generate output data prepared in a format or organization (projection) used by one or more output HUBs.
[0011] According to one embodiment, the system has a visual environment used in the above system, which is referred to as a pipeline editor or Lambda Studio IDE in some embodiments herein. It may include a graphical user interface that provides an environment and software development components. This includes providing real-time recommendations based on the understanding of the meaning or semantics associated with the data for the execution of semantic actions on the data accessed from the input HUB.
[0012] According to certain embodiments, this system can provide a service for recommending actions and conversions for input data based on patterns identified from the functional decomposition of the data flow of a software application, which includes determining possible conversions to the data flow in a later application. The data flow can be decomposed into a model that describes data conversions, predicates, and business rules applied to the data, and attributes used within the data flow.
[0013] According to certain embodiments, this system can identify the types of data and data sets or entities associated with this schema by performing ontological analysis of the schema definition, and generate or update a model from a reference schema that includes an ontology defined based on the relationships between the data sets or entities and their attributes. The data flow can be analyzed using a reference HUB that includes one or more schemas, and further classified or recommendations can be made, such as conversion, enrichment, filtering, or cross-entity data fusion of the input data.
[0014] According to certain embodiments, this system provides a programming interface that is in some embodiments referred to as an other-language function interface herein. Through this interface, a user or third party can extend the functionality of the system by declaratively defining services, functional and business types, semantic actions, and patterns or predefined composite data flows based on functional and business types.
[0015] According to certain embodiments, the system can provide a data governance function. This can include, for example, for each slice of data temporally related to a particular snapshot, provenance information (where the particular data came from), lineage (how the data was obtained / processed), security (who was responsible for the data), classification (what the data is related to), impact (how much impact the data has on the business), retention (how long the data should persist), and validity (whether the data should be excluded / included for analysis / processing). Thus, these can be used in lifecycle determination and data flow recommendations.
[0016] According to certain embodiments, the system can be implemented as a service, for example, as a cloud service provided within a cloud-based computing environment, and can serve as a single control point for the analysis of data used in design, simulation, deployment, development, operation, and software applications. This includes enabling data input from one or more data sources (e.g., an input HUB in certain embodiments), providing a graphical user interface that allows a user to specify an application for the data, and scaling the data according to the destination, use, or target of the intended data (e.g., an output HUB in certain embodiments).
[0017] According to certain embodiments, as used herein, the terms "input" and "output" when used in connection with a particular HUB are provided only as labels reflecting the obvious flow of data in a particular use case or example, and are not intended to limit the type or function of a particular HUB.
[0018] For example, according to one embodiment, an input HUB that functions as a data source can also function as an output HUB or as a target that receives this data or other data, either simultaneously or at another time, and vice versa.
[0019] In addition, although some of the examples described herein for illustrative purposes show the use of input and output HUBs, according to one embodiment, in actual implementation examples, a data integration or other computing environment may include multiple such HUBs. At least some of these function as both input HUBs and / or output HUBs.
[0020] According to one embodiment, this system enables rapid development of software applications in large-scale, for example, cloud-based computing environments. In such an environment, the data model will evolve rapidly. Also, features such as search, recommendation, or suggestion are useful business conditions. In such an environment, by combining artificial intelligence (AI) and semantic search, users are given the authority to achieve more using their existing systems. For example, integrated interactions such as attribute-level mapping can be recommended based on an understanding of the user's interaction with metadata, data, and the system.
[0021] According to one embodiment, using this system can also propose complex cases. It is an interesting dimensional edge, for example, that can be used for information analysis and enables users to discover hitherto unknown facts within that data.
[0022] In some embodiments, the system provides a graphical user interface. This graphical user interface enables the automation of manual tasks (such as recommendations or suggestions), and by leveraging the integration of machine learning and probabilistic knowledge, it provides useful context for the user and enables discovery and semantic-driven solutions. It can be, for example , the creation of data warehouses, the scaling of services, the creation and enrichment of data, and the design and monitoring of software applications.
[0023] According to various embodiments, the system can include or utilize some or all of the following features.
[0024] Design-time system: A computing environment that enables the design, creation, monitoring, and management of software applications (such as data flow applications, pipelines, or Lambda applications), including the use of a data AI subsystem that provides, for example, machine learning capabilities, according to one embodiment.
[0025] Runtime system: A computing environment that enables the execution of software applications (such as data flow applications, pipelines, or Lambda applications), receives inputs from the design-time system, and provides recommendations to the design-time system.
[0026] Pipeline: Declaration means for defining a processing pipeline having multiple stages or semantic actions according to an embodiment. Each of the multiple stages or semantic actions corresponds to a function such as filtering, combining, enriching, transforming, or fusing the input data to create output data. For example, a data flow software application or a data flow application representing a data flow in DFML. According to an embodiment, the system can use the same codebase (e.g., together with the Spark runtime platform) for both batch (historical) data processing and real-time (streaming) data processing, and supports the construction of pipelines or applications that can operate on real-time data streams for real-time data analysis. Reprocessing of data due to pipeline design changes can be processed through rolling upgrades of the deployed pipeline. According to an embodiment, the pipeline can be provided as a Lambda application that can handle real-time data and batch data processing within different batch layers and real-time layers. Port, and also supports the construction of pipelines or applications that can operate on real-time data streams for real-time data analysis. Reprocessing of data due to pipeline design changes can be processed through rolling upgrades of the deployed pipeline. According to an embodiment, the pipeline can be provided as a Lambda application that can handle real-time data and batch data processing within different batch layers and real-time layers.
[0027] HUB: A data source or target (cloud or on-premises) containing a dataset or entity according to an embodiment. A data source that can be introspected, from which data can be consumed, or to which data can be published. The data source contains a dataset or entity, which has attributes, semantics, or relationships with other datasets or entities. Examples of HUBs include streaming data, telemetry, batch-based, structured or unstructured, or other types of data sources. Data can be received from a HUB associated with a source dataset or entity and mapped to a target dataset or entity of the same or another HUB.
[0028] System HUB: According to one embodiment, the System HUB can function as a knowledge source that stores other metadata, profile information, and data sets or entities that can be associated with other HUBs in the other metadata, and can also function like a regular HUB as a source or recipient of data to be processed. For example, in DFML, it is a central repository where metadata and system states are managed.
[0029] Data Set (Entity): A data structure including attributes (e.g., columns) according to one embodiment. This can be, for example, a database table, view, file, or API that is owned by or can be associated with one or more HUBs. According to one embodiment, it is one or more business entities, such as customer records, and can function as a semantic business type and be stored as a data component, such as a table within a HUB. A data set or entity can have relationships with other data sets or entities, along with attributes (e.g., columns within a table) and data types (e.g., string or integer). According to one embodiment, the system supports schema - independent processing of all types of data (e.g., structured, semi - structured, or unstructured data) during operations such as enrichment, preparation, conversion, model training, or scoring.
[0030] Data AI Sub - system: A component of a system, such as a data AI system according to one embodiment, responsible for machine learning and semantic - related functions. This includes one or more of providing search, profiling, recommendation engines, or support for automatic mapping. The Data AI Sub - system is connected to a design - time system, such as a software development component like Lambda Studio, via an event coordinator. It can support the operation of the [[[ID=]]], and provide recommendations based on the continuous processing of data by data flow applications (such as pipelines, Lambda applications), for example, recommend modifications to existing pipelines (such as pipelines) to utilize the data being processed. The data AI subsystem can analyze the amount of input data and continuously update the domain knowledge model. During the processing of a data flow application (such as a pipeline), for example, each stage of the pipeline can process the updated domain model and input from the user based on the recommended alternatives or options provided by the data AI subsystem, for example, accept or reject the recommended semantic actions.
[0031] Event coordinator: An event-driven architecture (EDA) component that operates between a design-time system and a runtime system, and coordinates events related to the design, creation, monitoring, and management of data flow applications (such as pipelines, Lambda applications). For example, the event coordinator receives a published notification of data (such as new data conforming to a known data type) from a HUB, normalizes the data from this HUB, and provides the normalized data for use by a set of subscribers, for example, by a pipeline or other downstream consumer. Also, the event coordinator can receive notifications of state transactions within the system for use in lineage tracking or logging, including the creation of temporary slices and schema evolution. It coordinates events related to the design, creation, monitoring, and management of data flow applications (such as pipelines, Lambda applications). For example, the event coordinator receives a published notification of data (such as new data conforming to a known data type) from a HUB, normalizes the data from this HUB, and provides the normalized data for use by a set of subscribers, for example, by a pipeline or other downstream consumer. Also, the event coordinator can receive notifications of state transactions within the system for use in lineage tracking or logging, including the creation of temporary slices and schema evolution.
[0032] Profiling: An operation that extracts a sample of data from a HUB, profiles the data provided by this HUB, as well as the data sets or entities and attributes within this HUB, determines metrics related to the sampling of the HUB, and updates the metadata related to the HUB to reflect the data profile in this HUB.
[0033] Software Development Component (Lambda Studio): A design-time system tool according to an embodiment that provides a graphical user interface, enabling a user to create, monitor, and manage the life cycle of a Lambda application or pipeline as a pipeline of semantic actions. For example, a graphical user interface (UI, GUI) or studio that allows a user to design pipelines and Lambda applications. By providing a graphical user interface, a user can create, monitor, and manage the life cycle of a Lambda application or pipeline as a pipeline of semantic actions. For example, a graphical user interface (UI, GUI) or studio that allows a user to design pipelines and Lambda applications. The interface (UI, GUI) or studio.
[0034] Semantic Action: According to an embodiment, a data transformation function, such as a relational algebra operation. An action that a data flow application (such as a pipeline, Lambda application) can perform on a data set or entity within a HUB for projection to another entity. A semantic action functions as a different model or a higher-order function available to the HUB that can receive a data set input and generate a data set output. A semantic action may include a mapping. A semantic action may include a mapping that can be continuously updated, for example, by a data AI subsystem, in response to the processing of data as part of a pipeline or Lambda application.
[0035] Mapping: A recommended mapping of semantic actions between a first (e.g., source) dataset or entity, such as provided by a data AI subsystem in certain embodiments and accessible through a design-time system, e.g., via Lambda Studio which is a software development component, and another (e.g., target) dataset or entity. For example, the data AI subsystem can provide automatic mapping as a service. The automatic mapping can be driven by metadata, schema, and statistical profiling of datasets based on machine learning analysis of metadata associated with the HUB or data inputs. It is a recommended mapping of semantic actions between a first (e.g., source) dataset or entity, such as provided by a data AI subsystem in certain embodiments and accessible through a design-time system, e.g., via Lambda Studio which is a software development component, and another (e.g., target) dataset or entity. For example, the data AI subsystem can provide automatic mapping as a service. The automatic mapping can be driven by metadata, schema, and statistical profiling of datasets based on machine learning analysis of metadata associated with the HUB or data inputs.
[0036] Pattern: A pattern of semantic actions that a data flow application (e.g., pipeline, Lambda application) can execute in certain embodiments. By using templates, it is possible to provide a definition of a pattern that can be reused by other applications. The logical flow of data and related conversions associated with normal business semantics and processes.
[0037] Policy: A set of policies in certain embodiments that control how a data flow application, e.g., a pipeline, Lambda application, is scheduled, which users or components can access which HUBs and semantic actions, and how data should be aged or other considerations. Configuration settings that define how, for example, a pipeline should be scheduled, executed, or accessed.
[0038] Application design services: According to one embodiment, it provides services such as data flow (e.g., pipeline), services for Lambda applications (e.g., verification, compilation, packaging), and other services (e.g., DFML services such as UI, system facade) including deployment to these services. It verifies pipelines (e.g., Lambda applications) in software development components (e.g., Lambda Studio, e.g., its input and output), persists the pipelines, controls the deployment of the pipelines and Lambda applications to the system (e.g., to a Spark cluster) for execution, and then is a design-time system component that can be used for the management of the application life cycle or state. Edge layer: A layer that transfers data to a scalable input / output layer as a data collection (e.g., storage and transfer) layer according to one embodiment. It can receive data through, for example, a gateway accessible to the Internet, and is a runtime system component that includes one or more nodes with security and other features that support, for example, secure access to a data AI system.
[0039]
[0040] Compute layer: An application execution and data processing layer (e.g., Spark) according to one embodiment. It functions as a distributed processing component, such as a Spark cloud service, a cluster of compute nodes, an aggregate of virtual machines, or other components or nodes, and is a runtime system component used, for example, for the execution of pipelines and Lambda applications. In a multi-tenant environment, the nodes within the compute layer can be allocated to tenants for use in the execution of pipelines or Lambda applications by these tenants.
[0041] Scalable Input / Output (I / O) Layer: According to an embodiment, it provides a scalable data persistence and access layer configured as topics and partitions (e.g., Kafka). It enables data to be moved within the system and shared among various components of the system, providing a queue or other logical storage, such as a runtime system component in a Kafka environment. In a multi-tenant environment, the scalable I / O layer can be shared among multiple tenants.
[0042]
[0043] Data Lake: A repository for the persistence of information from a system HUB or other components according to an embodiment. Usually, it is a repository of data, such as in DFML, which is typically normalized or processed by pipelines, Lambda applications, etc., and consumed by other pipelines, Lambda applications, or publish layers.
[0044]
[0045] Data Flow Machine Learning (DFML): A data integration and data flow management system that assists in building composite data flow applications by leveraging machine learning (ML) according to an embodiment.
[0046] Metadata: Definitions and descriptions underlying datasets or entities, attributes, and their relationships according to an embodiment. It can also be descriptive data about artifacts in DFML, for example.
[0046] Data: Application data represented by datasets or entities according to an embodiment. These can be in batch or stream form. For example, customers, orders, or products.
[0047] System Facade: A unified API layer for accessing the functions of, for example, a DFML event-driven architecture, according to one embodiment.
[0048] Data AI Subsystem: Provides artificial intelligence (AI) services, according to one embodiment, including but not limited to, for example, search, auto-map, recommendation, or profiling.
[0049] Streaming Entity: A continuous input of data and near-real-time processing and output conditions that can support an emphasis on the speed of data, according to one embodiment.
[0050] Batch Entity: An ingestion of data that can be characterized by an emphasis on volume and that is performed scheduled or in response to a request, according to one embodiment. Chon.
[0051] Data Slice: A partition of data that is typically marked by time, according to one embodiment.
[0052] Rule: Represents a directive that affects artifacts in, for example, DFML, according to one embodiment, such as a data rule, relationship rule, metadata rule, and composite or hybrid rule.
[0053] Recommendation (Data AI): A recommended series of actions, according to one embodiment, typically represented by one or more semantic actions or fine-grained directives to assist in the design of, for example, a pipeline, Lambda application.
[0054] Search (Data AI): A semantic search in, for example, DFML, characterized by context and the user's intent, to return relevant artifacts, according to one embodiment.
[0055] Auto Map (Data AI): A certain type of recommendation in some embodiments that shortlists candidate source or target data sets or entities to be tried in a data flow.
[0056] Data Profiling (Data AI): An aggregate of several metrics, such as minimum value, maximum value, interquartile range, or scatter, that characterize data in attributes belonging to a data set or entity, in some embodiments.
[0057] Action Parameter: A reference to a data set that is the execution target of a semantic action, in some embodiments. For example, parameters for an equi-join in a pipeline, Lambda application, etc., as an example.
[0058] Foreign Language Function Interface: A mechanism that registers and invokes services (and semantic actions), for example, as part of a DFML Lambda application framework, in some embodiments. Can be used to extend functions or transformation vocabularies in DFML, for example.
[0059] Service: An aggregate of semantic actions that can be characterized by data integration stages (such as preparation, discovery, transformation, or visualization), for example, unaddressed artifacts in DFML, in some embodiments.
[0060] Service Registry: A repository of services, their semantic actions, and other instance information, in some embodiments.
[0061] Data Lifecycle: A stage in the use of data within DFML, for example, that starts with ingestion and ends with publication, in some embodiments.
[0062] Metadata Harvesting: Collecting metadata and sample data for profiling, typically after registration of the HUB, according to an embodiment.
[0063] Pipeline Normalization: Normalizing data in a specific format that facilitates consumption by, for example, pipelines, Lambda applications, according to an embodiment.
[0064] Monitoring: Identifying, measuring, and evaluating the execution of, for example, pipelines, Lambda applications, according to an embodiment.
[0065] Ingest: Taking in data through an edge layer into, for example, DFML, according to an embodiment.
[0066] Publish: Writing data from, for example, DFML to a target endpoint, according to an embodiment.
[0067] Data AI System FIG. 1 is a diagram showing a system for providing data flow artificial intelligence according to an embodiment.
[0068] As shown in FIG. 1, according to an embodiment, a system, such as data AI system 150, can provide one or more services for processing and transforming data, such as business data, consumer data, and enterprise data, which includes the use of machine learning processing used with various computing assets such as databases, cloud data warehouses, storage systems, or storage services.
[0069] According to an embodiment, the computing asset can be cloud-based, enterprise-based, on-premises, or agent-based. The various elements of this system can be connected by one or more networks 130.
[0070] According to one embodiment, the system may include one or more input HUBs 110 (e.g., data sources, data sources) and an output HUB 180 (e.g., data targets, data targets).
[0071] According to one embodiment, each input HUB, e.g., HUB 111, may include a plurality of (source) data sets or entities 192.
[0072] According to one embodiment, examples of input HUBs may include a database management system (DB, DBMS) 112 (e.g., an on-line transaction processing system (OLTP), a business intelligence system, or an on-line analytical processing system (OLAP)). In such an example, the data provided by a source such as a data management system may be structured data or semi-structured data.
[0073] According to one embodiment, other examples of input HUBs may include a cloud store / object store 114 (e.g., AWS S3 or another object store), a data cloud 116 (e.g., a third-party cloud), a streaming data source 118 (e.g., AWS Kinesis or another streaming data source), or other input sources 119, which may be an object bucket or clickstream source having unstructured data.
[0074] According to one embodiment, the input HUB may include a data source. The data source receives data from, for example, an Oracle Big Data Prep (BDP) service.
[0075] According to certain embodiments, the system may include one or more output HUBs 180 (e.g., output destinations). Each output HUB, e.g., HUB 181, may include a plurality of (target) data sets or entities 194.
[0076] According to certain embodiments, examples of output HUBs may include public cloud 182, data cloud 184 (e.g., AWS and Azure), on-premises cloud 186, or other output targets 187. The data output provided by this system can be generated for data flow applications (e.g., pipelines, Lambda applications) accessible in the output HUB.
[0077] According to certain embodiments, an example of a public cloud may include, for example, Oracle Public Cloud. This may include, for example, Big Data Prep cloud service, Exadata cloud service, Big Data Discovery cloud service, and Business Intelligence cloud service.
[0078] According to certain embodiments, the system can be implemented as a unified platform for streaming and on-demand (batch) data processing that is delivered to users as a service (e.g., as software as a service), providing scalable multi-tenant data processing for multiple input HUBs. The data is analyzed in real time using machine learning techniques and visual insights and monitoring provided by a graphical user interface as part of the service. Data sets can be fused from multiple input HUBs for output to the output HUB. For example, through the data processing services provided by this system, data can be generated for a data warehouse and populated into one or more output HUBs.
[0079] According to one embodiment, the system can provide declarative and programming topologies for data transformation, enrichment, routing, classification, and blending, and can include a design-time system 160 and a runtime system 170. A user can create an application, such as a data flow application (e.g., a pipeline, a Lambda application) 190, which is designed to execute data processing.
[0080] According to one embodiment, the design-time system enables a user to design a data flow application, define a data flow, and define data for data flow processing. For example, the design-time system can provide a software development component 162 (referred to herein as Lambda Studio in one embodiment) that provides a graphical user interface for creating a data flow application. This is possible.
[0081] For example, according to one embodiment, a user can use the software development component to specify an input HUB and an output HUB for creating a data flow for an application. The graphical user interface can indicate an interface for services for data integration. Thereby, the user can create, operate, and manage a data flow for an application. This includes the function of dynamically monitoring and managing a data flow pipeline, which is, for example, observing a data lineage and performing forensic analysis.
[0082] According to one embodiment, the design-time system can also include an application design service 164 for deploying a data flow application to the runtime system.
[0083] According to an embodiment, the design-time system may also include one or more system hubs 166 (e.g., a metadata repository) for storing metadata to process data flows. The one or more system hubs can store samples of data, such as samples of data types including functional and business data types. The By using the information within the system hub, one or more of the techniques disclosed herein can be executed. The data lake 167 component can operate as a repository for the persistence of information from the system hub.
[0084] According to an embodiment, the design-time system may also include a data artificial intelligence (AI) subsystem 168 that performs operations for data artificial intelligence processing. The operations may include using ML techniques, such as search and retrieval. The data AI subsystem can sample data to generate metadata for the system hub.
[0085] According to an embodiment, the data AI subsystem can perform schema object analysis, metadata analysis, sample data, correlation analysis, and classification analysis for each input hub. By continuously performing on the input data, the data AI subsystem can provide rich data to data flow applications and can provide recommendations, insights, and type inference to, for example, pipelines, Lambda applications.
[0086] According to an embodiment, the design-time system enables a user to create policies, artifacts, and flows that define the functional requirements of a use case.
[0087] For example, according to one embodiment, at design time, the system can create a HUB, ingest data, and define an ingestion policy by providing a graphical user interface. The ingestion policy may be time-based or in response to requests from the associated data flow. When an input HUB is selected, the source can be profiled by sampling data from this input HUB. This can be, for example, executing a metadata query, obtaining samples, and obtaining user-defined inputs. The profile can be stored in the system HUB. The graphical user interface enables combining multiple sources to define a data flow pipeline. This can be done by creating a script or by using a guided editor. The guided editor allows visualizing the data at each step. The graphical user interface can provide access to a recommendation service that suggests how to modify, enrich, or combine, for example, a data cloud.
[0088] According to one embodiment, during design time, the application design service can suggest a structure suitable for analyzing the obtained content. The application design service can suggest treatments and related dimension hierarchies by using a knowledge service (functional classification). Once this is done, the design-time system can obtain the blended data from the previous pipeline and recommend the data flow necessary to populate the dimension target structure. Based on the dependency analysis, the target schema can also be loaded / refreshed by deriving and generating an orchestration flow. In the case of a forward engineering use case, the design-time system can also host the target structure and create the target schema by generating a HUB.
[0089] According to one embodiment, the runtime system can execute processing during the execution time of services for data processing.
[0090] According to one embodiment, in the runtime or operating mode, the policies and flow regulations created by the user are applied and / or executed. For example, such processing may include calling ingestion, transformation, model, and publish services to process data within a pipeline. and can include calling ingestion, transformation, model, and publish services to process data within a pipeline.
[0091] According to one embodiment, the runtime system may include an edge layer 172, a scalable input / output (I / O) layer 174, and a distributed processing system or computing layer 176. At runtime (e.g., when data is ingested from one or more input HUBs 110), the edge layer can receive data for events. Data is generated by this event.
[0092] According to one embodiment, the event coordinator 165 operates between the design-time system and the runtime system to coordinate events related to the design, creation, monitoring, and management of data flow applications (e.g., pipelines, Lambda applications).
[0093] According to one embodiment, the edge layer sends data to the scalable input / output layer. The scalable input / output layer routes this data to the distributed processing system or computing layer.
[0094] According to one embodiment, the distributed processing system or computing layer can process data for output by implementing the pipeline process (per tenant). The distributed processing system can be implemented using, for example, Apache Spark and Alluxio. Data can be sampled into a data lake and then output to an output HUB. The distributed processing system can communicate with a scalable input / output layer to initiate and process the data.
[0095] According to one embodiment, a data AI system including some or all of the above components can be provided on or executed by one or more computers including, for example, one or more processors (CPUs), memory, and a persistent storage device (198).
[0096] Event-driven architecture As described above, according to one embodiment, the system may include an event-driven architecture (EDA) component or an event coordinator. This operates between the design-time system and the run-time system to coordinate events related to the design, creation, monitoring, and management of data flow applications (e.g., pipelines, Lambda applications).
[0097] FIG. 2 is a diagram showing an event-driven architecture including an event coordinator used in a system according to one embodiment.
[0098] As shown in FIG. 2, according to one embodiment, the event coordinator may include an event queue 202 (e.g., Kafka), an event bootstrap service 204 (e.g., ExecutorService), and an event configuration publisher / event consumer 206 (e.g., DBCS).
[0099] According to one embodiment, events received by a system facade 208 (e.g., an event API extension) are sent to one or more event brokers 210, e.g., Kafka consumers It can be transmitted by the bus to various components of the system. For example, by way of example, data and / or events such as external data 212 (such as S3, OSCS, or OGG data), and / or inputs from the graphical user interface 214 (such as a browser or DFML UI) can be transmitted via the event coordinator to other components, such as the above-mentioned application runtime 216, data lake, system HUB, data AI subsystem, application design service, and / or ingest 220, publish 230, scheduling 240, or other components. For example, inputs from the graphical user interface 214 (such as a browser or DFML UI) can be transmitted via the event coordinator to other components, such as the above-mentioned application runtime 216, data lake, system HUB, data AI subsystem, application design service, and / or ingest 220, publish 230, scheduling 240, or other components.
[0100] According to one embodiment, the event broker may be configured as a consumer of stream events. The event bootstrap can process events on behalf of registered subscribers by starting a plurality of configured event brokers. For processing a given event, each event broker delegates the processing of the event to a registered callback endpoint. The event coordinator enables registration of event types, registration of eventing entities, registration of events, and registration of subscribers. Table 1 provides an example of various event objects including publish events and subscribe events.
[0101]
Table 1
[0102] Event Type According to an embodiment, an event type defines a state change of an event that is important to the system. It is, for example, the creation of a HUB, the modification of a data flow application such as a pipeline, a Lambda application, the ingestion of data for a dataset or entity, or the publishing of data to a target HUB, etc. An example of a data format and examples of various event types are shown below and in Table 2.
[0103] [Table 2]
[0104] Eventing entity According to an embodiment, an eventing entity may be a publisher and / or subscriber of an event. For example, an eventing entity can be registered to publish one or more events and / or can be a consumer of one or more events. This includes the registration of endpoints or callback URLs that are used to notify or send an acknowledgement for a publish and delegate the processing of subscribed events. Examples of eventing entities may include metadata services, ingestion services, system HUB artifacts, and pipelines, Lambda applications. An example of a data format and examples of various eventing entities are shown below and in Table 3.
[0105] [Table 3]
[0106] Event According to an embodiment, an event is an instance of an event type associated with an eventing entity registered as a publisher and may have subscribers (eventing entities). For example, a metadata service can register a HUB creation event for publishing, and publish one or more of the event instances (one instance per created HUB) for this event. Examples of various events are shown in Table 4.
[0107]
Table 4
[0108] Example According to an embodiment, the following example shows the creation of an event type, the registration of a publish event, the registration of subscribers, the publishing of an event, the retrieval of an event type, the retrieval of a publisher for an event type, and the retrieval of subscribers for an event type.
[0109]
Number
[0110] This returns a universally unique ID (UUID), for example, "8e8 7039b-a8b7-4512-862c-fdb05b9b8888". The eventing object is the system Events within the mu can be published or subscribed to. For example, service endpoints such as ingest services, metadata services, and application design services can be publish or subscribe events along with static endpoints for acknowledgement, notification, error, or processing. DFML artifacts (e.g., DFMLEntity, DFMLLambdaApp, DFMLHub) can also be registered as eventing objects, and instances of these types can publish or subscribe to events without registration as an eventing object.
[0111]
Number
[0112] The following example registers DFMLLambdaApps (type) as eventing objects.
[0113]
Number
[0114] For eventing entities of HUB, Entity, and LambdaApp types, <publisherurl>can be annotated to the REST endpoint URL, and the event-driven architecture derives the actual URL by replacing the DFML artifact instance URL. For example, if the notificationEndpointURL is http: / / den00tnk:9021 / <publisherurl>Registered as / notification and identified as part of the message If the publisher URL specified as part of the message is hubs / 1234 / entities / 3456, the URL called for notification will be http: / / den00tnk:9021 / hubs / 1234 / entities / 3456 / notification . The POST returns a UUID, for example, "185cb819-7599-475b-99a7-65e0bd2ab947".
[0115] Registration of Publish Events According to an embodiment, a publish event can be registered as follows.
[0116]
Number
[0117] The above eventType is the UUID returned for the registration of the event type DATA_INGESTED, and the above publishingEntity is the DFMLEntity type registered as the eventing object. This registration returns a UUID, for example, "2c7a4b6f-73ba-4247-a07a-806ef659def5".
[0118] Registration of Subscribers According to an embodiment, a subscriber can be registered as follows.
[0119]
Number
[0120] The UUID returned from the publish event registration is used as a path segment for the subscriber registration.
[0121]
Number
[0122] The above publisherURL and publishingObjectType are instances and types of publisher objects. Here, note that a data flow (e.g., Lambda) application subscribes to DATA_INGESTED events from entity / hubs / 1234 / entities / 3456 in the specification of URI / lambdaApps / 123456. This registration returns a UUID , for example, "1d542da1-e18e-4590-82c0-7fe1c55c5bc8".
[0123] Publishing of Events According to an embodiment, events can be published as follows.
[0124]
Number
[0125] The above publisherURL is used when the publish object is one of DFMLEntity, DFMLHub, or DFMLLambdaApps, and is used to check the instance of the eventing object that publishes the message that the subscriber participates in. Also, the publisher URL is used to derive the notification URL when the subscriber successfully processes the message. This publish returns the message body that was part of the published event.
[0126] Obtaining Event Types According to an embodiment, event types can be obtained as follows.
[0127]
Number
[0128] Obtain a publisher for the event type
[0129]
Number
[0130] Obtain a subscriber for the event type According to an embodiment, the subscriber for the event type can be obtained as follows.
[0131]
Number
[0132] The above description has been shown as an example for explaining a specific embodiment of an event coordinator, an event type, an eventing entity, and an event. According to other embodiments, by using other types of EDA to provide communication within a system that functions between a design-time system and a run-time system, events related to the design, creation, monitoring, and management of data flow applications can be coordinated, and other types of event types, eventing entities, and events can be supported.
[0133] Data Flow Machine Learning (DFML) Flow As described above, according to various embodiments, the present system can be used in a data integration or other computing environment that utilizes machine learning (ML, Data Flow Machine Learning, DFML) used for the management of data flow (data flow, DF) and the construction of composite data flow software applications (such as data flow applications, pipelines, Lambda applications).
[0134] FIG. 3 is a diagram showing steps within a data flow according to an embodiment. As shown in FIG. 3, according to an embodiment, the processing of the DFML data flow 260 may include a plurality of steps including an ingest step 262. In the ingest step, data can be ingested from various sources such as Salesforce (SFDC), S3, or DBaaS.
[0135] In the data preparation step 264, the ingested data can be prepared, for example, by deduplication, normalization, or enrichment.
[0136] In the transformation step 266, the system can transform the data by performing one or more merges, filters, or lookups of data sets.
[0137] In the model step 268, one or more models are generated along with the mappings for the models.
[0138] In the publish step 270, the system can publish the model, specify policies and schedules, and populate the target data structure.
[0139] According to an embodiment, the system supports the use of the search / recommend function 272 throughout each of its data preparation step, transformation step, and model step. The user can interact with the system through a set of well-defined services with a limited scope of functions within the data integration framework. This set of services defines the logical view of the system. For example, in the design mode, the user can create flows, artifacts, and policies that define the functional requirements of a particular use case.
[0140] FIG. 4 is a diagram showing an example of a data flow including a plurality of sources according to an embodiment.
[0141] As shown in data flow 280 as an example shown in FIG. 4, according to one embodiment, what is needed here is SFDC and FACS (Fusion Apps Cloud Service) content from a plurality of sources 282 shown, along with some files within OSCS (Oracle Storage Cloud Service), is imported, blended so that this information can be used for the analysis of the desired content, target cubes and dimensions are derived, the blended content is mapped to the target structure, and this content is made available for use in an Oracle Business Intelligence Cloud Service (BICS) environment with a dimensional model, which includes ingest, transform 266A / 266B, model, orchestrate 292, and deploy 294 steps.
[0142] The example shown is provided to illustrate the techniques described herein, and the functions described herein are not limited to use with these particular data sources.
[0143] According to one embodiment, in the ingest step, to access and ingest SFDC content, a HUB is created in the data lake to receive this content. This can be accomplished, for example, by selecting an SFDC adapter of the relevant access mode (JDBC, REST, SOAP), creating a HUB, naming it, and defining an ingest policy based on time or in response to the requirements of the relevant data flow.
[0144] According to one embodiment, a similar process can be performed for the other two sources. The difference is that in the case of the OSCS source, the schema may not be known at the start and instead can be obtained by some means (such as metadata query, sampling, or user definition).
[0145] According to one embodiment, the source of the data can be further investigated by optionally profiling it. This can help derive recommendations later in the integration flow.
[0146] According to one embodiment, the next step is to define how to combine the separate sources around the central item. This is typically the basis (facts) of the analysis and can be achieved by defining a data flow pipeline. This can be done directly by creating a pipeline domain specific language (DSL) script or using a guided editor. In the case of a guided editor, the user can know the effect on the data at each step and use a recommendation service that suggests how the data can be modified, enriched, combined, etc.
[0147] At this point, the user can request that the system suggest a structure suitable for analyzing the obtained content. For example, according to one embodiment, the system can suggest countermeasures and related dimensional hierarchies by using a knowledge service (functional classification). Once this is done, the system can recommend the data flow necessary to obtain the blended data from the previous pipeline and populate the dimensional target structure. This also loads / refresh the target schema by deriving and generating an orchestration flow based on dependency analysis.
[0148] According to one embodiment, the system then generates a HUB that hosts the target structure and associates this with a DBCS that generates the data definition language (DDL) necessary to create the target schema via the adapter, for example, it can deploy XDML or some form that BICS can use to generate the models necessary to access the newly created schema. This can be populated by executing an orchestration flow and triggering an exhaust service can be done.
[0149] FIG. 5 is a diagram showing an example of the use of a data flow having a pipeline according to an embodiment.
[0150] As shown in FIG. 5, according to one embodiment, the system allows the user to describe the processing of data when constructed / executed 304 as an application by defining a pipeline 302 that represents a data flow including pipeline steps S1 to S5 in this example.
[0151] For example, according to one embodiment, the user can call ingest, transform, model, and publish services, or other services such as policy 306, execution 310, or persistence service 312 to process the data within the pipeline. The user can also identify a unified flow that can integrate the relevant pipelines by defining a solution (i.e., control flow). Typically, a solution models the loading of a complete use case such as a sales cube and related dimensions.
[0152] Data AI System Component According to one embodiment, the adapter enables connection to various endpoints and ingestion of data from various endpoints and is specific to the application or source type.
[0153] According to one embodiment, the system may include a predetermined set of adapters. Some of these adapters can utilize other SOA adapters and enable additional adapters to be registered with the framework. There may be two or more adapters for a given connection type. In that case, the ingest engine selects the most suitable adapter based on the HUB's connection type configuration.
[0154] FIG. 6 is a diagram showing an example of the use of an ingest / publish engine and an ingest / publish service according to one embodiment.
[0155] As shown in FIG. 6, according to one embodiment, pipeline 334 can access ingest / publish engine 330 via ingest / publish service 332. In this example, pipeline 334 is designed to ingest data (e.g., sales data) from an input HUB (e.g., SFDC HUB1) at 336, transform the ingested data at 338, and publish this data to an output HUB (e.g., Oracle HUB) at 340.
[0156] According to one embodiment, the ingest / publish engine supports multiple connection types 331. One of these types, connection type 342, is associated with one or more adapters 344 that provide access to the HUB.
[0157] For example, as shown in the example of FIG. 6, according to one embodiment, the SFDC connection type 352 can be associated with the SFDC-Adp1 adapter 354 and the SFDC-Adp2 adapter 356 that provide access to the SFDC HUBs 358 and 359, the ExDaaS connection type 362 can be associated with the ExDaaS-Adp adapter 364 that provides access to the ExDaas HUB 366, and the Oracle connection type 372 can be associated with the Oracle Adp adapter 374 that provides access to the Oracle HUB 376.
[0158] Recommendation engine According to one embodiment, the system may include a recommendation engine or a knowledge service that operates as a specialized filtering system that predicts / suggests the most relevant of several possible actions that can be performed on the data.
[0159] According to one embodiment, by chaining recommendations, the user can be guided through these recommendations to more easily achieve a given end goal. For example, the recommendation engine can guide the user through a set of steps when converting a dataset into a data cube and publishing it to a target BI system.
[0160] According to an embodiment, the recommendation engine utilizes three aspects: (A) business type classification, (B) functional type classification, and (C) knowledge base. Ontology management and query / search functions for datasets or entities can be provided, for example, by a shared ontology derived from YAGO3 with query APIs, MRS, and audit repositories. Business entity classification can be provided, for example, by an ML pipeline-based classification that identifies business types. Functional type classification can be provided, for example, by a deductive rule-based functional type classification. Action recommendations can be provided, for example, by an inductive rule-based data preparation, transformation, model, dependencies, and related recommendations.
[0161] Classification service According to an embodiment, the system provides a classification service that can be categorized into business type classification and functional type classification, which will be further described below for each.
[0162] Business type classification According to an embodiment, the business type of an entity is its phenotype. The observable characteristics of the individual attributes within an entity are as important as the definition in identifying the business type of the entity. The classification algorithm uses a general definition of the dataset or entity, but can also classify the business type of the dataset or entity using a model built with data.
[0163] For example, according to an embodiment, a dataset ingested from a HUB can be classified as one of the existing business types (seeded from the main HUB) known to the system, or added as a new type if it cannot be classified into an existing business type.
[0164] According to one embodiment, the business type classification is utilized when making recommendations, either through inductive reasoning (from the transformations defined for similar business types within the pipeline) or based on simple propositions derived from the classification root entity.
[0165] Describe the overview. According to one embodiment, the classification process is described by the following set of steps. These steps include the step of ingesting and seeding from the main (training) hub, the step of building a model, calculating column statistics, and registering these for use in classification, the step of classifying a dataset or entity from a newly added hub, including creating a profile / calculating column statistics, the step of classifying a dataset or entity to provide a shortlist of entity models to use based on structure and column statistics, and the step of classifying a dataset or entity including multi-class classification and calculating / predicting using the model.
[0166] FIG. 7 is a diagram showing the process of ingestion and training from the HUB according to one embodiment.
[0167] As shown in FIG. 7, according to one embodiment, data from the HUB 382 (for example, the RelatedIQ source in this example) can be read by the recommendation engine 380 as a dataset 390 (for example, a Resilient Distributed Dataset :RDD). This dataset includes, in this example, an account dataset 391, an event dataset 392, a contact dataset 393, a list dataset 394, and a user dataset 395.
[0168] According to one embodiment, a plurality of type classification tools 400 can be used along with the ML pipeline 402. For example, GraphX 404, Wolfram / Yago 406, and / or MLlib Statistics 408 can be used to seed the knowledge graph 440 with entity metadata (training or seed data) when the HUB is first registered.
[0169] According to one embodiment, a dataset or entity metadata and data are ingested from the source HUB and stored in the data lake. During model generation 410, entity metadata (attributes and relationships with other entities) is used for the generation of the model 420 and the knowledge graph, for example, through FP-growth logistic regression 412. The knowledge graph represents all datasets or entities, which in this example represents events 422, accounts 424, contacts 426, and users 428. As part of the seed, a regression model is constructed using the dataset or entity data, and attribute statistics (minimum, maximum, average, or probability density) are calculated.
[0170] FIG. 8 is a diagram showing a model construction process according to an embodiment. As shown in FIG. 8, according to one embodiment, when running, for example, within the Spark environment 430, Spark MLlib Statistics can be used to calculate column statistics added as attribute properties to the knowledge graph. The calculated column statistics can be used together with other datasets or entity metadata to shortlist the entities for which the regression model is used in testing new entities for classification.
[0171] FIG. 9 is a diagram showing a process for classifying a dataset or entity from a newly added HUB according to an embodiment.
[0172] As shown in FIG. 9, according to one embodiment, when a new HUB, in this example Oracle HUB 442, is added, the data sets or entities provided from this HUB, such as party information 444 and customer information 446, are classified as parties 448 by the model based on the previously created training or seed data.
[0173] For example, according to one embodiment, column statistics are calculated from the data of a new data set or entity, and a set of predicates representing a subgraph of the entity is created using this information together with other metadata that can be used as part of the ingest.
[0174] According to one embodiment, the calculation of column statistics is useful in the maximum likelihood estimation (MLE) method, while the subgraph is useful in the regression model of the data set. A set of graph predicates generated for the new entity is used to shortlist candidate entity models for testing and classifying the new entity. FIG. 10 is a diagram further showing the process of classifying a data set or entity from a newly added HUB according to one embodiment.
[0175] As shown in FIG. 10, according to one embodiment, a predicate representing a subgraph of a new data set or entity to be classified is compared 450 with a similar subgraph representing a data set or entity that is already part of the knowledge graph. The ranking of matching entities based on the probability of match is used when shortlisting entity models for use in testing for classifying the new entity.
[0176] FIG. 11 is a diagram further showing the process of classifying a data set or entity from a newly added HUB according to one embodiment.
[0177] FIG. 11 is a diagram further showing the process of classifying a data set or entity from a newly added HUB according to one embodiment.
[0178] As shown in FIG. 11, according to one embodiment, a regression model of the matching data sets or entities included in the shortlist is used to test the data from new data sets or entities. By extending the ML pipeline to include additional classification methods / models, the accuracy of the process can be improved. The classification service classifies new entries 452 if there is a match within an acceptable threshold, for example, a probability higher than 0.8 in this example. Otherwise, this data set or entity can be added to the knowledge graph as a new business type. Also, the user can verify the classification by accepting or rejecting the results.
[0179] Functional classification According to one embodiment, the functional type of an entity is its genotype. The functional type can also be described as an interface through which transformation actions are defined. For example, a binding transformation or a filter is defined by a functional type such as a relational entity in this case. In summary, all transformations are defined by a functional type as a parameter.
[0180] FIG. 12 is a diagram showing an object diagram used for functional classification according to one embodiment.
[0181] As shown in FIG. 12 by object diagram 460, according to one embodiment, the system can describe a general case (dimensions, levels, or cubes in this example) through a set of rules that serve as a criterion when evaluating a data set or entity to identify its functional type.
[0182] For example, according to one embodiment, a multi-dimensional cube can be described by its measurement attributes and dimensions. These can each be defined by themselves by their types and other characteristics. The rule engine evaluates business-type entities and annotates their functional types based on the evaluation.
[0183] FIG. 13 is a diagram showing an example of a dimension functional classification according to an embodiment. As shown in the hierarchy of the functional classification 470 as an example shown in FIG. 13, according to one embodiment, a level can be defined, for example, by its dimension and level attributes.
[0184] FIG. 14 is a diagram showing an example of a cube functional classification according to an embodiment. As shown in the hierarchy of the functional classification 480 as an example shown in FIG. 14, according to one embodiment, a cube can be defined, for example, by its measurement attributes and dimensions.
[0185] FIG. 15 is a diagram showing an example of the use of a functional classification for evaluating the functional type of a business entity according to an embodiment.
[0186] As shown in FIG. 15, in this example 490, according to one embodiment, a sales dataset must be evaluated as a cube functional type by the rule engine. Similarly, products, customers, and time must be evaluated as dimensions and levels (e.g., age group, gender).
[0187] According to an embodiment, the function type of the entity and the rules for specifying the dataset or entity elements in this specific example are shown below. This includes several rules that can be specified for the evaluation of the same function type. For example, a column of type "Date" can be considered as a dimension regardless of whether there is a reference to the parent-level entity. Similarly, postal code, gender, and age may only require data rules to specify them as dimensions.
[0188]
Number
[0189] FIG. 16 is a diagram showing an object diagram used for function conversion according to an embodiment. As shown in FIG. 15, in this example 500, according to an embodiment, the conversion function can be defined by a function type. The business entity (business type) is annotated as a function type, which by default includes that the composite business type is a function type "entity".
[0190] FIG. 17 is a diagram showing the operation of a recommendation engine according to an embodiment. As shown in FIG. 17, according to an embodiment, the recommendation engine generates recommendations that are a set of actions defined by a business type. Each action is a command that requires applying a conversion to a dataset.
[0191] According to an embodiment, the recommendation context 530 includes metadata that extracts the source of the recommendation and identifies a set of propositions that generated the recommendation. With this context, the recommendation engine can learn and prioritize recommendations based on the user's response.
[0192] According to one embodiment, the target entity deduction / mapping unit 512 makes transformation recommendations that facilitate mapping the current data set to a target using the target definition (and a classification service that annotates data sets or entity and attribute business types). This is common when a user starts with a known target object (e.g., a sales cube) and constructs a pipeline to instantiate the cube.
[0193] According to one embodiment, the template (pipeline / solution) 514 defines a reusable set of pipeline steps and transformations to obtain a desired final result. For example, the template may include steps to enrich, transform, and publish to a data mart. In this case, the set of recommendations reflects the template design.
[0194] According to one embodiment, the classification service 516 identifies the business type of a data set or entity ingested from the HUB to the data lake. Recommendations for an entity can be made based on transformations applied to similar entities (business types) or in relation to the target entity deduction / mapping unit.
[0195] According to one embodiment, the functional service 518 annotates the functional types that a data set or entity can have based on defined rules. For example, when generating a cube from a given data set or joining it to a dimension table, it is important to evaluate whether the data set conforms to the rules that define the functional type of the cube.
[0196] According to an embodiment, through pattern inference from pipeline component 520, the recommendation engine can summarize the transformations performed based on a given business type in an existing pipeline definition of a similar context, and propose similar transformations as recommendations for the current context.
[0197] According to an embodiment, a recommendation context can be used to process recommendation 532, which includes action 534, transformation function 535, action parameter 536, function parameter 537, and business type 538.
[0198] Data Lake / Data Management Strategy As described above, according to an embodiment, the data lake provides a repository for the persistence of information from the system HUB or other components.
[0199] FIG. 18 is a diagram showing the use of a data lake according to an embodiment. As shown in FIG. 18, according to an embodiment, the data lake can be associated with one or more data access APIs 540, a cache 542, and a persistence store 544. By operating together, they receive the normalized ingested data for use in multiple pipelines 552, 554, 556.
[0200] According to an embodiment, a variety of data management strategies can be used to manage data (performance, scalability) and its life cycle in the data lake. These can be broadly classified as data-driven or process-driven.
[0201] FIG. 19 is a diagram showing the management of a data lake using a data-driven strategy according to an embodiment.
[0202] As shown in FIG. 19, according to an embodiment, in a data-driven approach, the management unit is derived based on the HUB or data server definition. For example, in this approach, data from Oracle1 HUB can be stored in the first data center 560 associated with this HUB, and data from SFHUB1 can be stored in the second data center 562 associated with this HUB.
[0203] FIG. 20 is a diagram showing the management of a data lake using a process-driven strategy according to an embodiment.
[0204] As shown in FIG. 20, according to an embodiment, in a process-driven approach, the management unit is derived based on the relevant pipeline for accessing data. For example, in this approach, data associated with the sales pipeline can be stored in the first data center 564 associated with this pipeline, and data from other pipelines (such as pipelines 1, 2, 3) can be associated with the second data center 566 associated with these other pipelines.
[0205] Pipeline According to an embodiment, the pipeline defines the transformation or processing to be performed on the ingested data. The processed data may be stored in the data lake or published to another endpoint such as DBCS, for example.
[0206] FIG. 21 is a diagram showing the use of a pipeline compiler according to an embodiment. As shown in FIG. 21, according to one embodiment, the pipeline compiler 582 operates between the design environment 570 and the execution environment 580. This includes receiving one or more pipeline metadata 572 and DSLs, such as Java (registered trademark) DSL 574, JSON DSL 576, Scala DSL 578, and providing the output used in the execution environment as, for example, Spark application 584 and / or SQL statement 586.
[0207] FIG. 22 is a diagram showing an example of a pipeline graph according to one embodiment. As shown in FIG. 22, according to one embodiment, the pipeline 588 includes a list of pipeline steps. Different types of pipeline steps represent different types of operations that can be performed within this pipeline. Each pipeline step can generally include a plurality of input datasets and a plurality of output datasets, which are described by pipeline step parameters. The processing order of operations within the pipeline is defined by combining the output pipeline step parameters from the previous pipeline step with the next pipeline step. In this way, the pipeline steps and the relationships between the pipeline step parameters form a directed acyclic graph (DAG).
[0208] According to one embodiment, if the pipeline includes one or more unique pipeline steps (signature pipelines) that represent the input and output pipeline step parameters of the pipeline, it can be reused in another pipeline. The surrounding pipeline is the pipeline that is reused in the (pipeline usage) pipeline step.
[0209] FIG. 23 is a diagram showing an example of a data pipeline according to one embodiment. As shown in data pipeline 600 as an example shown in FIG. 23, according to an embodiment, the data pipeline performs data conversion. The flow of data within the pipeline is represented as the combination of pipeline step parameters. Various types of pipeline steps are supported for different conversion operations. This includes, for example, entities (retrieving data from the data lake or publishing processed data to the data lake / other HBUs), and joins (fusion of multiple sources).
[0210] FIG. 24 is a diagram showing another example of a data pipeline according to an embodiment. As shown in data pipeline 610 as an example shown in FIG. 24, according to an embodiment, data pipeline P1 can be reused in another data pipeline P2.
[0211] FIG. 25 is a diagram showing an example of an orchestration pipeline according to an embodiment.
[0212] As shown in orchestration pipeline 620 as an example shown in FIG. 25, according to an embodiment, when using the orchestration pipeline, the pipeline steps can be used to represent tasks or jobs that need to be executed throughout the orchestration flow. All pipeline steps within the orchestration pipeline are assumed to have one input pipeline step parameter and one output pipeline step parameter. The execution dependencies between tasks can be represented as the combination between pipeline step parameters.
[0213] According to one embodiment, parallel execution of tasks can be scheduled if the pipeline steps depend on the same previous pipeline step without conditions (i.e., fork). If a pipeline step depends on multiple previous paths, this pipeline step waits for all of the multiple paths to complete prior to its own execution (i.e., join). However, this does not always mean that the tasks are executed in parallel. The orchestration engine can determine whether to execute the tasks serially or in parallel depending on the available resources.
[0214] In the example shown in FIG. 25, according to one embodiment, pipeline step 1 is executed first. If pipeline step 2 and pipeline step 3 are executed in parallel, pipeline step 4 is executed after both pipeline step 2 and pipeline step 3 are completed. The orchestration engine can also execute this orchestration pipeline serially (pipeline step 1, pipeline step 2, pipeline step 3, pipeline step 4) or (pipeline step 1, pipeline step 3, pipeline step 2, pipeline step 4) as long as the dependencies between the pipeline steps are satisfied.
[0215] FIG. 26 is a diagram further showing an example of an orchestration pipeline according to one embodiment.
[0216] As shown in pipeline 625 as an example shown in FIG. 26, according to one embodiment, each pipeline step can return a status 630, such as a success or error status, according to its own semantics. The dependency between two pipeline steps may be conditional based on the return status of the pipeline step. In the example shown, pipeline step 1 is executed first, and if it ends successfully, pipeline step 2 is executed, otherwise pipeline step 3 is executed. After either pipeline step 2 or pipeline step 3 is executed, pipeline step 4 is executed.
[0217] According to one embodiment, by nesting an orchestration pipeline, one orchestration pipeline may be able to reference another orchestration pipeline through pipeline usage. An orchestration pipeline can also reference a data pipeline as pipeline usage. The difference between an orchestration pipeline and a data pipeline is that an orchestration pipeline references a data pipeline that does not include signature pipeline steps, while a data pipeline can reuse another data pipeline that includes signature pipeline steps.
[0218] According to one embodiment, depending on the type of pipeline step and code optimization, a data pipeline can be generated as a single Spark application running within a Spark cluster, or as multiple SQL statements running within DBCS, or as a mix of SQL and Spark code. In the case of an orchestration pipeline, it can be generated to run within an underlying execution engine or within a workflow scheduling component such as Oozie. For execution.
[0219] Coordination fabric According to one embodiment, a coordination fabric or fabric controller provides the tools necessary to deploy and manage framework components (service providers) and an (user-designed) application, manages the execution of the application and resource requests / allocations, and provides an integration framework (messaging bus) to facilitate interactions between various components.
[0220] FIG. 27 is a diagram showing the use of a coordination fabric including a messaging system according to an embodiment.
[0221] As shown in FIG. 27, according to one embodiment, a messaging system (e.g., Kafka) 650 coordinates interactions between a resource manager 660 (e.g., Yarn / Mesos), a scheduler 662 (e.g., Chronos), an application scheduler 664 (e.g., Spark), and a plurality of nodes here shown as nodes 652, 654, 656, 658. are shown.
[0222] According to one embodiment, a resource manager is used to manage the life cycle of data computing tasks / applications. This includes scheduling, monitoring, application execution, resource mediation and allocation, load balancing, managing the configuration and deployment of components (which are message creators and consumers) within a message-driven component integration framework, upgrading components (services) without downtime, and upgrading the infrastructure with minimal or no disruption to services.
[0223] FIG. 28 is a diagram further showing the use of a coordination fabric including a messaging system according to an embodiment.
[0224] As shown in FIG. 28, according to one embodiment, dependencies between components within a coordination fabric are shown by a simple data-driven pipeline execution use case. (c) represents a consumer and (p) represents a producer.
[0225] According to the embodiment shown in FIG. 28, the scheduler (p) starts the process (1) by initiating the ingestion of data into the HUB. The ingestion engine (c) processes the request (2) and ingests the data from the HUB to the data lake. After the ingestion process ends, the ingestion engine (p) starts the pipeline processing (3) by communicating the end status. If the scheduler supports data-driven execution, it can automatically start the pipeline process to be executed (3a). The pipeline engine (c) calculates the pipelines waiting for data execution (4). The pipeline engine (p) conveys a list of pipeline applications to be scheduled for execution (5). The scheduler obtains the execution schedule request for the pipeline (6) and starts the execution of the pipeline (6a). The application scheduler (e.g., Spark) negotiates with the resource manager for resource allocation (7) and executes the pipeline. The application scheduler sends the pipeline for execution to the execution part within the allocated node (8).
[0226] On-premises agent According to one embodiment, an on-premises agent facilitates access to local data and, in a limited way, facilitates distributed pipeline execution. The on-prem ises agent is provisioned and configured to communicate with, for example, cloud DI services and processes data access and remote pipeline execution requests.
[0227] FIG. 29 is a diagram showing an on-premises agent used in a system according to one embodiment.
[0228] As shown in FIG. 29, according to an embodiment, the cloud agent adapter 682 provisions an on-premises agent 680 (1) and configures an agent adapter endpoint for communication.
[0229] The ingest service initiates a local data access request to HUB1 through the messaging system (2). The cloud agent adapter operates as a mediator between the on-premises agent and the messaging system by providing access to requests initiated through the ingest service (3), writes data from the on-premises agent to the data lake, and notifies the messaging system of the completion of the task.
[0230] The on-premises agent polls the cloud agent adapter (4) for processing of data access requests or uploads data to the cloud. The cloud agent adapter writes the data to the data lake (5) and notifies the pipeline through the messaging system.
[0231] DFML Flow Process FIG. 30 is a diagram showing a data flow process according to an embodiment.
[0232] As shown in FIG. 30, according to an embodiment, in the ingest step 692, data is ingested from various sources such as SFDC, S3, or DBaaS.
[0233] In the data preparation step 693, the ingested data is prepared, for example, by deduplication, standardization, or enrichment.
[0234] In the transformation step 694, the system transforms the data by performing merges, filters, or lookups on the data sets.
[0235] In model step 695, one or more models are generated along with the mapping for the model.
[0236] In publish step 696, the system can publish the model, specify policies and schedules, and populate the target data structure.
[0237] Metadata and Data-Driven Automatic Mapping According to an embodiment, the system can support the automatic mapping of complex data structures, data sets, or entities between one or more data sources or targets (referred to as HUBs in some embodiments herein). The automatic mapping can be driven by metadata, schema, and statistical profiling of the data set. By using the automatic mapping, the source data set or entity associated with the input HUB can be mapped to the target data set or entity, or vice versa, at one or more output HUBs to generate output data prepared in the format or organization (projection) used.
[0238] For example, according to an embodiment, a user may desire to select data mapped from a source or input data set or entity in an input HUB to a target or output data set or entity in an output HUB in order to implement (e.g., construct) a data flow, pipeline, or Lambda application.
[0239] According to one embodiment, for a very large set of HUBs and datasets or entities, manually creating a map of data from an input HUB to an output HUB can be a very time-consuming and inefficient task. Therefore, by providing recommendations for mapping data to the user through automatic mapping, the user can focus on simplifying data flow applications such as pipelines and Lambda applications.
[0240] According to one embodiment, the data AI subsystem can receive an automatic mapping request for an automatic mapping service via a graphical user interface (e.g., Lambda Studio Integrated Development Environment (IDE)). It can do so.
[0241] According to one embodiment, this request may include the file specified for the application to be executed by the automatic mapping service, along with information identifying the input HUB, the dataset or entity, and one or more attributes. The application file may contain information about the data for the application. The data AI subsystem can process the application file to extract entity names and other shape characteristics of the entity, including attribute names and data types. The automatic mapping service can use this in a search to discover a possible set of candidates for mapping.
[0242] According to an embodiment, the system can access data for conversion to a HUB such as a data warehouse, for example. The accessed data can include various types of data including semi-structured and structured data. The data AI subsystem can perform metadata analysis on the accessed data, including identifying one or more shapes, features, or structures of the data. For example, metadata analysis can identify data types (e.g., business types and function types), and the column shape of the data.
[0243] According to an embodiment, based on the metadata analysis of the data, one or more samples of the data can be identified, and by applying a machine learning process to the sampled data, the categories of the data within the accessed data can be determined and the model can be updated. The category of the data can indicate a related part of the data, such as a fact table within the data, for example.
[0244] According to an embodiment, machine learning can be implemented using, for example, a logistic regression model or other types of machine learning models that can be implemented for machine learning. According to an embodiment, the data AI subsystem can analyze the relationships of one or more data items within the data based on the category of the data. This relationship indicates one or more fields within the data for the category of the data.
[0245] According to an embodiment, the data AI subsystem can perform a process for feature extraction, which includes identifying a statistical profile of the data randomly sampled for the attributes of the accessed data, the data type, and one or more metadata.
[0246] For example, according to one embodiment, the data AI subsystem can generate a profile of the accessed data based on the category of the data. This profile can be generated to convert the data to the output HUB and can be displayed, for example, on a graphical user interface.
[0247] According to one embodiment, as a result of creating such a profile, the model can support recommendations with a certain degree of confidence regarding the similarity of candidate data sets or entities to the input data set or entity. The recommendations can be provided to the user via a graphical user interface after filtering and sorting.
[0248] According to one embodiment, the automatic map service can dynamically suggest recommendations based on the stage at which the user is building a data flow application, such as a pipeline or a Lambda application.
[0249] An example of a recommendation at the entity level can include, according to one embodiment, a recommendation of an attribute, such as a column of an entity that is automatically mapped to another attribute or another entity. The service can continuously provide recommendations and guide the user based on the user's past activities.
[0250] According to one embodiment, the recommendations can be mapped, for example, from a source data set or entity associated with the input HUB to a target data set or entity associated with the output HUB, using an application programming interface (API) (such as a REST API) provided by the automatic map service. The recommendations can indicate a projection of the data, such as an attribute, a data type, and a representation, and the above representation may be a mapping of an attribute with respect to a data type.
[0251] According to an embodiment, the system can provide a graphical user interface for selecting an output HUB for the conversion of data accessed based on recommendations. For example, through the graphical user interface, the user can select a recommendation for converting data to an output HUB.
[0252] Automatic mapping According to an embodiment, the automatic mapping function can be mathematically defined, and the entity set E is defined as follows.
[0253]
Number
[0254] Wherein, the shape set S includes metadata, data type, and statistical profiling dimension. The objective is to find j such that the probability of similarity between e i and e j is maximized.
[0255]
Number
[0256] At the dataset or entity level, the problem is a binary problem. That is, whether the datasets or entities are similar or not. f s , f t , h(f s , f t ) represents a set of source, target features, and interactive features between the source and the target. Therefore, the objective is to estimate the probability of similarity.
[0257]
Mathematics
[0258] The log-likelihood function is defined as follows
[0259]
Mathematics
[0260] Therefore, in the logistic regression model, the unknown coefficients can be estimated as follows.
[0261]
Mathematics
[0262] According to an embodiment, the automatic mapping service can be triggered, for example, by receiving an HTTP POST request from a system facade service. The system facade API sends data flow applications such as pipelines, Lambda application JSON files from the UI to the automatic mapping REST API, and the parser module processes the application JSON file and extracts the entity names and shapes of the dataset or entities including the attribute names and data types.
[0263] According to an embodiment, the automatic mapping service quickly discovers a set of strong candidates for mapping by using search. The candidate set is a highly relevant set Since it is necessary, this can be achieved using special indexes and queries. This special index incorporates a special search field in which all attributes of the entity are stored and all are tokenized in combinations of N-grams. At query time, the search query builder module constructs a special query using both the entity name and the attribute name of a given entity, for example, by leveraging fuzzy search features based on the Levenshtein distance, and sorts the results by relevance in terms of string similarity by leveraging the search boost function.
[0264] According to an embodiment, the recommendation engine presents to the user a selection of a plurality of relevant results, for example, often the top N results.
[0265] According to an embodiment, to obtain high accuracy, the machine learning model compares the source and the target and determines a score for the similarity of entities based on the extracted features. Feature extraction includes statistical profiles, data types, and metadata of randomly sampled data for each attribute.
[0266] According to an embodiment, the description herein generally describes learning an example of automatic mapping obtained from Oracle Business Intelligence (OBI) system mapping data using a logistic regression model, but other supervised machine learning models can be used instead.
[0267] According to an embodiment, the output of the logistic regression model represents, in a statistical sense, the overall confidence level of the degree of similarity of a candidate data set or entity to an input data set or entity. To discover accurate mappings, one or more other models can be used to calculate the similarity between source and target attributes using similar features.
[0268] Finally, according to one embodiment, the recommendations are filtered and sorted and sent back to the system facade and then sent to the user interface. The automatic mapping service dynamically presents recommendations based on where the user is in the data flow application, e.g., pipeline or Lambda application design. The service can continuously provide recommendations and guide the user based on the user's past activities. The automatic mapping can be performed either in the forward engineering direction or the reverse engineering direction.
[0269] FIG. 31 is a diagram showing an automatic mapping of data types according to an embodiment. As shown in FIG. 31, according to one embodiment, the system facade 701 and the automatic mapping API 702 can receive a data flow application, e.g., pipeline or Lambda application, from a software development component, e.g., Lambda Studio. The parser 704 processes the JSON file of the application and extracts entity names and shapes including attribute names and data types.
[0270] According to one embodiment, a set of strong candidate data sets or entities for mapping is discovered by supporting primary research 710 using a search index 708. The search query builder module 706 seeks a selected data set or entity 712 by constructing a query using both the entity name and the attribute name of a given entity.
[0271] According to one embodiment, a machine learning (ML) model is used to compare source and target pairs and determine a similarity score for data sets or entities based on the extracted features. Feature extraction 714 includes statistical profiles of randomly sampled data, data types, and metadata for each attribute.
[0272] According to one embodiment, the logistic regression model 716 provides, as output, the overall confidence level of the degree of similarity of candidate entities to input entities. To find a more accurate mapping, the column mapping model 718 is used to further evaluate the similarity between source attributes and target attributes.
[0273] According to one embodiment, next, as the automatic mapping 720, the recommendations are sorted in order to return them to a software development component such as Lambda Studio. The automatic map service dynamically presents recommendations based on which stage the user is in during the design of a data flow application such as a pipeline or a Lambda application. The service can continuously provide recommendations and guide the user based on the user's past activities.
[0274] FIG. 32 is a diagram showing an automatic map service for generating mappings according to one embodiment.
[0275] As shown in FIG. 32, according to one embodiment, an automatic map service can be provided for generating mappings. This includes receiving a UI query 728, sending it to a query understanding engine 729, and then sending it to a query decomposition 730 component.
[0276] According to one embodiment, by performing the primary research 710 using the data HUB 722, candidate data sets or entities 731 are obtained for use in subsequent metadata and statistical profile processing 732.
[0277] According to one embodiment, the results are obtained by the Get Stats profile 734 component, the data It is sent to the AI system 724 and feature extraction 735. The results are used for synthesis 736, final confidence merging and ranking 739 according to the model 723, and given to recommendations and related confidence 740.
[0278] Example of automatic mapping FIG. 33 is a diagram showing an example of a mapping between a source schema and a target schema according to an embodiment.
[0279] As shown in FIG. 33, this example 741 shows an example of a simple automatic mapping according to an embodiment, for example, based on (a) hypernym, (b) synonym, (c) equality, (d) Soundex, and (e) fuzzy matching.
[0280] FIG. 34 is a diagram showing another example of a mapping between a source schema and a target schema according to an embodiment.
[0281] As shown in FIG. 34, an approach based solely on metadata fails when this information is irrelevant. According to an embodiment, FIG. 34 shows an example 742 when there is no information in the source attribute name and the target attribute name. When there are no metadata features, the system can discover similar entities using a model that includes statistical profiling of features.
[0282] Automatic mapping process FIG. 35 is a diagram showing a process for providing an automated mapping of data types according to an embodiment.
[0283] As shown in FIG. 35, in step 744, according to an embodiment, by processing the accessed data, metadata analysis of the accessed data is performed.
[0284] In step 745, one or more samples of the accessed data are identified. In step 746, by applying a machine learning process, the category of data in the accessed data is discriminated.
[0285] In step 748, a profile of the accessed data is generated based on the discriminated category of the data for use in automatic mapping of the accessed data.
[0286] Dynamic recommendations and simulations According to an embodiment, the system may include a software development component (referred to as Lambda Studio in some embodiments herein) and a graphical user interface that provides the visual environment used in the system (referred to as a pipeline editor or Lambda Studio IDE in some embodiments herein). This includes providing real-time recommendations based on an understanding of the meaning or semantics associated with the data for the execution of semantic actions on the data accessed from the input HUB. For example, according to an embodiment, the graphical user interface can provide real-time recommendations for performing an action (also referred to as a semantic action) on the data accessed from the input HUB. This includes partial data and the shape or other characteristics of this data. Semantic actions can be executed on the data based on the meaning or semantics associated with this data. The meaning of the data can be used to select the semantic actions that can be performed on this data.
[0287]
[0288] According to one embodiment, a semantic action can represent an operator on one or more datasets and can refer to base semantic actions or functions defined declaratively within the system. One or more processed datasets can be generated by executing the semantic action. The semantic action can be defined by parameters associated with a particular function type or business type. These represent the specific upstream datasets to be processed. The graphical user interface may be metadata-driven, and the graphical user interface may be dynamically generated to dynamically provide recommendations based on the metadata specified within the data.
[0289] FIG. 36 is a diagram showing a system that displays one or more semantic actions enabled for accessed data according to an embodiment.
[0290] As shown in FIG. 36, according to one embodiment, a query for a semantic action enabled for accessed data is sent to the system's knowledge source using a graphical user interface 750 having a user input area 752. This query indicates the classification of the accessed data.
[0291] According to one embodiment, a response to the query is received from the knowledge source. This response indicates one or more semantic actions enabled for the accessed data and identified based on the classification of the data.
[0292] According to one embodiment, the selected semantic action among the semantic actions enabled for the accessed data is displayed for selection and for use with the accessed data. This includes automatically providing or updating a list of the selected semantic actions or recommendations 758 among the semantic actions 756 enabled for the accessed data during the processing of the accessed data.
[0293] According to one embodiment, recommendations can be provided dynamically rather than pre-computed based on static data. For example, the system can provide recommendations in real time based on data accessed in real time, considering information such as, for example, a user profile or the user's experience level. The recommendations provided by the system for real-time data can be notable, relevant, and accurate for generating data flow applications such as pipelines, Lambda applications. Recommendations can be provided based on the user's actions with respect to data associated with specific metadata. The system can recommend semantic actions for the information.
[0294] For example, according to one embodiment, the system can ingest, transform, integrate, and publish data to any system. The system can recommend analyzing some of its numerical criteria using interesting analysis techniques with entities, pivoting this data in various dimensions to show which are the interesting dimensions, summarizing the data with respect to the dimension hierarchy, and enriching the data with more insights.
[0295] According to one embodiment, recommendations can be provided based on an analysis of the data using techniques such as, for example, metadata analysis of the data.
[0296] According to one embodiment, metadata analysis may include determining a classification of data, such as the shape, features, and structure of the data. Metadata analysis can determine the data type (e.g., business type and functional type). Metadata analysis can also indicate the column shape of the data. According to one embodiment, by comparing data with a metadata structure (e.g., shape and features), the data type and the attributes associated with the data can be determined. The metadata structure can be defined in the system HUB (e.g., knowledge source) of the system.
[0297] According to one embodiment, the system can use metadata analysis to query the system HUB to identify semantic actions based on the metadata. The recommendation may be a semantic action determined based on the analysis of the metadata of the data accessed from the input HUB. Specifically, the semantic action can be mapped to the metadata. For example, the semantic action can be mapped to the metadata for which these actions are permitted and / or applicable. The semantic action may be defined by the user and / or may be defined based on the structure of the data.
[0298] According to one embodiment, the semantic action can be defined based on the conditions associated with the metadata. The system HUB can be modified so that the semantic action is modified, deleted, or augmented.
[0299] According to one embodiment, examples of semantic actions may include building a cube, filtering data, grouping data, aggregating data, or other actions that can be performed on data. By defining semantic actions based on metadata, there is no need for a mapping or schema to determine the semantic actions permitted on the data. Semantic actions may be defined when new and different metadata structures are discovered. Thus, the system can dynamically obtain recommendations based on the identification of semantic actions, using the metadata parsed for the data received as input.
[0300] According to one embodiment, a third party may define semantic actions and be able to supply data, for example data that defines one or more semantic actions associated with metadata. The system can determine the semantic actions available for the metadata by dynamically querying a system HUB. Thus, the system HUB may be modified and the system may then determine the semantic actions permitted based on such modifications. The system can process data obtained from a third party by performing operations (such as filtering, detecting, and registering). At this time, the data can define semantic actions and the semantic actions can be made available based on the semantic actions identified by the above processing.
[0301] FIG. 37 and FIG. 38 are diagrams showing a graphical user interface that displays one or more semantic actions enabled for accessed data according to one embodiment.
[0302] As shown in FIG. 37, according to one embodiment, a software development component (such as Lambda Studio) is a graphical user interface (such as a pipeline An editor or Lambda Studio IDE) 750 can be provided. This is the output HU It is possible to display recommendation semantic actions used when processing input data for projection to output HU B or simulating the processing of input data.
[0303] For example, according to one embodiment, through the interface of FIG. 37, a user can display options 752 associated with a data flow application such as a pipeline, a Lambda application. This includes, for example, input HUB specifications 754.
[0304] According to one embodiment, during the creation of a data flow application such as a pipeline, a Lambda application, or during the simulation of a data flow application such as a pipeline, a Lambda application with respect to input data, one or more semantic actions 756 or other recommendations 758 can be displayed in a graphical user interface for user consideration.
[0305] According to one embodiment, in simulation mode, a software development component (such as Lambda Studio) provides a sandbox environment. Through this environment, a user can immediately view the results of executing various semantic actions on the output. This includes automatically updating a list of semantic actions suitable for the accessed data during the processing of the accessed data.
[0306] For example, as shown in FIG. 38, according to one embodiment, a user searching for some information From the user's starting point, the system can recommend actions 760 for the information. For example, it is recommended to analyze some of the numerical criteria using entities with interesting analysis methods, pivot this data in various dimensions to show which are interesting dimensions, summarize the data with respect to the dimension hierarchy, and enrich the data with more insights.
[0307] According to an embodiment, in the example shown, both the source and the dimension are recommended for analyzable entities within the system, and the task of constructing a multi-dimensional cube is generally one of point-and-click.
[0308] Typically, such operations require a lot of experience and domain-specific knowledge. By using machine learning to analyze both data characteristics and the user's behavior patterns with respect to general integration patterns, in combination with semantic search and recommendations from machine learning, it enables cutting-edge tooling for developing applications for building business-specific applications.
[0309] FIG. 39 is a diagram showing a process of displaying one or more semantic actions enabled for accessed data according to an embodiment.
[0310] As shown in FIG. 39, at step 772, according to an embodiment, by processing the accessed data, a metadata analysis of the accessed data is performed. The metadata analysis includes determining the classification of the accessed data.
[0311] At step 774, a query for semantic actions enabled for the accessed data is sent to the system's knowledge source. This query indicates the classification of the accessed data.
[0312] In step 775, a response to the query is received from the knowledge source. This response indicates one or more semantic actions that are validated against the accessed data and identified based on the classification of the data.
[0313] In step 776, a semantic action selected from among the semantic actions validated against the accessed data is displayed on a graphical user interface for selection and for use with the accessed data. This includes automatically providing or updating a list of semantic actions selected from among the semantic actions validated against the accessed data during processing of the accessed data.
[0314] Functional Decomposition of Data Flow According to one embodiment, the system can provide a service for recommending actions and transformations for input data based on patterns identified from a functional decomposition of the data flow of a software application, which includes determining possible transformations to the data flow in a later application. The data flow can be decomposed into a model that describes data transformations, predicates, and business rules applied to the data, and attributes used within the data flow.
[0315] FIG. 40 shows supporting pattern detection and inductive learning by evaluating a pipeline, a Lambda application, to determine its components, according to one embodiment.
[0316] As shown in FIG. 40, according to one embodiment, the functional decomposition logic 800, i.e., the software component, can be provided as software or program code executable by a computer system or other processing device, and for a display 805 (e.g., within a pipeline editor or Lambda Studio IDE), the function It can be used to provide solutions 802 and recommendations 804. For example, this system can provide a service for recommending actions and transformations on data based on patterns / templates identified from the functional decomposition of data flow applications such as pipelines, Lambda application data flows, i.e., through the functional decomposition of data flows, patterns can be observed for determining possible transformations to the data flow in subsequent applications.
[0317] According to one embodiment, this service can be implemented by a framework that can decompose or classify a data flow into a model that describes data transformations, predicates, and business rules applied to the data, and attributes used in the data flow.
[0318] Conventionally, the data flow of an application can represent a series of transformations on data, and the type of transformation applied to the data is highly contextual. In most data integration frameworks, usually, the process lineage is limited or non-existent regarding how the data flow is persisted, analyzed, and generated. According to one embodiment, this system can derive context-related patterns from a flow or graph based on semantically rich entity types, further learn data flow grammars and models, and use these to generate composite data flow graphs when given a predetermined similar context.
[0319] According to one embodiment, the system can generate one or more data structures that define patterns and templates based on the design specifications of the data flow. By decomposing the data flow into a data structure that defines a functional expression, patterns and templates can be determined. Using the data flow, functional expressions for determining patterns for data transformation recommendations can be predicted and generated. The recommendations are based on the decomposed data flow and models derived from inductive learning of the unique patterns, and can be made more granular (e.g., recommend a scalar transformation for a specific attribute, or use one or more attributes in a predicate for filtering or joining).
[0320] According to one embodiment, with a data flow application such as a pipeline, Lambda application, a user can generate complex data transformations based on semantic actions on the data. The system can store the data transformation as one or more data structures that define the flow of data for the pipeline, Lambda application.
[0321] According to one embodiment, by utilizing the decomposition of the data flow of a data flow application such as a pipeline, Lambda application, pattern analysis of the data can be determined and a functional expression can be generated. The decomposition can be performed on semantic actions and transformations and predicates, or business rules. Each semantic action of the previous application can be identified through the decomposition. Using the process of induction, business logic can be extracted from the data flow including its context elements (business type and functional type).
[0322] According to one embodiment, a model can be generated for a process, and based on induction, rich context normative data flow design recommendations can be generated It can be done. These recommendations may be based on patterns inferred from the model, and each recommendation may correspond to a semantic action that can be performed on the data for the application.
[0323] According to an embodiment, the system can execute a process of inferring a pattern of data transformation based on function decomposition. The system can access the data flow of one or more data flow applications, such as pipelines, Lambda applications. By processing the data flow, one or more functional expressions can be obtained. The functional expressions can be generated based on actions, predicates, or business rules identified within the data flow. By using actions, predicates, or business rules, a pattern of transformation for the data flow can be identified (e.g., inferred). The inference of the transformation pattern may be a passive process.
[0324] According to an embodiment, the pattern of transformation can be determined by a crowdsourcing method based on passive analysis of the data flows of different applications. This pattern can be determined using machine learning (e.g., deep reinforcement learning).
[0325] According to an embodiment, the pattern of transformation can be identified for the functional expressions generated for data flow applications, such as pipelines, Lambda applications. By decomposing one or more data flows, a pattern of data transformation can be inferred.
[0326] According to an embodiment, the system can use this pattern to recommend one or more data transformations for a new data flow application, such as a pipeline or a Lambda application's data flow. In an example of a data flow for processing data for currency exchange, the system can identify a pattern of transformation for the data. Also, the system can recommend one or more transformations for a new data flow of an application, where the data flow includes data for similar currency exchange. The transformation can be performed in a similar manner according to the pattern so as to produce a similar currency exchange by modifying the new data flow according to the transformation.
[0327] FIG. 41 is a diagram showing means for identifying a pattern of transformation within a data flow for one or more functional expressions generated for each of one or more applications according to an embodiment.
[0328] As described above, according to an embodiment, by a pipeline, such as a Lambda application, a user can define complex data transformations based on semantic actions corresponding to operators in a relational algorithm. The data transformation is typically persisted as a directed acyclic graph or a query, or as nested functions in the case of DFML. By decomposing and serializing a data flow application, such as a pipeline or a Lambda application, into nested functions, it enables pattern analysis of the data flow and induces a data flow model that can be used to generate functional expressions for extracting complex transformations for a data set in a similar context.
[0329] According to one embodiment, the nested function decomposition is performed not only at the level of semantic actions (rows or dataset operators), but also on scalar transformations and predicate structures. This enables deep lineage capabilities for complex dataflows. Recommendations based on inductive models can be highly granular (e.g., recommend a scalar transformation for a particular attribute, or one or more attributes in a predicate for filtering or joining). be used).
[0330] According to one embodiment, the elements included in the function decomposition are generally as follows. The application represents a top-level dataflow transformation.
[0331] An action represents an operator on one or more datasets (unique dataframes).
[0332] An action refers to a base semantic action or function defined declaratively in the system. An action can have one or more action parameters, each of which can have a specific role (in, out, in / out) and type, can return one or more processed datasets, and can be embedded or nested to several levels of depth.
[0333] An action parameter is owned by an action, has a specific functional or business type, and represents a specific upstream dataset to be processed. A join parameter represents a dataset or entity within the HUB used for the transformation. A value parameter represents an intermediate or temporary data structure processed in the context of the current transformation.
[0334] A scope resolver can derive a process lineage for a dataset used in the overall dataflow or an element within the dataset.
[0335] FIG. 42 is a diagram showing an object diagram used to identify patterns of transformations within a data flow for one or more functional expressions generated for each of one or more applications according to an embodiment.
[0336] As shown in FIG. 42, according to an embodiment, using function decomposition logic, the data flow of a data flow application, such as a pipeline, a Lambda application, can be decomposed or classified into a model that describes data transformations, predicates, and business rules applied to the data, and attributes used within the data flow. This can be decomposed, for example, into a pattern or template 812 (a pipeline if the template is associated with a Lambda application), a service 814, a function 816, a function parameter 818, and a function type 820.
[0337] According to an embodiment, each of these function components can be further decomposed into, for example, a task 822 or an action 824 that reflects a data flow application such as a pipeline, a Lambda application.
[0338] According to an embodiment, using a scope resolver 826, references to specific attributes or embedded objects can be resolved through their scope. For example, as shown in FIG. 42, the scope resolver resolves references to attributes or embedded objects through their adjacent scopes. For example, a join function that uses the output of a filter and another table has a reference to its scope resolver and, by using it in combination with the InScopeOf operation, can resolve to the leaf node itself to the root node.
[0339] FIG. 43 is a diagram showing a process for identifying patterns of transformations within a data flow for one or more functional expressions generated for each of one or more applications according to an embodiment.
[0340] As shown in FIG. 43, according to an embodiment, in step 842, for each of one or more software applications, access the data flow.
[0341] In step 844, by processing the data flows of the one or more software applications, generate one or more functional expressions representing the data flows. The one or more functional expressions are generated based on semantic actions and business rules identified within the data flows.
[0342] In step 845, for each of the one or more functional expressions generated for each of the one or more software applications, identify a pattern of transformations within the data flow. Use semantic actions and business rules to identify the pattern of transformations within the data flow.
[0343] In step 847, using the pattern of transformations identified within the data flow, provide one or more data transformation recommendations for the data flow of another software application.
[0344] Ontology Learning According to an embodiment, the system can execute ontology analysis of schema definitions to identify the data and the types of data sets or entities associated with the schema, and generate or update a model from a reference schema including an ontology defined based on the relationships between entities and their attributes. Analyze the data flow using a reference HUB including one or more schemas, and further classify or make recommendations such as transformation, enrichment, filtering, or cross-entity data fusion of input data.
[0345] According to an embodiment, the system can determine the ontology of the types of data and entities in the reference schema by performing ontology analysis of the schema definition. In other words, this system can generate a model from a schema that includes an ontology defined based on the relationship between entities and their attributes. The reference schema may be provided by the system, a default reference schema, or alternatively, supplied by the user or a third-party reference schema.
[0346] Some of the data integration frameworks may reverse engineer the engineer metadata from well-known system source types, but do not provide analysis of the metadata for constructing a functional system that can be used for pattern definition and entity classification. Also, metadata harvesting is limited in scope and is not extended to data profiling for the extracted dataset or entity. There is currently no functionality that allows a user to specify a reference schema for functional system ontology learning for use in complex processes (business logic) and integration patterns in addition to entity classification (in a similar topological space).
[0347] According to an embodiment, one or more schemas can be stored in the reference HUB. This itself can be provided within or as part of the system HUB. Similar to the reference schema, the reference HUB may also be provided by the user or a third-party reference HUB, or in a multi-tenant environment, may be accessed through, for example, a data flow API associated with a specific tenant.
[0348] According to an embodiment, use the reference HUB to analyze the data flow and further Classification or recommendations can be performed. It can be, for example, transformation, enrichment, filtering, or cross-entity data fusion.
[0349] For example, according to one embodiment, the system can receive an input that defines a reference HUB as a schema for ontology analysis. The reference HUB may import and obtain entity definitions (attribute definitions, data types, and relationships between data sets or entities, constraints, or business rules). Some metrics of the data may be derived by extracting sample data (e.g., attribute vectors such as column data as an example) in the reference HUB for all data sets or entities and the profiled data.
[0350] According to one embodiment, the type system can be instantiated based on the glossary of the reference schema. The system can derive an ontology (e.g., a set of rules) that describes the types of the data by performing ontology analysis. Ontology analysis can determine data rules. The data rules are defined for profiled data (e.g., attribute or composite value) metrics and describe business type elements (e.g., UOM, ROIL, or currency type) together with their data profiles. Ontology analysis can determine relationship rules and composite rules. The relationship rules define the correspondence between data sets or entities and attribute vectors (constraints or references imported from the reference schema), and the composite rules can be derived from a combination of data rules and relationship rules. Then, the type system can be defined based on the rules obtained through metadata harvesting and data sampling.
[0351] According to one embodiment, based on a type system instantiated using ontology analysis, patterns and templates from the system HUB can be utilized. Then, the system can execute data flow processing using the type system.
[0352] For example, according to one embodiment, the classification and type annotation of a dataset or entity can be specified by the type system of the registered HUB. Using the type system, rules can be defined for functional and business types derived from a reference schema. Using the type system, actions such as blend, enrichment, and transformation recommendations can be performed on entities identified within the data flow based on the type system.
[0353] FIG. 44 is a diagram showing a system for generating functional rules according to one embodiment.
[0354] As shown in FIG. 44, according to one embodiment, rule induction logic 850 or a software component provided as software or program code executable by a computer system or other processing device can associate rule 851 with functional system 852.
[0355] FIG. 44 is a diagram showing a system for generating functional rules according to one embodiment.
[0356] As shown in FIG. 45, according to one embodiment, HUB1 can function as a reference ontology. This can be used for type tagging, comparison, classification, or otherwise evaluation of metadata schemas or ontologies provided by other (e.g., newly registered) HUBs such as HUB2 and HUB3, and the data AI system can create appropriate rules to use.
[0357] FIG. 46 is a diagram showing an object diagram used for generating functional rules according to an embodiment.
[0358] According to an embodiment, as shown in FIG. 46 for example, the rule induction logic can associate a rule with a functional system having a set of functional 853 (e.g., HUB, dataset or entity, and attribute), and store it in a registry for use in creating a data flow application such as a pipeline, Lambda application. This includes being able to associate each functional 854 with a functional rule 856 and a rule 858. Each rule can be associated with a rule parameter 860.
[0359] According to an embodiment, first a reference schema is processed and an ontology containing a set of rules suitable for this schema can be created.
[0360] According to an embodiment, then a new HUB or a new schema is evaluated, its dataset or entity is compared with the existing ontology and the created rules, and it can be used for the analysis of the new HUB / schema and its entity, as well as for further learning of the system.
[0361] Metadata harvesting in a data integration framework may be limited to reverse engineering entity definitions (attributes and their data types, and possibly relationships), but according to certain embodiments, the approach provided by the systems described herein differs in the following ways. That is, the schema definition can be used as a reference ontology from which business types and function types can be derived, and data profiling metrics for data sets or entities in the reference schema can be derived. Then, by using this reference HUB to analyze business entities within other HUBs (data sources), further classification or recommendations (such as blends or enrichments) can be made.
[0362] According to certain embodiments, the system uses the following set of steps for ontology learning using a reference schema.
[0363] The user specifies an option to use a newly registered HUB as a reference schema.
[0364] Entity definitions (e.g., attribute definitions, data types, relationships between entities, constraints or business rules are imported).
[0365] Sample data is extracted for all data sets or entities, and the data is profiled to derive some metrics about the data.
[0366] A type system is instantiated based on the glossary of the reference schema (function types and business types).
[0367] A set of rules describing the business type is derived. Data rules are defined with respect to the profiled data metrics and describe the nature of the business type elements (e.g., the UOM, ROI, or currency type of the data) (which can be defined as a business element along with the profile).
[0368] Relationship rules are generated that define associations in the elements (constraints or references imported from the reference schema).
[0369] Composite rules are generated that can be derived through combinations of data and relationship rules. The type system (functional and business) is defined based on rules derived through metadata harvesting and data sampling.
[0370] Then, patterns or templates can define complex business logic using types instantiated based on the reference schema.
[0371] Then, the HUB registered in the system can be analyzed in the context of the reference schema.
[0372] In the newly registered HUB, classification and type annotation of datasets or entities can be performed based on rules for the functional and business types derived from the reference schema.
[0373] Based on the type annotation, blending, enrichment, and transformation recommendations can be executed on the dataset or entity.
[0374] FIG. 47 is a diagram showing a process of generating a functional system based on one or more generated rules according to an embodiment.
[0375] As shown in FIG. 47, according to an embodiment, at step 862, an input defining the reference HUB is received.
[0376] In step 863, by accessing the reference HUB, obtain one or more entity definitions associated with a dataset or entity provided by the reference HUB.
[0377] In step 864, generate sample data for one or more datasets or entities from the reference HUB.
[0378] In step 865, by profiling this sample data, determine one or more metrics associated with this sample data.
[0379] In step 866, generate one or more rules based on the entity definitions. In step 867, generate a functional system based on the one or more generated rules.
[0380] In step 868, persist the functional system and the profile of the sample data for use in processing data inputs.
[0381] Foreign language function interface According to an embodiment, this system provides a (in some embodiments herein referred to as a foreign language function interface) programming interface. Through this interface, a user or third party can extend the functionality of the system by declaratively defining services, functional and business types, semantic actions, and patterns or predetermined composite data flows based on functional and business types.
[0382] As described above, the interfaces that the current data integration system can provide are limited, there is no support for types, and there is no interface clearly defined for object composition and pattern definition. Due to such shortcomings, complex functions such as cross-service recommendations for calling semantic actions across the entire service for extending the framework or a unified application design platform are not currently provided.
[0383] According to an embodiment, through a multilingual function interface, a user can extend the functions of the system by declaratively providing definitions or other information (for example, from a customer to another third party).
[0384] According to an embodiment, the system is metadata-driven, obtains metadata by processing the definitions received through the multilingual function interface, discriminates the classification of the metadata, for example, data types (for example, function types and business types), and can compare these data types (both functions and business) with the existing metadata to determine whether there is a type match.
[0385] According to an embodiment, the metadata received through the multilingual function interface may be stored in the system HUB so that the system can access the metadata to process the data flow. For example, by accessing the metadata, semantic actions can be determined based on the type of the input data set. The system can determine the permitted semantic actions for the type of data provided through this interface.
[0386] According to one embodiment, by providing a common declarative interface, the system can enable a user to map service-native types and actions to platform-native types and actions. This enables a unified application design experience through the discovery of types and patterns. This also facilitates purely declarative data flow definitions and designs that require components of various services that extend the platform, and the generation of native code for each semantic action.
[0387] According to one embodiment, metadata received through a multi-language function interface can be automatically processed and the objects or artifacts described therein (such as data types or semantic actions) can be used in the operation of the data flow processed by the system. One or more function types and business types that define a service, one or more semantic actions that display, or one or more patterns / templates that display may be displayed using metadata information received from one or more third-party systems.
[0388] For example, according to one embodiment, the classification of accessed data, such as the function type and business type of data, can be determined. This classification can be identified based on information about the data received with the information. By receiving data from one or more third-party systems and expanding the functionality of the system, data integration of the data flow can be performed based on information received from a third party (such as a service, semantic action, or pattern).
[0389] According to an embodiment, by updating the metadata in the system HUB, it is possible to include information identified regarding the data. For example, services and patterns / templates can be executed based on the information (e.g., semantic actions) identified in the metadata received through the multi-language function interface by updating. In this way, the system can be enhanced through the multi-language function interface without interfering with the processing of the data flow.
[0390] According to an embodiment, subsequent data flows can be processed using the metadata in the system HUB after its update. Metadata analysis can be performed on data flows of data flow applications, such as pipelines, Lambda application data flows. Next, the system HUB can be used to determine transformation recommendations considering the definitions provided via the other-language function interface. The transformation can be determined based on the patterns / templates used to define semantic actions to execute the service, and the semantic actions can also consider the definitions provided via the other-language function interface.
[0391] FIG. 48 is a diagram showing a system for identifying a pattern used in providing recommendations regarding a data flow based on information provided via a multi-language function interface according to an embodiment.
[0392] As shown in FIG. 48, according to an embodiment, the service registry 902, function and business type registry 904, or pattern / template 906 in the system HUB can be updated using the definitions received via the other-language function interface 900.
[0393] According to one embodiment, by using the updated information in a data AI subsystem including a rule engine 908, for example, to identify a HUB, dataset or entity, or attribute 910 annotated with a type in the system HUB, and to provide these datasets or entities for use in providing recommendations for data flow applications (such as pipelines, Lambda applications) via a software development component (such as Lambda Studio), the recommendation can be provided to a recommendation engine 912.
[0394] FIG. 49 is a diagram further showing a pattern used in providing recommendations for a data flow based on information provided via a foreign language function interface according to one embodiment.
[0395] As shown in FIG. 49, according to one embodiment, third-party metadata 920 can be received through a foreign language function interface.
[0396] FIG. 50 is a diagram further showing a pattern used in providing recommendations for a data flow based on information provided via a foreign language function interface according to one embodiment.
[0397] As shown in FIG. 50, according to one embodiment, the functions of the system can be extended using third-party metadata received through a multilingual function interface.
[0398] According to one embodiment, this system enables the extension of the framework through a clearly defined interface. Through this interface, the service It is possible to register, together with the native type of the service, the semantic actions realized by the service, and typed parameters, patterns or templates that can extract predefined algorithms that can be used, in particular, as part of the service.
[0399] According to an embodiment, by providing a common declarative programming paradigm, a pluggable service architecture enables mapping service-native types and actions to platform-native types and actions. This enables a unified application design experience through the discovery of types and patterns. This also facilitates purely declarative data flow definitions and designs that require components of various services to extend the platform and the generation of native code for each semantic action.
[0400] According to an embodiment, a pluggable service architecture also defines a compile, generate, deploy, and runtime execution framework (unified application design service) for plugins. The recommendation engine can perform machine learning and inference on the semantic actions and patterns of all plugged-in services and can make cross-service semantic action recommendations regarding distributed composite data flow design and development.
[0401] FIG. 51 is a diagram showing a process for identifying a pattern used in providing recommendations regarding a data flow based on information provided via a foreign language function interface according to an embodiment.
[0402] As shown in FIG. 51, according to an embodiment, in step 932, one or more definitions of metadata used in processing data are received via a foreign language function interface.
[0403] In step 934, by processing the metadata received via the foreign language function interface, identify information regarding the received metadata that includes one or more of a classification, a semantic action, a template that defines a pattern, or a service as defined by the received metadata.
[0404] In step 936, store the metadata received via the foreign language function interface in the system HUB. The system HUB, when updated, includes information regarding the received metadata and extends the functionality of the system to include the supported types, semantic actions, templates, and services of the system.
[0405] In 938, identify a pattern for providing a recommendation regarding a data flow based on the information updated in the system HUB via the foreign language function interface.
[0406] Policy-based Lifecycle Management According to an embodiment, this system can provide a data governance function. This is, for example, historical information (where the specific data came from), lineage (how this data was acquired / processed), security (who was responsible for this data), classification (what the data is related to), impact (how much impact this data has on the business), retention time (how long this data should exist), and validity (whether this data should be excluded / included for analysis / processing) for each slice of data temporally related to a specific snapshot. These can be used in lifecycle decisions and data flow recommendations. be
[0407] Current approaches to data lifecycle management do not include governance-related functions based on changes in overall data characteristics in temporary partitions or tracking the evolution of data (changes in data profiles or drifts). Data characteristics (classification, frequency of change, type of change, or use in a process) observed or derived by the system are not used for lifecycle decisions or recommendations about the data (retention time, security, validity, acquisition interval).
[0408] According to one embodiment, the system can provide a graphical user interface that can display the lifecycle of a data flow based on lineage tracking. The lifecycle can indicate where the data was processed and whether an error occurred during the processing of that data, and can be presented as a timeline diagram of the data (e.g., number of data sets, volume of data sets, and use of data sets). The interface can provide a snapshot of the data at a given point in time and can provide that visual indicator during the processing of the data. Thus, this interface enables a complete audit of the data or a system snapshot of the data based on the lifecycle (e.g., performance metrics or resource usage).
[0409] According to one embodiment, the system can determine the lifecycle of data based on sample data (periodically sampled from the ingested data) and data obtained for processing by user-specified applications. Some aspects of data lifecycle management are similar across the categories of ingested data, i.e., streaming data and batch data (reference and incremental). For incremental data, the system can use a scheduled log collection event-driven approach to obtain a temporary slice of the data and manage the allocation of slices across application instances covering the following functions.
[0410] According to one embodiment, in the case of data loss, the system can reconstruct data using a lineage spanning the entire layer from the metadata managed within the system HUB.
[0411] For example, according to one embodiment, incremental data can be obtained by specifying incremental data attribute columns or user configuration settings, and high and low watermarks are maintained across the entire data ingest. Queries or APIs and corresponding parameters (timestamp or ID column) can be associated with the ingested data.
[0412] According to one embodiment, the system can manage lineage information across the entire layer. This includes, for example, query or log metadata within the edge layer, topic / partition offsets per ingest within the scalable I / O layer, slices (file partitions) within the data lake, references to the process lineage of subsequent downstream processed data sets using this data (specific execution instances of an application that generate data and associated parameters), topics / partitions for "marked" data sets published to a target endpoint and corresponding data slices within the data lake, and offsets within the partitions published and processed by a job execution instance and published to a target endpoint.
[0413] According to one embodiment, in the event of a failure of a layer (e.g., edge, scalable I / O, data lake or publish), the data can be reconstructed from the upstream layer or obtained from the source.
[0414] According to one embodiment, the system can perform other lifecycle management functions.
[0415] For example, according to one embodiment, security is enhanced and audited at each of these layers of the data slice. The data slice can be excluded from or (if already excluded) made the subject of processing or access. This can prevent spurious or corrupted data slices from being processed. The retention policy can be enforced on slices of data through a sliding window. Analyze the impact on slices of data (e.g., analyze the ability to tag slices of a given window as impactful in the context of a data mart built for quarterly reporting).
[0416] According to one embodiment, data is classified by tagging it with a function type or business type defined within the system (e.g., tagging a data set with a function type along with a business type (such as order, customer, product, or time) as a cube or dimension or hierarchical data).
[0417] According to one embodiment, the system can execute a method that includes accessing data from one or more HUBs. This data can be sampled, and the system determines and manages a temporary slice of the data. This includes accessing the system HUB of the system and obtaining metadata regarding the sampled data. The sampled data can be managed for lineage tracking across one or more layers within the system.
[0418] According to one embodiment, incremental data and parameters regarding sample data can be managed with respect to the ingested data. The data can be classified by tagging it with the type of data associated with the sample data.
[0419] FIG. 52 is a diagram showing the management of sampled or accessed data for system tracking across one or more layers, according to an embodiment.
[0420] For example, as shown in FIG. 52, according to an embodiment, using this system, data can be received from HUB952, in this example an Oracle database, and HUB954, in this example S3 or other environment. Data received from the input HUB in the edge layer can be provided to the scalable I / O layer as one or more topics for use by a data flow application such as a pipeline, Lambda application (each topic can be provided as a distributed partition).
[0421] According to an embodiment, the ingested data, typically represented by an offset to a partition, can be normalized 964 by the compute layer and written to the data lake as one or more temporary slices spanning the layers of the system.
[0422] According to an embodiment, this data is then used by a data flow application such as a pipeline, Lambda applications 966, 968, and finally published 970 to one or more additional topics 960, 962, and then, in this example, can be published to a target endpoint (e.g., a table) at one or more output HUBs such as a DBCS environment.
[0423] As shown in FIG. 52, according to an embodiment, initially, the data reconstruction and system tracking information can include, for example, provenance information (Hub 1, S3), lineage (Source Entity in Hub 1), security (Connection Credential used), or other information regarding the ingestion of the data.
[0424] FIG. 53 further shows the management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. As shown in FIG. 53, data reconstruction and lineage tracking information can then be updated to include, for example, updated history information (→T1), lineage (→T1 (Ingest Process)), or other information.
[0425] FIG. 54 further shows the management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. As shown in FIG. 54, data reconstruction and lineage tracking information can then be further updated to include, for example, updated history information (→E1), lineage (→E1 (Normalize)), or other information. can be made.
[0426] According to an embodiment, a temporary slice 972 used by one or more data flow applications, such as pipelines, Lambda applications, can be created to span the layers of the system. In the event of a failure, such as a failure in writing to a data lake, the system can identify one or more unprocessed data slices and complete the processing of those data slices, either as a whole or incrementally.
[0427] FIG. 55 is a diagram further showing the management of sampled or accessed data for system tracking across one or more layers according to an embodiment. As shown in FIG. 55, thereafter, by further updating the data reconstruction and system tracking information and creating additional temporary slices, it is possible to include information such as, for example, updated history information (→E11(App1)), security (Role Executing App 1), or other information.
[0428] FIG. 56 is a diagram further showing the management of sampled or accessed data for system tracking across one or more layers according to an embodiment. As shown in FIG. 56, thereafter, by further updating the data reconstruction and system tracking information and creating additional temporary slices, it is possible to include information such as, for example, updated system (→E12(App2)), security (Role Executing App 2), or other information.
[0429] FIG. 57 is a diagram further showing the management of sampled or accessed data for system tracking across one or more layers according to an embodiment. As shown in FIG. 57, thereafter, by further updating the data reconstruction and system tracking information and creating additional temporary slices, it is possible to include information such as, for example, updated system (→T2(Publish)), security (Role Executing Publish to I / O Layer), or other information.
[0430] FIG. 58 is a diagram further showing the management of sampled or accessed data for system tracking across one or more layers according to an embodiment. As shown in FIG. 58, thereafter, by further updating the data reconstruction and system tracking information, the output of data to the target endpoint 976 can be reflected.
[0431] Data Lifecycle Management According to an embodiment, the data lifecycle management based on the above system tracking is directed to several functional areas. Some of these areas can be configured by the user (access control, retention time, validity), some are derived (history information, lineage), and others use machine learning algorithms (classification, influence). For example, data management is applied to both sample data (periodically sampled from the ingested data) and data obtained for processing by user-defined applications. Some aspects of data lifecycle management are similar across the entire category of ingested data, i.e., streaming data and batch data (reference and incremental). In the case of incremental data, DFML uses a scheduled log collection event-driven approach to obtain a temporary slice of the data and manages the allocation of slices across the entire application instance covering the following functions.
[0432] In the case of data loss, the data is reconstructed using the lineage across the entire layer from the metadata managed in the system HUB.
[0433] Incremental data is obtained by identifying incremental data attribute columns or user configurations, and high and low watermarks are maintained across the entire ingest.
[0434] For each ingest, associate a query or API and the corresponding parameters (timestamp or ID column).
[0435] Maintaining lineage information across tiers. Query or log metadata in the edge layer. Topic / partition offset per ingest in the scalable I / O layer. Slices (file partitions) in the data lake. A reference to the process lineage (the specific execution instance of the application that produced the data and its associated parameters) of all subsequent downstream processed datasets that use this data. Topic / partition offsets of datasets that are "marked" to be published to target endpoints and the corresponding data slices in the data lake. The offset within the partition that a job execution instance publishes and processes and publishes to the target endpoint.
[0436] In case of a layer failure, data can be reconstructed from an upstream layer or retrieved from the source. Security is enforced and audited at each of these layers for data slices. Data slices can be excluded from processing or access, or made eligible (if already excluded), so that spurious or corrupted data slices are not processed. Retention time policies can be enforced for data slices through sliding windows. Slices of data are analyzed for impact (e.g., the ability to tag slices for a given window as impactful in the context of a data mart built for quarterly reporting).
[0437] Classifying data by functional or business type tagging prescribed within the system (e.g. tagging a dataset by functional type (cube or dimensional or hierarchical data) along with business type (e.g. order, customer, product or time)).
[0438] FIG. 59 illustrates a process for managing sampled or accessed data for lineage tracking across one or more layers, according to an embodiment.
[0439] As shown in FIG. 59, according to an embodiment, in step 982, data is accessed from one or more HUBs.
[0440] In step 983, the accessed data is sampled. In step 984, a temporary slice is identified for the sampled data and the accessed data.
[0441] In step 985, by accessing the system HUB, metadata regarding the sampled data or the accessed data represented by the temporary slice is obtained.
[0442] In step 986, classification information is determined for the sampled data or the accessed data represented by the temporary slice.
[0443] In step 987, the sampled data or the accessed data represented by the temporary slice is managed for lineage tracking across one or more layers within the system.
[0444] Embodiments of the present invention can be implemented using a general-purpose or special-purpose digital computer, computing device, machine, or microprocessor, including one or more processors, memories, and / or computer-readable storage media programmed according to the teachings of this disclosure. Skilled programmers can readily create appropriate software coding based on the teachings of this disclosure, which will be apparent to those skilled in the software art.
[0445] In some embodiments, the present invention includes a computer program product that is a non-transitory computer-readable medium (media) storing instructions that can be used to cause a computer to perform any of the processes of the present invention. Examples of storage media can include, but are not limited to, floppy (registered trademark) disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), or other types of storage media or devices suitable for non-transitory storage of instructions and / or data.
[0446] The foregoing description of the present invention has been provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations will be apparent to practitioners skilled in the art.
[0447] For example, some of the above embodiments show performing various calculations by using products such as Wolfram, Yago, Chronos, and Spark, and using data sources such as BDP, SFDC, and S3 as data sources or targets, but the embodiments described herein can also be used with other types of products and data sources that provide similar types of functionality.
[0448] In addition, some of the above embodiments show components, layers, objects, logic, or other features of various embodiments, but such features can be provided as software or program code executable by a computer system or other processing device.
[0449] Embodiments are selected and described with the aim of enabling those skilled in the art to understand the invention, with various embodiments and various modifications, which accompany the principles of the invention and its actual applications, suitable for the intended particular uses. Changes and modifications include any combination of the related disclosed features. It is intended that the scope of the present invention be defined by the following claims and their equivalents.< / publisherurl> < / publisherurl>
Claims
1. A method for use in a computing environment, comprising: one or more processors ingesting data from one or more input hubs including a dataset into a first topic of a scalable input / output layer; the one or more processors writing to a data lake a temporary slice generated by normalizing the data ingested into the first topic by a normalization application of a computing layer; the one or more processors writing to the data lake each temporary slice processed by each of one or more applications of a data flow in the computing layer; the one or more processors publishing the temporary slice written to the data lake to a second topic of the scalable input / output layer by a publish application of the computing layer; the one or more processors publishing the temporary slice published to the second topic to an output hub; the one or more processors updating data reconstruction and lineage tracking information including at least lineage and security for each process between ingestion and publish of data from the input hub; and the lineage indicating how data was acquired and processed.
2. The method of claim 1, wherein metadata received through a foreign language function interface is stored in a knowledge source accessible by the one or more processors for processing of the data flow.
3. The method of claim 1 or 2, wherein the life cycle of the data flow includes where the data was processed.
4. The method according to any one of claims 1 to 3, wherein the method is executed in a cloud or cloud-based computing environment.
5. A system for providing a data governance function for use in data integration or other computing environments, the system comprising a computer including one or more processors, the one or more processors: ingesting data from one or more input hubs including a dataset into a first topic of a scalable input / output layer; Write a temporary slice generated by normalizing the data ingested into the first topic by the normalization application of the computing layer to the data lake. Write each temporary slice processed by each of one or more applications of the data flow in the computing layer to the data lake, where the temporary slice was written to the data lake. Publish the temporary slice written to the data lake to the second topic of the scalable input / output layer by the publish application of the computing layer. Publish the temporary slice published to the second topic to the output HUB. For each process from ingestion to publishing of the data from the input HUB, update data reconstruction and lineage tracking information including at least lineage and security. The lineage is a system that indicates how data was acquired and processed. **Claim 6** A computer-readable program for causing one or more processors to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Declarative language and visualization system for recommended data transformations and repairs
WO2016049460A1