System and method for dynamic lineage tracking, reconstruction, and lifecycle management
Patent Information
- Application Number
- JP2024016373
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2016-08-22
- Filing Date
- 2024-02-06
- Publication Date
- 2025-08-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Modern computing environments face challenges in integrating data from distributed applications with varying configurations due to differences in data types, execution environments, and lifecycle management, requiring significant resource-intensive efforts and domain expertise for software design and integration.
A system leveraging machine learning (ML) for data flow management, providing automated mapping, data governance, and lifecycle management through a graphical user interface, enabling dynamic lineage tracking and reconfiguration of complex data structures across multiple sources and targets.
Facilitates rapid development of scalable software applications by automating data integration, reducing manual effort, and enhancing data governance with real-time recommendations and insights, allowing users to discover new data patterns and improve data flow efficiency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Copyright Notice A portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of this patent document or the patent disclosure in the Patent and Trademark Office patent file or records, but otherwise reserves all copyrights whatsoever.
[0002] Claiming priority This application is a direct sequel to U.S. Provisional Patent Application No. 2016-010363, filed on August 22, 2016, and entitled "SYSTEM AND METHOD FOR AUTOMATED MAPPING OF DATA TYPES BETWEEN CLOUD AND DATABASE SERVICES." No. 62 / 378,143, filed on August 22, 2016, entitled "SYSTEM AND METHOD FOR DYNAMIC, INCREMENTAL RECOMMENDATIONS WITHIN REAL-TIME VISUAL SIMULATION" U.S. Provisional Patent Application No. 62 / 378,146, filed August 22, 2016, entitled "SYSTEM U.S. Provisional Patent Application No. 62 / 378,147, filed August 22, 2016, entitled "AND METHOD FOR INFERENCING OF DATA TRANSFORMATIONS THROUGH PATTERN DECOMPOSITION" “SYSTEM AND METHOD FOR ONTOLOGY INDUCTION THROUGH STATISTICAL PROFILING AND No. 62 / 378,150, filed on Aug. 22, 2016, entitled “SYSTEM AND METHOD FOR METADATA-DRIVEN EXTERNAL INTERFACE GENERATION OF APPLICATION PROGRAMMING INTERFACES” and U.S. Provisional Patent Application No. 62 / 378,150, filed on Aug. 22, 2016, entitled “SYSTEM AND METHOD FOR METADATA-DRIVEN EXTERNAL INTERFACE GENERATION OF APPLICATION PROGRAMMING INTERFACES.” No. 62 / 378,151, and U.S. Provisional Patent Application No. 62 / 378,152, filed August 22, 2016, entitled “SYSTEM AND METHOD FOR DYNAMIC LINEAGE TRACKING AND RECONSTRUCTION OF COMPLEX BUSINESS ENTITIES WITH HIGH-LEVEL POLICIES.” No. 60 / 339,933 filed on Oct. 23, 2003, and claims priority based on, and each of the above applications is incorporated herein by reference.
[0003] FIELD OF THEINVENTION FIELD OF THE DISCLOSURE Embodiments of the present invention relate generally to methods for integrating data from various sources, and more particularly to supporting dynamic lineage tracking, reconfiguration, and lifecycle management. [Background technology]
[0004] background Many modern computing environments require the ability to share large amounts of data among various types of software applications. However, distributed applications may differ significantly in their configuration, for example, due to differences in the respective data types supported or the respective execution environments. Applications may differ according to, for example, their application programming interfaces, runtime environments, deployment techniques, life cycle management, or security management.
[0005] Software design tools intended to be used to develop such distributed applications tend to be resource intensive and often require the services of domain model experts to manage the application and data integration. As a result, application developers faced with the challenge of building complex, scalable distributed applications that are used to integrate different kinds of data in different kinds of execution environments typically must expend a great deal of effort to design, build, and configure these applications. Summary of the Invention
[0006] overview According to various embodiments, the present specification provides a system (data artificial intelligence system) for use in data integration or other computing environments that utilizes machine learning (ML, dataflow machine learning, DFML) for managing data flows (dataflows, DF) and building composite dataflow software applications (dataflow applications, pipelines). According to an embodiment, the system can provide data governance functionality, which can include providing provenance (where does a particular piece of data come from), lineage (how is this data obtained / processed) for each slice of data that is temporally related to a particular snapshot. data), security (who was responsible for this data), classification (what is this data related to), impact (how much impact does this data have on the business), retention (how long should this data last) and validity (how long should this data last) Efficacy (should this data be excluded / included for analysis / processing or not?). These can be used in lifecycle decisions and data flow recommendations. [Brief description of the drawings]
[0007] [Figure 1] FIG. 1 illustrates a system for providing dataflow artificial intelligence, according to an embodiment. [Diagram 2] FIG. 2 illustrates an event-driven architecture including an event coordinator for use in a system according to an embodiment. [Diagram 3] FIG. 2 illustrates steps in a data flow according to an embodiment. [Figure 4] FIG. 2 illustrates an example of a data flow including multiple sources according to an embodiment. [Diagram 5] FIG. 2 illustrates an example of using dataflow with a pipeline according to an embodiment. [Figure 6] FIG. 2 illustrates an example of using an ingest / publish engine and ingest / publish services with a pipeline according to an embodiment. [Figure 7] FIG. 1 illustrates a process of ingesting and training from a HUB, according to an embodiment. [Figure 8] FIG. 1 illustrates a process for building a model, according to an embodiment. [Figure 9] FIG. 1 illustrates a process for classifying a dataset or entities from a newly added HUB, according to one embodiment. [Figure 10] FIG. 13 further illustrates the process of classifying a dataset or entities from a newly added HUB according to one embodiment. [Figure 11] FIG. 13 further illustrates the process of classifying a dataset or entities from a newly added HUB according to one embodiment. [Figure 12] FIG. 2 illustrates an object diagram for use in functional classification according to one embodiment. [Figure 13] FIG. 2 illustrates an example of a dimensional function type classification, according to an embodiment. [Figure 14] FIG. 1 illustrates an example of cube function classification according to an embodiment. [Figure 15] FIG. 1 illustrates an example of the use of functional type classification to evaluate the functional type of a business entity according to an embodiment. [Figure 16] FIG. 2 illustrates an object diagram for use in a function transformation according to one embodiment. [Figure 17] FIG. 2 illustrates the operation of a recommendation engine according to an embodiment. [Figure 18] FIG. 1 illustrates the use of a data lake, according to an embodiment. [Figure 19] FIG. 1 illustrates managing a data lake using a data-driven strategy according to an embodiment. [Figure 20] FIG. 1 illustrates managing a data lake using a process-driven strategy according to an embodiment. [Figure 21] FIG. 2 illustrates the use of a pipeline compiler, according to an embodiment. [Figure 22] FIG. 2 illustrates an example of a pipeline graph according to an embodiment. [Figure 23] FIG. 2 illustrates an example of a data pipeline, according to an embodiment. [Figure 24] FIG. 2 illustrates another example of a data pipeline according to an embodiment. [Diagram 25] FIG. 2 illustrates an example of an orchestration pipeline, according to an embodiment. [Figure 26] FIG. 2 further illustrates an example of an orchestration pipeline, according to an embodiment. [Figure 27] FIG. 2 illustrates the use of a coordination fabric that includes a messaging system according to an embodiment. [Figure 28] FIG. 13 further illustrates the use of a coordination fabric that includes a messaging system according to an embodiment. [Figure 29]FIG. 2 illustrates an on-premise agent for use with the system, according to an embodiment. [Diagram 30] FIG. 2 illustrates a data flow process according to an embodiment. [Diagram 31] FIG. 2 illustrates automatic mapping of data types according to an embodiment. [Diagram 32] FIG. 2 illustrates an automated map service for generating mappings, according to an embodiment. [Diagram 33] FIG. 2 illustrates an example of a mapping between a source schema and a target schema according to an embodiment. [Diagram 34] FIG. 13 illustrates another example of a mapping between a source schema and a target schema according to an embodiment. [Diagram 35] FIG. 2 illustrates a process for providing automatic mapping of data types according to an embodiment. [Diagram 36] FIG. 1 illustrates a system for displaying one or more semantic actions enabled for accessed data, according to one embodiment. [Figure 37] FIG. 2 illustrates a graphical user interface displaying one or more semantic actions enabled for accessed data, according to an embodiment. [Figure 38] FIG. 10 further illustrates a graphical user interface displaying one or more semantic actions enabled for accessed data, according to an embodiment. [Figure 39] FIG. 1 illustrates a process for displaying one or more semantic actions enabled for accessed data according to an embodiment. [Diagram 40] FIG. 2 illustrates a means for identifying patterns of transformations in a data flow for one or more functional expressions generated for each of one or more applications, according to one embodiment. [Diagram 41]FIG. 2 illustrates an example of identifying patterns of transformations in a data flow for one or more function expressions according to an embodiment. [Diagram 42] FIG. 2 illustrates an object diagram for use in identifying patterns of transformations in data flows for one or more function expressions generated for each of one or more applications, according to one embodiment. [Diagram 43] FIG. 2 illustrates a process for identifying patterns of transformations in data flows for one or more function expressions generated for each of one or more applications, according to one embodiment. [Diagram 44] FIG. 1 illustrates a system for generating functional rules according to an embodiment. [Diagram 45] FIG. 2 further illustrates a system for generating functional rules according to an embodiment. [Figure 46] FIG. 2 illustrates an object diagram used to generate functional rules according to one embodiment. [Figure 47] FIG. 2 illustrates a process for generating a functional system based on one or more generated rules, according to one embodiment. [Figure 48] FIG. 1 illustrates a system for identifying patterns for use in providing recommendations regarding data flows based on information provided via a multilingual function interface, according to one embodiment. [Figure 49] FIG. 1 illustrates identifying patterns for use in providing recommendations regarding data flows based on information provided via a foreign language function interface, according to one embodiment. [Figure 50] FIG. 13 further illustrates identifying patterns for use in providing recommendations regarding data flows based on information provided via a foreign language function interface, according to an embodiment. [Figure 51] FIG. 1 illustrates a process for identifying patterns for use in providing recommendations regarding data flows based on information provided via a foreign language function interface, according to one embodiment. [Figure 52] FIG. 1 illustrates management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. [Figure 53] FIG. 13 further illustrates management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. [Figure 54] FIG. 13 further illustrates management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. [Figure 55] FIG. 13 further illustrates management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. [Figure 56] FIG. 13 further illustrates management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. [Figure 57] FIG. 13 further illustrates management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. [Figure 58] FIG. 13 further illustrates management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. [Figure 59] FIG. 1 illustrates a process for managing sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] Detailed Description The above description, together with other embodiments and features thereof, will become apparent by reference to the following description, including the specification and claims, and the accompanying drawings. In the following description, specific details for the purpose of explanation are set forth in order to provide a thorough understanding of various embodiments of the present invention. However, it will be apparent that various embodiments can be practiced without these specific details. The following description, including the specification and claims, and the accompanying drawings are not intended to be limiting.
[0009] Introduction According to various embodiments, the present specification provides a method for managing data flows (DFs) and This paper describes systems (data artificial intelligence systems, data AI systems) used in data integration or other computing environments that leverage machine learning (ML, data flow machine learning, DFML) to build complex data flow software applications (data flow applications, pipelines).
[0010] According to an embodiment, the system can provide support for automatic mapping of complex data structures, datasets or entities between one or more data sources or data targets, referred to in some embodiments herein as HUBs. Automatic mapping can be driven by metadata, schemas, and statistical profiling of datasets, and can be used to map source datasets or entities associated with an input HUB to target datasets or entities, or vice versa, to generate output data prepared in a format or organization (projection) used by one or more output HUBs.
[0011] According to one embodiment, the system includes a visual environment used in the system, referred to herein in some embodiments as a Pipeline Editor or Lambda Studio IDE. The system may include a graphical user interface that provides an environment for developing and executing software applications, including providing real-time recommendations for performing semantic actions on data accessed from an input hub based on an understanding of the meaning or semantics associated with the data.
[0012] According to an embodiment, the system can provide a service for recommending actions and transformations on input data based on patterns identified from a functional decomposition of a software application's data flow, including determining possible transformations of the data flow in a subsequent application. The data flow can be decomposed into models that describe transformations of the data, predicates, and business rules that are applied to the data, and attributes used within the data flow.
[0013] According to an embodiment, the system can perform ontology analysis of the schema definition to determine the type of data and each data set or entity associated with the schema, and generate or update a model from a reference schema that includes an ontology defined based on the relationship between the data set or entity and their attributes. The reference HUB that includes one or more schemas can be used to analyze the data flow and further classify or make recommendations, such as, for example, transforming, enriching, filtering, or cross-entity data fusion of input data.
[0014] According to an embodiment, the system provides a programmatic interface, referred to in some embodiments herein as a foreign language functional interface, that allows users or third parties to extend the functionality of the system by declaratively specifying services, functional and business types, semantic actions, and patterns or predefined composite data flows based on functional and business types.
[0015] According to one embodiment, the system can provide data governance capabilities, such as providing provenance (where does a particular piece of data come from), lineage, and time-related information for each slice of data that is temporally related to a particular snapshot. (how was this data obtained / processed?), security (who was responsible for this data?), classification (what is this data related to?), impact (how much impact does this data have on the business?), retention (how long does this data last?) These are the lifecycle decisions (how long should this data persist), and validity (should this data be excluded / included for analysis / processing or not). These can then be used in lifecycle decisions and data flow recommendations.
[0016] According to an embodiment, the system can be implemented as a service, e.g., a cloud service offered within a cloud-based computing environment, and can act as a single control point for the design, simulation, deployment, development, operation, and analysis of data for use in software applications, including enabling data input from one or more data sources (e.g., an input HUB in one embodiment), providing a graphical user interface that allows a user to specify an application for this data, and scaling the data according to the intended data destination, use, or target (e.g., an output HUB in one embodiment).
[0017] In accordance with certain embodiments, the terms "input" and "output" when used herein in connection with a particular HUB are provided merely as labels reflecting the apparent flow of data in a particular use case or example, and are not intended to limit the type or functionality of a particular HUB.
[0018] For example, according to one embodiment, an input HUB that functions as a source of data can simultaneously or at another time also function as an output HUB or as a target receiving this data or other data, and vice versa.
[0019] Additionally, while for illustrative purposes some of the examples described herein show the use of input and output hubs, in actual implementations, according to certain embodiments, a data integration or other computing environment may include multiple such hubs, at least some of which function as both input and / or output hubs.
[0020] According to an embodiment, the system allows for rapid development of software applications in large-scale, e.g., cloud-based computing environments. In such environments, data models will evolve rapidly. Also, features such as search, recommendation, or suggestion are useful business terms. In such environments, combining artificial intelligence (AI) with semantic search empowers users to achieve more with their existing systems. For example, integrated interactions such as attribute-level mapping can be recommended based on understanding metadata, data, and user interactions with the system.
[0021] According to an embodiment, the system can also be used to suggest complex cases, such as interesting dimensional edges that can be used for information analysis and allow users to discover previously unknown facts in their data.
[0022] In some embodiments, the system provides a graphical user interface that can enable automation of manual tasks (e.g., recommendations or suggestions) and leverage the integration of machine learning and probabilistic knowledge to provide useful context for users and enable discovery and semantic-driven solutions, such as: , creating data warehouses, scaling services, creating and enriching data, and designing and monitoring software applications.
[0023] According to various embodiments, the system may include or utilize some or all of the following features:
[0024] Design-Time System: According to one embodiment, a computing environment that enables the design, creation, monitoring, and management of software applications (e.g., data flow applications, pipelines, or Lambda applications), including the use of a Data AI subsystem that provides, for example, machine learning capabilities.
[0025] Runtime System: According to one embodiment, a computing environment that enables the execution of software applications (e.g., dataflow applications, pipelines, or Lambda applications) and receives input from and provides recommendations to the design-time system.
[0026] Pipeline: A declarative means for specifying a processing pipeline having multiple stages or semantic actions according to an embodiment. Each of the multiple stages or semantic actions corresponds to a function, such as one or more of filtering, combining, enriching, transforming, or fusing input data to produce output data. A data flow software application or data flow application that represents data flows, for example in DFML. According to an embodiment, the system supports a declarative pipeline design that allows the same code base to be used (e.g., with the Spark runtime platform) for both batch (historical) and real-time (streaming) data processing. The pipeline also supports building pipelines or applications that can work against real-time data streams for real-time data analytics. Reprocessing of data due to changes in the pipeline design can be handled through rolling upgrades of deployed pipelines. According to an embodiment, the pipeline can be provided as a Lambda application that can accommodate processing of real-time and batch data in different batch and real-time layers.
[0027] HUB: According to an embodiment, a data source or target (cloud or on-premise) that contains a dataset or entity. A data source that can be introspected and from which data can be consumed or to which data can be published. A data source contains a dataset or entity, which has attributes, semantics or relationships to other datasets or entities. Examples of HUBs include streaming data, telemetry, batch-based, structured or unstructured, or other types of data sources. Data can be received from a HUB that is associated with a source dataset or entity and can be mapped to a target dataset or entity in the same or another HUB.
[0028] System HUB: According to an embodiment, the System HUB can act as a knowledge source that stores other metadata and profile information and data sets or entities that can be associated with other HUBs, and can also act like a regular HUB as a source or recipient of data to be processed. For example, in DFML, it is a central repository where metadata and system state are managed.
[0029] Dataset (Entity): According to an embodiment, a data structure that includes attributes (e.g., columns). This can be, for example, a database table, view, file, or API that can be owned by or associated with one or more HUBs. According to an embodiment, one or more business entities, for example, customer records, can function as semantic business types and are stored as data components, for example, tables in a HUB. A dataset or entity can have relationships with other datasets or entities, along with attributes, for example, columns in a table, and data types, for example, strings or integers. According to an embodiment, the system supports schema-agnostic processing of all types of data (e.g., structured, semi-structured, or unstructured data), for example, during enrichment, preparation, transformation, model training, or scoring operations.
[0030] Data AI Subsystem: According to an embodiment, a component of a system, e.g., a Data AI System, that is responsible for machine learning and semantic related functions. This may include one or more of searching, profiling, providing a recommendation engine, or supporting automatic mapping. The Data AI Subsystem may be configured to communicate with software development components, such as a design-time system, e.g., Lambda Studio, via an event coordinator. The data AI subsystem can support the behavior of a data flow application (e.g., pipeline, Lambda application) by providing recommendations based on the continuous processing of data by the data flow application (e.g., pipeline, Lambda application), e.g., recommend modifications to an existing pipeline to leverage the data being processed. The Data AI subsystem can analyze the volume of input data and continuously update the domain knowledge model. During the processing of a data flow application (e.g., pipeline), each stage of the pipeline can, e.g., accept or reject a recommended semantic action by processing the updated domain model and input from a user based on the recommended alternatives or choices provided by the Data AI subsystem.
[0031] Event Coordinator: According to one embodiment, an event-driven architecture (EDA) coordinator that operates between a design-time system and a run-time system. The event coordinator is a component of the data flow system that coordinates events related to designing, creating, monitoring, and managing data flow applications (e.g., pipelines, Lambda applications). For example, the event coordinator receives published notifications of data (e.g., new data conforming to known data types) from the HUB, normalizes the data from the HUB, and provides the normalized data to a set of subscribers, e.g., for consumption by pipelines or other downstream consumers. The event coordinator can also receive notifications of state transactions in the system for use in lineage tracking or logging, including temporary slice creation and schema evolution.
[0032] Profiling: In one embodiment, the operation of extracting a sample of data from a HUB, profiling the data provided by the HUB, as well as the data sets or entities and attributes within the HUB, determining metrics associated with the sampling of the HUB, and updating metadata associated with the HUB to reflect the data profile in the HUB.
[0033] Software Development Component (Lambda Studio): In one embodiment, A design-time system tool that allows users to create, monitor, and manage the lifecycle of a Lambda application or pipeline as a pipeline of semantic actions by providing a graphical user interface. For example, a graphical user interface that allows users to design a pipeline or Lambda application. The interface (UI, GUI) or studio.
[0034] Semantic Action: According to an embodiment, a data transformation function, e.g., a relational algebra operation. An action that a dataflow application (e.g., a pipeline, a Lambda application) can perform on a dataset or entity in a HUB for projection to another entity. Semantic actions act as higher-order functions available to different models or HUBs that can receive dataset inputs and generate dataset outputs. Semantic actions can include mappings. Semantic actions can include mappings that can be continuously updated, e.g., by a data AI subsystem, in response to processing of data, e.g., as part of a pipeline or Lambda application.
[0035] Mapping: According to one embodiment, a first (e.g., source) data set or entity provided, for example, by the Data AI Subsystem and made accessible through the design-time system, for example, via the software development component Lambda Studio. A recommended mapping of semantic actions between an entity and another (e.g., target) dataset or entity. For example, the Data AI subsystem can provide automatic mapping as a service. Automatic mapping can be driven by metadata, schemas, and statistical profiling of datasets, based on machine learning analysis of metadata or data inputs mapped to the HUB.
[0036] Pattern: According to one embodiment, a pattern of semantic actions that a data flow application (e.g., a pipeline, a Lambda application) can perform. Templates can be used to provide a definition of a pattern that can be reused by other applications. The logical flow of data and associated transformations that are typically associated with business semantics and processes.
[0037] Policy: According to one embodiment, a set of policies that control how a dataflow application, e.g., a pipeline, a Lambda application, is scheduled, which users or components can access which HUBs and semantic actions, and how data should be aged, or other considerations. Configuration settings that dictate how, e.g., a pipeline, should be scheduled, executed, or accessed.
[0038] Application Design Services: According to an embodiment, data flow, e.g., pipeline, services for Lambda applications, e.g., validation, compilation, packaging, and deployment to other, e.g., DFML services (e.g., UI, system facades), etc. Software development components, e.g., Lambda Studio (e.g., Inspect the pipeline (e.g., its inputs and outputs), persist the pipeline, and run the pipeline and the Lambda application. to control the deployment to a system (e.g., a Spark cluster) for A design-time system component that can be used to manage the lifecycle or state of an application.
[0039] Edge Layer: According to one embodiment, a layer that collects data and forwards it to a scalable input / output layer, e.g., as a store and forward layer. A runtime system component that includes one or more nodes that can receive data, e.g., via a gateway, that is accessible to the Internet, and that includes security and other features that support, e.g., secure access to the data AI system.
[0040] Compute Layer: According to one embodiment, the application execution and data processing layer (e.g., Spark). A runtime system component that functions as a distributed processing component, e.g., Spark cloud service, a cluster of compute nodes, a collection of virtual machines, or other components or nodes, and is used to execute pipelines, Lambda applications, for example. In a multi-tenant environment, nodes in the compute layer can be allocated to tenants for use in the execution of pipelines or Lambda applications by those tenants.
[0041] Scalable Input / Output (I / O) Layer: According to one embodiment, a scalable data persistence and access layer organized as topics and partitions (e.g., Kafka) is provided. Data can be moved around the system and accessed by various A runtime system component that provides a queue or other logical storage that can be shared among multiple components, such as a Kafka environment. Thus, the scalable I / O layer can be shared among multiple tenants.
[0042] Data Lake: According to one embodiment, a repository of persistence of information from a system HUB or other components, typically a repository of data, e.g., in DFML, that is normalized or processed by, e.g., pipelines, Lambda applications, and consumed by other pipelines, Lambda applications, or publishing layers.
[0043] Registry: According to one embodiment, one or more information repositories for storing, e.g., functional and business types used to decompose pipelines, Lambda applications, into their functional components.
[0044] Dataflow Machine Learning (DFML): According to one embodiment, a data integration, dataflow management system that leverages machine learning (ML) to assist in building composite dataflow applications.
[0045] Metadata: The underlying definition, description of a dataset or entities and attributes and their relationships, according to one embodiment. It can also be descriptive data about an artifact, for example in DFML.
[0046] Data: According to an embodiment, application data represented by a data set or entity, which may be a batch or stream. For example, a customer, an order, or a product.
[0047] System Facade: A unified API layer for accessing the functionality of, for example, the DFML event-driven architecture, according to one embodiment.
[0048] Data AI Subsystem: According to an embodiment, provides artificial intelligence (AI) services including, but not limited to, search, automap, recommendation, or profiling.
[0049] Streaming Entity: A continuous input of data and near real-time processing and output requirements that may support acceleration of the rate of data, according to an embodiment.
[0050] Batch Entity: A scheduled or on-demand data ingestion event that can be characterized by an emphasis on volume, according to one embodiment. Chong.
[0051] Data Slice: A partition of data, typically marked by time, according to one embodiment.
[0052] Rule: According to one embodiment, illustratively represents a directive that governs an artifact in DFML, such as data rules, relationship rules, metadata rules, and compound or hybrid rules.
[0053] Recommendation (Data AI): According to one embodiment, a recommended set of actions, typically represented by one or more semantic actions or fine-grained directives to aid in the design of, e.g., a pipeline, Lambda application.
[0054] Search (Data AI): According to one embodiment, a semantic search, for example in DFML, characterized by context and user intent, to return relevant artifacts.
[0055] AutoMap (Data AI): A type of recommendation that shortlists candidate source or target datasets or entities to be tried in a dataflow, according to one embodiment.
[0056] Data Profiling (Data AI): According to one embodiment, the collection of several metrics, e.g., minimum, maximum, interquartile range, or spread, that characterize the data in an attribute belonging to a dataset or entity.
[0057] Action Parameter: According to one embodiment, a reference to a dataset on which a semantic action is performed, such as a parameter for an equi-join in a pipeline or Lambda application, for example.
[0058] Foreign Function Interface: According to one embodiment, a mechanism for registering and invoking services (and semantic actions) as part of, for example, the DFML Lambda application framework, which can be used to extend, for example, the functionality or transformation vocabulary in DFML.
[0059] Service: According to one embodiment, a collection of semantic actions that can be characterized by a data integration stage (eg, preparation, discovery, transformation, or visualization), e.g., an unsupported artifact in DFML.
[0060] Service Registry: A repository of services, their semantic actions and other instance information, according to one embodiment.
[0061] Data Lifecycle: The stages in the use of data, for example in DFML, starting with ingestion and ending with publishing, according to one embodiment.
[0062] Metadata Harvesting: Collecting metadata and sample data for profiling, typically after registration of a HUB, according to one embodiment.
[0063] Pipeline Normalization: According to one embodiment, standardizing data into a specific format that facilitates consumption by a pipeline, e.g., a Lambda application.
[0064] Monitoring: Identifying, measuring, and evaluating the execution of, for example, a pipeline, a Lambda application, according to one embodiment.
[0065] Ingest: According to one embodiment, taking in data through an edge layer, for example in DFML.
[0066] Publish: Writing data to a target endpoint, for example from DFML, according to one embodiment.
[0067] Data AI System FIG. 1 is a diagram illustrating a system for providing dataflow artificial intelligence according to one embodiment.
[0068] As shown in FIG. 1, according to an embodiment, a system, e.g., data AI system 150, can provide one or more services for processing and transforming data, e.g., business data, consumer data, and enterprise data, including the use of machine learning processing used in conjunction with various computational assets, e.g., databases, cloud data warehouses, storage systems, or storage services.
[0069] According to an embodiment, the computing assets may be cloud-based, enterprise-based, on-premise or agent-based. The various elements of the system may be connected by one or more networks 130.
[0070] According to an embodiment, the system may include one or more input HUBs 110 (eg, a source of data, data source) and output HUBs 180 (eg, a target of data, data target).
[0071] According to one embodiment, each input HUB, for example HUB 111, may include multiple (source) data sets or entities 192.
[0072] According to one embodiment, an example of an input hub is a database management system (DB, DBMS) 112 (e.g., an on-line transaction processing system (OLTP), a business intelligence system, or an online In such instances, the data provided by a source, such as a data management system, may be structured or semi-structured.
[0073] According to an embodiment, other examples of input HUBs may include a cloud store / object store 114 (e.g., AWS S3 or another object store), which may be an object bucket or clickstream source having unstructured data, a data cloud 116 (e.g., a third-party cloud), a streaming data source 118 (e.g., AWS Kinetics or another streaming data source), or other input source 119.
[0074] According to one embodiment, the input hub may include a data source, for example, receiving data from an Oracle Big Data Prep (BDP) service.
[0075] According to an embodiment, the system may include one or more output HUBs 180 (e.g., output destinations). Each output HUB, e.g., HUB 181, may include multiple (target) data sets or entities 194.
[0076] According to one embodiment, examples of output HUBs include public clouds 182, data clouds 184 (e.g., AWS and Azure), on-premise clouds 186, or other The system may include output targets 187. The data outputs provided by the system may be generated for data flow applications (e.g., pipelines, Lambda applications) accessible at the output HUB.
[0077] According to an embodiment, an example of a public cloud may include, for example, Oracle Public Cloud, which includes, for example, Big Data Prep cloud services, Exadata These may include cloud services, Big Data Discovery cloud services, and Business Intelligence cloud services.
[0078] According to an embodiment, the system can be realized as a unified platform for streaming and on-demand (batch) data processing delivered to users as a service (e.g., software as a service), providing scalable multi-tenant data processing for multiple input HUBs. Data is analyzed in real time using machine learning techniques and visual insights and monitoring provided by a graphical user interface as part of the service. Data sets can be fused from multiple input HUBs for output to an output HUB. For example, through the data processing services provided by the system, data can be generated for a data warehouse and populated into one or more output HUBs.
[0079] According to an embodiment, the system provides a declarative and programmatic topology for transforming, enriching, routing, classifying, and blending data and may include a design-time system 160 and a run-time system 170. Users can create applications, such as data flow applications (e.g., pipelines, Lambda applications) 190, designed to perform data processing.
[0080] According to one embodiment, the design-time system can enable a user to design a dataflow application, specify the dataflow, and specify the data for dataflow processing. For example, the design-time system can provide a software development component 162 (referred to in one embodiment herein as Lambda Studio) that provides a graphical user interface for the creation of dataflow applications. Cut.
[0081] For example, according to one embodiment, a user can use a software development component to specify input and output hubs to create a data flow for an application. A graphical user interface can present an interface for services for data integration, allowing a user to create, manipulate, and manage data flows for an application. This includes the ability to dynamically monitor and manage data flow pipelines, such as observing data lineage and performing method analysis.
[0082] According to an embodiment, the design-time system may also include application design services 164 for deploying dataflow applications to a run-time system.
[0083] According to an embodiment, the design-time system may also include one or more system HUBs 166 (e.g., metadata repositories) for storing metadata for processing the data flows. The one or more system HUBs may store samples of data, such as data types including functional and business data types. The information in the system HUB can be used to perform one or more of the techniques disclosed herein. The data lake 167 component can act as a repository for persistence of information from the system HUB.
[0084] According to an embodiment, the design-time system may also include a data artificial intelligence (AI) subsystem 168 that performs operations for data artificial intelligence processing. Such operations may include using ML techniques, such as search and retrieval. The data AI subsystem may sample data to generate metadata for the system HUB.
[0085] According to an embodiment, the data AI subsystem can perform schema object analysis, metadata analysis, sample data, correlation analysis, and classification analysis for each input HUB. The data AI subsystem can provide rich data to data flow applications by continuously running on the input data, and can provide recommendations, insights, and type induction to pipelines, Lambda applications, for example.
[0086] According to one embodiment, the design-time system enables users to create policies, artifacts, and flows that specify the functional requirements of a use case.
[0087] For example, according to an embodiment, the design-time system can provide a graphical user interface to create a HUB to ingest data and define an ingest policy. The ingest policy can be time-based or on demand from the associated data flow. Once an input HUB is selected, data can be sampled from the input HUB to profile the source, for example by running metadata queries, taking samples, and taking user-defined inputs. The profile can be stored in the system HUB. The graphical user interface allows multiple sources to be combined to define a data flow pipeline. This can be done by creating a script or by using a guided editor. The guided editor can visualize the data at each step. The graphical user interface can provide access to a recommendation service that suggests how the data cloud can be modified, enriched, or combined, for example.
[0088] According to an embodiment, during design time, the application design service can suggest suitable structures to parse the resulting content. The application design service can suggest actions and associated dimensional hierarchies by using knowledge services (functional classification). Once this is completed, the design time system can recommend the data flows required to take the blended data from the previous pipeline and populate the dimensional target structures. Based on the dependency analysis, the orchestration flows can also be derived and generated to load / refresh the target schema. For forward engineering use cases, the design time system can also generate a HUB to host the target structures and create the target schema.
[0089] According to an embodiment, the runtime system can perform operations during the run-time of a service for processing data.
[0090] According to one embodiment, during run-time or operational mode, user-created policy and flow specifications are applied and / or executed. For example, such processing may include, but is not limited to, the following: This may include invoking ingest, transform, model, and publish services to process the data within the pipeline.
[0091] According to an embodiment, the runtime system may include an edge layer 172, a scalable input / output (I / O) layer 174, and a distributed processing system or computation layer 176. At runtime (e.g., as data is ingested from one or more input HUBs 110), the edge layer may receive data due to events that generate data.
[0092] According to one embodiment, an event coordinator 165 operates between the design-time and run-time systems to coordinate events related to the design, creation, monitoring, and management of dataflow applications (e.g., pipelines, Lambda applications).
[0093] According to one embodiment, the edge layer sends data to a scalable I / O layer, which routes the data to a distributed processing system or compute layer.
[0094] According to one embodiment, a distributed processing system or computation layer can process the data for output by implementing a pipeline process (per tenant). The distributed processing system can be implemented, for example, using Apache Spark and Alluxio. , it can sample data into the data lake and then output this data to an output HUB. The distributed processing system can communicate with the scalable input / output layer to launch and process the data.
[0095] According to one embodiment, a data AI system including some or all of the above components can be provided on or executed by one or more computers including, for example, one or more processors (CPUs), memory, and persistent storage (198).
[0096] Event-Driven Architecture As previously mentioned, according to one embodiment, the system may include an event-driven architecture (EDA) component or event coordinator that operates between the design-time and run-time systems to coordinate events related to the design, creation, monitoring, and management of dataflow applications (e.g., pipelines, Lambda applications).
[0097] FIG. 2 is a diagram illustrating an event-driven architecture including an event coordinator for use in a system according to one embodiment.
[0098] As shown in FIG. 2, according to one embodiment, the event coordinator may include an event queue 202 (e.g., Kafka), an event bootstrapper service 204 (e.g., ExecutorService), and an event configuration publisher / event consumer 206 (e.g., DBCS).
[0099] According to one embodiment, events received at the system façade 208 (e.g., an event API extension) are forwarded to one or more event brokers 210, e.g., a Kafka consumer. The data and / or events can be communicated to various components of the system by the data and / or events, such as, for example, external data 212 (e.g., S3, OSCS, or OGG data) or a graphical user interface 214 (e.g., Input from a data source (e.g., a browser or DFML UI) can be communicated via the event coordinator to other components, such as the application runtime 216, data lake, system hub, data AI subsystem, application design services, and / or ingest 220, publish 230, scheduling 240, or other components described above.
[0100] According to an embodiment, an event broker may be configured as a consumer of stream events. An event bootstrapper can start multiple configured event brokers to process events on behalf of registered subscribers. For processing a given event, each event broker delegates the processing of the event to a registered callback endpoint. An event coordinator enables registration of event types, registration of eventing entities, registration of events, and registration of subscribers. Table 1 provides an example of various event objects, including published and subscribed events.
[0101] [Table 1]
[0102] Event Type According to an embodiment, the event type defines an event state change that is significant to the system, such as the creation of a HUB, modification of a dataflow application, such as a pipeline, a Lambda application, ingesting data for a dataset or entity, or publishing data to a target HUB. An example data format and examples of various event types are shown below and in Table 2.
[0103] [Table 2]
[0104] Eventing Entities According to an embodiment, an eventing entity may be a publisher and / or subscriber of events. For example, an eventing entity may register to publish one or more events and / or be a consumer of one or more events. This includes registering an endpoint or callback URL that is used to notify or send acknowledgments for publications and delegate processing of subscribed events. Examples of eventing entities may include metadata services, ingest services, system HUB artifacts and pipelines, Lambda applications. An example data format and examples of various eventing entities are shown below and in Table 3.
[0105] [Table 3]
[0106] event According to an embodiment, an event is an instance of an event type associated with an eventing entity registered as a publisher and may have subscribers (eventing entities). For example, a metadata service may register for a HUB creation event for publishing and may publish one or more event instances for this event (one instance per HUB created). Examples of various events are shown in Table 4.
[0107] [Table 4] TIFF2024054219000006.tif167170
[0108] example According to one embodiment, the following example shows creating an event type, registering a publish event, registering a subscriber, publishing an event, getting an event type, getting a publisher for the event type, and getting a subscriber for the event type.
[0109]
number
[0110] This is a universally unique ID (UUID), e.g. "8e8 The eventing object returns "7039b-a8b7-4512-862c-fdb05b9b8888". Eventing objects can publish or subscribe to events within a system. For example, service endpoints such as ingest services, metadata services, and application design services may publish or subscribe events, along with static endpoints for acknowledgment, notification, error, or handling. DFML artifacts (e.g., DFMLEntity, DFMLLambdaApp, DFMLHub) can also be registered as eventing objects, and instances of these types can publish or subscribe to events without being registered as eventing objects.
[0111]
number
[0112] The following example registers DFMLLambdaApps (type) as an eventing object.
[0113]
number
[0114] For eventing entities of type HUB, Entity and LambdaApp, <publisherurl>can be annotated to REST endpoint URLs, allowing for event-driven The architecture derives the actual URL by substituting the DFML artifact instance URL. For example, if the notificationEndpointURL is http: / / den00tnk:9021 / <publisherurl> / notification and specified as part of the message If the publisher URL given is hubs / 1234 / entities / 3456, then it will be called for notification. The URL would be http: / / den00tnk:9021 / hubs / 1234 / entities / 3456 / notification The POST returns a UUID, e.g. "185cb819-7599-475b-99a7-65e0bd2ab947".
[0115] Registering for a publish event According to one embodiment, a publish event can be registered as follows:
[0116]
number
[0117] The eventType above is the UUID returned for the registration of the event type DATA_INGESTED, and the publishingEntity above is the DFMLEntity type being registered as the eventing object. This registration returns a UUID, e.g. "2c7a4b6f-73ba-4247-a07a-806ef659def5".
[0118] Registering a Subscriber According to one embodiment, a subscriber can register as follows.
[0119]
number
[0120] The UUID returned from the publish event registration is used as a path segment for the subscriber registration.
[0121]
number
[0122] The publisherURL and publishingObjectType above are the instance and type of the publisher object. Here, a dataflow (e.g. Lambda) application is interested in identifying the URI / lambdaApps / 123456 and subscribing to DATA_INGESTED events from the entity / hubs / 1234 / entities / 3456. This subscription is made using the UUID , which returns, for example, "1d542da1-e18e-4590-82c0-7fe1c55c5bc8".
[0123] Publishing an Event According to one embodiment, an event can be published as follows:
[0124]
number
[0125] The above publisherURL is used when the publishing object is one of the following: DFMLEntity, DFMLHub, or DFMLLambdaApps, and is used to send messages in which subscribers subscribe. It is used to check for an instance of the eventing object to publish, and the publisher URL is used to derive a notification URL when a subscriber successfully processes the message. This publish returns the message body that was part of the published event.
[0126] Get the event type According to one embodiment, the event type can be determined as follows:
[0127]
number
[0128] Get the publisher for an event type
[0129]
number
[0130] Get the subscribers for an event type According to one embodiment, subscribers for an event type can be found as follows:
[0131]
number
[0132] The above description is provided by way of example to illustrate specific embodiments of event coordinators, event types, eventing entities, and events. According to other embodiments, other kinds of EDAs can be used to coordinate events for designing, creating, monitoring, and managing data flow applications by providing communication within the system that operates between the design-time and run-time systems, and other kinds of event types, eventing entities, and events can be supported.
[0133] Data Flow Machine Learning (DFML) Flows As previously mentioned, according to various embodiments, the system can be used in data integration or other computing environments that leverage machine learning (ML, Dataflow Machine Learning, DFML) for managing the flow of data (Dataflow, DF) and building composite dataflow software applications (e.g., dataflow applications, pipelines, Lambda applications).
[0134] FIG. 3 is a diagram illustrating steps in a data flow according to one embodiment. 3, according to one embodiment, processing of a DFML data flow 260 may include multiple steps, including an ingest step 262. In the ingest step, data may be ingested from various sources, for example, Salesforce (SFDC), S3, or DBaaS.
[0135] In a data preparation step 264, the ingested data may be prepared, for example, by de-duplication, standardization, or enrichment.
[0136] In a transformation step 266, the system may transform the data by performing one or more of a merge, filter, or lookup of a dataset.
[0137] In a model step 268, one or more models are generated along with mappings for the models.
[0138] In a publishing step 270, the system can publish the model, specify policies and schedules, and populate target data structures.
[0139] According to an embodiment, the system supports the use of search / recommendation functionality 272 throughout each of its data preparation, transformation, and model steps. Users can interact with the system through a set of well-defined services whose breadth of functionality is encapsulated within a data integration framework. This set of services defines a logical view of the system. For example, in design mode, users can create flows, artifacts, and policies that define the functional requirements of a particular use case.
[0140] FIG. 4 is a diagram illustrating an example of a data flow including multiple sources, according to an embodiment.
[0141] As shown in the example data flow 280 in FIG. 4, according to one embodiment, what is needed is a cloud service called SFDC and FACS (Fusion Apps Cloud Service). Content from multiple sources 282 shown can be sent to OSCS (Oracle Storage Cloud The goal of the Oracle Business Intelligence Cloud Service (BICS) is to ingest data from Oracle Business Intelligence Cloud Service (BIS) along with several files in the BICS environment, blend this information so that it can be used to analyze the desired content, derive target cubes and dimensions, map the blended content to the target structure, and make this content available in the BICS environment along with the dimensional model, which includes the Ingest, Transform 266A / 266B, Model, Orchestrate 292, and Deploy 294 steps.
[0142] The illustrated examples are provided to illustrate the techniques described herein, and the functionality described herein is not limited to use with these particular data sources.
[0143] According to an embodiment, in the ingest step, a HUB is created in the data lake to receive SFDC content in order to access and ingest this content. This can be done, for example, by selecting an SFDC adapter with the relevant access mode (JDBC, REST, SOAP), creating the HUB, giving it a name, and defining an ingest policy that is time-based or according to the requirements of the relevant data flow.
[0144] According to an embodiment, a similar process can be performed for the other two sources, the difference being that for the OSCS source, the schema may not be known at the start, but instead can be obtained by some means (e.g., metadata query, sampling, or user-defined).
[0145] According to an embodiment, the source of the data can be further investigated by optionally profiling it, which can help derive recommendations later in the integration flow.
[0146] According to one embodiment, the next step is to specify how the separate sources are joined around a central item. This is typically the basis (facts) of the analysis and can be achieved by specifying a data flow pipeline. This can be done directly, by creating a pipeline Domain Specific Language (DSL) script, or by using a guided editor, where the user can see the effect on the data at each step and can use a recommendation service that suggests how the data can be, for example, modified, enriched, or combined.
[0147] At this point, the user can request that the system suggest suitable structures to analyze the resulting content. For example, according to one embodiment, the system can suggest measures and associated dimensional hierarchies by using knowledge services (functional classification). Once this is completed, the system can recommend the data flows required to take the blended data from the previous pipeline and populate the dimensional target structures. It also derives and generates orchestration flows based on dependency analysis to load / refresh the target schema.
[0148] According to one embodiment, the system then creates a HUB to host the target structure and maps it, via an adapter, to the DBCS which generates the data definition language (DDL) required to create the target schema, and can deploy, for example, XDML or some form of data that BICS can use to generate the models required to access the newly created schema. This is populated by executing an orchestration flow and triggering an exhaust service. It is possible.
[0149] FIG. 5 is a diagram illustrating an example of the use of a dataflow with a pipeline, according to one embodiment.
[0150] As shown in FIG. 5, in accordance with one embodiment, the system allows a user to describe the processing of data as it is constructed and executed 304 as an application by specifying a pipeline 302 that represents a data flow, in this example including pipeline steps S1-S5.
[0151] For example, according to one embodiment, a user can invoke ingest, transform, model, and publish services or other services, such as policy 306, execution 310, or persistence service 312, to process the data in the pipeline. A user can also specify a unified flow that can integrate related pipelines by specifying a solution (i.e., a control flow). Typically, a solution models a complete use case, such as the loading of a sales cube and related dimensions.
[0152] Data AI System Components According to an embodiment, adapters allow for connecting to and ingesting data from a variety of endpoints and are application or source type specific.
[0153] According to an embodiment, the system may include a set of predefined adapters. Some of these adapters may leverage other SOA adapters, allowing additional adapters to be registered with the framework. There may be more than one adapter for a given connection type. In that case, the ingest engine selects the most suitable adapter based on the connection type configuration of the HUB.
[0154] FIG. 6 illustrates an example of the use of an ingest / publish engine and an ingest / publish service according to an embodiment.
[0155] 6, according to one embodiment, the pipeline 334 can access the ingest / publish engine 330 via the ingest / publish service 332. The pipeline 334, in this example, is designed to ingest 336 data (e.g., sales data) from an input HUB (e.g., SFDC HUB1), transform 338 the ingested data, and publish 340 the data to an output HUB (e.g., Oracle HUB).
[0156] According to one embodiment, the ingest / publish engine supports multiple connection types 331. One of these types, connection type 342, is associated with one or more adapters 344 that provide access to the HUB.
[0157] For example, as shown in the example of FIG. 6 , according to one embodiment, SFDC connection type 352 can be associated with SFDC-Adp1 adapter 354 and SFDC-Adp2 adapter 356 that provide access to SFDC HUBs 358 and 359, ExDaaS connection type 362 can be associated with ExDaaS-Adp adapter 364 that provides access to ExDaas HUB 366, and Oracle connection type 372 can be associated with Oracle Adp adapter 374 that provides access to Oracle HUB 376.
[0158] Recommendation Engine According to one embodiment, the system may include a recommendation engine or knowledge service that acts as an expert filtering system that predicts / suggests the most relevant of several possible actions that can be taken on the data.
[0159] According to an embodiment, recommendations can be linked to help users follow these recommendations to achieve a given end goal. For example, a recommendation engine can guide a user through a set of steps in transforming a dataset into a data cube and publishing it to a target BI system.
[0160] According to an embodiment, the recommendation engine utilizes three aspects: (A) business type classification, (B) functional classification, and (C) knowledge base. Ontology management and query / search capabilities for datasets or entities can be provided by, for example, a shared ontology derived from YAGO3 with the query API, MRS, and audit repository. Business entity classification can be provided by, for example, an ML pipeline-based classification that identifies business types. Functional classification can be provided by, for example, a deductive rule-based functional classification. Action recommendations can be provided by, for example, an inductive rule-based data preparation, transformation, model, dependency, and associated recommendations.
[0161] Classification Services According to one embodiment, the system provides a classification service that can be categorized into business type classifications and functional type classifications, each of which is further described below.
[0162] Business Type Classification According to one embodiment, the business type of an entity is its phenotype. The observable characteristics of individual attributes within an entity are as important as the definition in identifying the business type of the entity. Classification algorithms use a general definition of a data set or entity, but can also utilize models built with the data to classify the business type of the data set or entity.
[0163] For example, according to one embodiment, a dataset ingested from a HUB can be classified as one of the existing business types known to the system (seeded from the main HUB), or it can be added as a new type if it cannot be classified into an existing business type.
[0164] According to one embodiment, the business type classification is utilized in making recommendations based on inductive reasoning (from transformations defined for similar business types in the pipeline) or simple propositions derived from the classification root entities.
[0165] In summary, according to one embodiment, the classification process is described by the following set of steps: ingest and seed from the main (training) hub, build models and calculate column statistics and register them for use in classification, classify datasets or entities from newly added hubs including creating profiles / calculating column statistics, classify datasets or entities to provide a shortlist of entity models to use based on structure and column statistics, and classify datasets or entities including multi-class classification and compute / predict with models.
[0166] FIG. 7 is a diagram illustrating the process of ingesting and training from a HUB according to one embodiment.
[0167] As shown in FIG. 7, according to one embodiment, data from a HUB 382 (e.g., RelatedIQ source in this example) is fed to a recommendation engine 380 via a dataset 390 (e.g., a Resilient Distributed Dataset :RDD), which in this example includes an accounts dataset 391, an events dataset 392, a contacts dataset 393, a list dataset 394, and a users dataset 395.
[0168] According to an embodiment, multiple type classifiers 400 can be used in conjunction with the ML pipeline 402. For example, GraphX 404, Wolfram / Yago 406, and / or MLlib statistics 408 can be used in seeding the knowledge graph 440 with entity metadata (training or seed data) when the HUB is initially registered.
[0169] According to one embodiment, dataset or entity metadata and data are ingested from source HUBs and stored in the data lake. During model generation 410, the entity metadata (attributes and relationships to other entities) are used to generate a model 420 and a knowledge graph, for example through FP-growth logistic regression 412. The knowledge graph is a set of all datasets or entities, in this example represents events 422, accounts 424, contacts 426, and users 428. As part of the seeding, a regression model is built using the dataset or entity data to calculate attribute statistics (minimum, maximum, average, or probability density).
[0170] FIG. 8 is a diagram illustrating a model building process, according to one embodiment. 8, according to one embodiment, for example when running within Spark environment 430, Spark MLlib statistics can be used to compute column statistics that are added as attribute properties to the knowledge graph. The computed column statistics can be used along with other datasets or entity metadata to shortlist entities whose regression models are used in testing new entities for classification.
[0171] FIG. 9 illustrates a process of classifying a data set or entities from a newly added HUB according to one embodiment.
[0172] As shown in FIG. 9, according to one embodiment, when a new HUB is added, in this example Oracle HUB 442, the data set or entities provided by the HUB, such as party information 444 and customer information 446, are classified by the model as parties 448 based on previously created training or seed data.
[0173] For example, according to one embodiment, column statistics are computed from the data of a new data set or entity, and a set of predicates that represent a subgraph of the entity are computed. This information, along with other metadata available as part of the ingest, is used to create the
[0174] According to one embodiment, the calculation of column statistics is useful in maximum likelihood estimation (MLE) methods, while the subgraphs are useful in regression models of a data set. The set of graph predicates generated for the new entity is used to shortlist candidate entity models for testing and classification of the new entity.
[0175] FIG. 10 further illustrates the process of classifying a data set or entities from a newly added HUB according to one embodiment.
[0176] 10, according to one embodiment, predicates representing a subgraph of a new dataset or entity to be classified are compared with similar subgraphs representing datasets or entities that are already part of the knowledge graph 450. A ranking of matching entities based on the probability of match is used to shortlist entity models to be used in testing for classification of the new entity.
[0177] FIG. 11 further illustrates the process of classifying a data set or entities from a newly added HUB according to one embodiment.
[0178] As shown in FIG. 11, according to one embodiment, the regression models of the shortlisted matching datasets or entities are used to test data from the new dataset or entity. The ML pipeline can be expanded to encompass additional classification methods / models to improve the accuracy of the process. If there is a match within an acceptable threshold, e.g., probability higher than 0.8 in this example, the classification service classifies 452 the new entry. Otherwise, this dataset or entity can be added to the knowledge graph as a new business type. The user can also accept or reject the results. The classification can be verified by accepting or rejecting it.
[0179] Functional Classification According to an embodiment, the functional type of an entity is its genotype. A functional type can also be described as an interface through which transformation actions are specified. For example, a join transformation or a filter is specified on a functional type, in this case a relational entity. In summary, all transformations are specified with a functional type as a parameter.
[0180] FIG. 12 is a diagram illustrating an object diagram for use in functional classification according to one embodiment.
[0181] As shown in FIG. 12 via object diagram 460, according to one embodiment, the system can describe a general case (in this example, a dimension, level, or cube) through a set of rules against which a data set or entity can be evaluated to identify its functional type.
[0182] For example, according to one embodiment, a multidimensional cube can be described by its measurement attributes and dimensions, each of which can itself be defined by their type and other characteristics. A rules engine evaluates business type entities and annotates their function types based on the evaluation.
[0183] FIG. 13 is a diagram illustrating an example of a dimensional function type classification, according to an embodiment. As shown in the example hierarchy of functional classifications 470 illustrated in FIG. 13, according to one embodiment, levels may be defined, for example, by their dimensions and level attributes.
[0184] FIG. 14 illustrates an example of cube function classification according to one embodiment. According to one embodiment, a cube may be defined, for example, by its measurement attributes and dimensions, as shown in the example hierarchy of functional classifications 480 depicted in FIG.
[0185] FIG. 15 illustrates an example of using a functional type classification to evaluate the functional type of a business entity according to an embodiment.
[0186] 15, in this example 490, according to one embodiment, the sales data set must be evaluated by the rules engine as a cube function. Similarly, products, customers, and time must be evaluated as dimensions and levels (e.g., age group, gender).
[0187] According to an embodiment, the rules for identifying the entity function type and data set or entity element in this example are shown below. This includes several rules that can be specified for evaluation of the same function type. For example, a column of type "Date" can be considered as a dimension regardless of whether there is a reference to a parent level entity. Similarly, zip code, gender, and age may only require data rules to identify them as dimensions.
[0188]
number
[0189] FIG. 16 is a diagram illustrating an object diagram for use in function transformation according to one embodiment. 15, in this example 500, according to one embodiment, the transformation function can be specified with a functional type. Business entities (business types) are annotated as functional types, which implies that by default, composite business types are of functional type "entity".
[0190] FIG. 17 is a diagram illustrating the operation of a recommendation engine according to one embodiment. As shown in Figure 17, according to one embodiment, the recommendation engine generates recommendations that are a set of actions defined by business types. Each action is a directive that calls for applying a transformation to a dataset.
[0191] According to one embodiment, the recommendation context 530 includes metadata that abstracts the source of the recommendation and identifies the set of propositions that generated the recommendation. This context allows the recommendation engine to learn and prioritize recommendations based on user responses.
[0192] According to one embodiment, the target entity deduction / mapping unit 512 uses the target definition (and the classification service that annotates the dataset or entity and attribute business types) to make transformation recommendations that facilitate mapping the current dataset to the target. This is typical when a user starts with a known target object (e.g., a sales cube) and builds a pipeline to instantiate the cube.
[0193] According to one embodiment, a template (pipeline / solution) 514 defines a reusable set of pipeline steps and transformations to achieve a desired end result. For example, a template may include steps to enrich, transform, and publish to a data mart. The set of recommendations in this case reflects the template design.
[0194] According to one embodiment, the classification service 516 identifies the business type of a data set or entity ingested from the HUB into the data lake. Recommendations can be made based on transformations applied to similar entities (business types) or in conjunction with the target entity deduction / mapping part.
[0195] According to one embodiment, the function type service 518 annotates the function types that a dataset or entity can have based on defined rules. For example, to generate a cube from a given dataset or join it to a dimension table, it is important to evaluate whether the dataset complies with the rules that define the function type of the cube.
[0196] According to one embodiment, pattern inference from the pipeline component 520 enables the recommendation engine to summarize the transformations performed based on a given business type in existing pipeline definitions of similar contexts and suggest similar transformations as recommendations for the current context.
[0197] According to an embodiment, a recommendation context can be used to process a recommendation 532 , which includes an action 534 , a conversion function 535 , action parameters 536 , function parameters 537 , and a business type 538 .
[0198] Data Lake / Data Management Strategy As previously mentioned, according to one embodiment, the data lake provides a repository of persistence of information from system HUBs or other components.
[0199] FIG. 18 is a diagram illustrating the use of a data lake according to one embodiment. 18, according to one embodiment, a data lake can be associated with one or more data access APIs 540, a cache 542, and a persistence store 544 that work together to receive normalized ingested data for use by multiple pipelines 552, 554, 556.
[0200] According to an embodiment, a variety of different data management strategies can be used to manage the data (performance, scalability) and its lifecycle in the data lake, which can be broadly categorized as data-driven or process-driven.
[0201] FIG. 19 illustrates managing a data lake using a data-driven strategy according to one embodiment.
[0202] As shown in Fig. 19, according to an embodiment, in a data-driven approach, management units are derived based on HUB or data server definitions. For example, in this approach, data from Oracle1 HUB can be stored in the first data center 560 associated with this HUB, and data from SFHUB1 can be stored in the second data center 562 associated with this HUB.
[0203] FIG. 20 illustrates managing a data lake using a process-driven strategy according to one embodiment.
[0204] 20, according to one embodiment, in a process-driven approach, management units are derived based on the associated pipeline from which they access data. For example, in this approach, data associated with a sales pipeline may be stored in a first data center 564 associated with that pipeline, and data from other pipelines (e.g., pipelines 1, 2, 3) may be associated with a second data center 566 associated with those other pipelines.
[0205] Pipeline According to one embodiment, the pipeline specifies the transformation or processing to be performed on the ingested data. The processed data may be stored in the data lake or published to another endpoint, such as DBCS.
[0206] FIG. 21 is a diagram illustrating the use of a pipeline compiler according to one embodiment. 21, according to one embodiment, a pipeline compiler 582 operates between a design environment 570 and an execution environment 580. It includes receiving one or more pipeline metadata 572 and a DSL, e.g., Java DSL 574, JSON DSL 576, Scala DSL 578, and providing output for use in the execution environment, e.g., as a Spark application 584 and / or SQL statements 586.
[0207] FIG. 22 illustrates an example pipeline graph according to an embodiment. As shown in FIG. 22, according to one embodiment, a pipeline 588 includes a list of pipeline steps. Different types of pipeline steps represent different types of operations that can be performed within the pipeline. Each pipeline step may include multiple input data sets and multiple output data sets, typically described by pipeline step parameters. The processing order of operations within the pipeline is defined by binding the output pipeline step parameters from the previous pipeline step to the next pipeline step. In this way, the pipeline steps and the relationships between the pipeline step parameters form a directed acyclic graph (DAG).
[0208] According to an embodiment, a pipeline can be reused in another pipeline if the pipeline contains one or more unique pipeline steps (signature pipelines) that represent the input and output pipeline step parameters of the pipeline. The enclosing pipeline is the pipeline that is reused in the pipeline steps (pipeline usage).
[0209] FIG. 23 is a diagram illustrating an example of a data pipeline, according to an embodiment. According to one embodiment, a data pipeline performs data transformations, as shown in the example data pipeline 600 in FIG. 23. The flow of data through the pipeline is represented as a combination of pipeline step parameters. Various types of pipeline steps are supported for different transformation operations, including, for example, entity (taking data out of the data lake or publishing processed data to the data lake / other HBUs) and join (blending multiple sources).
[0210] FIG. 24 illustrates another example of a data pipeline, according to an embodiment. As shown in the example data pipeline 610 shown in FIG. 24, according to one embodiment, a data pipeline P1 can be reused in another data pipeline P2.
[0211] FIG. 25 is a diagram illustrating an example of an orchestration pipeline, according to an embodiment.
[0212] As shown in the example orchestration pipeline 620 in FIG. 25, according to one embodiment, when an orchestration pipeline is used, the pipeline An in-step can be used to represent a task or job that needs to be executed in the overall orchestration flow. Every pipeline step in an orchestration pipeline shall have one input pipeline step parameter and one output pipeline step parameter. Execution dependencies between tasks can be represented as bonds between pipeline step parameters.
[0213] According to an embodiment, parallel execution of tasks can be scheduled when a pipeline step depends unconditionally on the same previous pipeline step (i.e., fork). If a pipeline step depends on multiple previous paths, the pipeline step waits for all the multiple paths to complete before executing itself (i.e., join). However, this does not always mean that tasks are executed in parallel. The orchestration engine can decide whether to execute tasks serially or in parallel depending on the available resources.
[0214] In the example shown in FIG. 25, according to one embodiment, pipeline step 1 is executed first. If pipeline step 2 and pipeline step 3 are executed in parallel, pipeline step 4 is executed after both pipeline step 2 and pipeline step 3 are completed. The orchestration engine can also execute this orchestration pipeline serially as (pipeline step 1, pipeline step 2, pipeline step 3, pipeline step 4) or (pipeline step 1, pipeline step 3, pipeline step 2, pipeline step 4) as long as the dependencies between the pipeline steps are satisfied.
[0215] FIG. 26 further illustrates an example of an orchestration pipeline, according to an embodiment.
[0216] As shown in the example pipeline 625 in FIG. 26, according to an embodiment, each pipeline step can return a status 630, e.g., a success or error status, depending on its semantics. A dependency between two pipeline steps may be conditional based on the return status of the pipeline steps. In the example shown, pipeline step 1 is executed first, and if it completes successfully, pipeline step 2 is executed, otherwise pipeline step 3 is executed. After either pipeline step 2 or pipeline step 3 is executed, pipeline step 4 is executed.
[0217] According to an embodiment, nesting of orchestration pipelines may allow one orchestration pipeline to reference another orchestration pipeline through a pipeline usage. An orchestration pipeline may also reference a data pipeline as a pipeline usage. The difference between an orchestration pipeline and a data pipeline is that an orchestration pipeline references a data pipeline that does not contain a signature pipeline step, whereas a data pipeline can reuse another data pipeline that does contain a signature pipeline step.
[0218] According to an embodiment, depending on the type of pipeline steps and code optimization, a data pipeline can be generated as a single Spark application running in a Spark cluster, as multiple SQL statements running in DBCS, or as a mix of SQL and Spark code. In the case of an orchestration pipeline, the execution is done in the underlying execution engine or in a web application such as Oozie. A workflow can be generated for execution within the workflow schedule component.
[0219] Coordination Fabric According to one embodiment, the coordination fabric or fabric controller provides the tools necessary to deploy and manage framework components (service providers) and (user-designed) applications, manages application execution and resource requests / allocation, and facilitates interaction between the various components by providing an integration framework (messaging bus).
[0220] FIG. 27 illustrates the use of a coordination fabric with a messaging system according to one embodiment.
[0221] As shown in FIG. 27, according to one embodiment, a messaging system (e.g., Kafka) 650 includes a resource manager 660 (e.g., Yarn / Mesos), a scheduler 662 (e.g., Chronos), an application scheduler 664 (e.g., Spark), and a set of application servers, shown here as nodes 652, 654, 656, and 658. Coordinates interactions between multiple nodes in a network.
[0222] According to one embodiment, a resource manager is used to manage the lifecycle of data computing tasks / applications, including scheduling, monitoring, application execution, resource arbitration and allocation, load balancing, managing the configuration and deployment of components (that are message producers and consumers) within a message-driven component integration framework, upgrading components (services) with no downtime, and upgrading infrastructure with minimal or no disruption to services.
[0223] FIG. 28 is a diagram further illustrating the use of a coordination fabric that includes a messaging system, according to one embodiment.
[0224] As shown in Figure 28, according to one embodiment, dependencies between components in a coordination fabric are illustrated by a simple data-driven pipeline execution use case, where (c) denotes a consumer and (p) denotes a producer.
[0225] According to the embodiment shown in FIG. 28, the scheduler (p) starts the process by starting the ingestion of data into the HUB (1). The ingest engine (c) processes the request (2) and ingests the data from the HUB into the data lake. After the ingest process is completed, the ingest engine (p) starts the pipeline processing by communicating the completion status (3). If the scheduler supports data-driven execution, it can automatically start the pipeline process to execute (3a). The pipeline engine (c) calculates (4) the pipelines waiting to execute the data. The pipeline engine (p) communicates (5) a list of pipeline applications to schedule for execution. The scheduler gets (6) an execution schedule request for the pipeline and starts the execution of the pipeline (6a). The application scheduler (e.g. Spark) arbitrates with the resource manager for resource allocation (7) and executes the pipeline. The application scheduler sends the pipeline to the executor in the assigned node for execution (8).
[0226] On-Premises Agent According to one embodiment, the on-premise agent facilitates access to local data and, in a limited way, facilitates distributed pipeline execution. The agent is provisioned and configured to communicate with, for example, a cloud DI service to handle data access and remote pipeline execution requests.
[0227] FIG. 29 illustrates an on-premise agent for use in the system according to one embodiment.
[0228] As shown in FIG. 29, according to one embodiment, a cloud agent adapter 682 provisions an on-premise agent 680 (1) and configures an agent adapter endpoint for communication.
[0229] The ingest service initiates local data access requests to HUB1 through the messaging system (2). The cloud agent adapter acts as an intermediary between the on-premises agent and the messaging system by providing access to requests initiated through the ingest service (3), writing data from the on-premises agent to the data lake, and notifying the completion of the task through the messaging system.
[0230] The premises agent polls the cloud agent adapter for processing data access requests (4) or uploading data to the cloud. The cloud agent adapter writes the data to the data lake (5) and notifies the pipeline through a messaging system.
[0231] DFML Flow Process FIG. 30 is a diagram illustrating a data flow process according to one embodiment.
[0232] As shown in FIG. 30, according to one embodiment, in an ingest step 692, data is ingested from various sources, for example, SFDC, S3, or DBaaS.
[0233] In a data preparation step 693, the ingested data is prepared, for example, by de-duplication, normalization, or enrichment.
[0234] In a transformation step 694, the system transforms the data by merging, filtering, or performing a lookup of the data sets.
[0235] In a model step 695, one or more models are generated along with mappings for the models.
[0236] In a publishing step 696, the system can publish the model, specify policies and schedules, and populate target data structures.
[0237] Metadata and data-driven auto-mapping According to an embodiment, the system may provide support for automatic mapping of complex data structures, datasets, or entities between one or more data sources or targets (referred to in some embodiments herein as HUBs). Automatic mapping may be driven by metadata, schema, and statistical profiling of the datasets. Using automatic mapping, source datasets or entities associated with an input HUB may be mapped to target datasets or entities, or vice versa, to generate a consistent, unified view of the data in one or more output HUBs. It is possible to generate output data prepared in the format or projection used.
[0238] For example, according to one embodiment, a user may want to select data to be mapped from a source or input dataset or entity in an input HUB to a target or output dataset or entity in an output HUB to implement (e.g., build) a data flow, pipeline, or Lambda application.
[0239] According to one embodiment, since manually mapping data from input HUBs to output HUBs for very large sets of HUBs and datasets or entities can be a very time-consuming and inefficient task, auto-mapping allows users to focus on simplifying their data flow applications, e.g., pipelines, Lambda applications, by providing users with recommendations for mapping data.
[0240] According to one embodiment, the data AI subsystem receives an automap request for the automap service via a graphical user interface (e.g., the Lambda Studio Integrated Development Environment (IDE)). It is possible.
[0241] According to one embodiment, the request may include a file specified for the application for which the automap service is to run, along with information identifying the input hub, the data set or entity, and one or more attributes. The application file may contain information about the data for the application. The data AI subsystem may process the application file to extract entity names and other geometric characteristics of the entity, including attribute names and data types. The automap service may use this in a search to discover possible candidate sets for mapping.
[0242] According to an embodiment, the system can access data for transformation into a HUB, such as a data warehouse. The accessed data can include various types of data, including semi-structured and structured data. The data AI subsystem can perform metadata analysis on the accessed data, including determining one or more shapes, characteristics, or structures of the data. For example, metadata analysis can determine data types (e.g., business and functional) and column shapes of the data.
[0243] According to an embodiment, one or more samples of the data can be identified based on metadata analysis of the data, and a machine learning process can be applied to the sampled data to determine categories of data within the accessed data and update the model. The categories of data can indicate relevant portions of the data, such as, for example, fact tables within the data.
[0244] According to an embodiment, the machine learning can be implemented using, for example, a logistic regression model or other types of machine learning models that can be implemented for machine learning. According to an embodiment, the data AI subsystem can analyze a relationship of one or more data items in the data based on a category of the data. The relationship indicates one or more fields in the data for the category of the data.
[0245] According to an embodiment, the data AI subsystem can perform a process for feature extraction, which includes determining a statistical profile of the randomly sampled data, a data type, and one or more metadata about attributes of the accessed data.
[0246] For example, according to one embodiment, the data AI subsystem can generate a profile of the accessed data based on the category of the data, which can be generated for transforming the data into an output HUB and displayed, for example, in a graphical user interface.
[0247] According to one embodiment, as a result of creating such a profile, the model can support recommendations with a degree of confidence regarding the similarity of the candidate data sets or entities to the input data sets or entities, which, after filtering and sorting, can be provided to the user via a graphical user interface.
[0248] According to one embodiment, the automap service can dynamically suggest recommendations based on the stage at which a user is building a dataflow application, e.g., a pipeline, a Lambda application.
[0249] An example of an entity-level recommendation may include, according to an embodiment, attribute recommendations, such as a column of an entity that is automatically mapped to another attribute or another entity. The service can continually provide recommendations and guide the user based on the user's past activity.
[0250] According to an embodiment, the recommendations can be mapped from, for example, a source dataset or entity associated with an input HUB to a target dataset or entity associated with an output HUB using an application programming interface (API) (e.g., a REST API) provided by the automated map service. The recommendations can represent a projection of data, such as, for example, attributes, data types, and representations, where the representations can be a mapping of attributes with respect to data types.
[0251] According to an embodiment, the system can provide a graphical user interface to select an output hub for conversion of the accessed data based on the recommendation. For example, the graphical user interface can allow a user to select a recommendation for converting data to an output hub.
[0252] Automatic Mapping According to one embodiment, the automatic mapping function can be mathematically defined, where an entity set E is defined as follows:
[0253]
number
[0254] where the shape set S includes metadata, data type and statistical profiling dimension. The purpose is to i and e j The goal is to find j that maximizes the probability of similarity between
[0255]
number
[0256] At the dataset or entity level, the problem is a binary problem: datasets or entities are either similar or dissimilar. s , f t , h(f s ,f t Let,(,i,) denote the set of source,target features and interactive features between source and target.,Hence, the goal is to estimate the probability of,similarity.
[0257]
number
[0258] The log-likelihood function is defined as
[0259]
number
[0260] Therefore, in the logistic regression model, the unknown coefficients can be estimated as follows:
[0261]
number
[0262] According to one embodiment, the automap service can be triggered, for example, by receiving an HTTP POST request from the system facade service. The system facade API sends dataflow application, e.g., pipeline, Lambda application JSON files from the UI to the automap REST API, and a parser module processes the application JSON files and extracts the entity names and shapes of the datasets or entities, including attribute names and data types.
[0263] According to one embodiment, the automated map service uses a search to quickly find a set of potential candidates for mapping. A candidate set is a set of highly related objects. Since there is a need, this can be achieved using a special index and queries. The special index introduces a special search field where all attributes of an entity are stored and all tokenized with N-gram combinations. At query time, the search query builder module constructs a special query using both the entity name and attribute names of a given entity, leveraging fuzzy search features, for example based on Levenshtein distance, and leverages the search boost feature to sort the results by relevance in terms of string similarity.
[0264] According to one embodiment, the recommendation engine presents a number of relevant results, often a selection of the top N results, to the user.
[0265] According to one embodiment, to achieve high accuracy, the machine learning model compares the source and target and scores the entity similarity based on the extracted features. The feature extraction includes statistical profiles, data types, and metadata of the randomly sampled data for each attribute.
[0266] In accordance with one embodiment, the description herein generally describes using a logistic regression model to learn automated mapping examples from Oracle Business Intelligence (OBI) lineage mapping data, although other supervised machine learning models may be used instead.
[0267] According to one embodiment, the output of the logistic regression model represents an aggregate confidence in the degree of similarity of the candidate dataset or entity to the input dataset or entity in a statistical sense. To find an accurate mapping, one or more other models can be used to calculate the similarity between the source and target attributes using similar features.
[0268] Finally, according to one embodiment, the recommendations are filtered and sorted and sent back to the system façade for a user interface. The automap service dynamically presents recommendations based on which stage the user is at during dataflow application design, e.g., a pipeline or Lambda application. The service can continuously provide recommendations and guide the user based on the user's past activities. Automapping can be performed in either a forward engineering or reverse engineering direction.
[0269] FIG. 31 is a diagram illustrating automatic mapping of data types according to one embodiment. As shown in FIG. 31, according to one embodiment, a system façade 701 and an automap API 702 allow a dataflow application, e.g., a pipeline or a Lambda application, to be created from a software development component, e.g., Lambda Studio. The parser 704 processes the application's JSON file and extracts entity names and shapes, including attribute names and data types.
[0270] According to one embodiment, a search index 708 is used to support a primary search 710 to discover a likely candidate set of datasets or entities for mapping. A search query builder module 706 seeks selected datasets or entities 712 by constructing a query using both the entity name and attribute names of a given entity.
[0271] According to one embodiment, a machine learning (ML) model is used to compare source and target pairs and generate a similarity score for the datasets or entities based on the extracted features. Feature extraction 714 includes a statistical profile, data type, and metadata of the randomly sampled data for each attribute.
[0272] According to one embodiment, the logistic regression model 716 provides as output an overall confidence of the degree of similarity of the candidate entity to the input entity. To find a more accurate mapping, a column mapping model 718 is used to further evaluate the similarity between the source and target attributes.
[0273] According to one embodiment, the recommendations are then sorted for return to a software development component, such as Lambda Studio, as an automatic mapping 720. The dynamic map service dynamically surfaces recommendations based on which stage the user is in while designing a dataflow application, such as a pipeline or a Lambda application. The service can continuously provide recommendations and guide the user based on the user's past activities.
[0274] FIG. 32 illustrates an automated map service for generating mappings according to one embodiment.
[0275] As shown in Figure 32, according to one embodiment, an automated map service can be provided for the generation of a mapping, which involves receiving a UI query 728 and passing it to a query understanding engine 729 and then to a query decomposition 730 component.
[0276] According to one embodiment, a primary search 710 is performed using the data hub 722 to find candidate data sets or entities 731 for use in subsequent metadata and statistical profiling processing 732 .
[0277] According to one embodiment, the results are stored in the Get Stats Profile 734 component, The results are then fed to the AI system 724 and feature extraction 735. The results are used for synthesis 736, merging and ranking 739 the final confidence according to the model 723, and giving a recommendation and associated confidence 740.
[0278] Automap Example FIG. 33 is a diagram illustrating an example of a mapping between a source schema and a target schema according to an embodiment.
[0279] As shown in FIG. 33, this example 741 illustrates an example of simple automatic mapping based, for example, on (a) hypernyms, (b) synonyms, (c) equality, (d) Soundex, and (e) fuzzy matching, according to one embodiment.
[0280] FIG. 34 illustrates another example of a mapping between a source schema and a target schema according to an embodiment.
[0281] As shown in Figure 34, an approach based solely on metadata fails when this information is irrelevant. According to one embodiment, Figure 34 shows an example 742 where there is no information at all in the source and target attribute names. In the absence of metadata features, the system can discover similar entities using a model that includes statistical profiling of features.
[0282] Automating the Map Process FIG. 35 illustrates a process for providing automated data type mapping according to one embodiment. FIG.
[0283] As shown in FIG. 35, in step 744, according to one embodiment, the accessed data is processed to perform a metadata analysis of the accessed data.
[0284] In step 745, one or more samples of the accessed data are identified. In step 746, a machine learning process is applied to determine categories of data within the accessed data.
[0285] In step 748, a profile of the accessed data is generated based on the determined data categories for use in automatic mapping of the accessed data.
[0286] Dynamic Recommendations and Simulations According to one embodiment, the system includes a software development component (referred to in some embodiments herein as Lambda Studio) and a and a graphical user interface (referred to in some embodiments herein as a Pipeline Editor or Lambda Studio IDE) that provides a visual environment. This includes providing real-time recommendations for performing semantic actions on data accessed from an input HUB based on an understanding of the meaning or semantics associated with the data.
[0287] For example, according to one embodiment, the graphical user interface can provide real-time recommendations to perform operations (also referred to as semantic actions) on data accessed from the input HUB, including the portion of data and the shape or other characteristics of the data. Semantic actions can be performed on data based on the meaning or semantics associated with the data. The meaning of the data can be used to select semantic actions that can be performed on the data.
[0288] According to an embodiment, a semantic action may represent an operator on one or more data sets and may reference base semantic actions or functions declaratively defined in the system. One or more processed data sets may be generated by executing a semantic action. A semantic action may be defined by parameters that are associated with a particular functional or business type, which represents a particular upstream data set to be processed. A graphical user interface may be metadata-driven, dynamically generated to provide recommendations based on metadata identified in the data.
[0289] FIG. 36 illustrates a system for displaying one or more semantic actions enabled for accessed data according to one embodiment.
[0290] 36, according to one embodiment, a query for semantic actions enabled on the accessed data is sent to the knowledge sources of the system using a graphical user interface 750 having a user input area 752. The query indicates a classification of the accessed data.
[0291] According to one embodiment, a response to the query is received from the knowledge source, the response being validated against the accessed data and specific based on a classification of the data. Indicates one or more semantic actions that have been
[0292] According to an embodiment, selected ones of the semantic actions enabled for the accessed data are displayed for selection and use with the accessed data, including automatically providing or updating the list of selected semantic actions or recommendations 758 from the semantic actions enabled for the accessed data 756 during processing of the accessed data.
[0293] According to an embodiment, recommendations can be provided dynamically, rather than pre-computed based on static data. For example, the system can provide recommendations in real-time based on data accessed in real-time, taking into account information such as user profile or user experience level. Recommendations provided by the system for real-time data can be noteworthy, relevant, and accurate to generate dataflow applications, e.g., pipelines, Lambda applications. Recommendations can be provided based on user behavior with respect to data associated with specific metadata. The system can recommend semantic actions for the information.
[0294] For example, according to one embodiment, the system can ingest, transform, integrate and publish data to any system, the system can recommend using entities to analyze some of their metrics in interesting analytical ways, pivot this data on different dimensions, indicate which dimensions are interesting, summarize the data with respect to dimension hierarchies and enrich the data with more insights.
[0295] According to one embodiment, recommendations may be provided based on analysis of the data using techniques such as metadata analysis of the data.
[0296] According to an embodiment, metadata analysis may include determining a classification of the data, such as, for example, the shape, characteristics, and structure of the data. Metadata analysis may determine a data type (e.g., business and functional). Metadata analysis may also indicate the column shape of the data. According to an embodiment, data may be compared to a metadata structure (e.g., shape and characteristics) to determine the data type and attributes associated with the data. The metadata structure may be defined in a system HUB (e.g., a knowledge source) of the system.
[0297] According to an embodiment, the system can identify semantic actions based on metadata by querying the system HUB using metadata analysis. The recommendation may be a semantic action determined based on analysis of metadata of data accessed from the input HUB. Specifically, the semantic actions can be mapped to metadata. For example, the semantic actions can be mapped to metadata where these actions are permitted and / or applicable. The semantic actions may be defined by a user and / or based on the structure of the data.
[0298] According to an embodiment, semantic actions can be defined based on conditions associated with the metadata. The system HUB can be modified so that semantic actions can be modified, deleted, or augmented.
[0299] According to an embodiment, examples of semantic actions may include building cubes, filtering data, grouping data, aggregating data, or other actions that can be performed on the data. By defining semantic actions based on metadata, no mapping or scheme is required to determine the semantic actions allowed on the data. Semantic actions may be defined as new and different metadata structures are discovered. Thus, the system can dynamically determine recommendations based on the identification of semantic actions using the metadata analyzed for the data received as input.
[0300] According to an embodiment, semantic actions may be defined by a third party, and the third party may provide data, e.g., data defining one or more semantic actions associated with metadata. The system may dynamically query the system HUB to determine the semantic actions available for the metadata. Thus, the system HUB may be modified, and the system may determine the semantic actions allowed at that time based on such modifications. The system may process the data obtained from the third party by performing operations (e.g., filtering, detection, and registration). The data may then define semantic actions and make the semantic actions available based on the semantic actions identified by the process.
[0301] 37 and 38 are diagrams illustrating a graphical user interface displaying one or more semantic actions enabled for accessed data, according to one embodiment.
[0302] As shown in FIG. 37, according to one embodiment, a software development component (e.g., Lambda Studio) may include a graphical user interface (e.g., a pipeline Editor or Lambda Studio IDE) 750 can be provided. This is the output HU You can display the recommended semantic actions to use when processing the input data or simulating the processing of the input data for projection into B.
[0303] For example, according to one embodiment, the interface of Figure 37 allows a user to view options 752 associated with a data flow application, e.g., a pipeline, a Lambda application, which may include, e.g., an input HUB specification 754.
[0304] According to one embodiment, during creation of a dataflow application, e.g., a pipeline, a Lambda application, or during simulation of a dataflow application, e.g., a pipeline, a Lambda application, on input data, one or more semantic actions 756 or other recommendations 758 may be displayed in a graphical user interface for review by a user.
[0305] According to one embodiment, in a simulation mode, the software development component (e.g., Lambda Studio) provides a sandbox environment that allows a user to immediately see the results of the execution of various semantic actions on the output, including automatically updating the list of semantic actions appropriate for the accessed data as the accessed data is being processed.
[0306] For example, as shown in FIG. 38, according to one embodiment, a user who is looking for some information From the user's starting point, the system can recommend actions 760 on the information, such as using the entities to analyze some of their metrics in interesting analytical ways, pivoting this data on different dimensions and indicating which dimensions are interesting, summarizing the data with respect to dimensional hierarchies, and enriching the data with more insights.
[0307] In accordance with one embodiment, in the example shown, both sources and dimensions are recommended for analysable entities in the system, making the task of building a multidimensional cube largely one of pointing and clicking.
[0308] Typically, such work requires a lot of experience and domain-specific knowledge. Using machine learning to analyze both data characteristics and user behavior patterns for common integration patterns, along with a combination of semantic search and machine learning recommendations, enables advanced tooling for application development to build business-specific applications.
[0309] FIG. 39 illustrates a process for displaying one or more semantic actions enabled for accessed data according to one embodiment.
[0310] 39, in step 772, according to one embodiment, the accessed data is processed to perform a metadata analysis of the accessed data. The metadata analysis includes determining a classification of the accessed data.
[0311] At step 774, a query of semantic actions enabled for the accessed data is sent to the knowledge sources of the system, the query indicating a classification of the accessed data.
[0312] At step 775, a response to the query is received from the knowledge source, the response indicating one or more semantic actions enabled for the accessed data and identified based on the classification of the data.
[0313] At step 776, the selected semantic actions from among the semantic actions enabled for the accessed data are displayed in a graphical user interface for selection and use on the accessed data, including automatically providing or updating the list of the selected semantic actions from among the semantic actions enabled for the accessed data during processing of the accessed data.
[0314] Functional Decomposition of Data Flow According to an embodiment, the system can provide a service for recommending actions and transformations on input data based on patterns identified from a functional decomposition of a software application's data flow, including determining possible transformations of the data flow in a subsequent application. The data flow can be decomposed into models that describe transformations of the data, predicates, and business rules that are applied to the data, and attributes used within the data flow.
[0315] FIG. 40 illustrates an embodiment of a pipeline that supports evaluating a Lambda application for its components to facilitate pattern detection and inductive learning.
[0316] As shown in FIG. 40, according to one embodiment, function decomposition logic 800, i.e., software components, may be provided as software or program code executable by a computer system or other processing device, and may be displayed on a display 805 (e.g., in a pipeline editor or Lambda Studio IDE) displaying the function decomposition logic 800. It can be used to provide solutions 802 and recommendations 804. For example, the system can provide a service to recommend actions and transformations on data based on patterns / templates identified from functional decomposition of data flows of a data flow application, e.g., a pipeline, a Lambda application, i.e., through functional decomposition of the data flows, patterns can be observed to determine possible transformations to the data flows in subsequent applications.
[0317] According to one embodiment, this service can be realized by a framework that can decompose or classify data flows into models that describe the transformations, predicates, and business rules that are applied to the data, and attributes used in the data flows.
[0318] Traditionally, the data flow of an application can represent a set of transformations on data, and the types of transformations applied to the data are highly contextual. In most data integration frameworks, process lineage is usually limited or non-existent with respect to how data flows are persisted, parsed, and generated. According to an embodiment, the system can derive context-related patterns from flows or graphs based on semantic-rich entity types, and can further learn data flow grammars and models and use them to generate composite data flow graphs given a given similar context.
[0319] According to an embodiment, the system can generate one or more data structures that define patterns and templates based on the design specifications of the data flows. The patterns and templates can be determined by decomposing the data flows into data structures that define functional expressions. The data flows can be used to predict and generate functional expressions to determine patterns for data transformation recommendations. The recommendations are based on the decomposed data flows and models derived by inductive learning of the inherent patterns, and can be fine-grained (e.g., recommending scalar transformations for specific attributes, or using one or more attributes in predicates for filtering or joining).
[0320] According to one embodiment, a data flow application, e.g., a pipeline, Lambda application, allows users to create complex data transformations based on semantic actions on data. The system can store the data transformations as one or more data structures that define the flow of data for the pipeline, Lambda application.
[0321] According to an embodiment, decomposition of the data flow of a data flow application, e.g., a pipeline, a Lambda application, can be used to determine pattern analysis of the data and generate functional expressions. The decomposition can be performed on semantic actions, transformations and predicates, or business rules. Each of the semantic actions of the above applications can be identified through the decomposition. Using a process of induction, the business logic can be extracted from the data flow including its context elements (business and functional).
[0322] According to an embodiment, a model can be generated for a process and, based on induction, context-rich prescriptive dataflow design recommendations can be generated. These recommendations may be based on patterns inferred from the model, and each recommendation may correspond to a semantic action that can be performed on the data for the application.
[0323] According to an embodiment, a system can execute a process to infer patterns of data transformation based on functional decomposition. The system can access data flows of one or more data flow applications, e.g., pipelines, Lambda applications. The data flows can be processed to determine one or more function expressions. The function expressions can be generated based on actions, predicates, or business rules identified in the data flows. Using the actions, predicates, or business rules, patterns of transformations can be identified (e.g., inferred) for the data flows. Inferring patterns of transformations can be a passive process.
[0324] According to one embodiment, the patterns of transformations can be determined in a crowdsourced manner based on a passive analysis of the data flows of different applications, which can be determined using machine learning (e.g., deep reinforcement learning).
[0325] According to an embodiment, patterns of transformations can be identified for function expressions generated for data flow applications, e.g., pipelines, Lambda applications. By decomposing one or more data flows, patterns of data transformations can be inferred.
[0326] According to an embodiment, the system can use the pattern to recommend one or more data transformations for a data flow of a new data flow application, e.g., a pipeline, a Lambda application. In the example of a data flow of processing on data for a monetary exchange, the system can identify a pattern of transformations on the data. The system can also recommend one or more transformations for a new data flow of the application, which includes data for a similar monetary exchange. The transformations can be performed in a similar manner according to the pattern, such that the new data flow is modified according to the transformation to produce a similar monetary exchange.
[0327] FIG. 41 illustrates a means for identifying patterns of transformations in data flows for one or more function expressions generated for each of one or more applications, according to one embodiment.
[0328] As mentioned above, according to one embodiment, a pipeline, e.g., a Lambda application, allows a user to specify complex data transformations based on semantic actions corresponding to operators in a relational calculus. Data transformations are typically persisted as directed acyclic graphs or queries, or in the case of DFML, as nested functions. By decomposing and serializing a dataflow application, e.g., a pipeline, a Lambda application, as nested functions, we enable pattern analysis of the dataflow and induct a dataflow model that can be used to generate functional expressions that extract complex transformations on data sets in similar contexts.
[0329] According to an embodiment, nested function decomposition is performed not only at the level of semantic actions (row or dataset operators) but also at scalar transformations and predicate structures, enabling deep lineage capabilities of complex data flows. Recommendations based on inductive models can be highly granular (e.g., recommending scalar transformations for a particular attribute, or one or more attributes in a predicate for filtering or joining). ).
[0330] According to one embodiment, the elements involved in the functional decomposition are roughly as follows: An application represents a top-level dataflow transformation.
[0331] An action represents an operator on one or more datasets (specific data frames).
[0332] An action refers to a base semantic action or function declaratively defined in the system. Actions can have one or more action parameters, each of which can have a specific role (in, out, in / out) and type, can return one or more processed data sets, and can be embedded or nested to several levels deep.
[0333] Action parameters are owned by an action, have a specific functional or business type, and represent a specific upstream dataset to be processed. Binding parameters represent datasets or entities in the HUB that are used for the transformation. Value parameters represent intermediate or temporary data structures to be processed in the context of the current transformation.
[0334] The scope resolver allows process lineage to be derived for datasets or elements within datasets that are used in the overall data flow.
[0335] FIG. 42 illustrates an object diagram for use in identifying patterns of transformations in data flows for one or more function expressions generated for each of one or more applications, according to one embodiment.
[0336] 42, according to an embodiment, a dataflow of a dataflow application, e.g., a pipeline, a Lambda application, can be decomposed or categorized using function decomposition logic into models that describe the transformations, predicates, and business rules applied to the data and attributes used within the dataflow, which can be decomposed into, for example, patterns or templates 812 (pipelines if the template is associated with a Lambda application), services 814, functions 816, function parameters 818, and function types 820.
[0337] According to one embodiment, each of these functional components can be further decomposed into, for example, tasks 822 or actions 824 that reflect a dataflow application, for example a pipeline, a Lambda application.
[0338] According to one embodiment, a scope resolver 826 can be used to resolve references to a particular attribute or embedded object through its scope. For example, as shown in Figure 42, a scope resolver resolves references to attributes or embedded objects through their adjacent scopes. For example, a filter and a join function that uses the output of another table may have a reference to its scope resolver and may be used in conjunction with the InScopeOf operation to resolve references to a leaf node or its It can be resolved to the root node.
[0339] FIG. 43 illustrates a process for identifying patterns of transformations in data flows for one or more function expressions generated for each of one or more applications, according to one embodiment.
[0340] As shown in FIG. 43, according to one embodiment, in step 842, a data flow is accessed for each of one or more software applications.
[0341] In step 844, the data flow of one or more software applications is processed to generate one or more functional expressions that represent the data flow, the one or more functional expressions being generated based on semantic actions and business rules identified in the data flow.
[0342] At step 845, for the one or more function expressions generated for each of the one or more software applications, patterns of transformations in the data flow are identified. Semantic actions and business rules are used to identify patterns of transformations in the data flow.
[0343] In step 847, the pattern of transformations identified in the data flow is used to provide one or more data transformation recommendations for the data flow of another software application.
[0344] Ontology Learning According to an embodiment, the system can perform ontology analysis of the schema definition to determine the type of data and data sets or entities associated with the schema, and generate or update a model from a reference schema that includes an ontology defined based on the relationships between the entities and their attributes. A reference HUB that includes one or more schemas can be used to analyze data flows and further classify or make recommendations, such as transforming, enriching, filtering, or cross-entity data fusion of input data.
[0345] According to an embodiment, the system can determine the ontology of the types of data and entities in the reference schema by performing an ontology analysis of the schema definition. In other words, the system can generate a model from a schema that includes an ontology defined based on the relationships between entities and their attributes. The reference schema can be a system-provided or default reference schema, or alternatively, a user-supplied or third-party reference schema.
[0346] Some data integration frameworks may reverse engineer metadata from known system source types, but do not provide analysis of the metadata to build functional systems that can be used for pattern definition and entity classification. Also, metadata harvesting is limited in scope and does not extend to data profiling for the extracted data sets or entities. Features that allow users to specify reference schemas for functional system ontology learning to use in complex process (business logic) and integration patterns in addition to entity classification (in a similar topological space) are not currently available.
[0347] According to an embodiment, one or more schemas can be stored in a reference hub, which itself can be provided in or as part of a system hub. Like the reference schema, the reference hub can be a user-provided or third-party reference hub, or in a multi-tenant environment, can be associated with a particular tenant and accessed, for example, through a data flow API.
[0348] According to one embodiment, a reference HUB is used to analyze the data flow, and It can perform classification or recommendation, for example transformation, enrichment, filtering or cross-entity data fusion.
[0349] For example, according to one embodiment, the system can receive an input that specifies a reference HUB as a schema for ontology analysis. The reference HUB can be imported to obtain entity definitions (attribute definitions, data types, and relationships between datasets or entities, constraints, or business rules). Sample data (e.g., attribute vectors such as column data, for example) in the reference HUB can be extracted for all datasets or entities and profiled data to derive some metrics of the data.
[0350] According to an embodiment, the type system can be instantiated based on the glossary of the reference schema. The system can derive an ontology (e.g., a set of rules) that describes the type of data by performing ontology analysis. The ontology analysis can determine data rules. The data rules are prescribed for the profiled data (e.g., attributes or composite values) metrics and describe the business type elements (e.g., UOM, ROIL, or currency type) with their data profile. The ontology analysis can determine relationship rules and composition rules. The relationship rules prescribe the correspondence between the data sets or entities and the attribute vectors (constraints or references imported from the reference schema), and the composition rules can be derived from a combination of the data rules and the relationship rules. The type system can then be prescribed based on the rules obtained through metadata harvesting and data sampling.
[0351] According to one embodiment, patterns and templates from the System HUB can be utilized based on a type system instantiated using ontology analysis, and the system can then use the type system to perform data flow processing.
[0352] For example, according to one embodiment, classification and type annotations of a dataset or entity can be specified by a type system of a registered HUB. The type system can be used to specify rules for functional and business types derived from a reference schema. The type system can be used to perform actions, such as blending, enrichment, and transformation recommendations, on the entities identified in the data flow based on the type system.
[0353] FIG. 44 illustrates a system for generating functional rules according to an embodiment.
[0354] As shown in FIG. 44, according to one embodiment, rules 851 can be mapped to a functional system 852 by rule induction logic 850 or software components provided as software or program code executable by a computer system or other processing device.
[0355] FIG. 44 illustrates a system for generating functional rules according to an embodiment.
[0356] As shown in FIG. 45, according to one embodiment, HUB1 can serve as a reference ontology that can be used to type tag, compare, classify, and otherwise evaluate metadata schemas or ontologies provided by other (e.g., newly registered) HUBs, such as HUB2 and HUB3, and the data AI system can use the Create the appropriate rules to use.
[0357] FIG. 46 illustrates an object diagram used to generate functional rules according to one embodiment.
[0358] According to an embodiment, as shown for example in Figure 46, the rule induction logic can map rules to a functional system having a set of function types 853 (e.g., HUB, dataset or entity, and attributes) and store them in a registry for use in creating data flow applications, e.g., pipelines, Lambda applications. This includes that each function type 854 can be mapped to a functional rule 856 and a rule 858. Each rule can be mapped to rule parameters 860.
[0359] According to one embodiment, a reference schema may first be processed to create an ontology that includes a set of rules appropriate to this schema.
[0360] According to one embodiment, the new HUB or new schema is then evaluated and its data set or entities are compared to the existing ontology and the created rules, which can be used to analyze the new HUB / schema and its entities and further train the system.
[0361] While metadata harvesting in a data integration framework may be limited to reverse engineering entity definitions (attributes and their data types, and possibly relationships), according to an embodiment, the system described herein provides a different approach: it allows a schema definition to be used as a reference ontology from which business and functional types can be derived, and data profiling metrics can be derived for data sets or entities in the reference schema. This reference HUB can then be used to analyze business entities in other HUBs (data sources) for further classification or recommendation (e.g. blending or enrichment).
[0362] According to one embodiment, the system uses the following set of steps for ontology learning using a reference schema.
[0363] The user specifies an option to use the newly registered HUB as a reference schema.
[0364] Entity definitions (for example attribute definitions, data types, relationships between entities, constraints or business rules are imported).
[0365] For every dataset or entity, sample data is extracted and the data is profiled to derive some metrics about the data.
[0366] A type system is instantiated based on the vocabulary of reference schemas (functional and business types).
[0367] A set of rules is derived that describes the business types. Data rules are defined in terms of the profiled data metrics and describe the nature of the business type element (e.g., UOM, ROI, or currency type) that the data (Can be specified as a business type element along with the profile).
[0368] Relationship rules are generated that govern the associations on the elements (constraints or references imported from referenced schemas).
[0369] Complex rules are generated that can be derived through the combination of data and relationship rules. The type system (functions and businesses) is defined based on rules derived through metadata harvesting and data sampling.
[0370] A pattern or template can then specify complex business logic with types instantiated based on the reference schema.
[0371] HUBs registered in the system can then be parsed in the context of the reference schema.
[0372] In newly registered HUBs, classification and type annotation of datasets or entities can be performed based on rules for functional and business types derived from the reference schema.
[0373] Based on type annotations, blending, enrichment, and transformation recommendations can be performed on datasets or entities.
[0374] FIG. 47 illustrates a process for generating a functional system based on one or more generated rules, according to one embodiment.
[0375] As shown in FIG. 47, according to one embodiment, in step 862, input is received defining a reference HUB.
[0376] In step 863, the reference HUB is accessed to obtain one or more entity definitions associated with the data set or entity provided by the reference HUB.
[0377] In step 864, sample data for one or more data sets or entities is generated from the reference HUB.
[0378] In step 865, the sample data is profiled to determine one or more metrics associated with the sample data.
[0379] In step 866, one or more rules are generated based on the entity definition. In step 867, a functional system is generated based on the generated one or more rules.
[0380] In step 868, the functional system and sample data profiles are persisted for use in processing the data input.
[0381] Foreign language function interface According to an embodiment, the system comprises (as referred to in some embodiments herein) It provides a programmatic interface (called the Multilingual Function Interface) that allows users or third parties to express services, functional and business types, semantic actions, and patterns or expressions based on functional and business types. The functionality of the system can be extended by declaratively specifying predefined complex data flows.
[0382] As mentioned above, current data integration systems can provide limited interfaces, no support for types, and no well-defined interfaces for object composition and pattern definition. Due to these shortcomings, complex features such as cross-service recommendations or a unified application design platform for invoking semantic actions across services that extend the framework are currently not provided.
[0383] According to one embodiment, a multilingual function interface allows users to extend the functionality of the system by providing definitions or other information declaratively (eg, from a customer to other third parties).
[0384] According to one embodiment, the system is metadata-driven and can derive metadata by processing definitions received through a multilingual function interface, determine the classification of the metadata, e.g., data type (e.g., functional and business), and compare these data types (both functional and business) with existing metadata to determine whether there is a type match.
[0385] According to an embodiment, metadata received through the multilingual function interface may be stored in the system HUB, allowing the system to access the metadata to process the data flow. For example, the metadata may be accessed to determine semantic actions based on the type of data set received as input. The system may determine the semantic actions that are allowed for the type of data provided through this interface.
[0386] According to an embodiment, by providing a common declarative interface, the system can enable users to map service-native types and actions to platform-native types and actions. This enables a unified application design experience through type and pattern discovery. It also facilitates purely declarative data flow definition and design, which requires the generation of native code for each semantic action and components of various services that extend the platform.
[0387] According to an embodiment, metadata received through the multilingual function interface can be automatically processed and the objects or artifacts (e.g., data types or semantic actions) described therein can be used in the operations of data flows processed by the system. Metadata information received from one or more third party systems can be used to define a service, to represent one or more functional and business types, to represent one or more semantic actions, or to represent one or more patterns / templates.
[0388] For example, according to an embodiment, a classification of the accessed data, such as functional and business types of data, can be determined. This classification can be determined based on information about the data received with the information. Receiving data from one or more third party systems can extend the functionality of the system to perform data integration of data flows based on information (e.g., services, semantic actions, or patterns) received from the third parties.
[0389] According to an embodiment, metadata in the system HUB can be updated to include information specified about the data. For example, services and patterns / templates can be updated to execute based on information (e.g., semantic actions) specified in the metadata received through the multilingual function interface. In this way, the system can be enhanced through the multilingual function interface without disrupting the processing of the data flow.
[0390] According to an embodiment, subsequent data flows can be processed using the metadata in the system HUB after it has been updated. Metadata analysis can be performed on data flows of data flow applications, such as pipelines, Lambda applications. The system HUB can then be used to determine transformation recommendations taking into account the definitions provided via the multilingual function interface. The transformations can be determined based on patterns / templates used to specify semantic actions to execute the service, and the semantic actions can also take into account the definitions provided via the multilingual function interface.
[0391] FIG. 48 illustrates a system for identifying patterns for use in providing recommendations regarding data flows based on information provided via a multilingual function interface according to an embodiment.
[0392] As shown in FIG. 48, in one embodiment, definitions received via a multilingual function interface 900 can be used to update a service registry 902, a function and business type registry 904, or a pattern / template 906 in the system HUB.
[0393] According to one embodiment, the updated information is used by a data AI subsystem including a rules engine 908 to, for example, determine type-annotated hubs, data sets or entities, or attributes 910 in the system hub, and to generate recommendations for use in providing recommendations for data flow applications, e.g., pipelines, Lambda applications, via a software development component (e.g., Lambda Studio). The information can be provided to the consultation engine 912.
[0394] FIG. 49 is a diagram further illustrating identifying patterns for use in providing recommendations regarding data flows based on information provided via a foreign language function interface, according to an embodiment.
[0395] As shown in FIG. 49, according to one embodiment, third party metadata 920 can be received in a multilingual function interface.
[0396] FIG. 50 is a diagram further illustrating identifying patterns for use in providing recommendations regarding data flows based on information provided via a multilingual function interface, according to an embodiment.
[0397] As shown in FIG. 50, in accordance with one embodiment, third party metadata received through a multilingual function interface can be used to extend the functionality of the system.
[0398] According to one embodiment, the system allows the framework to be extended through a well-defined interface. In one embodiment, a service can be registered along with its native type, the semantic actions realized by the service, along with typed parameters, patterns or templates that abstract predefined algorithms that can be utilized as part of the service.
[0399] According to one embodiment, by providing a common declarative programming paradigm, the pluggable services architecture allows for mapping service-native types and actions to platform-native types and actions. This enables a unified application design experience through type and pattern discovery. It also facilitates purely declarative data flow definition and design, which requires the generation of native code for each semantic action and components of various services that extend the platform.
[0400] According to an embodiment, the pluggable services architecture also defines a compilation, generation, deployment, and run-time execution framework (unified application design services) for plug-ins. The recommendation engine can perform machine learning and inference of semantic actions and patterns of all plugged-in services and can make cross-service semantic action recommendations for distributed composite dataflow design and development.
[0401] FIG. 51 illustrates a process for identifying patterns for use in providing recommendations regarding data flows based on information provided via a multilingual function interface, according to one embodiment.
[0402] As shown in FIG. 51, in accordance with one embodiment, in step 932, one or more definitions of metadata for use in processing the data are received via a foreign language function interface.
[0403] In step 934, the received metadata is processed via the foreign language function interface to identify information about the received metadata including one or more of a classification, a semantic action, a template defining a pattern, or a service defined by the received metadata.
[0404] In step 936, the metadata received via the foreign language function interface is stored in the system HUB. The system HUB is updated to include information about the received metadata and to extend the functionality of the system, including the system's supported types, semantic actions, templates, and services.
[0405] At 938, a pattern is identified for providing recommendations regarding data flow based on the information updated at the system HUB via the foreign language function interface.
[0406] Policy-Based Lifecycle Management According to an embodiment, the system can provide data governance capabilities such as historical information (where did this data come from), lineage (how was this data acquired / processed), security (who was responsible for this data), classification (what is this data related to), impact (how much impact does this data have on the business), retention time (how long should this data persist), and validity (should this data be excluded / included for analysis / processing) for each slice of data that is temporally related to a particular snapshot. These can be used in lifecycle decisions and data flow recommendations. do.
[0407] Current approaches to data lifecycle management do not include governance-related functionality or tracking of data evolution (changes in data profile or drift) based on changes in data characteristics across temporal partitions. System observed or derived data characteristics (classification, frequency of change, type of change, or use in a process) are not used in lifecycle decisions or recommendations about data (retention time, security, validity, retrieval intervals).
[0408] According to an embodiment, the system can provide a graphical user interface that can display the life cycle of the data flow based on lineage tracking. The life cycle can show where the data was processed and if any errors occurred during the processing of the data, and can be shown as a timeline diagram of the data (e.g., number of datasets, volume of datasets, and usage of datasets). The interface can provide a point-in-time snapshot of the data and can provide visual indicators of the data as it is being processed. Thus, the interface allows a full audit of the data or a system snapshot of the data based on its life cycle (e.g., performance metrics, or resource usage).
[0409] According to an embodiment, the system can determine the lifecycle of data based on sample data (sampled periodically from ingested data) and data acquired for processing by user-defined applications. Some aspects of data lifecycle management are similar across categories of ingested data, i.e. streaming data and batch data (reference and incremental). For incremental data, the system can use a scheduled log collection event-driven method to acquire temporary slices of data and manage allocation of slices across application instances covering the following functions:
[0410] According to one embodiment, in the event of data loss, the system can reconstruct the data using lineage across layers from metadata managed in the system HUB.
[0411] For example, according to an embodiment, incremental data can be obtained by specifying incremental data attribute columns or user configuration settings to maintain high and low watermarks throughout data ingestion. A query or API and corresponding parameters (timestamp or ID column) can be associated with the ingested data.
[0412] According to an embodiment, the system can manage lineage information across tiers, such as query or log metadata in the edge layer, topic / partition offsets per ingest in the scalable I / O layer, slices (file partitions) in the data lake, references to process lineage (specific execution instances of the application that generated the data and its associated parameters) of subsequent downstream processed datasets that used this data, topics / partitions for datasets that are "marked" for publication to target endpoints and their corresponding data slices in the data lake, and offsets within partitions where job execution instances are published and processed and published to target endpoints.
[0413] According to one embodiment, layers (e.g., edge, scalable I / O, data lake, In case of a failure of one of the following (either the sender or the publisher), the data can be reconstructed from the upstream layer or retrieved from the source.
[0414] According to some embodiments, the system can perform other life cycle management functions.
[0415] For example, according to an embodiment, security is enforced and audited at each of these layers of data slices. Data slices can be excluded or included (if already excluded) from processing or access. This allows spurious or corrupted data slices to not be processed. Retention policies can be enforced for slices of data through sliding windows. Slices of data are analyzed for impact (e.g., the ability to tag slices of a given window as impactful in the context of a data mart built for quarterly reporting).
[0416] According to one embodiment, data is classified by tagging it with a functional or business type defined within the system (e.g., tagging a data set with a functional type (as cube or dimension or hierarchical data) along with a business type (e.g., order, customer, product, or time)).
[0417] According to an embodiment, a system can perform a method that includes accessing data from one or more HUBs. The data can be sampled, and the system determines a temporal slice of the data and manages the slices, including accessing a system HUB of the system and obtaining metadata about the sampled data. The sampled data can be managed for lineage tracking across one or more tiers in the system.
[0418] According to an embodiment, incremental data and parameters relating to the sample data can be managed for the ingested data. The data can be classified by tagging the type of data associated with the sample data.
[0419] FIG. 52 illustrates management of sampled or accessed data for lineage tracking across one or more layers according to an embodiment.
[0420] For example, as shown in Figure 52, according to one embodiment, the system can be used to receive data from HUB 952, in this example an Oracle database, and HUB 954, in this example S3 or other environments. Data received from the input HUB at the edge layer can be provided to the scalable I / O layer as one or more topics (each of which can be provided as a distributed partition) for use by a data flow application, such as a pipeline, a Lambda application.
[0421] According to one embodiment, the ingested data, typically represented by an offset into a partition, can be normalized 964 by the compute layer and written to the data lake as one or more temporary slices that span the tiers of the system.
[0422] According to one embodiment, this data is then consumed by a dataflow application, e.g., a pipeline, Lambda application 966, 968, and ultimately published 970 to one or more additional topics 960, 962, which then In this example, data can be published to a target endpoint (eg, a table) at one or more output hubs, such as in a DBCS environment.
[0423] As shown in FIG. 52, according to one embodiment, initially, data reconstruction and lineage tracking information may include, for example, provenance (Hub 1, S3), lineage (Source Entity in Hub 1), security (Connection Credential used), or other information regarding the ingestion of data.
[0424] FIG. 53 further illustrates management of sampled or accessed data for lineage tracking across one or more tiers, according to an embodiment. As shown in FIG. 53, the data reconstruction and lineage tracking information can then be updated to include information such as, for example, updated history information (→T1), lineage (→T1 (Ingest Process)), or other information.
[0425] FIG. 54 further illustrates management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. As shown in FIG. 54, the data reconstruction and lineage tracking information can then be further updated to include information such as, for example, updated history information (→E1), lineage (→E1(Normalize)), or other information. do.
[0426] According to an embodiment, temporary slices 972 used by one or more data flow applications, e.g., pipelines, Lambda applications, can be created across tiers of the system. In the event of a failure, e.g., a failure writing to the data lake, the system can determine one or more data slices that are not yet processed and complete the processing of the data slices, either in whole or incrementally.
[0427] Figure 55 further illustrates management of sampled or accessed data for lineage tracking across one or more tiers, according to an embodiment. As shown in Figure 55, the data reconstruction and lineage tracking information can then be further updated to create additional temporary slices to include information such as updated historical information (→E11(App1)), security (Role Executing App 1), or other information.
[0428] Figure 56 further illustrates management of sampled or accessed data for lineage tracking across one or more tiers, according to an embodiment. As shown in Figure 56, the data reconstruction and lineage tracking information can then be further updated to create additional temporary slices to include information such as updated lineage (→E12(App2)), security (Role Executing App 2), or other information.
[0429] Figure 57 further illustrates management of sampled or accessed data for lineage tracking across one or more layers, according to an embodiment. As shown in Figure 57, the data reconstruction and lineage tracking information can then be further updated to create additional temporal slices to include information such as updated lineage (→T2(Publish)), security (Role Executing Publish to I / O Layer), or other information.
[0430] 58 is a diagram further illustrating management of sampled or accessed data for lineage tracking across one or more tiers, according to an embodiment. As shown in FIG. 58, data reconstruction and lineage tracking information can then be further updated to reflect output of data to a target endpoint 976.
[0431] Data Lifecycle Management According to an embodiment, the data lifecycle management based on the lineage tracking is directed to several functional areas. Some of these areas can be configured by the user (access control, retention time, validity), some are derived (historical information, lineage), and others use machine learning algorithms (classification, impact). For example, data management applies to both sample data (sampled periodically from ingested data) and data obtained for processing by user-defined applications. Some aspects of data lifecycle management are similar across the categories of ingested data, i.e. streaming data and batch data (reference and incremental). For incremental data, DFML uses a scheduled log collection event-driven method to obtain temporary slices of data and manages the allocation of slices across application instances covering the following functions:
[0432] In the event of data loss, the data is reconstructed using lineage across layers from metadata managed at the system hub.
[0433] Get incremental data and maintain high and low watermarks across ingests by specifying incremental data attribute columns or user configuration settings.
[0434] For each ingest, map the query or API and the corresponding parameters (timestamp or ID column).
[0435] Managing lineage information across tiers. Query or log metadata in the edge layer. Topic / partition offsets per ingest in the scalable I / O layer. Slices (file partitions) in the data lake. A reference to the process lineage (the specific execution instances of the application that generated the data and its associated parameters) of all subsequent downstream processed datasets that used this data. Topic / partition offsets of datasets that are "marked" to be published to a target endpoint and the corresponding data slices in the data lake. The offsets within the partitions that the job execution instances publish and are processed and published to the target endpoint.
[0436] In case of a layer failure, data can be reconstructed from an upstream layer or retrieved from the source. Security is enforced and audited at each of these layers for data slices. Data slices can be excluded or targeted (if already excluded) from processing or access. This allows spurious or corrupted data slices to not be processed. Retention time policies can be enforced for slices of data through sliding windows. Slices of data are analyzed for impact (e.g., the ability to tag slices for a given window is analyzed as impactful in the context of a data mart built for quarterly reporting).
[0437] Categorize data by functional or business type tagging defined within the system (e.g. tag a data set with its functional type (cube or dimensional or hierarchical data) along with its business type (e.g. order, customer, product or time).
[0438] FIG. 59 illustrates a process for managing sampled or accessed data for lineage tracking across one or more layers according to an embodiment.
[0439] As shown in FIG. 59, according to one embodiment, in step 982, data is accessed from one or more HUBs.
[0440] In step 983, the accessed data is sampled. In step 984, a temporal slice is identified for the sampled and accessed data.
[0441] In step 985, the system HUB is accessed to obtain metadata about the sampled or accessed data represented by the temporal slice.
[0442] In step 986, classification information is determined for the sampled or accessed data represented by the temporary slice.
[0443] In step 987, the sampled or accessed data represented by the temporary slices is managed for lineage tracking across one or more tiers in the system.
[0444] Embodiments of the present invention can be implemented using a general purpose or special purpose digital computer, computing device, machine, or microprocessor, including one or more processors, memory, and / or computer readable storage media, programmed according to the teachings of the present disclosure. A skilled programmer can readily produce appropriate software coding based on the teachings of the present disclosure, as will be apparent to one skilled in the software arts.
[0445] In some embodiments, the present invention includes a computer program product that is a non-transitory computer readable medium(s) having stored thereon instructions that can be used to program a computer to perform any of the processes of the present invention. Examples of storage media may include, but are not limited to, floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROM, RAM, EPROM, EEPROM, DRAM, VRAM, flash memory devices, magnetic or optical cards, memory systems (including molecular memory ICs), or other types of storage media or devices suitable for non-transitory storage of instructions and / or data.
[0446] The foregoing description of the present invention has been provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form disclosed. Numerous modifications and variations will be apparent to those skilled in the art.
[0447] For example, some of the above embodiments may involve using products such as Wolfram, Yago, Chronos, and Spark to perform various calculations, and may also involve the use of, for example, For example, you can use data sources such as BDP, SFDC, and S3 to find the source of the data. Although the embodiments described herein are illustrated as functioning as a data source or target, the embodiments described herein may also be used with other types of products and data sources that provide similar types of functionality.
[0448] In addition, while some of the above embodiments illustrate components, layers, objects, logic or other features of various embodiments, such features may be provided as software or program code executable by a computer system or other processing device.
[0449] The embodiments have been chosen and described in order to best explain the principles of the invention and its practical application, thereby enabling those skilled in the art to understand the invention with its various embodiments and various modifications suited to the particular uses intended. Modifications and variations include any pertinent combination of the disclosed features. It is intended that the scope of the invention be defined by the following claims and their equivalents.< / publisherurl> < / publisherurl>
Claims
1. 1. A method comprising: one or more processors normalizing the data ingested into the first topic of the scalable input / output layer using a normalization application in the compute layer to generate temporary slices and writing the resulting slices to the data lake; The one or more processors process the temporary slices written to the data lake by each of one or more applications of a data flow in the compute layer, and write each temporary slice to the data lake; The one or more processors publishing the temporary slices written to the data lake to a second topic in the scalable I / O layer via a publishing application in the compute layer; The one or more processors publish the temporary slices published to the second topic to an output HUB; and updating data reconstruction and lineage tracking information, including lineage and security, by the one or more processors for each operation between ingesting and publishing data to the first topic of the scalable input / output layer.
2. The method of claim 1 , further comprising the one or more processors ingesting data from one or more input hubs containing a data set into the first topic of the scalable input / output layer.
3. The method of claim 2 , wherein metadata received through a foreign language function interface is stored in a knowledge source accessed by the one or more processors for processing the data flow.
4. The method of claim 2 or 3, wherein the lifecycle of the data flow includes where the data is processed.
5. The method of any one of claims 1 to 4, wherein the lineage indicates how the data was acquired and processed.
6. The method of any one of claims 1 to 5, wherein the method is performed in a cloud or cloud-based computing environment.
7. a computer including one or more processors; The one or more processors: The data ingested into the first topic of the scalable input / output layer is normalized by the normalization application in the computation layer, and the resulting temporary slices are written to the data lake. the temporary slices written to the data lake are processed by each of one or more applications of a data flow in the compute layer, and each temporary slice is written to the data lake; publishing the temporary slices written to the data lake by a publishing application in the compute layer to a second topic in the scalable input / output layer; Publish the temporary slice published to the second topic to an output HUB; The system updates data reconstruction and lineage tracking information, including lineage and security, for each operation between ingesting and publishing data to the first topic of the scalable input / output layer.
8. A computer readable program for causing one or more processors to carry out the method of any one of claims 1 to 6.