Systems and methods for automatically enriching datasets with system knowledge data
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ORACLE INT CORP
- Filing Date
- 2023-09-22
- Publication Date
- 2026-04-23
AI Technical Summary
Existing enterprise software users face challenges in efficiently extracting and integrating data from horizontal and vertical business applications into data warehouses, which is both time and resource-intensive, limiting the effectiveness of data analytics and business intelligence.
A data analytics environment with a control plane and data plane architecture that automates the extraction, transformation, and loading of data from enterprise applications into a data warehouse, utilizing a data pipeline and transformation layer to convert data into a model format suitable for analysis, and a semantic layer for user understanding, supported by a query engine for federated queries.
Facilitates efficient and automated data enrichment, enabling users to derive meaningful insights from diverse data sources, enhancing strategic decision-making through improved data integration and visualization.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] Copyright Notice A portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to the copying by anyone of this patent document or the patent disclosure, as it appears in the U.S. Patent and Trademark Office patent file or records, but otherwise reserves all and any rights under copyright.
[0002] Priority claims This application claims the benefit of priority to U.S. provisional patent application Ser. No. 63 / 416,379, entitled "SYSTEM AND METHOD FOR AUTOMATICALLY ENRICHING DATASETS WITH SYSTEM KNOWLEDGE DATA," filed on October 14, 2022, and U.S. patent application Ser. No. 18 / 137,963, entitled "SYSTEM AND METHOD FOR AUTOMATICALLY ENRICHING DATASETS WITH SYSTEM KNOWLEDGE DATA," filed on April 21, 2023, the contents of which are incorporated herein by reference.
[0003] Technical Field FIELD OF THE INVENTION The embodiments described herein relate generally to computer-based methods of computer data analysis and providing business intelligence or other data, and particularly to systems and methods for automatically enriching datasets in a data analysis environment with system knowledge data. [Background technology]
[0004] background Data analytics allows for the computer-based examination of large amounts of data, for example, to derive conclusions or other information from the data. For example, business intelligence tools may be used to provide users with business intelligence that describes their enterprise data in a format that enables them to make strategic business decisions. Summary of the Invention
[0005] overview According to one embodiment, described herein are systems and methods for automatically enriching datasets in a data analytics environment with system knowledge data. The system may operate to automatically enrich datasets as the datasets are analyzed. Users of the data analytics environment, such as business users preparing data visualizations, may not be aware of additional data and system knowledge data available to improve the data visualization. The systems and methods described herein can provide automatic enrichment of data, for example from a knowledge repository, which can be delivered to data analytics customers using various delivery means. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 illustrates an exemplary data analysis environment, according to one embodiment. [Figure 2] FIG. 2 further illustrates an exemplary data analysis environment, according to one embodiment. [Figure 3] FIG. 2 further illustrates an exemplary data analysis environment, according to one embodiment. [Figure 4] FIG. 2 further illustrates an exemplary data analysis environment, according to one embodiment. [Figure 5] FIG. 2 further illustrates an exemplary data analysis environment, according to one embodiment. [Figure 6] FIG. 1 illustrates the use of the system to transform, analyze, or visualize data, according to one embodiment. [Figure 7]1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 8] 1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 9] 1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 10] 1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 11] 1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 12] 1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 13] 1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 14] 1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 15] 1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 16] 1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 17] 1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 18] 1A-1D illustrate various examples of user interfaces for use in a data analysis environment, according to one embodiment. [Figure 19] FIG. 1 illustrates a system for automatically enriching a dataset in a data analytics environment with system knowledge data, according to one embodiment. [Figure 20]1 is a flowchart of a method for automatically enriching a dataset with system knowledge data in a data analytics environment, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0007] Detailed Description Generally speaking, within an organization, data analytics enables the computer-based examination of large amounts of data, for example, to derive conclusions or other information from the data. For example, business intelligence (BI) tools may be used to provide users with business intelligence that describes their enterprise data in a format that enables them to make strategic business decisions.
[0008] Examples of such business intelligence tools / servers include Oracle Business Intelligence Applications (OBIA), Oracle Business Intelligence Enterprise Edition (OBIEE), or Oracle Business Intelligence Server (OBIS), which provide query, reporting, and analytical servers that can work in conjunction with databases to support functions such as data mining or analysis, and analytical applications.
[0009] Data analytics can increasingly be delivered within the context of enterprise software application environments, such as Oracle Fusion Applications environments, or within software-as-a-service (SaaS) or cloud environments, such as Oracle Analytics Cloud or Oracle Cloud Infrastructure environments, or other types of analytical applications or cloud environments.
[0010] Introduction According to one embodiment, a data warehouse environment or component, such as, for example, an Oracle Autonomous Data Warehouse (ADW), an Oracle Autonomous Data Warehouse Cloud (ADWC), or any other type of data warehouse environment or component adapted to store large amounts of data, can provide a central repository for storing data collected by one or more business applications.
[0011] For example, according to one embodiment, a data warehouse environment or component may be provided as a multidimensional database that utilizes online analytical processing (OLAP) or other techniques to generate business-related data from multiple disparate data sources. An organization may extract such business-related data from one or more vertical and / or horizontal business applications and inject the extracted data into a data warehouse instance associated with the organization.
[0012] As mentioned above, examples of horizontal business applications include ERP, HCM, CX, SCM, and EPM, which can provide a wide range of functions across various corporate organizations.
[0013] Vertical business applications are generally narrower in scope than horizontal business applications, but provide access to data further up or further down the data chain within a defined area or industry. Examples of vertical business applications might be medical software or banking software used within a particular organization.
[0014] While software vendors are increasingly offering enterprise software products or components as SaaS or cloud-oriented offerings, such as Oracle Fusion Applications, while other enterprise software products or components, such as Oracle ADWC, may be offered as one or more of SaaS, Platform-as-a-Service (PaaS), or hybrid subscriptions, enterprise users of traditional business intelligence applications and processes are typically faced with the task of extracting data from their horizontal and vertical business applications and introducing the extracted data into a data warehouse, a process that can be both time and resource intensive.
[0015] According to one embodiment, the analytical application environment enables customers (tenants) to develop computer-executable software analytical applications for use with a BI component, such as an OBIS environment, or other type of BI component adapted to explore large amounts of data sourced by the customers (tenants) themselves or from multiple third-party entities.
[0016] As another example, according to one embodiment, the analytical application environment may be used to pre-configure the reporting interface of a data warehouse instance with associated metadata that describes business-related data objects in the context of various business productivity software applications, for example, to include pre-defined dashboards, key performance indicators (KPIs), or other types of reports.
[0017] Data analysis Generally speaking, data analytics allows for the computer-based exploration or analysis of large amounts of data in order to derive conclusions or other information from that data, while business intelligence tools (BI) provide an organization's business users with information that describes their enterprise data in a format that enables them to make strategic business decisions.
[0018] Examples of data analytics environments and business intelligence tools / servers include Oracle Business Intelligence Server (OBIS), Oracle Analytics Cloud (OAC), and Fusion Analytics Warehouse (FAW), which support functions such as data mining or analysis and analytical applications.
[0019] FIG. 1 illustrates an exemplary data analysis environment, according to one embodiment. The exemplary embodiment shown in Figure 1 is provided for purposes of illustrating an example of a data analysis environment that may be used in connection with various embodiments described herein. According to other embodiments and examples, the techniques described herein may be used with other types of data analysis, database, or data warehouse environments. The components and processes shown in Figure 1 and as further described herein with respect to various other embodiments may be provided as software or program code executable by, for example, a cloud computing system or other appropriately programmed computer system.
[0020] As shown in FIG. 1, according to one embodiment, data analysis environment 100 may be provided by, or otherwise operate on, a computer system having computer hardware (e.g., processor, memory) 101 and including one or more software components that operate as a control plane 102 and a data plane 104 and provide access to a data warehouse, data warehouse instance 160 (database 161, or other type of data source).
[0021] According to one embodiment, the control plane operates to provide control of cloud products or other software products offered within the context of a SaaS or cloud environment, such as an Oracle Analytics Cloud environment, or other type of cloud environment. For example, according to one embodiment, the control plane may include a console interface 110 that allows access by customers (tenants) and / or cloud environments having provisioning components 111.
[0022] According to one embodiment, the console interface may allow access by customers (tenants) operating a graphical user interface (GUI) and / or a command line interface (CLI) or other interface, and / or may include an interface for use by a provider of a SaaS or cloud environment and its customers (tenants). For example, according to one embodiment, the console interface may provide an interface that allows customers to provision services for use within their SaaS environment and configure those services that have been provisioned.
[0023] According to one embodiment, a customer (tenant) can request provisioning of a customer schema within a data warehouse. The customer can also provide, via a console interface, a number of attributes associated with the data warehouse instance, including required attributes (e.g., login credentials) and optional attributes (e.g., size or speed). The provisioning component can then provision the requested data warehouse instance including the data warehouse customer schema and populate the data warehouse instance with the appropriate information provided by the customer.
[0024] According to one embodiment, the provisioning component may also be used to update or edit the data warehouse instance and / or ETL processes running on the data plane, for example, by changing or updating the requested frequency of ETL process execution for a particular customer (tenant).
[0025] According to one embodiment, the data plane may include a data pipeline or processing layer 120 and a data transformation layer 134 that together process operational or transactional data from an organization's enterprise software applications or data environments, for example, business productivity software applications provisioned in a customer's (tenant's) SaaS environment. The data pipeline or processing may include various functionality to extract transactional data from business applications and databases provisioned in the SaaS environment and then load the transformed data into a data warehouse.
[0026] According to one embodiment, the data transformation layer may include data models, such as knowledge models (KMs) or other types of data models, that the system uses to transform transactional data received from business applications and corresponding transactional databases provisioned in the SaaS environment into a model format understood by the data analytics environment. The model format may be provided in any data format suitable for storage in a data warehouse. According to one embodiment, the data plane may also include data and configuration user interfaces, and mapping and configuration databases.
[0027] According to one embodiment, the data plane is responsible for performing extract, transform, and load (ETL) operations, including extracting transactional data from an organization's enterprise software applications or data environment, such as business productivity software applications and corresponding transactional databases provided in a SaaS environment, transforming the extracted data into a model format, and loading the transformed data into customer schemas in a data warehouse.
[0028] For example, according to one embodiment, each customer (tenant) of the environment can be associated with its own customer tenancy in the data warehouse, which is associated with its own unique customer schema, and can further be provided read-only access to a data analysis schema, which can be updated periodically or on another basis by a data pipeline or process, such as an ETL process.
[0029] According to one embodiment, a data pipeline or process may be scheduled to run at intervals (e.g., hourly, daily, weekly) to extract transactional data from an enterprise software application or data environment, such as, for example, a business productivity software application and corresponding transactional database 106 provisioned in a SaaS environment.
[0030] According to one embodiment, the extraction process 108 may extract transaction data, during which a data pipeline or process may insert the extracted data into a data staging area, which may serve as a temporary staging area for the extracted data. Data quality and data protection components may be used to ensure the integrity of the extracted data. For example, according to one embodiment, a data quality component may perform validation of the extracted data while the data is temporarily held in the data staging area.
[0031] According to one embodiment, once the extraction process has completed its extraction, a transformation process can begin using a data transformation layer to convert the extracted data into a model format that will be loaded into a customer schema in the data warehouse.
[0032] According to one embodiment, the data pipeline or process may operate in combination with a data transformation layer to transform data into a model format. A mapping and configuration database may store metadata and data mappings that define the data model used by the data transformation. A data and configuration user interface (UI) may facilitate access and modification of the mapping and configuration database.
[0033] According to one embodiment, the data transformation layer can convert the extracted data into a format suitable for loading into a data warehouse customer schema, for example, according to a data model. During the transformation, the data transformation can perform dimension generation, fact generation, and aggregation generation as needed. Dimension generation can include generating dimensions or fields for loading into the data warehouse instance.
[0034] According to one embodiment, after transforming the extracted data, the data pipeline or process may execute a warehouse load procedure 150 to load the transformed data into a customer schema of a data warehouse instance. After loading the transformed data into the customer schema, the transformed data can be analyzed and used in a variety of additional business intelligence processes.
[0035] Different customers of the data analytics environment may have different requirements for how their data is classified, aggregated, or transformed for the purposes of providing data analysis or business intelligence data or developing software analytics applications. According to one embodiment, to support such different requirements, semantic layer 180 may include data that defines a semantic model of the customer's data, which helps to help users understand and access that data using commonly understood business terms. Semantic layer 180 may also provide custom content to presentation layer 190.
[0036] According to one embodiment, a semantic model can be specified, for example in an Oracle environment, as a BI repository (RPD) file having metadata that specifies logical schemas, physical schemas, physical-to-logical mappings, aggregate table navigation, and / or other components that implement various physical, business model and mapping, and presentation layer aspects of the semantic model.
[0037] According to one embodiment, a customer can make modifications to their data source model to support their specific requirements, for example, by adding custom facts or dimensions associated with the data stored in their data warehouse instance, and the system can extend the semantic model accordingly.
[0038] According to one embodiment, the presentation layer may enable access to data content using, for example, software analytics applications, user interfaces, dashboards, key performance indicators (KPIs), or other types of reports or interfaces, such as may be provided by products such as Oracle Analytics Cloud or Oracle Analytics for Applications.
[0039] Business Intelligence Server According to one embodiment, query engine 18 (e.g., an OBIS instance) operates in the manner of a federated query engine to respond to analytical queries or requests, for example from clients within an Oracle Analytics Cloud environment, directed to data stored in the database.
[0040] According to one embodiment, the OBIS instance can push down operations to supported databases according to a query execution plan 56, where a logical query can include Structured Query Language (SQL) statements received from a client, while a physical query includes database-specific statements that the query engine sends to the database to retrieve data when processing the logical query. In this way, the OBIS instance translates business user queries into the appropriate database-specific query language (e.g., Oracle SQL, SQL Server SQL, DB2 SQL, or Essbase MDX). The query engine (e.g., OBIS) may also support internal execution of SQL operators that cannot be pushed down to the database.
[0041] According to one embodiment, a user / developer can interact with a client computing device 10, which includes computing hardware 11 (e.g., processor, storage, memory), a user interface 12, and an application 14. A query engine or business intelligence server, such as OBIS, generally operates to process inbound requests to a database model, e.g., SQL requests, construct and execute one or more physical database queries, process the data appropriately, and then return the data in response to the request.
[0042] To accomplish this, according to one embodiment, a query engine or business intelligence server may include various components or functions, such as a logical or business model or metadata that describes the data available as the subject area of a query, a request generator that receives incoming queries and converts them into physical queries for use with connected data sources, and a navigator that receives incoming queries, navigates the logical model, and generates physical queries that optimally return the data needed for a particular query.
[0043] For example, according to one embodiment, a query engine or business intelligence server may utilize a logical model that is mapped to the data in the data warehouse by creating a simplified star-schema business model for various data sources so that users can query the data as if it were from a single source. Information can then be returned to the presentation layer as subject areas according to the mapping rules in the business model layer.
[0044] According to one embodiment, a query engine (e.g., OBIS) can process queries against a database according to a query execution plan, which can include various child (leaf) nodes, generally referred to in various embodiments herein as RqList, such as: Execution plan: [[ RqList < <191986> > [for database 0:0,0] D102.c1 as c1 [for database 0:0,0], sum(D102.c2 by [ D102.c1] ) as c2 [for database 0:0,0] Child Nodes (RqJoinSpec): < <192970> > [for database 0:0,0] RqJoinNode < <192969> > [] ( RqList < <193062> > [for database 0:0,0] D2.c2 as c1 [for database 0:0,0], D1.c2 as c2 [for database 0:0,0] Child Nodes (RqJoinSpec): < <193065> > [for database 0:0,0] RqJoinNode < <193061> > [] ( RqList < <192414> > [for database 0:0,118] T1000003.Customer_ID as c1 [for database 0:0,118], T1000003.TARGET as c2 [for database 0:0,118] Child Nodes (RqJoinSpec): < <192424> > [for database 0:0,118] RqJoinNode < <192423> > [] [users / administrator / dv_joins / multihub / input::##dataTarget] as T1000003 ) as D1 LeftOuterJoin (Eager) < <192381> > On D1.c1 = D2.c1; actual join vectors:
[0000] =
[0000] ( RqList < <192443> > [for database 0:0,0] D104.c1 as c1 [for database 0:0,0], nullifnotunique(D104.c2 by [ D104.c1] ) as c2 [for database 0:0,0] Child Nodes (RqJoinSpec): < <192928> > [for database 0:0,0] RqJoinNode < <192927> > [] ( RqList < <192852> > [for database 0:0,118] T1000006.Customer_ID as c1 [for database 0:0,118], T1000006.Customer_City as c2 [for database 0:0,118] Child Nodes (RqJoinSpec): < <192862> > [for database 0:0,118] RqJoinNode < <192861> > [] [users / administrator / dv_joins / my_customers / input::data] as T1000006 ) as D104 GroupBy: [ D104.c1] [for database 0:0,0] sort OrderBy: c1, Aggs:[ nullifnotunique(D104.c2 by [ D104.c1] ) ] [for database 0:0,0] ) as D2 ) as D102 GroupBy: [ D102.c1] [for database 0:0,0] sort OrderBy: c1 asc, Aggs:[ sum(D102.c2 by [ D102.c1] ) ] [for database 0:0,0] Within a query execution plan, each execution plan component (RqList) represents a block of queries within the query execution plan and is generally translated into a SELECT statement. An RqList may have nested child RqLists, similar to how a SELECT statement selects from nested SELECT statements.
[0045] According to one embodiment, the query engine can interact with different databases, for each of which it can use a data source-specific code generator. A typical strategy is to send as many SQL executions as possible to the database by submitting them as part of the physical query, thereby reducing the amount of information returned to the OBIS server.
[0046] According to one embodiment, during operation, the query engine or business intelligence server can create a query execution plan that can then be further optimized to perform, for example, aggregations of the data necessary to respond to the request, combine the data together, and apply further calculations before returning the results to the calling application, for example via an ODBC interface.
[0047] According to one embodiment, complex multi-pass requests involving multiple data sources may require the query engine or business intelligence server to decompose the query, determine which sources, multi-pass calculations, and aggregations are available, and generate a logical query execution plan that spans multiple databases and physical SQL statements, the results of which can then be returned and further joined or aggregated by the query engine or business intelligence server.
[0048] FIG. 2 further illustrates an exemplary data analysis environment, according to one embodiment. 2, according to one embodiment, the provisioning component may also include a provisioning application programming interface (API) 112, a number of workers 115, a metering manager 116, and a data plane API 118, as further described below. The console interface may communicate with the provisioning API by making API calls, for example, when commands, instructions, or other inputs are received at the console interface, to provision services within the SaaS environment or to make configuration changes to provisioned services.
[0049] According to one embodiment, a data plane API can communicate with the data plane. For example, according to one embodiment, provisioning and configuration changes directed to services provided by the data plane can be communicated to the data plane via the data plane API.
[0050] According to one embodiment, the metering manager may include various functionality for metering services provisioned through the control plane and usage of those services. For example, according to one embodiment, the metering manager may record usage over time of processors provisioned through the control plane for a particular customer (tenant) for billing purposes. Similarly, the metering manager may record the amount of data warehouse storage space partitioned for use by customers of a SaaS environment for billing purposes.
[0051] According to one embodiment, the data pipeline or processing provided by the data plane may include a monitoring component 122, a data staging component 124, a data quality component 126, and a data projection component 128, as further described below.
[0052] According to one embodiment, the data transformation layer may include a dimension generation component 136, a fact generation component 138, and an aggregation generation component 140, as described further below. The data plane may also include a data and configuration user interface 130 and a mapping and configuration database 132.
[0053] According to one embodiment, the data warehouse may include a default data analysis schema (referred to herein, according to some embodiments, as an analysis warehouse schema) 162 and a customer schema 164 for each customer (tenant) of the system.
[0054] According to one embodiment, to support multiple tenants, the system may enable the use of multiple data warehouses or data warehouse instances. For example, according to one embodiment, a first warehouse customer tenancy for a first tenant may include a first database instance, a first staging area, and a first data warehouse instance of the multiple data warehouses or data warehouse instances, while a second customer tenancy for a second tenant may include a second database instance, a second staging area, and a second data warehouse instance of the multiple data warehouses or data warehouse instances.
[0055] According to one embodiment, based on the data model defined in the mapping and configuration database, the monitoring component can determine the dependencies of several different datasets (data sets) to be converted. Based on the determined dependencies, the monitoring component can determine which of the several different datasets should be converted to the model format first.
[0056] For example, according to one embodiment, if a first model dataset does not include a dependency on any other model dataset, and if a second model dataset includes a dependency on the first model dataset, the monitoring component may determine to transform the first dataset before the second dataset to accommodate the dependency of the second dataset on the first dataset.
[0057] For example, according to one embodiment, a dimension may include categories of data, such as "Name," "Address," or "Age." Generating facts involves generating values, or "measurements," that the data can take. Facts can be associated with appropriate dimensions within the data warehouse instance. Generating aggregations involves creating data mappings that calculate aggregations of the transformed data to existing data in a customer schema of the data warehouse instance.
[0058] According to one embodiment, once any transformations have been performed (as specified by the data model), a data pipeline or process can read the source data, apply the transformations, and then push the data to a data warehouse instance.
[0059] According to one embodiment, data transformations can be expressed in rules, and once transformations have occurred, values can be held intermediately in a staging area where data quality and data projection components can validate and check the integrity of the transformed data before the data is uploaded to customer schemas in the data warehouse instance. Monitoring can be provided as the extract, transform, and load processes run, for example across multiple computing instances or virtual machines. Dependencies can also be maintained during the extract, transform, and load processes, and the data pipeline or processes can accommodate such sequencing decisions.
[0060] According to one embodiment, after transforming the extracted data, the data pipeline or process may execute a warehouse load procedure to load the transformed data into a customer schema of a data warehouse instance. After loading the transformed data into the customer schema, the transformed data can be analyzed and used in a variety of additional business intelligence processes.
[0061] FIG. 3 further illustrates an exemplary data analysis environment, according to one embodiment. As shown in FIG. 3, according to one embodiment, data may be sourced from a customer's (tenant's) enterprise software application or data environment (106), for example, using data pipeline processing, or may be sourced as custom data 109 from one or more customer-specific applications 107, and loaded into the data warehouse instance, in some examples including using object storage 105 for data storage.
[0062] In embodiments of an analytics environment, such as Oracle Analytics Cloud (OAC), users can create datasets that use tables from different connections and schemas, and the system uses the relationships defined between these tables to create relationships or joins within the dataset.
[0063] According to one embodiment, for each customer (tenant), the system pre-configures the customer's data warehouse instance based on an analysis of data in that customer's enterprise application environment and in the customer tenancy 117 using a data analysis schema maintained and updated by the system in the system / cloud tenancy 114. Thus, the data analysis schema maintained by the system enables data to be retrieved from the customer's environment by a data pipeline or process and loaded into the customer's data warehouse instance.
[0064] According to one embodiment, the system also provides, for each customer of the environment, a customer schema that is easily modifiable by the customer and allows the customer to supplement and utilize the data in their data warehouse instance. For each customer, the resulting data warehouse instance acts as a database, but its contents are partly controlled by the customer and partly controlled by the environment (system).
[0065] For example, according to one embodiment, a data warehouse (e.g., ADW) may include a data analysis schema and a customer schema for each customer / tenant sourced from their enterprise software applications or data environment. Data provisioned in the data warehouse tenancy (e.g., ADW cloud tenancy) is accessible only to that tenant while simultaneously allowing access to various functions of the shared environment, such as ETL-related or other functions.
[0066] According to one embodiment, to support multiple customers / tenants, the system enables the use of multiple data warehouse instances, where, for example, a first customer tenancy may include a first database instance, a first staging area, and a first data warehouse instance, and a second customer tenancy may include a second database instance, a second staging area, and a second data warehouse instance.
[0067] According to one embodiment, for a particular customer / tenant, upon extraction of its data, the data pipeline or process may insert the extracted data into the tenant's data staging area, which may serve as a temporary staging area for the extracted data. For example, data quality and data protection components may be used to ensure the integrity of the extracted data by performing validation of the extracted data while it is temporarily held in the data staging area. Once the extraction process completes its extraction, a data transformation layer may be used to initiate a transformation process to convert the extracted data into a model format that will be loaded into the customer schema of the data warehouse.
[0068] FIG. 4 further illustrates an exemplary data analysis environment, according to one embodiment. As shown in FIG. 4 , according to one embodiment, the process of extracting data, for example from a customer's (tenant's) enterprise software application or data environment using a data pipeline process as described above, or as custom data sourced from one or more customer-specific applications, and loading the data into a data warehouse instance or refreshing the data in a data warehouse generally involves three broad stages performed by ETP services 160 or processes performed by one or more computing instances 170, including one or more extraction services 163, transformation services 165, and load / publish services 167.
[0069] For example, according to one embodiment, a list of view objects for extraction can be sent to an Oracle BI Cloud Connector (BICC) component, e.g., via a REST call. The extracted files can be uploaded to an object storage component, e.g., an Oracle Storage Service (OSS) component, for data storage. A transform process receives data files from the object storage component (e.g., OSS) and applies business logic while loading them into a target data warehouse, e.g., an ADW database, that is internal to the data pipeline or process and not exposed to customers (tenants). A load / publish service or process receives data from, e.g., an ADW database or warehouse, and publishes it to a data warehouse instance accessible to customers (tenants).
[0070] FIG. 5 further illustrates an exemplary data analysis environment, according to one embodiment. As shown in FIG. 5, which illustrates the operation of a system with multiple tenants (customers) according to one embodiment, data can be sourced, for example, from each of the multiple customer's (tenant's) enterprise software applications or data environments using data pipeline processing as described above, and loaded into a data warehouse instance.
[0071] According to one embodiment, the data pipeline or process maintains a data analysis schema that is periodically updated for each of multiple customers (tenants), e.g., Customer A 180, Customer B 182, by a system that follows best practices for a particular analytical use case.
[0072] According to one embodiment, for each of multiple customers (e.g., Customer A, B), the system populates the customer's data warehouse instance based on an analysis of data within that customer's enterprise application environment 106A, 106B and within each customer's tenancy (e.g., Customer A's tenancy 181, Customer B's tenancy 183) using data analysis schemas 162A, 162B that are maintained and updated by the system such that data is obtained from the customer's environment via a data pipeline or process and loaded into the customer's data warehouse instance 160A, 160B.
[0073] According to one embodiment, the data analysis environment also provides, for each of the environment's multiple customers, a customer schema (e.g., customer A schema 164A, customer B schema 164B) that is easily modifiable by the customer, thereby enabling the customers to supplement and utilize the data in their own data warehouse instance.
[0074] As noted above, according to one embodiment, for each of multiple customers of the data analytics environment, the resulting data warehouse instance acts as a database whose contents are controlled in part by the customer and in part by the data analytics environment (system), including appearing to be pre-populated with appropriate data retrieved from the enterprise application environment to address various analytical use cases. Once the extraction process 108A, 108B for a particular customer has completed its extraction, a data transformation layer can be used to initiate a transformation process to convert the extracted data into a model format that is loaded into the customer schema of the data warehouse.
[0075] According to one embodiment, to address a customer's (tenant's) specific needs, activation plans 186 can be used to control the operation of data pipelines or processing services for that customer for specific functional areas.
[0076] For example, according to one embodiment, an activation plan may specify a number of extract, transform, and load (publish) services or steps to be performed in a certain order, at a certain time, and within a certain time frame.
[0077] According to one embodiment, each customer can be associated with its own activation plan. For example, an activation plan for a first customer A can determine the tables to be retrieved from that customer's enterprise software application environment (e.g., a Fusion Applications environment) or determine the manner in which services and their operations are to be executed in a certain order, while an activation plan for a second customer B can similarly determine the tables to be retrieved from that customer's enterprise software application environment or determine the manner in which services and their operations are to be executed in a certain order.
[0078] FIG. 6 illustrates the use of the system to transform, analyze, or visualize data, according to one embodiment.
[0079] 6, according to one embodiment, the systems and methods disclosed herein can be used to provide a data visualization environment 192 that enables insights to a user of the analytical environment regarding analytical artifacts and the relationships between them. The model can then be used to visualize, for example, via a user interface, the relationships between such analytical artifacts as network charts or visualizations of the relationships and lineage between artifacts (e.g., User, Role, DV Project, Dataset, Connection, Dataflow, Sequence, ML Model, ML Script).
[0080] According to one embodiment, the client application may be implemented as software or computer-readable program code executable by a computer system or processing device and have a user interface, such as a software application user interface or a web browser interface. The client application may obtain or access data via an internet / HTTP or other type of network connection to the analysis system, or in the example of a cloud environment, via cloud services provided by the environment.
[0081] According to one embodiment, the user interface may include or provide access to various data flow action types that enable self-service text analysis, including allowing a user to view a data set or manipulate the user interface to transform, analyze, or visualize data, for example, to generate graphs, charts, or other types of data analysis or data flow visualizations, as described in more detail below.
[0082] According to one embodiment, the analytics system allows for retrieving, receiving, or preparing datasets from one or more data sources, for example, via one or more data source connections. Examples of types of data that can be transformed, analyzed, or visualized using the systems and methods described herein include HCM, HR, or ERP data, email or text messages, or other free-form or unstructured text data provided in one or more databases, data storage services, or other types of data repositories or data sources.
[0083] For example, according to one embodiment, a request for data analysis or visualization information can be received via a client application and user interface such as those described above and communicated to an analytics system (via a cloud service in the example cloud environment). The system can retrieve a dataset appropriate to address the user / business context for use in generating and returning the requested data analysis or visualization information to the client. For example, the data analytics system can retrieve the dataset using, for example, a SELECT statement or a logical SQL instruction.
[0084] According to one embodiment, the system can create a model or dataflow that reflects an understanding of the dataflow or set of input data by applying various algorithmic processes to generate visualizations or other types of useful information associated with the data. The model or dataflow can be further modified in the dataset editor 193 by applying various processes or techniques to the dataflow or set of input data, for example, including one or more dataflow actions 194, 195 or steps that operate on the dataflow or set of input data. A user can interact with the system through a user interface to control the use of dataflow actions to generate data analysis, data visualizations 196, or other types of useful information associated with the data.
[0085] According to one embodiment, a dataset is a self-service data model that users can build to meet their data visualization and analysis requirements. A dataset includes data source connection information, tables and columns, data enrichment and transformations. Users can use a dataset in multiple workbooks and data flows.
[0086] According to one embodiment, when a user creates and builds a dataset, the user can, for example, select from many types of connections or spreadsheets, create a dataset based on data from multiple tables in a database connection, an Oracle data source, or a local subject area, or create a dataset based on data from a variety of connections and subject areas.
[0087] For example, according to one embodiment, a user can build a dataset that includes tables from an Autonomous Data Warehouse connection, tables from a Spark connection, and tables from a local subject area, specify joins between the tables, and transform and enrich columns in the dataset.
[0088] According to one embodiment, additional artifacts, features, and actions associated with the dataset may include, for example:
[0089] Viewing Available Connections: A dataset uses one or more connections to data sources to access data and provide it for analysis and visualization. A user's connection list includes connections that the user has established and connections that the user has permission to access and use.
[0090] Creating a dataset from a connection: When a user creates a dataset, the user can add tables, add joins, and enrich the data from one or more data source connections.
[0091] Adding multiple connections to a dataset: A dataset can contain two or more connections. Adding more connections allows the user to access and join all the tables and data needed to build the dataset. Users can add more connections to datasets that support multiple tables.
[0092] Creating Dataset Table Joins: Joins represent relationships between tables in a dataset. If a user is creating a dataset based on facts and dimensions, and if joins already exist in the source tables, joins are automatically created in the dataset. If a user is creating a dataset from multiple connections and schemas, the user can manually define joins between tables.
[0093] According to one embodiment, a user can use Dataflow to create a dataset by combining, organizing, and integrating data. Dataflow allows users to organize and integrate data to create curated datasets that can be visualized by either themselves or other users.
[0094] For example, according to one embodiment, a user may use a dataflow to create a dataset, combine data from different sources, aggregate data, and train or apply predictive machine learning models to the data.
[0095] According to one embodiment, a dataset editor such as that described above allows a user to add actions or steps, where each step performs a specific function, such as adding data, joining tables, merging columns, transforming data, or saving data. Each step is validated as the user adds or modifies it. Once the user has completed configuring the dataflow, the user can execute it to create or update a dataset.
[0096] According to one embodiment, users can curate data from datasets, subject areas, or database connections. Users can run dataflows individually or sequentially. Users can include multiple data sources in a dataflow and specify how they are joined. Users can save output data from a dataflow to either a dataset or a supported database type.
[0097] According to one embodiment, additional artifacts, features, and operations associated with a data flow may include, for example:
[0098] Add Column: Add a custom column to the target dataset. Add Data: Adds a data source to a dataflow. For example, if a user wants to merge two datasets, the user adds both datasets to the dataflow.
[0099] Aggregation: Applying an aggregation function, such as count, sum, or average, to create group totals.
[0100] Branch: Creating multiple outputs from a dataflow. Filter: Selects only the data that interests the user.
[0101] Joins: Combining data from multiple data sources using database joins based on common columns.
[0102] Graph analysis: Performing geospatial analysis, such as calculating the distance or number of hops between two vertices.
[0103] The above are provided as examples, and according to one embodiment, other types of steps can be added to the data flow to transform the data set or provide data analysis or visualization.
[0104] Dataset analysis and visualization According to one embodiment, the system provides functionality that allows a user to generate datasets, analyses, or visualizations for display within a user interface, for example, to explore datasets or data sourced from multiple data sources.
[0105] 7-18 show various examples of user interfaces for use in a data analysis environment, according to one embodiment.
[0106] The user interfaces and functionality illustrated in FIGS. 7-18 are provided as examples for purposes of illustrating the various features described herein, and alternative examples of user interfaces and functionality may be provided according to various embodiments.
[0107] As shown in FIGS. 7-8, according to one embodiment, a user may access a data analysis environment, for example, to submit analyses or queries against an organization's data.
[0108] For example, according to one embodiment, a user can choose from various types of connections to create a dataset based on data from tables in, for example, a database connection, an Oracle subject area, an Oracle ADW connection, or a spreadsheet, file, or other type of data source. In this way, the dataset acts as a self-service data model from which a user can build data analysis or visualizations.
[0109] As shown in Figures 9-10, according to one embodiment, the dataset editor can display a list of connections to which a user has access privileges and enable the user to create or edit datasets containing tables, joins, and / or enriched data. The editor can display the schemas and tables of the data source connections, from which the user can drag and drop onto the dataset diagram. If a particular connection does not itself provide a list of schemas and tables, the user can use manual queries for the appropriate tables. Adding a connection allows the user to access and join the associated tables and data to build a dataset.
[0110] According to one embodiment, a join diagram in the dataset editor displays the tables and joins in a dataset, as shown in Figures 11-12. Joins specified in a data source can be automatically created between tables in a dataset, for example, by creating joins based on column name matches found between tables.
[0111] According to one embodiment, when a user selects a table, a preview data area displays a sample of the table's data. Displayed join links and icons indicate which tables are joined and the type of join used. Users can create joins by dragging and dropping one table onto another, click on a join to view or update its configuration, or click on a column's type attribute to change its type, for example, from measure to attribute.
[0112] According to one embodiment, the system can generate source-specific optimized queries for visualizations, where the dataset is treated as a data model and only the tables necessary to satisfy the visualization are used in the query.
[0113] By default, the granularity of a dataset is determined by the table with the lowest granularity. Users can create measurements on any table in a dataset, but this can result in duplicate measurements on one side of a one-to-many or many-to-many relationship. To address this, according to the embodiment shown in Figure 13, users can maintain granularity by setting a table on one side of the cardinality to maintain that level of detail.
[0114] As shown in FIG. 14, according to one embodiment, dataset tables can be associated with data access settings that determine whether the system will load the table into cache or whether the table will receive its data directly from the data source.
[0115] According to one embodiment, when automatic caching mode is selected for a table, the system loads or reloads the table data into the cache, which allows for faster performance when the table's data is refreshed, for example from a workbook, and also exposes reload menu options at the table and dataset level.
[0116] According to one embodiment, when live mode is selected for a table, the system retrieves table data directly from the data source and the source system manages the data source queries for the table. This option is useful when data is stored in a high performance data warehouse, for example, Oracle ADW, and ensures that the most recent data is used.
[0117] According to one embodiment, when a dataset uses multiple tables, some tables may use automatic caching and some tables may contain live data. During a reload of multiple tables using the same connection, if a table's data fails to reload, any tables currently set to use automatic caching are switched to retrieving their data using live mode.
[0118] According to one embodiment, the system allows users to enrich and transform their data before it is available for analysis. When a workbook is created and a dataset is added to it, the system performs column-level profiling on a representative sample of the data. After profiling the data, the user can implement the transformation and enrichment recommendations provided for recognizable columns in the dataset, such as GPS enrichment with city latitude and longitude or zip code.
[0119] According to one embodiment, data transformation and enrichment changes applied to a dataset affect workbooks and dataflows that use that dataset. For example, when a user opens a workbook that shares a dataset, the user receives a message indicating that the workbook is using updated or refreshed data.
[0120] According to one embodiment, dataflows provide a means for organizing and integrating data to produce curated datasets that users can visualize. For example, users can use dataflows to create datasets, combine data from various sources, aggregate data, or train machine learning models or apply predictive machine learning models to that data.
[0121] 15, according to one embodiment, each step in a dataflow performs a specific function, for example, to add data, join tables, merge columns, transform data, or save data. Once configured, a dataflow can be executed to perform operations to create or update a dataset, including, for example, the use of SQL operators, conditional expressions, or functions such as BETWEEN®, LIKE, IN, etc.
[0122] According to one embodiment, data flows can be used to merge data sets, cleanse data, and output results to a new data set. Data flows can be executed individually or sequentially. If a data flow in a sequence fails, all changes made in that sequence are rolled back.
[0123] As shown in Figures 16-18, according to one embodiment, visualizations can be displayed within a user interface, for example, to explore and add insight into datasets or data sourced from multiple data sources.
[0124] For example, according to one embodiment, a user can create a workbook, add a dataset, and then drag and drop its columns onto a canvas to create a visualization. The system can automatically generate the visualization based on the contents of the canvas, automatically selecting one or more visualization types for the user to choose from. For example, if a user adds an income measure to the canvas, the data element can be placed in the Value area of the grammar panel and the tile visualization type can be selected. The user can continue to add data elements directly to the canvas to build the visualization.
[0125] According to one embodiment, the system can provide automatically generated data visualizations (automatically generated insights, auto-insights) by suggesting visualizations that are expected to provide the best insights for a particular data set. A user can see an automatically generated summary of the insights, for example, by hovering over the associated visualization in a workbook canvas.
[0126] Automatic enrichment of datasets According to one embodiment, a system and method are described herein for automatically enriching datasets in a data analytics environment with system knowledge data.
[0127] According to one embodiment, users of data analysis environments, particularly business users and users of data visualization, for example, often lack or are unaware of additional knowledge that may improve their data visualization.
[0128] According to one embodiment, the systems and methods described herein may provide a "no-click," automated way to enrich customer data with a knowledge repository. Such knowledge repository may be distributed to customers using a distribution system, such as, for example, the three distribution mechanisms described below.
[0129] First, according to one embodiment, the system and method can deliver static knowledge datasets that can be automatically combined on the fly with customer data, resulting in enriched customer data. Such static knowledge datasets can include, for example, basic calculations on the customer dataset, such as automatically calculating minimum, maximum, average, or standard deviation to provide suggestions.
[0130] Second, according to one embodiment, the system and method can readily deliver third-party integrations, for example, by making runtime REST calls to external services to pull dynamically enriched knowledge columns based on automatic join keys detected in customer data. The automatic join key detection algorithm is not limited to column name matching but can be more sophisticated by detecting the semantic type of columns based on sampled data profiling. Currency conversion can be one use case targeted for this type of integration. In one embodiment, such automatic joins are enabled by using a service, such as Oracle BI Server, that has the ability to combine disparate sources and present them as a single entity to end users.
[0131] Third, according to one embodiment, the system and method may provide a well-defined integration handshake for customers to implement customized integrations with either external third-party datasets and services or customer-hosted in-house datasets.
[0132] According to one embodiment, the system and method may provide a number of system key-value system data sets where the keys are common attributes that may be used elsewhere. For example, useful keys may be related to geography, or geography plus time. A key containing "zip code" or "country / state / city / year" is another good example. The value may reflect some domain knowledge.
[0133] According to one embodiment, when a user is using their own dataset to build a visualization, the system and method can analyze it on the fly to find columns that contain similar keys that can be used in joins, and in addition to the customer columns, the system and method can provide columns from a knowledge dataset (e.g., external to the user's dataset) to be used to build the visualization.
[0134] FIG. 19 illustrates a system for automatically enriching a dataset in a data analytics environment with system knowledge data, according to one embodiment.
[0135] As shown in FIG. 19, according to one embodiment, an architecture 1900 may be utilized that can make or suggest a number of enhancements to a user's dataset to produce one or more data visualization (DV) workbooks.
[0136] According to one embodiment, when a user accesses or uploads a dataset, the system and method may perform one or more of data pre-augmentation, OBIS function calculation, dataset insight statistics, and external enrichment. Note that each of these functions may optionally be presented to the user as optional enhancements to the dataset for use in their data visualization (e.g., providing the suggested enhancements as selectable options), or the enhancements may be automatically applied to the user's data visualization.
[0137] According to one embodiment, data pre-enrichment, OBIS function calculation, and dataset insight statistics may comprise numerous functions and extensions 1910 that are automatically applied to existing or newly uploaded client datasets. For example, as shown in FIG. 19 , the dataset “AccidentsDemoFinal” includes a knowledge base that includes city and state zip codes, among other data, such as vehicle speed limits or vehicle occupancy counts. Data pre-enrichment, OBIS function calculation, and dataset insight statistics may automatically calculate numerous statistics, such as minimum, maximum, mean, standard deviation (shown in FIG. 19 as total occupants, average occupants, minimum occupants, maximum occupants, and various percentiles). By automatically providing such functionality, this enables users who would otherwise be unaware or unable (due to lack of product knowledge) to view or use the join functionality to more quickly and easily generate better data visualizations.
[0138] According to one embodiment, in addition to the functions and calculations automatically applied to a user's dataset as described above, the systems and methods described herein can additionally incorporate data from sources external to the user's dataset. For example, as shown in Figure 19, the systems and methods can automatically incorporate census data upon detecting a zip code column in the user's dataset and automatically add or suggest joins between the census data and the user's dataset.
[0139] According to one embodiment, the above example utilizes census data, however, one skilled in the art will readily appreciate that the described embodiments may utilize additional data sources, such as demographic data, public health information, publicly available statistical data, etc.
[0140] According to one embodiment, for example, AccidentsDemoFinal is an uploaded dataset. The dataset includes a data column called "Zip Code." From this "Zip Code" data column in the customer's dataset, the system can automatically detect the data column, pull in external data, for example, from census data, and make this newly accessed data available and accessible. This is an example of on-the-fly non-persistent data. Once selected, the system creates a join between the customer's dataset and the external data.
[0141] According to one embodiment, data external to the user's dataset can be captured directly from a source (e.g., the U.S. Census website), or such data can be captured from another local source, such as, for example, a database containing previously captured census data.
[0142] FIG. 20 is a flowchart of a method for automatically enriching a dataset with system knowledge data, according to one embodiment.
[0143] As depicted in FIG. 20, according to one embodiment, in step 2010, the method may provide a data analysis environment in a computer having a microprocessor, operable to display data visualizations associated with the dataset.
[0144] According to one embodiment, in step 2020, the method may access a user data set via the data analysis environment.
[0145] According to one embodiment, in step 2030, the method can analyze at least a portion of the user data set by a data analysis environment.
[0146] According to one embodiment, in step 2040, the method can determine, by the data analysis environment, one or more data enrichment actions to be performed, the determination of the one or more data enrichment actions being based on an analysis of a portion of the user dataset.
[0147] According to one embodiment, in step 2050, the method may automatically perform one or more data enrichment actions determined by the analysis environment.
[0148] According to one embodiment, in step 2060, the method may create one or more data visualizations of the user dataset, and the one or more data enrichment actions determined and performed on the user dataset are reflected in the one or more data visualizations created.
[0149] According to various embodiments, the teachings herein may be implemented using one or more computers, computing devices, machines, or microprocessors, including one or more processors, memory, and / or computer-readable storage media, programmed according to the teachings herein. As will be apparent to those skilled in the software arts, appropriate software coding can be readily produced by skilled programmers based on the teachings of the present disclosure.
[0150] In some embodiments, the teachings herein may include a computer program product, which is a non-transitory computer-readable storage medium having stored thereon instructions that can be used to program a computer to perform any of the processes of the present teachings. Examples of such storage media may include, but are not limited to, a hard disk drive, a hard disk, a fixed disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, or any other type of storage medium or device suitable for non-transitory storage of instructions and / or data.
[0151] The foregoing description has been provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the scope of protection to the precise form disclosed. Further modifications and variations will be apparent to those skilled in the art.
[0152] These embodiments were chosen and described to best explain the principles of the teachings herein and their practical application, so that others skilled in the art can appreciate various embodiments, along with various modifications suited to the particular uses contemplated, the scope of which is intended to be defined by the following claims and their equivalents.
Claims
1. A system for data analysis that includes automated enrichment of datasets, A computer comprising a microprocessor and a data analysis environment provided on the microprocessor and capable of displaying data visualizations associated with a dataset, The user dataset specified by the user is accessed by the data analysis environment. The aforementioned data analysis environment is Perform an analysis on at least a portion of the user dataset, and determine one or more data enrichment actions to be performed on the user dataset based on the analysis of the user dataset. Automatically execute at least one of the determined data enrichment actions. Generate one or more data visualizations of the aforementioned user dataset, A system in which the one or more data enrichment actions decided upon and performed on the user dataset are reflected in the one or more generated data visualizations.
2. The system according to claim 1, wherein the user dataset includes one of a dataset triggered for uploading to the data analysis environment and an existing dataset in the data analysis environment.
3. The system according to claim 2, wherein at least one of the determined data enrichment actions performed includes performing an automatic calculation on at least one column of data in the dataset.
4. The system according to claim 2, wherein at least one of the determined data enrichment actions performed includes automatically performing a join between the user dataset and a column of data outside the user dataset.
5. The system according to claim 4, wherein the columns of data outside the dataset are identified based on at least the columns of data in the user dataset.
6. The system according to claim 5, wherein the columns of data outside the dataset are obtained from a source outside the data analysis environment.
7. The system according to claim 5, wherein the columns of data outside the dataset are obtained from sources inside the data analysis environment.
8. A method for automatically enriching datasets for use in data analysis, To provide a data analysis environment that can operate on a computer equipped with a microprocessor to display data visualizations associated with a dataset, The aforementioned data analysis environment accesses user datasets, The data analysis environment analyzes at least a portion of the user dataset, The data analysis environment determines one or more data enrichment actions to be performed. Includes, The determination of the one or more data enrichment actions is based on the analysis of the portion of the user dataset. The aforementioned method, The data analysis environment automatically executes the determined one or more data enrichment actions. To create one or more data visualizations of the aforementioned user dataset, It further includes, A method in which one or more data enrichment actions determined and performed on the user dataset are reflected in one or more data visualizations created.
9. The method according to claim 8, wherein the user dataset includes one of a dataset triggered for uploading to the data analysis environment and an existing dataset in the data analysis environment.
10. The method according to claim 9, wherein at least one of the determined data enrichment actions performed includes performing an automatic calculation on at least one column of data in the dataset.
11. The method according to claim 9, wherein at least one of the determined data enrichment actions performed includes automatically performing a join between the user dataset and a column of data outside the user dataset.
12. The method according to claim 11, wherein the columns of data outside the dataset are identified based on at least the columns of data in the user dataset.
13. The method according to claim 12, wherein the columns of data outside the dataset are obtained from a source outside the data analysis environment.
14. The method according to claim 12, wherein the columns of data outside the dataset are obtained from sources inside the data analysis environment.
15. A program that causes one or more computers to perform the method described in any one of claims 8 to 14.