Storage-independent semantic artifacts in cloud-based data warehousing environments

The integration of in-memory databases with lakehouse architectures using hyperscalers in data warehousing environments addresses the inefficiencies of traditional storage methods, offering flexible and cost-effective data management with real-time processing capabilities.

JP2026055099APending Publication Date: 2026-03-30エスアーペーエスエー
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-03-30

AI Technical Summary

Technical Problem

Data warehousing environments face challenges in managing large data volumes efficiently, particularly due to the high cost and difficulty of maintaining data in in-memory databases, while hyperscalers offer better scalability and cost-effectiveness, necessitating a storage-independent solution.

Method used

A data warehousing environment that integrates in-memory databases with lakehouse architectures, using hyperscalers for storage, enabling storage-independent data management through a single artifact that supports both in-memory and object storage, leveraging open data formats like Apache Parquet and Delta Lake for efficient data processing and management.

Benefits of technology

This integration provides flexible, scalable, and cost-effective data management, supporting both business intelligence and machine learning workloads with real-time data ingestion and processing, ensuring data quality and consistency across diverse data types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026055099000001_ABST
    Figure 2026055099000001_ABST
Patent Text Reader

Abstract

To provide storage-independent semantic artifacts in a cloud-based data warehousing environment. [Solution] In one exemplary embodiment, a data warehousing environment, or a similar architecture, is extended to enable data storage in either an in-memory database or a lakehouse architecture, leveraging one or more hyperscalers for the underlying storage. More specifically, a single artifact is defined in Datasphere, which stores data in either an in-memory database or a lakehouse architecture in a storage-independent manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Application No. 63 / 695,774, titled "DATA PLATFORM", filed on September 17, 2024, and U.S. Provisional Application No. 63 / 727,154, titled "EMBEDDED DATA LAKE CAPABILITIES", filed on December 2, 2024, the entire contents of both of which are incorporated herein by reference.

[0002] This disclosure relates to lakehouse environments, and more particularly, to storage - independent semantic artifacts in cloud - based lakehouse environments.

Background Art

[0003] A data warehousing environment is a specialized system designed to store, manage, and analyze large amounts of structured data, and in some cases semi - structured data. It serves as a central repository where data from various sources is integrated, transformed, and made available for querying and reporting. Data warehousing environments are often used to support business intelligence (BI), analytics, and decision - making processes within an organization. A data lake is a repository designed to store large amounts of data in its native format. A lakehouse environment is an architecture that combines a data warehouse and a data lake.

Summary of the Invention

Means for Solving the Problems

[0004] A system according to one aspect of the present invention includes at least one hardware processor and a non-temporary computer-readable medium that stores instructions causing the at least one hardware processor to perform an action when executed by the at least one hardware processor. The action includes receiving data from a data lake, the data lake storing the data in raw data storage format; deciding whether to store the data in an in-memory database or in a first form of object storage; creating a database artifact in a cloud database, the database artifact containing information about the data but stored in a format independent of whether the data is stored in an in-memory database or in a first form of object storage; generating a delta table containing the data in the first form of object storage based on the decision to store the data in the first form of object storage; generating a virtual table in the in-memory database determined by the delta table; and providing access to the data via the database artifact. [Brief explanation of the drawing]

[0005] [Figure 1] A block diagram showing a system according to one exemplary embodiment. [Figure 2] A block diagram showing an architecture including a data warehousing environment according to one exemplary embodiment. [Figure 3] This block diagram shows a provisioning workflow according to one exemplary embodiment. [Figure 4] This is a block diagram illustrating the relationships between various data structures in one exemplary embodiment. [Figure 5] This is a flowchart illustrating a method for storing data according to one exemplary embodiment. [Figure 6] This is a block diagram showing an example of a software architecture for a computing device. [Figure 7] This is a block diagram of a machine that is an exemplary form of a computer system in which instructions can be executed within the machine to cause the machine to perform one or more of the methods described herein. [Modes for carrying out the invention]

[0006] One example of a data warehousing environment is a cloud-based data management solution that helps organizations create a unified, intelligent data layer across their entire data landscape. It serves as the foundation for a "data fabric," which is a way to seamlessly connect, manage, and understand data from various sources, whether those sources are internal systems, third-party platforms, or on-premises databases.

[0007] One of the capabilities of a data warehousing environment is its ability to provide direct access to distributed data while presenting its original business meaning. Rather than simply moving raw data, it preserves critical elements such as business logic, semantics, and relationships. This is beneficial, in detail, for organizations that want to ensure consistency and governance across departments, enabling them to analyze data in context and act based on that data.

[0008] The platform supports both data replication and federated access, allowing companies to choose whether to move data or access it in real time from wherever it resides. The platform also features a user-friendly interface for data modeling and transformation, making it accessible to both technical users and business analysts. This helps bridge the gap between IT and business teams, making collaboration on data initiatives easier.

[0009] Security, governance, and regulatory compliance are built into the platform. Users can set controls on how data is accessed and shared, ensuring that sensitive information remains protected while further enabling secure collaboration with internal and external partners.

[0010] One technical issue that arises in data warehousing environments stems from where the data is stored. Data from various different sources is collected and stored in an in-memory database. An in-memory database (also known as an in-memory database management system) is a type of database management system that relies primarily on main memory for computer data storage. In-memory databases are contrasted with database management systems that use disk storage mechanisms. Because disk access is slower than memory access, in-memory databases are traditionally faster than disk storage databases. One exemplary in-memory database is the HANA® database provided by SAP SE in Waldorf, Germany. HANA also supports disk storage and may be adopted in data warehousing environments. As a result, a data warehousing environment has three storage layers: (1) HANA in-memory, (2) HANA disk, and (3) file storage (hyperscalar).

[0011] Storing data in an in-memory database results in extremely fast data access, but maintaining large amounts of data in an in-memory database can be difficult to manage and potentially costly. Other types of long-term storage offer better scalability and can be more cost-effective, for example, by using a hyperscaler. A hyperscaler is a large-scale cloud service provider that delivers highly scalable, distributed computing, storage, and networking infrastructure. These companies operate large data centers and deliver cloud-based services to businesses and individuals, allowing them to access computing resources on demand. Hyperscalers are known for their ability to dynamically scale resources, enabling users to handle workloads that change size efficiently and cost-effectively.

[0012] The term "hyperscaler" originates from the concept of "hyperscale computing," which refers to the ability to seamlessly scale computing infrastructure to meet the demands of large-scale applications such as big data analytics, artificial intelligence, machine learning, and enterprise-level workloads. Hyperscalers achieve this by leveraging advanced technologies, automation, and economies of scale to deliver reliable and flexible cloud services.

[0013] Examples of hyperscalers include Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform (GCP), and Alibaba Cloud. These providers offer a wide range of services, including virtual machines, object stores, databases, networking, analytics, and machine learning tools. Hyperscalers also provide global infrastructure, with data centers located in multiple regions, guaranteeing high availability and low latency for customers.

[0014] In one exemplary embodiment, a data warehousing environment, or a similar architecture, is extended to enable data storage in either an in-memory database or a lakehouse architecture, leveraging one or more hyperscalers for the underlying storage. More specifically, a single artifact is defined in a data warehousing environment that stores data in either an in-memory database or a lakehouse architecture, and does so in a storage-independent manner.

[0015] The Lakehouse architecture includes schema enforcement and evolutionary mechanisms, enabling users to define and modify data schemas as needed while maintaining data quality and consistency. The Lakehouse architecture supports both business intelligence and machine learning workloads, allowing organizations to perform a wide range of data analytics tasks on a single platform. The Lakehouse architecture decouples storage from computing resources, enabling independent scaling of each component, which provides flexibility in resource allocation and cost management.

[0016] Lakehouses often utilize open data formats such as Apache Parquet, Delta Lake, or Apache Iceberg to ensure compatibility with various data processing engines and tools. Lakehouses also support real-time data ingestion and processing, enabling timely insights and decision-making. By integrating the capabilities of data lakes and data warehouses, the Lakehouse architecture provides a comprehensive solution for managing large volumes of diverse data, supporting both batch and real-time processing, and enabling advanced analytics and machine learning applications.

[0017] OpenTables formats are data storage formats designed to facilitate efficient data management and processing in distributed environments, specifically in data lake and lakehouse architectures. These formats are characterized by their ability to support multi-engine read and write operations, allowing different data processing engines to access and manipulate data simultaneously. OpenTables formats are commonly used to store large datasets in object store systems, providing a flexible and scalable solution for data analysis.

[0018] SAP HANA Cloud, offered by SAP SE in Waldorf, Germany, is a cloud-based data platform that provides advanced data management and analytics capabilities. Designed to handle massive amounts of data in real time, SAP HANA Cloud offers features such as in-memory computing, data integration, and the processing of both transactional and analytical workloads.

[0019] Data lake files are digital files stored within a data lake, which is a centralized repository designed to store vast amounts of raw data in its native format. These files can include structured data such as databases and tables, semi-structured data such as JSON or XML files, and unstructured data such as text documents, images, and videos. Data lakes are designed to accommodate diverse data types and formats, enabling organizations to store data without needing to structure it beforehand. This flexibility allows data scientists and analysts to perform various types of data processing and analysis, including big data analytics, machine learning, and real-time data processing, directly on stored files that can be stored in their native format without additional structure.

[0020] The Hana Data Lake (HDL) files provided by SAP SE in Walldorf, Germany, provide a file and object store for structured, semi-structured, and unstructured data in HANA Cloud. The HDL files provide a single API that is independent of the infrastructure selection made by the customer during the provisioning of the HANA Cloud service. This system is available to customers as a feature set within the HANA Cloud data lake product.

[0021] The HDL files are implemented by SAP's Storage Gateway. The Storage Gateway is deployed in all HDL clusters of HANA Cloud. Tables can be represented and stored in a data lake, such as an HDL file, using a structured file format. The Open Table Format (OTF) is a category of structured file formats that represent tables and provide guarantees to users, such as ACID transactions. Assuming an increasing interest of HDL file users in OTF tables, the HDL files will be enhanced to provide the ability to manage OTF tables on top of their existing APIs for managing general objects.

[0022] An API may be introduced to enable users to compose OTF tables in the HDL file catalog. The API also provides a client library for integration with Apache Spark. Apache Spark is an open-source distributed computing system designed to efficiently process and analyze large-scale data. Apache Spark provides a unified framework for processing a wide range of data processing tasks, including batch processing, real-time streaming, machine learning, graph processing, and interactive querying.

[0023] Fundamentally, Spark is built on a distributed computing architecture that enables processing data across multiple nodes within a cluster. Spark achieves high performance by leveraging in-memory computing, which minimizes the need to read and write data to disk during processing. This approach significantly speeds up data processing tasks compared to traditional disk-based systems like Hadoop MapReduce.

[0024] In the context of a data warehousing environment, Spark is used as a compute layer for advanced data processing and transformation tasks. Spark is seamlessly integrated with HANA Data Lake files (HDL files) to enable complex workflows such as data transformation, aggregation, and machine learning on large datasets stored in an object store. Spark's ability to process both structured and unstructured data makes it a versatile tool for modern data management and analysis.

[0025] Furthermore, HDL files are enhanced to support Delta Sharing, an open protocol for securely sharing Delta tables. At the core of Delta Sharing is the idea that it allows a party, the data provider, to share access to specific datasets stored in cloud storage without actually moving or replicating the data. This is done by a Delta Sharing Server, which acts as a secure gateway, managing who can access which data and ensuring that all permissions and authentication are enforced.

[0026] When another organization, i.e., the recipient, wants to access this data, they can use a wide range of tools to read it directly from the cloud. The data is accessed in place using secure, time-limited URLs, eliminating the need to download or copy files. This setup keeps the data fresh and consistent, reduces data transfer costs, and ensures a high level of control over what is being shared.

[0027] Delta Sharing includes creating and managing Delta shares. A Delta share is a collection of objects, such as tables and table partitions, that will be shared with further recipients in different organizations.

[0028] In exemplary embodiments, an HDL file, or similar functionality, is integrated with a data warehousing environment, or similar functionality that provides a mechanism for users to choose whether to store data in an in-memory database or in a hyperscaler. This is all done in a seamless manner without requiring any changes to user-facing processes, despite the introduction of storage-independent storage artifacts such as tables.

[0029] Another aspect of the Lakehouse architecture is a large, scalable, and cost-effective object store layer that serves as inband history storage and fine-grained storage. It also provides, or enables, a wide variety of processes to run on top of the data to transform it into higher-quality datasets.

[0030] Figure 1 is a block diagram of a system 100 according to an exemplary embodiment. The data warehousing environment 102 includes a data warehousing environment user interface 104 and a data warehousing environment backend service 106. The data warehousing environment user interface 104 includes a data builder 108. The data builder is designed to help users model, transform, and prepare data for analysis and reporting. The data builder provides an intuitive interface that enables users to create and manage data models, define relationships between datasets, and perform data transformations. By providing both graphical and script-based methods, the data builder 108 bridges the gap between technical users, such as data engineers, and business users, such as analysts.

[0031] The tool allows users to design data models that define how data is structured and related, and supports schemas optimized for analysis. The tool also enables the transformation of raw data into meaningful datasets by applying business logic, performing calculations, and cleaning or enriching the data. These transformations can be performed using either graphical tools or scripting languages ​​such as Structured Query Language (SQL) or Python scripts, making the tool accessible to a wide range of users.

[0032] Data Builder 108 seamlessly integrates with a variety of data sources, including cloud platforms and on-premises databases. It supports the creation of virtual tables and views, enabling users to access and work with data without physically moving it. Furthermore, Data Builder 108 works with both in-memory cloud databases 110, such as HANA, and object stores 112, such as HANA Data Lake files, providing flexibility and cost-effectiveness in data management.

[0033] For datasets that support change data capture, Data Builder 108 incorporates incremental updates, allowing users to track and manage changes to data over time. This feature is particularly beneficial for maintaining up-to-date analytics in dynamic environments. The tool also supports collaborative workflows, enabling multiple users to work on data models and transformations while ensuring consistency and reusability across projects.

[0034] The data builder 108 may include a table editor 113, a tool that allows users to create, modify, and manage tables as part of their data modeling and transformation workflow. It provides an interface for defining the structure and properties of tables, allowing users to configure how data is stored, accessed, and processed.

[0035] Table Editor 113 allows users to define the schema of a table, including its columns, data types, and primary key. Table Editor 113 also enables the specification of additional table properties, such as partitioning configurations, which can optimize query performance and data organization. The tool also supports creating tables that leverage different storage types, such as in-memory storage or object stores, providing flexibility in how data is managed.

[0036] Furthermore, the table editor 113 provides options for managing table lifecycle operations, such as expanding, modifying, or deleting tables. The table editor 113 ensures that changes to the table structure are implemented in a controlled and consistent manner, even when data is already stored in the table. For example, the table editor 113 allows users to add new columns to tables with existing data, while imposing restrictions on actions that could compromise data integrity, such as deleting primary key columns.

[0037] The data warehousing environment user interface 104 further includes a data integration monitor 114. The data integration monitor 114 is a tool designed to provide visibility and control over the data integration process. The data integration monitor 114 enables users to monitor, manage, and troubleshoot data flows and transformations as they occur within the platform. By providing real-time insights into the status and performance of data integration activities, the data integration monitor 114 ensures that users can maintain the reliability and efficiency of their data pipelines.

[0038] This tool allows users to track the progress of data replication, transformation, and ingestion workflows. It provides detailed information about the execution of these processes, including their current status, estimated completion time, and any errors or warnings encountered. This visibility helps users quickly identify and resolve issues, minimizing disruptions to data operations.

[0039] The data integration monitor also supports the management of scheduled and on-demand data integration tasks. Users can view and control the execution of these tasks, including starting, stopping, or rescheduling them as needed. This flexibility ensures that the data integration process aligns with organizational requirements and priorities.

[0040] In addition to monitoring individual tasks, the data integration monitor 114 provides an aggregated view of data integration activity across the entire platform. This allows users to analyze trends, identify bottlenecks, and optimize the performance of their data workflows. The tool also integrates with other components of the data warehousing environment 102, enabling users to delve into specific data flows or transformations for further analysis.

[0041] The data integration monitor 114 includes a table monitor 116. The table monitor 116 is a specialized tool that provides users with detailed insights into the status and performance of operations related to tables. The table monitor 116 is designed to help users track, manage, and troubleshoot table-related activities, such as data ingestion, replication, and transformation, ensuring that table operations are performed efficiently and reliably.

[0042] The table monitor 116 allows users to view the current state of tables, including their data ingestion status, update frequency, and any associated transformation workflows. The table monitor 116 provides real-time information about the progress of data being written to or read from tables, enabling users to monitor the flow of data to and from the system. This is particularly useful for ensuring that the data pipeline is functioning as expected and that tables are populated with accurate and up-to-date information.

[0043] In addition to tracking data ingestion, the table monitor 116 provides visibility into table-specific behavior, such as merge tasks, optimization processes, and change data capture (CDC) activity. For example, the table monitor 116 can display the status of a merge task that integrates data from an inbound buffer into a target table, or show the results of an optimization task that improves query performance by compressing small files or reorganizing data partitions.

[0044] The tool also highlights any errors or warnings that occur during table operations, enabling users to quickly identify and address problems. For example, if a data ingestion task fails due to schema mismatch or connectivity issues, Table Monitor 116 provides diagnostic information to help users resolve the problem. This ensures that workflows involving tables remain consistent and reliable. Table Monitor 116 integrates seamlessly with other components of the Data Integration Monitor, allowing users to delve into specific table operations or link them to broader data integration workflows. Table Monitor 116 also supports lifecycle management tasks, such as monitoring new table deployments, tracking schema changes, or managing table deletions. By providing a centralized diagram of table activity, Table Monitor 116 helps users maintain control over their data assets and ensures that tables function as intended within the overall data landscape.

[0045] Next, referring to the data warehousing environment backend service 106, this component includes deployer middleware 118, which is responsible for coordinating the deployment and management of data models, transformations, and other artifacts within the data warehousing environment 102. The data warehousing environment backend service 106 acts as a middle tier to ensure the seamless execution of deployment tasks and maintains consistency and reliability throughout the platform's data management and integration processes.

[0046] The Deployer Middleware 118 facilitates the deployment of data models such as tables, views, and transformation flows by translating higher-level design specifications into executable instructions for the underlying infrastructure. The Deployer Middleware 118 ensures these deployments are controlled and performed in a consistent manner, adhering to defined configurations and dependencies. This includes managing the creation, modification, and deletion of database artifacts such as virtual tables, Delta tables, and associated metadata.

[0047] The deployer middleware 118 is responsible for transactional deployment operations. The deployer middleware 118 ensures that deployment tasks are executed atomically, meaning that all changes are either successfully applied or none are applied in case of errors. This guarantees the integrity of the deployed artifacts and prevents incomplete or inconsistent states. For example, when deploying local tables (files), this means storing the data on the hyperscaler along with the Change Data Capture (CDC) function, and the deployer middleware 118 coordinates the creation of both the virtual tables in the in-memory cloud database 110 and the delta tables in the object store 112, ensuring that all components are properly aligned.

[0048] The deployer middleware 118 also manages dependencies between different artifacts, ensuring that the deployment sequence is executed in the correct order. For example, the deployer middleware 118 ensures that after a Delta table is created in the object store 112, its corresponding virtual table is deployed to the in-memory cloud database 110. This dependency management helps maintain the logical consistency of the data landscape.

[0049] In addition to deployment, the deployer middleware 118 supports lifecycle management tasks, such as updating existing artifacts to reflect schema changes or reconfigurations. The deployer middleware 118 enforces rules to ensure changes are applied safely, such as restricting actions that could compromise data integrity, like deleting primary key columns. The deployer middleware 118 also handles cleanup actions, such as rolling back changes or removing artifacts during undeployment, to maintain a clean and consistent environment.

[0050] The data warehousing environment backend service 106 also includes a local table monitor backend 120 that provides backend functionality to the table monitor 116.

[0051] The in-memory cloud database 110 is where database artifacts are stored, but the underlying data related to the database artifacts can be stored in either the in-memory cloud database 110 or the object store 112. More specifically, the Open Table Format Structured Query Language Application Programming Interface (OTF SQL API) 124 stores data definition language (DDL) objects and procedures 126 to manage one or more virtual tables 128. The virtual tables 128 represent database artifacts for data stored in the object store 112. The virtual tables 128 are stubs that provide federated data access to data stored remotely, such as in an external system 142, via the in-memory database 110.

[0052] The file adapter 130 connects the virtual table 128 to the object store 112. The object store 112 may contain a container 132 that includes the file 134 and the delta table 136.

[0053] File 134 and Delta Table 136 work together to provide a robust and scalable storage solution for managing large amounts of data. File 134 serves as the foundational storage unit, capable of holding structured, semi-structured, and unstructured data in formats such as Apache Parquet. These formats are optimized for analysis and enable efficient querying and processing of data, even at large scales. File 134 is organized within Container 132, which acts as a logical grouping of related data assets, and File 134 can be accessed via API or integrated with other components of the data warehousing environment 102.

[0054] On the other hand, Delta Table 136 is built on top of the Delta Lake storage framework and provides advanced features such as ACID transactions, schema enforcement, time travel, and change data capture (CDC). These tables are stored as a collection of files, such as Parquet files, along with a transaction log that tracks all changes made to the tables. The transaction log ensures that changes, including insertions, updates, and deletions, are applied atomically and consistently. This enables features such as querying tables as they existed at a specific point in time and tracking gradual changes to the data.

[0055] File 134 and Delta Table 136 work together seamlessly. When data is written to Delta Table 136, the data is stored in Container 132 as a Parquet file, and the transaction log is updated to reflect these changes. Delta Table 136 provides a structured interface for querying the data stored in the underlying File 134, ensuring that queries return results that match the latest state of the table. Data ingestion into Object Store 112 can occur through various mechanisms such as replication flows, transformation flows, or direct API calls, and the ingested data is written as a new Parquet file and recorded in the transaction log.

[0056] Delta Table 136 also supports change data capture by recording changes at the record level in the transaction log. This allows users to efficiently extract incremental updates and synchronize data across the system. Over time, maintenance tasks such as vacuuming (deleting older versions of data) and compression (merging smaller files into larger ones) are performed to optimize storage and query performance. These tasks ensure that Delta Table 136 remains efficient and scalable.

[0057] The integration of Delta Table 136 with the data warehousing environment 102 allows users to create, manage, and query these tables through the data warehousing environment interface. The virtual table 128 in the in-memory cloud database 110 provides seamless access to Delta Table 136, enabling users to leverage in-memory computing for analysis while cost-effectively storing data in the object store 112. This integration ensures that users can benefit from the scalability of the object store 112 while maintaining Delta Lake's advanced capabilities for structured data management.

[0058] The Spark Engine 138 may be accessed via a Spark Adapter 140 on the in-memory cloud database 110, or may be managed by the in-memory cloud database 110. Lifecycle management and data processing of local table files may be handled by the Spark Engine 138.

[0059] In one exemplary embodiment, maintenance may be performed on the delta table 136 from time to time. Initially, the delta table retains historical versions of data for ACID transactions, for example, when it is necessary to perform a rollback or other "time travel" operation on the data. Over time, this can result in a large amount of old data being stored. Therefore, periodic vacuum commands may be executed to remove old versions of data and free up storage space.

[0060] Furthermore, if you don't need to query older versions, it's best to use the vacuum command to clean them up and remove them.

[0061] Furthermore, delta tables can store data from many small files, especially after multiple merge or update operations. This can cause performance issues because each file needs to be read or written during querying. The optimize command may be used to compress small files into larger ones, which improves file read / write performance by reducing file fragmentation.

[0062] Furthermore, if a delta table is partitioned, partition management may become necessary over time, especially if the partitions become large or there are too many partitions.

[0063] The data warehousing environment 102 uses a concept called space for secure isolation of data and for isolation of workloads. Space can be configured as in-memory database space or file space. In-memory database space is stored and computed in the in-memory cloud database 110, while file space is stored in the object store 112, and computations on them are performed by the Spark engine 138.

[0064] Provisioning each file space includes provisioning both object store 112 instances and spark engine 138 instances. Each local table (file) includes a virtual table in the in-memory cloud database 110 and a delta table in object store 112.

[0065] The deployment of local tables (files) involves creating and managing the necessary database artifacts to enable the functionality of the tables within the system. This process ensures that the tables are properly configured to store, access, and process data, regardless of whether the data is stored in object store 112 or in in-memory database 110. The deployment process is designed to maintain consistency, support advanced features such as Change Data Capture (CDC), and integrate seamlessly with broader systems.

[0066] When a local table (file) is expanded, the system creates two main entities: an entity for active records (if the "Change Data Feed" feature is enabled) and another entity for handling delta records. For tables with "Change Data Feed," the expansion process involves creating a virtual table in the cloud database to access active records stored in the delta table in the object store. Additionally, a "Delta History" table is created in the cloud database to temporarily store the change data captured from the delta table. This change data may be retrieved from an underlying data source (e.g., a data lake). A dedicated procedure is also expanded to populate the delta history table with the change data, and an extract view is created to query the delta history table based on parameters such as timestamps and subscriber IDs. For local tables (files) without "Change Data Feed," only a virtual table is expanded.

[0067] The deployment process begins with creating a Delta Table in the object store. This table is configured to store data in an open format such as Apache Parquet and is managed using the Delta Lake framework. Delta Tables support advanced features essential for ensuring data consistency and reliability, such as ACID transactions, schema enforcement, and time travel. The Delta Table is created with the necessary schema, including columns, data types, and partitioning configurations where applicable.

[0068] Once the Delta Table is created, the system then creates a virtual table in the cloud database. The virtual table serves as a logical representation of the Delta Table, allowing users to query and interact with the data stored in the object store using standard SQL commands. The virtual table is created in the corresponding schema of the cloud tenant database and is linked to the Delta Table through a remote source connection. This connection enables seamless access to the data stored in the object store while simultaneously leveraging HANA's in-memory computing capabilities for analysis and processing.

[0069] If the local table (file) contains a "Change Data Feed," additional artifacts are deployed to support change data capture. A "Delta History" table is created in the cloud database to store the results of the change data feed procedure. This table contains all fields from the local table (file), along with metadata columns for change type, commit version, and commit timestamp. A procedure for populating the Delta History table is also deployed, allowing the system to extract incremental changes from the Delta table and store them in the Delta History table. Extract views are created to provide a user-friendly interface for querying the Delta History table based on specific parameters.

[0070] Modeling using local tables (files) involves defining and managing the structure, relationships, and behavior of tables that utilize the object store as the underlying storage for those tables. This process seamlessly integrates into broader modeling frameworks within data warehousing environments, enabling users to create and manage local tables (files) while effectively leveraging the scalability and cost-effectiveness of the object store. The modeling process ensures that local tables (files) are optimized for analysis, support advanced features such as CDC, and integrate seamlessly with other components of the platform.

[0071] The modeling process begins with creating a local table (file) in the data builder. The user defines the table schema, including columns, data types, and primary keys, to appropriately structure the data according to their use case. The schema definition also includes additional properties, such as partitioning configurations, which can improve query performance and data organization. Partitioning is particularly useful for large datasets because it enables efficient data retrieval by segmenting data based on specific columns.

[0072] The local table (file) is modeled to support two main entities: Active Record and Delta Capture. The Active Record entity represents the current state of data stored in the Delta Table within the object store. This entity is linked to a virtual table in the cloud database, which provides a logical interface for querying the data using SQL. The Delta Capture entity, on the other hand, is used to track and manage changes to the data over time. This entity leverages the Delta Table's "Change Data Feed" feature, recording metadata such as change type, commit version, and commit timestamp. These two entities are defined during the modeling process to ensure that the local table (file) supports both real-time and historical data analysis.

[0073] The CSN of a local table (file) may have the following annotations to define the delta behavior and associated Active Records View: Active entity:

[0074]

number

[0075] Delta entities (for change data capture functionality):

[0076]

number

[0077] Additional annotations may be used to represent delta views that require parameters in delta extraction.

[0078]

number

[0079] During deployment, virtual tables are created in the corresponding schema of the cloud tenant database. After the creation of the database artifacts, the tables can be populated with data by performing replication flows and MERGE tasks.

[0080] As long as no data is stored in the virtual table, arbitrary changes can be made to the table model. If the virtual table is empty, the deployment will make changes using the DROP / CREATE command sequence.

[0081] When data is found in a virtual table, columns may be added. It is not possible to delete columns, change the column representing the primary key of the data, or change the data type, including Null / NotNull, default value specification, and length reduction.

[0082] Columns are also added when the local table (file) has data.

[0083] Cloud databases may provide dedicated DDL commands for creating local tables (files) where an HDL file remote source (or file adapter) is referred to as a Spark remote source. The following is an example of how a local table (file) may be generated in a cloud database.

[0084]

number

[0085] If delta changes need to be extracted into an SQL table, two procedures may be used to gather information about the change data feed and how to retrieve the change data. DELTA_LAKE_TABLE_VERSION: Returns the latest remote delta table version. DELTA_LAKE_TABLE_CHANGES: Stores the minimum and maximum range of CDC records in the result table.

[0086]

number

[0087] The following database artifacts are expanded to process the delta capture records. 1. The "Delta History" table ($CDC table) is used as "temporary" storage for the results of the procedure "DELTA_LAKE_TABLE_CHANGES". The table contains a list of fields present in the local table (file), plus three additional fields related to the Change Data Feed (_change_type, _commit-version, _commit_timestamp). example:

[0088]

number

[0089] 2. An SQL procedure is deployed to fetch the change data feed to the "Delta History Table". The SQL procedure is called when a delta consumer (such as HANA Transformation Flow) requests the next delta record.

[0090] 3. SQL views return either active records or delta records, based on the parameter (FULL, DELTA_AI) that defines the extraction mode. In active records, the view points to a virtual table, and in delta records, it points to the Delta log table ($CDC).

[0091] This view is deployed as an SQL view having the following characteristics:

[0092] Delta type "UPSERT" (same primary key as Active Record)

[0093] Two additional CDC columns: modeElement (change type) and datetimeElement (change time).

[0094] Three parameters: Extraction Mode (FULL, DELTA_AI), From date, and Till date (timestamp).

[0095] Furthermore, the modeling process also involves integrating local tables (files) with other components of the data warehousing environment, such as replication flows and transformation flows. Users can define local tables (files) as target tables in replication flows, allowing data to be directly ingested into the table from various sources. Similarly, local tables (files) can be used as source or target tables in transformation flows, enabling users to apply business logic, perform calculations, and transform data as part of their workflow.

[0096] Annotations are used to define the behavior and properties of local tables (files) and to represent the design time of local tables (files). For example, annotations specify whether a table supports delta capture and identify the relevant active records and delta capture entities. These annotations ensure that the table's behavior matches the requirements of the use case and that the table's behavior integrates seamlessly with the broader data warehousing environment ecosystem.

[0097] Once a local table (file) is modeled, it is deployed to the system. During deployment, the Delta Table is created in HDLF, and the virtual table is created in SAP HANA. If the table includes delta capture, additional artifacts such as delta history tables, change data feed procedures, and extract views are also deployed. These artifacts enable the table to support advanced features such as incremental updates and historical data analysis.

[0098] Modeling using local tables (files) also supports lifecycle management tasks such as schema evolution and partitioning. For example, users can add new columns to a table schema or define partitioning configurations to optimize data storage and retrieval. However, some changes, such as deleting columns or changing primary keys, are restricted to maintain data integrity.

[0099] The deployment process ensures transactional consistency by using a saga pattern that encapsulates individual actions. This means that if any step of the deployment process fails, the system can roll back the changes to maintain a consistent state. For example, if the creation of a virtual table in SAP HANA fails, the system can clean up the Delta Table in HDLF to ensure that no orphaned artifacts remain.

[0100] Once deployment is complete, the local table (file) becomes available for use. Data may be ingested into the table using replication flows, transformation flows, or direct API calls. The system supports lifecycle management operations, such as modifying table schemas or deleting tables, and ensures that these changes are applied in a controlled and consistent manner. Maintenance tasks, such as vacuuming and optimization, may also be performed to manage storage efficiency and query performance.

[0101] The Local Table Monitor Backend 120 is a backend service component responsible for managing and monitoring operations related to local tables (files), specifically operations stored in object stores or cloud databases. The Local Table Monitor Backend 120 provides the necessary infrastructure for tracking, executing, and managing tasks related to local tables (files), ensuring their efficient operation and integration within a broader data warehousing environment.

[0102] This backend service supports monitoring key activities related to local tables (files), including data ingestion, transformation, and lifecycle management. The backend service interacts with the local table monitor in the user interface to provide real-time updates on the status of table-related operations. For example, it tracks the progress of data ingestion tasks, merge operations, and optimization processes, ensuring users have visibility into the state of their tables and can address any issues that arise.

[0103] The local table monitor backend 120 also plays a role in managing change data capture (CDC) of local tables (files). The local table monitor backend 120 ensures that incremental updates to the data are accurately tracked and recorded, enabling users to maintain up-to-date data analysis and synchronize data across the entire system. This includes managing CDC-related metadata, such as timestamps and change types, and ensuring this information is accessible to downstream processes.

[0104] In addition to monitoring, the backend service supports lifecycle management tasks for local tables (files). This includes handling operations such as table creation, modification, and deletion. The backend service ensures that these tasks are performed in a controlled and consistent manner, adhering to rules and constraints defined for local tables (files). For example, the backend service imposes restrictions on schema changes to tables with existing data to maintain data integrity.

[0105] The backend service also facilitates the execution of maintenance tasks such as vacuuming and optimization, which are essential for managing the performance and storage efficiency of local tables (files). These tasks are triggered and managed through the local table monitor backend, ensuring that they are executed reliably and without interrupting other operations.

[0106] Through integration with other components of the data warehousing environment backend service, such as the deployer middleware 118 and the data integration monitor 114, the local table monitor backend 120 ensures that local tables (files) are seamlessly integrated into the overall data workflow. The local table monitor backend 120 provides the necessary backend support to enable users to efficiently monitor and manage their tables, whether they are stored in a cloud database or an object store.

[0107] Figure 2 is a block diagram showing an architecture 200 including a data warehousing environment 202 according to an exemplary embodiment. Here, the HDLF 204 expands the storage capacity of the data warehousing environment 202, allowing data to be stored in delta tables 206 and files 208 of the object store 210. The compute tier 212 provides data transformation and processing using the Spark Engine 214. The in-memory query compute engine 216 provides file persistence. Data from an external source 218 may then be ingested via an inbound replication flow 220 and persisted in the object store 210, or directly in the HANA cloud and elastic compute nodes (ECN) 222.

[0108] HANA Cloud & ECN222 may include model 224, view 226 stored in cache 228, transfer flow 230, and table 232.

[0109] The data warehousing environment 202 further includes outbound replication flows 234, Information Access (INA) / Open Data Protocol (ODATA) 236, and SQL 238. Furthermore, data storage is defined in Spaces & Catalogs 240, and tasks are defined in Task Chains 242.

[0110] Figure 3 is a block diagram showing a provisioning workflow 300 according to an exemplary embodiment. More specifically, a flexible tenant configuration 304 may be defined in the data warehousing environment 302, which may be used to update tenant licenses 306. The space management component 308 can call the file container provisioning service 310 to provision file containers. In doing so, the file container provisioning service 310 can check tenant licenses 306 based on tenant metadata 312. The space management component 308 can create or update service instances of the Spark provisioning service 314, which in turn can manage the Cloud Platform Management Service 316. The Cloud Platform Management Service 316 can create, update, or delete services on the HANA Cloud 318.

[0111] Figure 4 is a block diagram illustrating the relationships between various data structures in an exemplary embodiment. Here, a local table 400 with delta capture for a design-time entity is based on repository objects such as an inbound data buffer 402, an active record 404, and a delta capture (API) 406. The active record 404 relates to a virtual table 408 of a HANA database artifact, and the delta capture (API) 406 relates to a delta procedure 410 of a HANA database artifact. Both the virtual table 408 and the delta procedure 410 relate to a delta table with a change data feed 412 of an HDLF artifact. The delta table with the change data feed 412 and the inbound data buffer 402 relates to a file store (path) 414.

[0112] Figure 5 is a flowchart illustrating a method 500 for storing data according to an exemplary embodiment. The flowchart described in the figure illustrates a method for performing a series of operations related to data management in a cloud-based data warehousing environment. Although the operations in the flowchart are presented in order, those skilled in the art will understand that some or all of the operations may be performed in a different order, combined or omitted, or performed in parallel.

[0113] In operation 502, method 500 begins with receiving data from an external system, such as a data lake. The data lake stores data in a raw data storage scheme, which may include structured, semi-structured, or unstructured data. In some examples, data may be ingested from an external system, such as an enterprise resource planning (ERP) system, a customer relationship management (CRM) system, or another third-party platform. The data is extracted using an API or replication flow that writes the data to an inbound buffer associated with the data lake. The data may be received from an external system that either ingests data or "pushes" the data.

[0114] In operation 504, the system decides whether to store the data in an in-memory database or in data-first format object storage. This decision may be based on predefined rules, user settings, or system parameters. In some examples, the decision may be determined by factors such as the frequency of data access, the size of the dataset, or the cost-effectiveness of the storage option. The system evaluates these criteria and selects the appropriate storage type. In some embodiments, the first format object storage may be implemented as open data format object storage.

[0115] In operation 506, a database artifact is created in the cloud database. The database artifact contains information about the data, but is stored in a format that does not depend on whether the data is stored in an in-memory database or in a first form of object storage. In some examples, the database artifact may include metadata, schema definitions, and references to the underlying storage location. The artifact serves as a logical representation of the data, allowing users to query and interact with the data without being constrained by the physical storage type.

[0116] Assuming that in operation 508 the system decides that data will be stored in a first form of object storage, the system generates a Delta Table containing the data in the first form of object storage. The Delta Table is configured to store data in a format such as Apache Parquet and is managed using the Delta Lake framework. In some examples, the Delta Table supports features such as ACID transactions, schema enforcement, and time travel. The system creates the Delta Table with the necessary schema, including columns, data types, and partitioning configuration.

[0117] In operation 510, a virtual table is created in the in-memory database. The virtual table is defined by the Delta Table and serves as a logical representation of the data stored in object storage. In some examples, the virtual table is created in the corresponding schema of the cloud tenant database and linked to the Delta Table through a remote source connection. This connection enables seamless access to the data stored in object storage while simultaneously leveraging the database's in-memory computing capabilities for analysis and processing.

[0118] In operation 512, access to the data is provided through a database artifact. The database artifact allows the user to query and interact with the data using standard SQL commands. In some examples, the artifact may include a view, procedure, or API that extracts details of the underlying storage, allowing the user to focus on analytical tasks.

[0119] In view of the implementations of the subject matter described above, this application discloses the following list of examples, where one feature of an isolated example, or two or more features of an example, which are to be interpreted in combination, or in combination with one or more features of one or more further examples, are also further examples that fall within the scope of this application.

[0120] Example 1 is a system comprising at least one hardware processor and a non-temporary computer-readable medium that, when executed by at least one hardware processor, stores instructions causing at least one hardware processor to perform an action, wherein the action includes receiving data from a data lake, where the data lake stores the data in a raw data storage manner; deciding whether to store the data in an in-memory database or in a first form of object storage; creating a database artifact in a cloud database, where the database artifact contains information about the data but is stored in a form independent of whether the data is stored in an in-memory database or in a first form of object storage; generating a delta table containing the data in the first form of object storage based on the decision to store the data in the first form of object storage; generating a virtual table in the in-memory database determined by the delta table; and providing access to the data via the database artifact.

[0121] In Example 2, the subject of Example 1 further includes the operation of modifying a delta table based on changes made to the data, capturing change data about the changes from the delta table, and storing the change data in a cloud database as a delta history table.

[0122] In Example 3, the subject of Example 2 is that the first form of object storage is open data form object storage, and changes made to the data are received from the data lake.

[0123] In Example 4, the subject matter from Examples 2 and 3 is expanded to include the delta history table containing all fields of the delta table, as well as metadata columns for change type and timestamp.

[0124] In Example 5, the subject matter of Examples 1-4 is further expanded to include executing optimization commands on the delta table to combine multiple data points within the delta table into a single data point.

[0125] In Example 6, the subject matter of Examples 1-5 is further expanded to include the operation of executing a vacuum command to remove old data from the delta table.

[0126] In Example 7, the subject matter of Examples 1 through 6 is expanded to include the fact that a virtual table is a logical representation of a delta table, allowing users to query and interact with data stored in the delta table using Structured Query Language (SQL) commands.

[0127] Example 8 is a method, the method comprising the steps of: receiving data from a data lake, wherein the data lake stores the data in a raw data storage manner; determining whether to store the data in an in-memory database or in a first form of object storage; creating a database artifact in a cloud database, wherein the database artifact contains information about the data but is stored in a manner independent of whether the data is stored in an in-memory database or in a first form of object storage; generating a delta table containing the data in the first form of object storage based on the decision to store the data in the first form of object storage; generating a virtual table in the in-memory database determined by the delta table; and providing access to the data via the database artifact.

[0128] In Example 9, the subject of Example 8 includes the steps of modifying a delta table based on changes made to the data, capturing change data about the changes from the delta table, and storing the change data in a cloud database as a delta history table.

[0129] In Example 10, the subject of Example 9 is that the first form of object storage is open data form object storage, and changes made to the data are received from the data lake.

[0130] In Example 11, the subject matter of Examples 9 and 10 is expanded to include the delta history table containing all fields of the delta table, as well as metadata columns for change type and timestamp.

[0131] In Example 12, the themes from Examples 8 through 11 are expanded to include the step of executing optimization commands on the delta table to combine multiple data points within the delta table into a single data point.

[0132] In Example 13, the subject matter from Examples 8 through 12 is expanded to include the step of executing a vacuum command to remove old data from a delta table.

[0133] In Example 14, the subject matter of Examples 8 through 13 is expanded to include the idea that a virtual table is a logical representation of a delta table, allowing a user to query and interact with data stored in the delta table using Structured Query Language (SQL) commands.

[0134] Example 15 is a non-temporary machine-readable medium that, when executed by one or more processors, stores instructions causing one or more processors to perform an action, wherein the action includes receiving data from a data lake, where the data lake stores the data in a raw data storage manner; deciding whether to store the data in an in-memory database or in open data-format object storage; creating a database artifact in a cloud database, where the database artifact contains information about the data but is stored in a format independent of whether the data is stored in an in-memory database or in open data-format object storage; generating a delta table containing the data in open data-format object storage based on the decision to store the data in open data-format object storage; generating a virtual table in the in-memory database determined by the delta table; and providing access to the data via the database artifact.

[0135] In Example 16, the subject of Example 15 further includes the operation of modifying a delta table based on changes made to the data, capturing change data about the changes from the delta table, and storing the change data in a cloud database as a delta history table.

[0136] In Example 17, the subject of Example 16 is that the first form of object storage is open data form object storage, and changes made to the data are received from the data lake.

[0137] In Example 18, the subject of Example 17 is expanded to include a delta history table containing all fields of the delta table, as well as metadata columns for change type and timestamp.

[0138] In Example 19, the subject matter of Examples 15-18 is further expanded to include executing optimization commands on the delta table to combine multiple data points within the delta table into a single data point.

[0139] In Example 20, the subject matter of Examples 15-19 is further expanded to include the operation of executing a vacuum command to remove old data from the delta table.

[0140] Example 21 is at least one machine-readable medium that, when executed by a processing circuit, contains instructions causing the processing circuit to perform an action that performs one of Examples 1 to 20.

[0141] Example 22 is a device equipped with means for implementing any of Examples 1 to 20.

[0142] Example 23 is a system that implements one of Examples 1 through 20.

[0143] Example 24 is a way to implement any of Examples 1 through 20.

[0144] Figure 6 shows a block diagram 600 illustrating an example of a software architecture 602 for a computing device. The software architecture 602 may be used in conjunction with various hardware architectures, as described herein, for example. Figure 6 is merely a non-limiting example of a software architecture, and many other architectures may be implemented to facilitate the functions described herein. A typical hardware layer 604 is shown and can represent, for example, any of the referenced computing devices described above. In some examples, the hardware layer 604 may be implemented according to the architecture of the computer system 6 in Figure 6.

[0145] A typical hardware layer 604 includes one or more processing units 606 having associated executable instructions 608. The executable instructions 608 represent executable instructions of the software architecture 602, including implementations of methods, modules, subsystems, and components described herein, and may also include memory and / or storage modules 610, which also have executable instructions 608. The hardware layer 604 may also include other hardware 612 representing any other hardware of the hardware layer 604. Examples of other hardware 612 include the hardware components shown in Figure 7.

[0146] In the exemplary architecture of Figure 6, the software architecture 602 can be conceptualized as a stack of layers, each providing a specific function. For example, the software architecture 602 may include layers such as an operating system 614, libraries 616, framework / middleware 618, applications 620, and a presentation layer 644. Operationally, applications 620 and / or other components within the layer may invoke API calls 624 through the software stack and access responses, return values, etc., illustrated as messages 626 based on the API calls 624. The layers illustrated are actually representative, and not all software architectures have all of them. For example, some mobile or dedicated operating systems may not provide a framework / middleware 618 layer, while others may. Other software architectures may include additional or different layers.

[0147] The operating system 614 may manage hardware resources and provide common services. The operating system 614 may include, for example, a kernel 628, a service 630, and a driver 632. The kernel 628 may act as an abstraction layer between the hardware layer and other software layers. For example, the kernel 628 may be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, etc. The service 630 may provide other common services to other software layers. In some examples, the service 630 includes an interrupt service. When an interrupt is accessed, the interrupt service may detect the receipt of the interrupt and, accordingly, cause the software architecture 602 to pause its current processing and execute an interrupt service routine (ISR).

[0148] Driver 632 may be responsible for controlling or interacting with the underlying hardware. For example, driver 632 may include, depending on the hardware configuration, a display driver, camera driver, Bluetooth® driver, flash memory driver, serial communication driver (e.g., Universal Serial Bus (USB) driver), Wi-Fi® driver, NFC driver, audio driver, power management driver, etc.

[0149] Library 616 may provide a common infrastructure that may be used by Application 620 and / or other components and / or layers. Library 616 generally provides functions that enable other software modules to perform tasks in a way that is easier than directly interface with the functions of the underlying operating system 614 (e.g., kernel 628, services 630, and / or drivers 632). Library 616 may include a system library 634 (e.g., the C standard library) which may include functions such as memory allocation functions, string manipulation functions, and mathematical functions. In addition, Library 616 may include API libraries 636, such as a media library (e.g., a library that supports the presentation and manipulation of various media formats such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG), a graphics library (e.g., the OpenGL framework which may be used to render 2D and 3D graphic content to a display), a database library (e.g., SQLite which may provide various relational database functions), and a web library (e.g., WebKit which may provide web browsing functionality). Library 616 may also include a wide variety of other libraries 638 that provide many other APIs to application 620 and other software components / modules.

[0150] Framework / middleware 618 may provide a higher level of common infrastructure that can be utilized by application 620 and / or other software components / modules. For example, framework / middleware 618 may provide various graphical user interface (GUI) functions, higher-level resource management, higher-level services, etc. Framework / middleware 618 may also provide a wide range of other APIs that can be utilized by application 620 and / or other software components / modules, some of which may be specific to a particular operating system or platform.

[0151] Application 620 includes embedded applications 640 and / or third-party applications 642. Typical examples of embedded applications 640 may include, but are not limited to, contact applications, browser applications, book reader applications, location-based applications, media applications, messaging applications, and / or game applications. Third-party applications 642 may include embedded applications and other applications in a broader combination. In certain examples, a third-party application 642 (for example, an application developed by an entity other than the vendor of a particular platform using the Android® or iOS® Software Development Kit (SDK)) may be mobile software that runs on a mobile operating system such as iOS®, Android®, Windows® Phone, or another mobile computing device's operating system. In this example, a third-party application 642 may call API calls 624 provided by the mobile operating system, such as operating system 614, to facilitate the functionality described herein.

[0152] Application 620 may utilize built-in operating system functions (e.g., kernel 628, services 630 and / or drivers 632), libraries (e.g., system libraries 634, API libraries 636, and other libraries 638), and frameworks / middleware 618 to create a user interface for interacting with the system's users. Alternatively or additionally, in some systems, user interaction may occur via a presentation layer, such as presentation layer 644. In these systems, the application / module "logic" may be separated from the application / module aspects that interact with the user.

[0153] Some software architectures utilize virtual machines. In the example in Figure 6, this is represented by virtual machine 648. A virtual machine creates a software environment that allows applications / modules to run as if they were running on a hardware computing device. A virtual machine is hosted by a host operating system (operating system 614) and, though not always, generally has a virtual machine monitor 646, which manages the operation of virtual machine 648 and its interface with the host operating system (i.e., operating system 614). The software architecture runs within virtual machine 648, including the operating system 650, libraries 652, frameworks / middleware 654, applications 656, and / or presentation layer 658. These layers of the software architecture running within virtual machine 648 may be the same as or different from the corresponding layers mentioned above.

[0154] Modules, components, and logic A computer system may include logic, components, modules, mechanisms, or any preferred combination thereof. A module may constitute either a software module (e.g., code embodied (1) on a non-temporary machine-readable medium, or (2) in a transmitted signal) or a hardware implementation module. A hardware implementation module is a tangible unit capable of performing several operations and may be configured or arranged in a certain manner. One or more computer systems (e.g., a standalone, client, or server computer system) or one or more hardware processors may be configured by software (e.g., an application or application portion) as a hardware implementation module operating to perform some of the operations described herein.

[0155] Hardware implementation modules may be implemented mechanically or electrically. For example, a hardware implementation module may include dedicated circuitry or logic permanently configured to perform certain operations (e.g., as a dedicated processor such as a field-programmable gate array (FPGA) or application-specific integrated circuit (ASIC)). A hardware implementation module may also include programmable logic or circuitry temporarily configured by software to perform certain operations (e.g., contained within a general-purpose processor or other programmable processor). It will be understood that the decision of whether to implement a hardware implementation module mechanically with dedicated, permanently configured circuitry or with temporarily configured circuitry (e.g., configured by software) may be made considering cost and time.

[0156] Therefore, the term “hardware implementation module” should be understood to include tangible entities that are physically constructed or permanently configured (e.g., hardwired) or temporarily or transiently configured (e.g., programmed) to operate in a certain way and / or perform some of the operations described herein. Hardware implementation modules may be temporarily configured (e.g., programmed), and each hardware implementation module does not need to be permanently configured or instantiated. For example, if a hardware implementation module comprises a general-purpose processor configured using software, the general-purpose processor may be configured as each different hardware implementation module at different times. Thus, the software may configure the processor to form a particular hardware implementation module at one moment and a different hardware implementation module at another moment.

[0157] Hardware implementation modules can provide information to and receive information from other hardware implementation modules. Therefore, the hardware implementation modules described may be considered to be communicatively coupled. If multiple such hardware implementation modules exist simultaneously, communication may be achieved by signal transmission (e.g., through appropriate circuits and buses connecting the hardware implementation modules). Multiple hardware implementation modules may be configured or instantiated at different times. Communication between such hardware implementation modules may be achieved, for example, by storing information in a memory structure accessible to multiple hardware implementation modules and retrieving it. For example, one hardware implementation module may perform an operation and store the output of that operation in a memory device to which that hardware implementation module is communicatively coupled. Subsequently, further hardware implementation modules may access the memory device to retrieve and process the stored output. Hardware implementation modules may also initiate communication with input or output devices and operate on resources (e.g., information gathering).

[0158] Various operations of the exemplary methods described herein may be performed, at least in part, by one or more processors that are temporarily or permanently configured (e.g., by software) to perform the relevant operations. Whether temporarily or permanently configured, such processors may form a processor implementation module that operates to perform one or more operations or functions. The modules referred to herein may include processor implementation modules.

[0159] Similarly, the methods described herein may be processor implementations, at least in part. For example, at least part of the operation of the method may be performed by one or more processors or processor implementation modules. Some of the performance of the operation may be distributed among one or more processors located across several machines, as well as within a single machine. One or more processors may be located in a single location (e.g., in a home environment, an office environment, or a server farm), or the processors may be distributed across several locations.

[0160] One or more processors may also operate in a “cloud computing” environment or as “software as a service” (SaaS) to support the performance of the related operations. For example, at least part of the operations may be performed by a group of computers (as an example of machines including processors), and these operations may be accessible over a network (e.g., the Internet) and over one or more suitable interfaces (e.g., APIs).

[0161] Electronic devices and systems The systems and methods described herein may be implemented using digital electronic circuits, computer hardware, firmware, software, computer program products (for example, data processing devices, for example, programmable processors, computer programs tangibly embodied on information carriers, for example, on machine-readable media, for execution by one or more computers or for controlling their operation), or a suitable combination thereof.

[0162] Computer programs can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, such as as standalone programs or as modules, subroutines, or other units suitable for use in a computing environment. Computer programs can be deployed to run on a single computer, or on multiple computers located in one site or distributed across multiple sites (e.g., cloud computing) and interconnected by a communication network. In cloud computing, server-side functions may be distributed across multiple networked computers. Load balancers are used to distribute work among multiple computers. Thus, a cloud computing environment that implements the method is a system with multiple processors on multiple computers tasked with performing the operation of that method.

[0163] The operation may be performed by one or more programmable processors executing a computer program to perform operations on input data and generate an output. The operation of the method may also be performed by dedicated logic circuits, such as FPGAs or ASICs, and the devices of the system may be implemented as dedicated logic circuits, such as FPGAs or ASICs.

[0164] A computing system can include clients and servers. Clients and servers are generally geographically separated from each other and typically interact through a communication network. The client-server relationship arises from computer programs running on each computer and having a client-server relationship with each other. A programmable computing system may be deployed using a hardware architecture, a software architecture, or both. In detail, it will be understood that the choice of whether to implement a certain function with permanently configured hardware (e.g., ASICs), with temporarily configured hardware (e.g., a combination of software and a programmable processor), or with a combination of permanently and temporarily configured hardware may be a design choice. Below are some example hardware (e.g., machine) and software architectures that may be deployed.

[0165] Exemplary machine architecture and machine-readable media Figure 7 shows a block diagram of an exemplary machine of a computer system 700 in which instructions 724 may be executed internally to cause the machine to perform one or more of the methods described herein. The machine may operate as a standalone device or be connected to other machines (e.g., networked). In a networked deployment, the machine may operate as a server or client in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a web appliance, a network router, a switch or bridge, or any machine capable of executing instructions (sequential or non-sequential) that specify actions to be performed by that machine. Furthermore, although only a single machine is illustrated, the term “machine” shall be interpreted to include any set of machines that individually or collectively execute a set (or set of) instructions to perform one or more of the methods described herein.

[0166] An exemplary computer system 700 includes a processor 702 (for example, a central processing unit (CPU), a graphics processing unit (GPU), or both), main memory 704, and static memory 706, which communicate with each other via a bus 708. The computer system 700 may further include a video display unit 710 (for example, a liquid crystal display (LCD) or a cathode ray tube (CRT)). The computer system 700 also includes an alphanumeric input device 712 (for example, a keyboard or a touch-sensitive display screen), a user interface (UI) navigation (or cursor control) device 714 (for example, a mouse), a storage unit 716, a signal generating device 718 (for example, a speaker), and a network interface device 720.

[0167] Machine-readable media The storage unit 716 includes a machine-readable medium 722 in which one or more sets of data structures and instructions 724 (e.g., software) that embody or are utilized by any one or more of the methods or functions described herein are stored. The instructions 724 may also be entirely or at least partially in the main memory 704 and / or the processor 702 during their execution by the computer system 700, and the main memory 704 and the processor 702 also constitute the machine-readable medium 722.

[0168] Although the machine-readable medium 722 is shown as a single medium in Figure 7, the term “machine-readable medium” may include a single or multiple mediums (for example, a centralized or distributed database, and / or associated caches and servers) that store one or more instructions 724 or data structures. The term “machine-readable medium” shall also be interpreted as including any tangible medium that can store, encode, or carry instructions 724 for machine execution, or that can store, encode, or carry data structures that are utilized by or associated with the instructions 724, or that cause a machine to execute one or more of the methods of this disclosure. Accordingly, the term “machine-readable medium” shall be interpreted as including, but not limited to, solid-state memory, as well as optical and magnetic media. Specific examples of machine-readable media include non-volatile memory, such as semiconductor memory devices, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and compact disc read-only memory (CD-ROM) and digital versatile disc read-only memory (DVD-ROM) disks. Machine-readable media are not transmission media.

[0169] transmission medium Instructions 724 may further be transmitted or received through a communication network 726 using a transmission medium. Instructions 724 may be transmitted using a network interface device 720 and one of several well-known transport protocols (e.g., Hypertext Transport Protocol (HTTP)). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, cellular networks, legacy telephone (POTS) networks, and wireless data networks (e.g., WiFi and WiMAX networks). The term “transmission medium” shall be interpreted as including any intangible medium that can store, encode, or carry instructions 724 for execution by a machine, and which contains digital or analog communication signals, or other intangible medium that facilitates the communication of such software.

[0170] While this specification describes specific examples, it will be apparent that various modifications and changes can be made to these examples without departing from the broader intent and scope of this disclosure. Therefore, this specification and its drawings should be considered in an illustrative rather than restrictive sense. The accompanying drawings, which form part of this specification, illustrate, not restrictively, specific examples in which the subject matter may be put into practice. These illustrative examples are described in sufficient detail so that those skilled in the art can put the teachings disclosed herein into practice.

[0171] Some parts of the subject matter described herein may be presented in relation to algorithms or symbolic representations of operations on data stored as bits or binary digital signals in machine memory (e.g., computer memory). Such algorithms or symbolic representations are examples of techniques used by those skilled in the art in data processing to communicate the content of their operations to others skilled in the art. As used herein, “algorithm” is a self-consistent set of operations or similar processes that lead to a desired result. In this context, algorithms and operations include the physical manipulation of physical quantities. While not always the case, generally such quantities may take the form of electrical, magnetic, or optical signals that can be stored, accessed, transferred, combined, compared, or otherwise manipulated by machines. It is sometimes convenient to refer to such signals using words such as “data,” “content,” “bit,” “value,” “element,” “symbol,” “character,” “term,” “digit,” and “numerical value,” primarily for reasons of common usage. However, these words are merely convenient labels and should be associated with appropriate physical quantities.

[0172] Unless otherwise specified, descriptions herein using words such as “process,” “calculate,” “calculate,” “determine,” “present,” and “display” may refer to the actions or processes of a machine (e.g., a computer) that manipulates or transforms data expressed as physical (e.g., electronic, magnetic, or optical) quantities in one or more memories (e.g., volatile memory, non-volatile memory, or a preferred combination thereof), registers, or other machine components that receive, store, transmit, or display information. Furthermore, unless otherwise specified, the terms “a” and “an,” while common in patent literature, are used herein to include one or more instances. Finally, unless otherwise specified, the conjunction “or” as used herein refers to a non-exclusive “or.” [Explanation of Symbols]

[0173] 100 Systems 102 Data warehousing environment 104 Data Warehousing Environment User Interface 106 Data warehousing environment backend services 108 Data Builder 110 In-memory cloud databases 112 Object Store 114 Data Integration Monitor 116 Table Monitors 118 Deployers and Middleware 120 Local Table Monitor Backend 124 Open Table Format Structured Query Language Application Programming Interface (OTF SQL API) 126 Data Definition Language (DDL) Objects and Procedures 128 virtual tables 130 File Adapters 132 containers 134 files 136 Delta Table 138 Spark Engine 140 Spark Adapter 142 External Systems 200 Architectures 202 Data Warehousing Environment 204 HDLF 206 Delta Table 208 files 210 Object Store 212 compute layers 214 Spark Engine 216 In-memory query compute engine 218 External Sources 220 Inbound Replication Flow 222 HANA Cloud and Elastic Compute Nodes 224 Model 226 views 228 Cache 230 Transfer Flows 232 Tables 234 Outbound Replication Flow 236 Information Access (INA) / Open Data Protocol (ODATA) 238 SQL 240 Spaces & Catalog

Claims

1. At least one hardware processor, When executed by the at least one hardware processor, a non-temporary computer-readable medium stores instructions that cause the at least one hardware processor to perform an action. The operation includes, Receiving data from a data lake, wherein the data lake stores and receives the data in raw data storage format. The decision of whether to store the aforementioned data in an in-memory database or in a first type of object storage, Creating a database artifact in a cloud database, wherein the database artifact contains information about the data, but is stored in a format that does not depend on whether the data is stored in the in-memory database or in the first form of object storage. Based on the decision to store the data in the first type of object storage, A delta table containing the aforementioned data is generated in the first type of object storage, To generate a virtual table determined by the delta table in the in-memory database, To provide access to the data via the aforementioned database artifact. A system that includes this.

2. The aforementioned operation, Modifying the delta table based on the changes made to the aforementioned data, From the aforementioned delta table, capture change data related to the aforementioned change, The aforementioned change data is stored in the cloud database as a delta history table. The system according to claim 1, further comprising:

3. The system according to claim 2, wherein the first form of object storage is an open data form of object storage, and the changes made to the data are received from the data lake.

4. The system according to claim 2, wherein the delta history table includes all fields of the delta table, as well as metadata columns for change type and timestamp.

5. The aforementioned operation, To combine multiple data points within the aforementioned delta table into a single data point, an optimization command is executed on the delta table. The system according to claim 1, further comprising:

6. The aforementioned operation, Execute a vacuum command to remove old data from the aforementioned delta table. The system according to claim 1, further comprising:

7. The system according to claim 1, wherein the virtual table is a logical representation of the delta table, and enables a user to query the data stored in the delta table and interact with the data using structured query language (SQL) commands.

8. A step of receiving data from a data lake, wherein the data lake stores the data in a raw data storage format. The steps include determining whether to store the aforementioned data in an in-memory database or in a first type of object storage, A step of creating a database artifact in a cloud database, wherein the database artifact contains information about the data, but is stored in a format that does not depend on whether the data is stored in the in-memory database or in the first type of object storage, Based on the decision to store the data in the first type of object storage, The steps include generating a delta table containing the aforementioned data in the first type of object storage, The steps include generating a virtual table determined by the delta table in the in-memory database, The steps include providing access to the data via the aforementioned database artifact. Methods that include...

9. The steps include modifying the delta table based on the changes made to the aforementioned data, The steps include capturing change data related to the change from the delta table, The steps include storing the aforementioned change data as a delta history table in the cloud database. The method according to claim 8, further comprising:

10. The method according to claim 9, wherein the first form of object storage is an open data form of object storage, and the changes made to the data are received from the data lake.

11. The method according to claim 9, wherein the delta history table includes all fields of the delta table, as well as metadata columns for change type and timestamp.

12. The step of executing an optimization command on the delta table in order to combine multiple data points in the delta table into a single data point. The method according to claim 8, further comprising:

13. Step 1: Execute a vacuum command to remove old data from the delta table. The method according to claim 8, further comprising:

14. The method according to claim 8, wherein the virtual table is a logical representation of the delta table, and enables a user to query the data stored in the delta table and interact with the data using structured query language (SQL) commands.

15. A non-temporary machine-readable medium that stores instructions for causing one or more processors to perform an action, when executed by one or more processors, Receiving data from a data lake, wherein the data lake stores and receives the data in raw data storage format. The decision of whether to store the aforementioned data in an in-memory database or in open data object storage, Creating a database artifact in a cloud database, wherein the database artifact contains information about the data, but is stored in a format that does not depend on whether the data is stored in the in-memory database or in the open data format object storage. Based on the decision to store the aforementioned data in the open data format object storage, A delta table containing the aforementioned data is generated in the open data format object storage, To generate a virtual table determined by the delta table in the in-memory database, To provide access to the data via the aforementioned database artifact. Non-temporary machine-readable media, including [specific examples of such media].

16. The aforementioned operation, Modifying the delta table based on the changes made to the aforementioned data, From the aforementioned delta table, capture change data related to the aforementioned change, The aforementioned change data is stored in the cloud database as a delta history table. A non-temporary machine-readable medium according to claim 15, further comprising:

17. The non-temporary machine-readable medium according to claim 16, wherein the first form of object storage is an open data form of object storage, and the changes made to the data are received from the data lake.

18. The non-temporary machine-readable medium according to claim 17, wherein the delta history table includes all fields of the delta table, as well as metadata columns for change type and timestamp.

19. The aforementioned operation, To combine multiple data points within the aforementioned delta table into a single data point, an optimization command is executed on the delta table. A non-temporary machine-readable medium according to claim 15, further comprising:

20. The aforementioned operation, Execute a vacuum command to remove old data from the aforementioned delta table. A non-temporary machine-readable medium according to claim 15, further comprising: