Database load unit for replication log replay

By capturing the loading of database objects in a replay log and dynamically adjusting load units based on memory constraints, the secondary database system efficiently manages large data sets, addressing the memory limitations challenge in database systems.

JP2025090515AActive Publication Date: 2025-06-17エスアーペーエスエー
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024189991
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-10-29
Publication Date
2025-06-17
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

Existing database systems face challenges in efficiently managing large data sets due to memory limitations, where primary databases often load data in column-loadable formats, which may not be optimal for secondary databases with smaller memory capacities.

Method used

The implementation involves a primary database system loading database objects into an in-memory store and capturing this process in a replay log, which is then sent to a secondary database system. The secondary database system checks a log replay configuration parameter to determine the format in which to load the database objects, allowing it to switch between column-loadable and page-loadable formats based on its memory constraints.

Benefits of technology

This approach allows the secondary database system to efficiently manage its memory by dynamically adjusting the load units of database objects, ensuring optimal performance even with limited memory capacity, without the need for rewriting the persistent store.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025090515000001_ABST
    Figure 2025090515000001_ABST
Patent Text Reader

Abstract

To provide a database load unit for replication log replay.SOLUTION: A process for using different load units comprises: loading a database object into a primary in-memory store; capturing it in a replay log according to a predetermined format and transmitting it to a secondary database system; checking, by the secondary database system, the value of a log replay configuration parameter in response to reception of the replay log; if the configuration parameter is a first value, replaying the replay log to load the corresponding database objects into a secondary in-memory store according to a first format; if the configuration parameter is a second value, replaying the log and loading the object according to the second format; and if the configuration parameter is a third value, replaying the log and loading the object in the same format that was used by the primary database system.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to database processing.

Background Art

[0002] A database is a collection of organized data. A database typically organizes data to correspond to a logical arrangement of the data. This facilitates operations on the data, such as retrieving values within the database, adding data to the database, sorting data within the database, or summarizing related data within the database. A database management system (DBMS) mediates the interaction between the database, users, and applications to organize, create, update, capture, analyze, and otherwise manage the data within the database.

[0003] To process queries efficiently, a database is typically configured to perform in-memory operations on the data. In an in-memory database, the data required for query execution and response is loaded into memory, and the query is executed against that in-memory data. However, many applications have large data stores, and due to memory limitations, it may be difficult for these applications to load all the data they need into memory. The amount of data processed by a database system continues to increase at a pace faster than memory devices evolve to store more data.

Summary of the Invention

Means for Solving the Problems

[0004] In some implementations, the primary database system loads database objects into the primary in-memory store according to a given format determined in the primary database system. The primary database system captures the loading of the database objects in a replay log according to the given format. The primary database system sends the replay log to the secondary database system. In response to receiving the replay log, the secondary database system checks the value of a log replay configuration parameter. If the log replay configuration parameter is a first value, the secondary database system replays the replay log and loads the corresponding database objects into the secondary in-memory store according to a first format, where the first format may or may not be different from the given format used by the primary database system. If the log replay configuration parameter is a second value, the secondary database system replays the first log and loads the database objects into the secondary in-memory store according to a second format, where the second format may or may not be different from the given format used by the primary database system. If the log replay configuration parameter is a third value, the secondary database system replays the first log and loads at least one database object into the secondary in-memory store in the same format (i.e., the given format) as used by the primary database system.

[0005] A non-transitory computer program product (i.e., a physically embodied computer program product) will also be described. When executed by one or more data processors of one or more computing systems, the non-transitory computer program product stores instructions that cause at least one data processor to perform the operations described herein. Similarly, a computer system that may include one or more data processors and a memory coupled to the one or more data processors will be described. The memory may temporarily or persistently store instructions that cause at least one processor to perform one or more of the operations described herein. Further, the method can be implemented by one or more data processors that are distributed within a single computing system or across two or more computing systems. Such computing systems can be connected via one or more connections including connections via a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), direct connections between one or more of a plurality of computing systems, etc., to exchange data and / or commands, or other instructions, etc.

[0006] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims.

[0007] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate specific aspects of the subject matter disclosed herein and, together with the description, serve to explain some of the principles associated with the disclosed implementations.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 11

Figure 12

[0009] Data structures are typically created and imported in memory, but once imported, the data structures may be persisted in persistent storage. When entering persistent storage, the data structures can be deleted from memory when they are no longer needed. Then, if the data structures are needed again in memory in the future, the data structures can be reconstructed from the information persisted in persistent storage. Loading a data structure refers to reconstructing the data structure in memory from the information persisted in persistent storage. Although the representation of the data structure in the persistent store may not match the representation in memory, the information stored in persistent storage is sufficient to fully reconstruct the data structure in memory.

[0010] In a database, a database object is a data structure used to store or reference data. A common type of database object is a table. Other types of database objects include columns, indexes, stored procedures, sequences, views, and the like. The database may store each database object as a plurality of substructures that collectively form and represent the database object. Note that in this specification, the terms "substructure" and "subcomponent" may be used interchangeably. In the case of a column, the substructure may include a dictionary, a database vector, and an index. The dictionary may associate each unique data value with a corresponding value identifier (ID). The value IDs may be numbered sequentially. The database vector may include value IDs that map to the actual values of the database object. The index may include a contiguous list of each unique value ID and one or more positions within the database vector that includes the value ID.

[0011] When taking a database object from a persistent location into an in-memory location, the database can take the database object into memory using a plurality of different formats. For example, one format is called a column-loadable format. In the case of a column-loadable format, when a column is queried, the entire column is fully loaded into memory. Fully loading the entire column into memory means taking into memory the entire subcomponents of the column (e.g., data, dictionary, and index). Another type of format that can be adopted is called a page-loadable format. In the case of a page-loadable format, only the pages that include the portions of the columns relevant to the query are loaded into memory. In other embodiments, other types of formats may be adopted.

[0012] When loading data into memory, a load unit may be used to specify the format in which the data is loaded and the granularity of the data loaded from the persistent store. In other words, a "load unit" defines the granularity of the data loaded from the persistent store into the in-memory store. Data Definition Language (DDL) statements can change the load units of various database objects such as tables, table partitions or partition sets, table columns or column sets, or composite indexes defined on a table. As an example, different types of load units that can be specified include page loadable, column loadable, and default loadable. In the case of the default loadable format, database objects recursively inherit the parent's load unit.

[0013] The format or load unit specifies how columns are loaded into memory, but the format does not necessarily determine how columns are persisted. In some database deployments, the load units for column loadable and page loadable formats may use different persistences based on the load unit definition. For example, a column loadable column can be persisted by serializing the column's data, dictionary, and index into a single page chain. In the case of a page loadable column, the pages containing the column portions required for a query can be easily identified and these columns can be persisted using a paging scheme so that they can be fetched into memory for access. In the case of a page loadable column, individual page chains are persisted for each of the sub-components of the column's data, dictionary, and index.

[0014] In some embodiments, when importing a database object from a persistent store location to an in-memory location, the database may determine whether to load the database object in a column-loadable format or a page-loadable format according to the current workload, corresponding attributes, and / or one or more other operating conditions. In some cases, it may be desirable to store the database object in a column-loadable format, while in other cases, it may be desirable to store the database object in a page-loadable format. The mechanisms and data layouts for storing database objects in a column-loadable format are different from the configurations for storing database objects in a page-loadable format. Therefore, switching between the two configurations may sometimes require completely rewriting the data persistence. The rewriting consumes a significant amount of memory and processing resources. Therefore, switching the in-memory format can be expensive in terms of the memory and processing resources utilized for the necessary rewriting. Therefore, there is a need for an improved technique that does not require rewriting the persistent store for the converted load unit.

[0015] As an example, if a given database object has a column-loadable load unit and the processing logic determines not to convert the load unit, the given database object is loaded into the in-memory store in a column-loadable format. When loading a column into the in-memory store in a column-loadable format, the processing logic loads the entire column into the in-memory store. If a given database object has a page-loadable load unit and the processing logic determines not to convert the load unit, the given database object is loaded into the in-memory store in a page-loadable format. In the case of a page-loadable format, the processing logic loads only the relevant pages into the in-memory store.

[0016] As an example, less frequently used tables are loaded into memory in page-loadable form, even if those tables have column-loadable load units. For these less frequently used tables, only the pages relevant to the queries being executed are loaded into memory. In some embodiments, more frequently used tables are loaded into memory in column-loadable form, even if those tables have page-loadable load units. For these more frequently used tables, entire columns are loaded into memory. In some embodiments, when storing from an in-memory store to a persistence store, the processing logic stores database objects in a single "unified persistence form", regardless of their in-memory form.

[0017] The unified persistence form allows each database object to be loaded into memory from the same persistence in either column-loadable form or page-loadable form. The unified persistence form also allows for the conversion of the load unit of a given database object without rewriting the given database object within the persistence store. The unified persistence form includes individual composite page chains for the data, dictionary, and index sub-components of a given database object. The composite page chain may have one or more sub-page chains.

[0018] Another type of persistence format is the serial persistence format that stores sub-components in sequence like a single page chain. The serial persistence format can be suitable for scenarios where data components are relatively small and can be efficiently stored together in a single chain. The serial persistence format is often used for smaller datasets and has advantages such as compact storage and fast access to the dataset. On the other hand, the unified persistence format is more suitable for larger and more complex datasets where the sizes of sub-components can vary significantly. With the unified persistence format, each sub-component can be stored optimally and independently of each other, taking into account the specific size and characteristics of each sub-component. The integrated persistence offers advantages such as performance improvement and the ability to load persistent data in a column-loadable or page-loadable format, which cannot be utilized if the dataset is stored in the serial persistence format.

[0019] Up-to-date database systems may provide a data replication service that mirrors data to improve performance. Data replication involves copying data from a database on one server (i.e., the primary database) to another database on another server or client (i.e., the secondary database). In some implementations, the secondary database is configured to be an exact copy in terms of processing resources, memory capacity, storage capacity, etc. However, while this configuration simplifies the management of the primary and secondary databases, it can be a costly configuration to maintain.

[0020] Accordingly, some organizations may implement a secondary database, which is a reduced version of the primary database, for cost reduction. In such cases, data replication during a failure or downtime and switching from the primary database to the secondary database become more complex. For example, if the memory capacity of the secondary database is half that of the primary database, the load units used by the primary database to load database objects may not be optimal when applied to the secondary database. Further elaborating, if the primary database uses a column-loadable format for all or most database objects to be loaded into the in-memory store, these database objects are fully loaded into the in-memory store. This approach may not be executable in the secondary database because the capacity of the in-memory store is smaller. Therefore, to account for the differences in the physical configuration of the primary database compared to the secondary database, the secondary database may use different load units for the same database objects loaded by the primary database. As an example, the secondary database may use a page-loadable format for database objects that are loaded in a column-loadable format in the primary database. As described above, the page-loadable format loads only the pages relevant to the query into the in-memory store, reducing memory usage.

[0021] Next, referring to FIG. 1, a system diagram is shown that illustrates an example of a database system 100 according to some exemplary embodiments. In FIG. 1, database system 100 may include one or more client devices 110, a database execution engine 130, and one or more databases 140. Note that the terms "database execution engine", "processor", and "processing logic" may be used interchangeably herein. It is shown that database 140 includes a table 150 that represents any number and type of database objects stored by database 140.

[0022] One or more client devices 110, database execution engine 130, and one or more databases 140 may be communicatively coupled via network 120. One or more databases 140 may include various relational databases, such as, for example, in-memory databases, column-based databases, row-based databases, and the like. One or more client devices 110 may include processor-based devices, such as, for example, mobile devices, wearable devices, personal computers, workstations, Internet of Things (IoT) appliances, and the like. Network 120 may be a wired network and / or a wireless network, including, for example, a public land mobile network (PLMN), a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), the Internet, and the like.

[0023] Next, referring to FIG. 2, a diagram is shown that illustrates an example of a system 200 that conforms to an implementation form of the subject matter of the present invention. System 200 includes applications 240A-N that interact with an in-memory database 210. Applications 240A-N represent any number and type of applications that are executed within system 200. For example, applications 240A-N can include analytical services, transaction processing services, report services, dashboard services, and the like. Applications 240A-N interact with in-memory database 210 using structured queries or other expressions. System 200 also includes one or more platforms 245. In some embodiments, a given platform 245 can be a data warehouse service that exposes online analytical processing (OLAP) services to users or applications. A database management service 260 provides administrative services and access to the configuration of in-memory database 210. A user interface (UI) 270 mediates the interaction between database management service 260 and a user or application, for example, to configure in-memory database 210 or to set user preferences for sorting operations.

[0024] Within the in-memory database 210, a data management service 250 manages transactions with the main storage 230 and the delta storage 235. The data management service 250 can provide a calculation and planning engine, a modeling service, a real-time replication service, a data integration service, and / or other services. The main storage 230 supports fast read access to the data of the in-memory database 210 within the data store 225. The read operation accesses both the main storage 230 and the delta storage 235, and the delta storage 235 contains any recently changed data that has not yet been incorporated into the main storage 230. The data within the main storage 230 can be backed up to the persistent storage 280. The persistent storage 280 may be disk storage or other suitable type of memory device. The changed data within the delta storage 235 is also backed up to the persistent storage 280 so that changes are retained even if events such as database failures or downtimes occur. The buffer cache 233 refers to a portion of the memory that provides temporary storage for data loaded from the persistent storage 280. For example, if a query or database operation requires a column or a portion of a database table, the corresponding page needed to respond to the query can be identified and loaded into the buffer cache 233.

[0025] In some embodiments, the persistent storage 280 stores database objects in a single unified persistent format regardless of the object's load unit. The unified persistent format will be described in more detail throughout the remainder of the present disclosure. It should be understood that the exemplary architecture of the system 200 merely shows what may be employed in some embodiments. In other embodiments, the system 200 may be structured in other suitable ways using different arrangements of components.

[0026] Next, referring to FIG. 3, a block diagram of an example of a system 300 for implementing load unit determination is shown. When a table is persisted to a persistence store 325 using a unified persistence format, each column of the table has an individual page chain for each of the data, dictionary, and index sub-components. These are shown as a data page chain 330, a dictionary page chain 340, and an index page chain 350. When loading a portion of the table from the persistence store 325 into memory, a load unit determination 320 regarding the intended structure of the table in the in-memory store is made dynamically. When the load unit determination 320 uses a page load attribute 315 (i.e., a load unit capable of page loading), the table is loaded into memory using a page loadable format. The page loadable format has individual page chains for the data, dictionary, and index portions of the table. By having individual page chains for the data, dictionary, and index portions of the table, only the relevant pages rather than the entire column can be loaded into memory. When the determination of the load unit uses an in-memory load attribute 310, the table is loaded into memory using a column loadable load unit, also referred to as an in-memory load format or a column loadable format.

[0027] When columns of a table are loaded into memory using a column loadable format, the entire data is read from the persistent store 325, and an index vector, a dictionary, and a reverse index vector are created as in-memory structures. Once the data is read from the persistent store 325 and the in-memory structures are created, the corresponding pages from the persistent store 325 are no longer accessed. Using the in-memory index vector, direct access can be made to a specific offset within the in-memory store.

[0028] When the columns of a table are loaded into memory using a page-loadable format, only the relevant data is read from the persistence store 325 and loaded into memory. As an example, in a page-loadable format, the processor first calculates which page the offset is on, then that page is loaded, and then data is read from that page. Within that page, additional calculations may be required from the header to reach a particular offset of a particular row.

[0029] Next, referring to FIG. 4, another example of a system 400 for implementing a load unit decision is shown. As illustrated, system 400 includes a persistence store 410 that stores data in a unified persistence format. In the case of a table stored in a unified persistence format, each component of each column (i.e., attribute) of the table has an individual page chain. For example, data components are stored as data page chain 420, dictionary components are stored as dictionary page chain 430, and index components are stored as index page chain 440.

[0030] If it is necessary to load a given portion of the table from the persistent store 410 into the in-memory store, a load unit determination 450 is made to determine in what form the given portion is to be stored in memory. As an example, the determination of the form in which to load a given portion into memory is based at least on attributes associated with the table portion. For example, the attribute may be a page load attribute 460 or the attribute may be an in-memory load attribute 470. As another example, the determination of the form in which to load a given portion into memory is also based on real-time operating conditions (e.g., memory utilization rate, frequency at which the table is accessed). If the real-time operating conditions indicate that a load unit conversion should not be performed based on memory utilization rate, access frequency, and other factors, the load unit determination 450 can revert to the attributes associated with the table. In the case of a query that references a column with a page load attribute 460, if no load unit conversion is implemented, the data is loaded into memory in a page-loadable form. In the case of a query that references a column with an in-memory load attribute 470, if no load unit conversion is implemented, the data is loaded into memory in a column-loadable form.

[0031] When a load unit conversion is implemented, columns with a page load attribute 460 are loaded into the in-memory store in a column loadable format. Or, if a column has an in-memory load attribute 470 and a load unit conversion is implemented, the column is loaded into the in-memory store in a page loadable format. As an example, a load unit conversion can be implemented by converting a page loadable load unit to a column loadable load unit when the memory usage rate is below a threshold. Alternatively, or additionally, a load unit conversion can be implemented by converting a column loadable load unit to a page loadable load unit when the memory usage rate exceeds a threshold. Alternatively, or additionally, a load unit conversion can be implemented by converting a page loadable load unit to a column loadable load unit when the access frequency of a column exceeds a threshold. Alternatively, or additionally, a load unit conversion can be implemented by converting a column loadable load unit to a page loadable load unit when the access frequency of a column is below a threshold. Other conditions for performing a load unit conversion are possible and contemplated.

[0032] When the table is loaded into the in-memory store and changes are made to the table in a page loadable format, the delta component 480 is merged with the unchanged data and stored again in the persistence store 410 in a unified persistence format. Similarly, when changes are made to a table in a column loadable format within the in-memory store, the delta component 485 is merged with the unchanged data and stored again in the persistence store 410 in a unified persistence format. In other words, in this example, regardless of which load unit is used to load the table into memory, the changed data is merged again and stored in the persistence store 410 in a unified persistence format.

[0033] Next, referring to FIG. 5, a system 500 is shown that includes a primary database 510 and a secondary database 520. As shown, data is replicated from the primary database 510 to the secondary database 520. Also, REDO logs can be sent from the primary database 510 to the secondary database 520. The REDO logs include modifications executed in the primary database 510, and the REDO logs are sent to the secondary database 520 so that the same modifications can be applied to the secondary database 520. The secondary database 520 is prepared to be used if a failure occurs in the primary database 510. If a failure occurs in the primary database 510, the system 500 can switch so that the secondary database 520 is used instead of the primary database 510, and columns can be used in the secondary database 520 in page form before the columns are fully loaded into memory.

[0034] In the primary database 510, columns are loaded according to the in-memory format, and the entire column is loaded into memory. This is shown by the entire sub-components of the column data, dictionary, and index being loaded into the in-memory store. Since different load units can be used in the unified persistence format, the secondary database 520 uses a page structure to load columns into memory in a page-loadable form. In the secondary database 520, the load unit can be determined for each column. As an example, in the case of the secondary database 520, if a given column is not fully loaded into memory, the given column points to the corresponding page chain within the persistence store. During access, the relevant data is loaded into memory page by page. For example, when a query is executed, only a portion of the data required for the query is loaded from the persistence store into memory.

[0035] In some embodiments, system 500 may include various types of configuration parameters. By way of example, a first configuration parameter, referred to as a "load unit parameter during log playback" or a "log playback configuration parameter", may be any of three values: default, page loadable, or column loadable. When the log playback configuration parameter is set to default, the tables on the secondary database 520 match the format of the tables on the primary database 510. When the log playback configuration parameter is set to page loadable, a particular table becomes page loadable in the secondary database 520 even if that particular table is loaded in a column loadable format in the primary database 510. Similarly, when the log playback configuration parameter is set to column loadable, a given table becomes column loadable in the secondary database 520 even if that given table is loaded in a page loadable format in the primary database 510.

[0036] By way of example, another configuration parameter, referred to as a "load unit parameter after failover", may have one of two values: continue as log playback or reload and reset to primary. When the load unit parameter after failover is set to continue as log playback, the load unit continues in the secondary database 520 in the same load format used in the primary database 510. When the load unit parameter after failover is set to reload and reset to primary, in the secondary database 520, the load unit is switched to be the same as the primary database 510. Thereby, to identify the loaded columns, the procedures of an unload script or a reload script are explicitly executed and then unloaded and reloaded again in the defined load unit. By way of example, this script is integrated with a failover script. As another example, this script becomes effective when the load unit configuration parameter after failover is set.

[0037] Next, referring to FIG. 6, another example of a system 600 with a primary database 610 and a secondary database 620 is shown. System 600 has a reverse configuration compared to system 500, in that the primary database 610 loads columns into memory in a page-loadable format, while the secondary database 620 loads columns into memory in a column-loadable format. This is the flexibility enabled by the unified persistence format, and the primary database 610 and the secondary database 620 can load data from the same persistence in either a page-loadable format or a column-loadable format without rewriting the persistence.

[0038] As an example, the configuration of system 600 can be adopted immediately after a failure occurs in the primary database 610 and the secondary database 620 takes over as the new primary database. In this example, a recovery operation is performed on the primary database 610 to set it as the new secondary database. The recovery operation can include, as a non-limiting example, executing a previously replicated REDO log.

[0039] Next, referring to FIG. 7, a process for dynamically determining a load unit for loading a database object and / or a database object subcomponent into memory is shown. At the start of method 700, an operation to start loading a column into memory is shown (block 705). For example, when a query targeting a column is received, loading of the column into memory may be started. Next, the database determines whether the column needs to be paged (conditional block 710). If the column needs to be paged (conditional block 710, "yes" leg), a paged data structure is created and the relevant pages of the column are loaded into memory (block 715). If the column does not need to be paged (conditional block 710, "no" leg), the entire column is loaded into memory and no paged data structure is created (block 720). After blocks 715 and 720, method 700 may end.

[0040] Next, referring to FIG. 8, a process for using different load units when loading database objects into memory in a primary database system and a secondary database system is shown. At the start of method 800, at least one database object is loaded into a primary in-memory store in the primary database system according to a first format (block 805). As an example, the first format is a column loadable format. The term "primary in-memory store" refers to an in-memory store in the primary database system. The primary database system refers to the main or active database system that is actively used by one or more users, organizations, or other entities. As an example, the primary database system replicates data to a secondary database system or an inactive database system, and the secondary database system is intended to take over the primary database system in the event of a failure or other event or situation.

[0041] Next, the operation of loading at least one database object into the primary in-memory store according to the first format is captured in the first log (block 810). Next, the first log is sent to the secondary database system (block 815). In response to receiving the first log, the secondary database system replays the first log (block 820). When replaying the first log, the secondary database system checks the value of the log replay configuration parameter (conditional block 825). If the log replay configuration parameter is the first value (conditional block 825, "first" leg), the secondary database system loads at least one database object into the secondary in-memory store according to the second format, where the second format is different from the first format (block 830). The term "secondary in-memory store" refers to the in-memory store in the secondary database system. After block 830, method 800 may end. As an example, the second format is a page-loadable format. Otherwise, if the log replay configuration parameter is the second value (conditional block 825, "second" leg), the secondary database system loads at least one database object into the secondary in-memory store according to the first format (block 835). After block 835, method 800 may end.

[0042] Next, referring to FIG. 9, a process for executing a failover load script in a secondary database system is shown. At the start of method 900, a failover state is detected in the primary database system (block 905). In response to detecting the failover state, the secondary database system executes the failover load script (block 910). As an example, the failover load script may be integrated within an overall failover script executed by the secondary database system. In this example, the failover load script is executed only if a reset-to-primary configuration parameter is set.

[0043] As part of the execution of the failover load script, the secondary database system identifies one or more database objects to be loaded into the in-memory store according to a first format (block 915). By way of example, the first format is a page-loadable format. Next, the secondary database system unloads the one or more database objects (block 920). Then, the secondary database system reloads one or more database objects corresponding to the primary-defined load units into the in-memory store according to a second format (block 925). By way of example, the second format is a column-loadable format. After block 925, method 900 ends.

[0044] In some implementations, the subject matter of the present invention may be configured to be implemented in a system 1000, as shown in FIG. 10A. The system 1000 may include a processor 1010, a memory 1020, a storage device 1030, and an input / output device 1040. Each of the components (e.g., 1010, 1020, 1030, and 1040) may be interconnected using a system bus 1050. The processor 1010 may be configured to process instructions for execution within the system 1000. In some implementations, the processor 1010 may be a single-threaded processor. In alternative implementations, the processor 1010 may be a multi-threaded processor. The processor 1010 may be further configured to process instructions stored in the memory 1020 or the storage device 1030, including receiving or transmitting information through the input / output device 1040. The memory 1020 may store information within the system 1000. In some implementations, the memory 1020 may be a computer-readable medium. In alternative implementations, the memory 1020 may be a volatile memory unit. In still some implementations, the memory 1020 may be a non-volatile memory unit. The storage device 1030 may be capable of providing mass storage for the system 1000. In some implementations, the storage device 1030 may be a computer-readable medium. In alternative implementations, the storage device 1030 may be a floppy disk device, a hard disk device, an optical disk device, a tape device, a non-volatile solid state memory, or any other type of storage device. The input / output device 1040 may be configured to provide input / output operations for the system 1000. In some implementations, the input / output device 1040 may include a keyboard and / or a pointing device. In alternative implementations, the input / output device 1040 may include a display unit for displaying a graphical user interface.

[0045] FIG. 10B shows an exemplary implementation of a database 140 that provides database services. The database 140 may include physical resources 1080 such as at least one hardware server, at least one storage, at least one memory, at least one network interface, etc. The database 140 may also include an infrastructure that may include at least one operating system 1082 for the physical resources and at least one hypervisor 1084 (which may create and run at least one virtual machine 1086). For example, each multi-tenant application may run on a corresponding virtual machine.

[0046] Next, referring to FIG. 11, a process for determining a load unit when loading a database object into memory in a secondary database system is shown. At the start of method 1100, a first log is sent from the primary database system to the secondary database system, and the first log captures the loading of at least one database object into the primary in-memory store in the primary database system (block 1105). In response to receiving the first log, the secondary database system replays the first log (block 1110). When replaying the first log, the secondary database system checks the value of the log replay configuration parameter (conditional block 1115). If the log replay configuration parameter is a first value (conditional block 1115, "first" leg), the secondary database system loads at least one database object into the secondary in-memory store according to a first format (block 1120). As an example, the first format is a page-loadable format, and when the log replay configuration parameter is set to the first value, it specifies or indicates that the database object needs to be loaded in a page-loadable format. In this case, the secondary database system can override the determination of the load unit made in the primary database system to load at least one database object into the primary in-memory store. After block 1120, method 1100 may end.

[0047] When the log playback configuration parameter is a second value (conditional block 1115, "second" leg), the secondary database system loads at least one database object into the secondary in-memory store according to a second format (block 1125). It should be understood that the second value is different from the first value, and the second value can be indicated, specified, or encoded in any suitable way to indicate that it is different from the first value. By way of example, the second format is a column-loadable format, and when the log playback configuration parameter is set to the second value, it is specified or indicated that the database object needs to be loaded in a column-loadable format. Similarly, in this case, the secondary database system can also override any load unit determination made in the primary database system to load at least one database object into the primary in-memory store. In other words, regardless of the load unit determination made in the primary database system to load at least one database object into the primary in-memory store, because the log playback configuration parameter has a second value, the secondary database system loads at least one database object into the secondary in-memory store according to the second format. After block 1125, method 1100 can end.

[0048] When the log playback configuration parameter is a third value (conditional block 1115, "third" leg), the secondary database system loads at least one database object into the secondary in-memory store according to the same format used by the primary database system to load at least one database object into the primary in-memory store (block 1130). As an example, in block 1130, if at least one database object was loaded into the primary in-memory store in a page-loadable format, the secondary database system loads at least one database object into the secondary in-memory store in a page-loadable format. In this example, in block 1130, if at least one database object was loaded into the primary in-memory store in a column-loadable format, the secondary database system loads at least one database object into the secondary in-memory store in a column-loadable format. Note that the third value is different from the first and second values. In some cases, the third value may be a default value or be called "default", and the secondary database system is default-configured to the same load unit as the primary database system. In other words, when the log playback configuration parameter is set to default, the secondary uses the load unit format that matches what was used in the primary. After block 1130, method 1100 may end.

[0049] Next, referring to FIG. 12, a process for adjusting the buffer cache size in a secondary database system based at least on log playback configuration parameters is shown. When the secondary database system replays the logs generated in the primary database system, it determines which playback load unit is used to load database objects into the in-memory store in the secondary database system (block 1205). Also, the secondary database system determines buffer cache size configuration parameters based at least on the playback load unit (block 1210). As an example, when the playback load unit is a page-loadable load unit, the buffer cache size configuration parameter is set to a maximum value. When the playback load unit is set to be page-loadable (conditional block 1215, "page-loadable" leg), the secondary database system increases the size of the buffer cache (e.g., buffer cache 233 in FIG. 2) according to the buffer cache size configuration parameter (block 1220). Alternatively, in block 1220, the secondary database system may set the size of the buffer cache to a first capacity that is a relatively high capacity. Otherwise, when the playback load unit is set to be column-loadable (conditional block 1215, "column-loadable" leg), the secondary database system decreases the size of the buffer cache according to the buffer cache size configuration parameter (block 1225). Alternatively, in block 1225, the secondary database system may set the size of the buffer cache to a second capacity that is a relatively low capacity, and the second capacity is smaller than the first capacity used in block 1220. After blocks 1220 and 1225, method 1200 may end.

[0050] The systems and methods disclosed herein can be embodied in a variety of forms including, for example, a data processor such as a computer including databases, digital electronic circuits, firmware, software, or combinations thereof. Further, the above features as well as other aspects and principles of the implementations of the present disclosure can be implemented in a variety of environments. Such environments and related applications can also be specially constructed to perform various processes and operations in accordance with the disclosed implementations, or can include a general-purpose computer or computing platform that is selectively activated or reconfigured by code to provide the required functionality. The processes disclosed herein are not inherently related to any particular computer, network, architecture, environment, or other apparatus, and can be implemented by a suitable combination of hardware, software, and / or firmware. For example, various general-purpose machines can be used with programs written in accordance with the teachings of the disclosed implementations, although in some cases it may be convenient to construct a dedicated device or system to perform the required methods and techniques.

[0051] Ordinal numbers such as first, second, etc. are, in some cases, related to order, but when used in a document, ordinal numbers do not necessarily mean order. For example, ordinal numbers can be used simply to distinguish one item from another. For example, a first event is distinguished from a second event, but it is not necessary to imply a chronological order or a fixed reference system (thus, the first event in one paragraph of the description may be different from the first event in another paragraph of the description).

[0052] The foregoing description is not intended to limit the scope of the invention, but rather to describe the scope of the invention, which is defined by the appended claims. Other implementations are within the scope of the following claims.

[0053] These computer programs, also referred to as programs, software, software applications, applications, components, or code, contain program instructions (i.e., machine instructions) for a programmable processor and can be implemented in high-level procedural programming languages and / or object-oriented programming languages, and / or assembly language / machine language. As used herein, the term "machine-readable medium" refers to any computer program product, apparatus, and / or device, such as a magnetic disk, optical disk, memory, programmable logic device (PLD), etc., used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the program instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium can store such program instructions non-transitorily, for example, in a non-transitory solid-state memory or a magnetic hard drive, or any equivalent storage medium. Alternatively, or additionally, a machine-readable medium can store such machine instructions in a transient manner, similar to a processor cache or other random access memory associated with one or more physical processor cores.

[0054] To provide interaction with a user, the subject matter described herein can be implemented on a computer having, for example, a display device such as a cathode ray tube (CRT) or liquid crystal display (LCD) monitor for displaying information to the user, a keyboard by which the user can provide input to the computer, and a pointing device such as a mouse or trackball. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback such as, for example, visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form such as acoustic, voice, or tactile input.

[0055] The subject matter described herein can be implemented in a computing system that includes back-end components such as, for example, one or more data servers, or in a computing system that includes middleware components such as, for example, one or more application servers, or in a computing system that includes front-end components such as, for example, one or more client computers having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described herein, or in any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication such as, for example, a communication network. Examples of communication networks include, but are not limited to, local area networks ("LANs"), wide area networks ("WANs"), and the Internet.

[0056] A computing system can include a client and a server. The client and the server are generally separated from each other, but not necessarily so, and typically interact via a communication network. The relationship between the client and the server is created by computer programs that are executed on respective computers and have a client-server relationship with each other.

[0057] In the above description and claims, a concatenated list of elements or features may be used following phrases such as "at least one" or "one or more". The term "and / or" may also occur in a list of two or more elements or functions. Such phrases are intended to mean either individually any of the listed elements or functions, or in combination any of the listed elements or functions with any of the other listed elements or functions, unless implicitly or explicitly inconsistent with the context in which they are used. For example, the phrases "at least one of A and B", "one or more of A and B", and "A and / or B" are each intended to mean "only A, only B, or both A and B". A similar interpretation is intended for lists containing three or more items. For example, the phrases "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, and / or C" are each intended to mean "only A, only B, only C, both A and B, both A and C, both B and C, or all of A, B, and C". The use of the term "based on" in the above and in the claims is intended to mean "at least partially based on", and features or elements not recited are also permitted.

[0058] In view of implementations of the above-described subject matter, the present application discloses the following list of examples. Here, one feature of an example, alone or in combination with a plurality of features of said example, optionally in combination with one or more features of one or more additional examples, results in a further example that is within the scope of the disclosure of the present application.

[0059] Example 1: A method including the steps of receiving, by a secondary database system, data replicated from a primary database system; receiving, by the secondary database system, a first log capturing at least one database object loaded into a primary in-memory store in the primary database system according to a first format; determining a value of a log replay configuration parameter; and in response to the log replay configuration parameter having a first value, replaying the first log on the secondary database system to load at least one database object into a secondary in-memory store according to a second format.

[0060] Example 2: The method according to Example 1, further including the steps of detecting a failover state; in response to detecting the failover state, starting a first failover reload script in the secondary database system; unloading at least one database object from the secondary in-memory store; and reloading at least one database object into the secondary in-memory store according to the first format.

[0061] Example 3: The method according to any one of Examples 1 to 2, wherein the second format is different from the first format.

[0062] Example 4: The method according to any one of Examples 1 to 3, wherein the first format is a format loadable by columns.

[0063] Example 5: The method according to any one of Examples 1 to 4, wherein the second format is a format loadable by pages.

[0064] Example 6: The method according to any one of Examples 1 to 5, wherein in the case of a format loadable by pages, only the page related to the corresponding query is loaded into the in-memory store.

[0065] Example 7: The method according to any one of Examples 1 to 6, wherein when in a column loadable format, the entire corresponding database object is loaded into the in-memory store.

[0066] Example 8: The method according to any one of Examples 1 to 7, wherein the column loadable format includes having data sub-components, dictionary sub-components, and index sub-components of the corresponding database object serialized in a single page chain, and the page loadable format includes having individual page chains for each of the data sub-components, dictionary sub-components, and index sub-components of the first database object.

[0067] Example 9: The method according to any one of Examples 1 to 8, further comprising the step of replaying a first log on a secondary database system and loading at least one database object into a secondary in-memory store according to a first format in response to the log replay configuration parameter having a second value.

[0068] Example 10: The method according to any one of Examples 1 to 9, further comprising the step of replaying a second log on a secondary database system and loading at least a second database object into a secondary in-memory store according to the same format as used in the primary in-memory store in response to the log replay configuration parameter having a third value.

[0069] Example 11: A system comprising at least one processor and at least one memory containing program instructions, which, when executed by the at least one processor, cause the secondary database system to receive data replicated from the primary database system, receive a first log captured by the secondary database system of at least one database object loaded into the primary in-memory store in the primary database system according to a first format, determine a value of a log replay configuration parameter, and in response to the log replay configuration parameter having a first value, replay the first log on the secondary database system to load at least one database object into the secondary in-memory store according to a second format.

[0070] Example 12: The system according to Example 11, wherein when the program instructions are executed by the at least one processor, the system further performs operations including detecting a failover state, starting a first failover reload script in the secondary database system in response to detecting the failover state, unloading at least one database object from the secondary in-memory store, and reloading at least one database object into the secondary in-memory store according to the first format.

[0071] Example 13: The system according to any one of Examples 11 to 12, wherein the second format is different from the first format.

[0072] Example 14: The system according to any one of Examples 11 to 13, wherein the first format is a column loadable format.

[0073] Example 15: The system according to any one of Examples 11 to 14, wherein the second format is a page loadable format.

[0074] Example 16: The system according to any one of Examples 11 to 15, wherein when in a page-loadable format, only the page related to the corresponding query is loaded into the in-memory store.

[0075] Example 17: The system according to any one of Examples 11 to 16, wherein when in a column-loadable format, the entire corresponding database object is loaded into the in-memory store.

[0076] Example 18: The system according to any one of Examples 11 to 17, including that the column-loadable format has data sub-components, dictionary sub-components, and index sub-components of the corresponding database object serialized in a single page chain.

[0077] Example 19: The system according to any one of Examples 11 to 18, including that the page-loadable format has individual page chains for each of the data sub-components, dictionary sub-components, and index sub-components of the first database object.

[0078] Example 20: A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause the secondary database system to receive data replicated from the primary database system, receive a first log captured by the secondary database system of at least one database object loaded into the primary in-memory store in the primary database system according to a first format, determine a value of a log replay configuration parameter, and in response to the log replay configuration parameter having a first value, replay the first log on the secondary database system to load at least one database object into the secondary in-memory store according to a second format.

[0079] The implementation forms described in the foregoing description do not represent all implementation forms that conform to the subject matter described in this specification. Instead, the implementation forms described in the foregoing description are merely some examples that conform to aspects related to the described subject matter. Although several modification examples have been described in detail above, other modifications or additions are also possible. In particular, in addition to what is described in this specification, further features and / or modification examples can be provided. For example, the above-described implementation forms can be directed to various combinations and sub-combinations of the disclosed features, and / or combinations and sub-combinations of some further features disclosed above. Furthermore, the logic flows depicted in the accompanying drawings and / or described in this specification do not necessarily require the specific order or sequential order shown in order to achieve the desired results. Other implementation forms may be included within the scope of the following claims.

Explanation of Signs

[0080] 100 Database System 110 Client Device 120 Network 130 Database Execution Engine 140 Database 150 Table 200 System 210 In-Memory Database 225 Data Store 230 Main Storage 233 Buffer Cache 235 Delta Storage 240 A-N Application 245 Platform 250 Data Management Service 260 Database Management Service 270 User Interface (UI) 280 Persistent Storage 300 System 310 In-Memory Load Attribute 315 Page Load Attributes 320 Load Unit Determination 325 Persistent Store 330 Data Page Chain 340 Dictionary Page Chain 350 Index Page Chain 400 System 410 Persistent Store 420 Data Page Chain 430 Dictionary Page Chain 440 Index Page Chain 450 Load Unit Determination 460 Page Load Attributes 470 In-Memory Load Attributes 480 Delta Component 485 Delta Component 500 System 510 Primary Database 520 Secondary Database 600 System 610 Primary Database 620 Secondary Database 1000 System 1010 Processor 1020 Memory 1030 Storage Device 1040 Input / Output Device 1050 System Bus 1080 Physical Resource 1082 Operating System 1084 Hypervisor 1086 Virtual Machine

Claims

1. receiving, by a secondary database system, the replicated data from the primary database system; receiving, by the secondary database system, a first log capturing at least one database object that has been loaded into a primary in-memory store at the primary database system according to a first format; determining values ​​for log replay configuration parameters; in response to the log replay configuration parameter having a first value, replaying the first log on the secondary database system to load the at least one database object into a secondary in-memory store according to a second format; A method comprising:

2. detecting a failover condition; initiating a first failover reload script at the secondary database system in response to detecting the failover condition; unloading the at least one database object from the secondary in-memory store; reloading the at least one database object into the secondary in-memory store according to the first format; The method of claim 1, further comprising:

3. The method of claim 2 , wherein the second format is different from the first format.

4. The method of claim 1 , wherein the first format is a column-loadable format.

5. The method of claim 4 , wherein the second format is a page-loadable format.

6. The method of claim 5 , wherein for the page loadable format, only pages relevant to a corresponding query are loaded into the in-memory store.

7. The method of claim 6 , wherein for the column loadable format, the entire corresponding database object is loaded into the in-memory store.

8. 8. The method of claim 7, wherein the column loadable format includes having a data subcomponent, a dictionary subcomponent, and an index subcomponent of a corresponding database object serialized into a single page chain, and the page loadable format includes having a separate page chain for each of the data subcomponent, the dictionary subcomponent, and the index subcomponent of a first database object.

9. 2. The method of claim 1, further comprising, in response to the log replay configuration parameter having a second value, replaying the first log on the secondary database system to load the at least one database object into the secondary in-memory store according to the first format.

10. 2. The method of claim 1, further comprising, in response to the log replay configuration parameter having a third value, replaying a second log on the secondary database system to load at least a second database object into the secondary in-memory store according to a same format used in the primary in-memory store.

11. At least one processor; at least one memory having program instructions stored therein; Equipped with The program instructions, when executed by the at least one processor, receiving, by a secondary database system, replicated data from the primary database system; receiving, by the secondary database system, a first log capturing at least one database object that has been loaded into a primary in-memory store at the primary database system according to a first format; Determining values ​​for log replay configuration parameters; in response to the log replay configuration parameter having a first value, replaying the first log on the secondary database system to load the at least one database object into a secondary in-memory store according to a second format; A system that performs an operation including

12. The program instructions, when executed by the at least one processor, Detecting a failover condition; initiating a first failover reload script at the secondary database system in response to detecting the failover condition; unloading the at least one database object from the secondary in-memory store; reloading the at least one database object into the secondary in-memory store according to the first format; The system of claim 11 , further comprising:

13. The system of claim 11 , wherein the second format is different from the first format.

14. The system of claim 11 , wherein the first format is a column-loadable format.

15. 15. The system of claim 14, wherein the second format is a page-loadable format.

16. 16. The system of claim 15, wherein for the page loadable format, only pages relevant to a corresponding query are loaded into the in-memory store.

17. 17. The system of claim 16, wherein for the column loadable format, the entire corresponding database object is loaded into the in-memory store.

18. 20. The system of claim 17, wherein the column loadable format includes having a data subcomponent, a dictionary subcomponent, and an index subcomponent of a corresponding database object serialized into a single page chain.

19. 20. The system of claim 18, wherein the page loadable format includes having a separate page chain for each of the data subcomponent, the dictionary subcomponent, and the index subcomponent of a first database object.

20. A non-transitory computer-readable storage medium having instructions stored thereon, comprising: The instructions, when executed by at least one data processor, receiving, by a secondary database system, replicated data from the primary database system; receiving, by the secondary database system, a first log capturing at least one database object that has been loaded into a primary in-memory store at the primary database system according to a first format; Determining values ​​for log replay configuration parameters; in response to the log replay configuration parameter having a first value, replaying the first log on the secondary database system to load the at least one database object into a secondary in-memory store according to a second format; A non-transitory computer-readable storage medium that causes operations to be performed, including:

Citation Information

Patent Citations

  • Parallel Replication Across Formats

    US20190325055A1

  • Metadata converter and memory management system

    US20210311949A1