Operating the data center
The described system addresses data center failures by offloading queries to a secondary database system, ensuring minimal downtime and automated recovery, thus enhancing data center resilience and reducing manual intervention.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- INTERNATIONAL BUSINESS MACHINE CORPORATION
- Filing Date
- 2022-12-07
- Publication Date
- 2026-04-17
AI Technical Summary
Existing data synchronization and recovery methods in data centers do not effectively address data recovery when synchronization occurs between peripheral devices, leading to prolonged downtime and manual intervention during failures.
A system and method for a data center that includes a primary and secondary database system with synchronized data, where queries are offloaded to the secondary system upon primary failure, enabling seamless data replication and query acceleration with minimal downtime and manual intervention.
Minimizes downtime and eliminates manual intervention in data center failures by ensuring high availability and rapid recovery through automated query offloading and replication, reducing Recovery Time Objective (RTO) and maintaining continuous data processing.
Smart Images

Figure 0007847648000001 
Figure 0007847648000002 
Figure 0007847648000003
Abstract
Description
Technical Field
[0001] The present invention relates to digital computer systems, and more particularly to an approach for operating a data center.
Background Art
[0002] Replication is the process of maintaining a defined set of data at two or more locations. Replication may involve copying specified changes from one source location to a target location and synchronizing the data at both locations. The source and target can be within the same machine or different machines in a distributed network within a logical server. There are several approaches for processing data and moving it from one system to another.
[0003] U.S. Patent Application Publication No. 2019 / 0391740(A1) states the following: "The computing system includes a first storage unit located at a first computing site. The first storage unit stores work data and units of data synchronously replicated from a first server cluster at a second computing site. The computing system further includes a second server cluster located at the first computing site, which is a proxy node of the first server cluster. The computing system further includes a second storage unit located at the first computing site, which asynchronously stores work data and units of data from the first storage unit in the second storage unit. The computing system further includes a third server cluster located at the first computing site, which processes units of work data asynchronously replicated to the second storage unit." (Abstract, U.S. Patent Application Publication 2019 / 0391740(A1)). However, such an approach does not address data recovery in the specific environment described in the embodiments of the present invention when data synchronization occurs between peripheral devices. [Overview of the project]
[0004] Various embodiments of methods, computer systems, and computer program products for operating a data center, as described by the subject matter of the independent claims, are provided. Preferred embodiments are described in the dependent claims. The embodiments of the present invention can be freely combined with one another, provided they are not mutually exclusive.
[0005] In one embodiment, an embodiment of the present invention relates to a computer execution method and includes providing a primary data center comprising a primary source database system and a primary target database system. Activation of functions in the primary data center causes the primary target database system to have a copy of the data from the primary source database system, to receive parsing queries from the primary source database system, and to execute said parsing queries on the data. In response to detecting a failure in the primary source database system, the processor offloads queries targeting the primary source database system to a secondary source database system, the secondary source database system of the secondary data center further comprises a secondary target database system, the secondary source database system has a copy of the data, and the above functions are deactivated in the secondary data center. The processor, in response to the primary target database system being available, receives parsed queries processed by the secondary source database system for queries offloaded by the primary target database system, and copies the data to the secondary target database system. The processor then ensures that the functionality is activated in the secondary data center. This approach has the advantage of minimizing downtime and reducing or eliminating manual intervention to reactivate replication functionality.
[0006] In another embodiment, embodiments of the present invention relate to a computer program product comprising one or more computer-readable storage media and program instructions stored collectively in one or more computer-readable storage media, wherein the program instructions are program instructions for providing a primary data center, the primary data center comprising a primary source database system and a primary target database system, the primary data center having a copy of the data of the primary source database system, receiving parsing queries from the primary source database system, and causing the primary target database system to execute parsing queries on the data, and the program instructions for providing a primary data center are activated in the primary data center. In response to detecting a failure in the primary source database system, the processor further includes program instructions for offloading queries targeting the primary source database system to a secondary source database system, where the secondary source database system in the secondary data center further includes a secondary target database system, and the secondary source database system has a copy of the data, and the functionality is deactivated in the secondary data center. In response to the primary target database system being available, the primary target database system further includes program instructions for receiving the parsed queries processed by the secondary source database system for the offloaded queries, and the primary target database system further includes program instructions for copying the data to the secondary target database system. The functionality is further activated in the secondary data center. Such an approach has the advantage of minimizing downtime and reducing or eliminating manual intervention to restore replication functionality.
[0007] In another embodiment, embodiments of the present invention relate to a computer system comprising one or more computer processors, one or more computer-readable storage media, and program instructions collectively stored in one or more computer-readable storage media for execution by at least one of the one or more computer processors, wherein the program instructions are program instructions for providing a primary data center, the primary data center comprising a primary source database system and a primary target database system, the primary data center having a copy of the data of the primary source database system, receiving parsing queries from the primary source database system, and executing parsing queries on the data, the functions of which are activated in the primary data center. In response to detecting a failure in the primary source database system, the system offloads queries targeting the primary source database system to a secondary source database system, further including program instructions for the secondary source database system in a secondary data center to perform the offload, which further includes a secondary target database system, the secondary source database system having a copy of the data, and the functionality being deactivated in the secondary data center. In response to the primary target database system being available, the system further includes program instructions for the primary target database system to receive the parsed queries processed by the secondary source database system for the offloaded queries, and for the primary target database system to copy the data to the secondary target database system. The system further includes program instructions for ensuring that the functionality is activated in the secondary data center. Such an approach has the advantage of minimizing downtime and reducing or eliminating manual intervention to restore replication functionality.
[0008] Embodiments of the present invention may further include transmitting log locations from the primary target database system to the secondary target database system until the data is replicated from the primary source database system to the primary target database system. Here, the log locations are used to replicate any further changes that occur in the secondary source database system after a failure. Such an approach makes it possible to use log locations to replicate any further changes that occur after a failure.
[0009] Hereinafter, embodiments of the present invention will be described in more detail, merely as examples, with reference to the drawings. [Brief explanation of the drawing]
[0010] [Figure 1] This is a diagram of a data center according to an embodiment of the present invention. [Figure 2] This is a diagram of a computer system according to an embodiment of the present invention. [Figure 3] This is a flowchart of an approach to operating a data center according to an embodiment of the present invention. [Figure 4] This is a flowchart of the recovery approach according to an embodiment of the present invention. [Modes for carrying out the invention]
[0011] Various embodiments of the present invention will be presented for illustrative purposes, but are not intended to be exhaustive or limit to the embodiments disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the embodiments described. The terminology used herein has been chosen to best describe the principles, practical applications, or technical improvements to the technologies available on the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
[0012] A data center may be a data processing system. As used herein, a data center may refer to a primary data center or a secondary data center. The source and target database systems of a data center may initially have the same dataset. Subsequently, the dataset may be modified in the source database system. The data center may be configured to replicate the changes from the source database system to the target database system, so that the target database system applies the same changes and therefore has the same contents as the source database system. The data center may enable such replication, for example, by activating the replication function. That is, if the replication function is deactivated, the data center cannot replicate changes made in the source database system. The data center may enable a hybrid transaction and analytic processing (HTAP) environment to process data in a database (e.g., Db2) using different types of queries. The source database system in the data center may enable transactional queries, and the target database system (also called an accelerator) may enable the execution of complex queries, for example, complex queries may include costly SQL operations such as grouping and aggregation. The source database system may identify the complex queries and forward them to the target database system. This execution of complex queries may be enabled by activating acceleration features in the data center. That is, if the acceleration features are deactivated, the source database system may not be able to forward complex queries to the target database system and therefore may not be executed. The data center may allow both features to be activated and deactivated.For example, the replication acceleration function may include the replication function and the acceleration function, such that activating the replication acceleration function includes activating the replication function and the acceleration function, and deactivating the replication acceleration function includes deactivating the replication function and the acceleration function. When the replication acceleration function is activated, the data center is considered active.
[0013] The source database system may be, for example, a transaction engine. The target database system may be, for example, an analysis engine. In certain combinations, such as those implemented by the IBM Db2 Analytic Accelerator for z / OS data center, the source database system may be a relational DBMS optimized for OLTP (row-first storage), and the target database system may be a relational DBMS optimized for analysis (column-first storage). However, this subject is not limited to the combination of OLTP (online transaction processing) and OLAP (online analytical processing). Other combinations may include OLTP / graph store or OLAP / key-value store. The processing resources of the source database system may be less than those of the target database system. The source database system may be designed with an emphasis on high-speed processing because the source database may be frequently read, written to, and updated. The target database system can enable complex queries on large amounts of data and therefore may have more processing resources than the source database system. A combination of relational database systems may be advantageous for executing various types of queries, such as HTAP queries. The system can provide optimal response times for both parsing and transactional queries on the same data, without the need to copy data or convert it to a schema best suited to the use case.
[0014] Following a primary data center failure, this approach can provide an optimal disaster recovery approach to enable the recovery or continuation of data processing. This subject enables a reactive and proactive process in cases where a sudden disaster damages or renders the primary center's source database system inoperable. This can improve the Recovery Time Objective (RTO), which represents the amount of time an application can be down without causing significant business damage. In practice, data center accelerators can typically run two types of application workloads: reporting applications for internal enterprise use and external customer applications that directly interact with clients and generate revenue. Both types of applications may be high-priority applications and, due to their direct customer interaction, may require extremely low RTOs (nearly zero RTOs). This means that long recovery times can cause serious financial impacts for the enterprise and dissatisfaction with its customers. To address this, this approach can utilize redundant systems and software. This ensures high availability, prevents downtime and data loss, and eliminates single points of failure. In particular, embodiments of this approach provide a secondary data center that can be used as a fallback solution in the event of a primary data center failure. Embodiments of this approach may be implemented in actual disaster recovery situations or in annual disaster recovery testing scenarios.
[0015] The combination of a primary and secondary data center can provide a passive-standby system architecture, as the secondary data center may be in passive and standby mode while the primary data center is operational. In practice, the hardware and software in the secondary data center are installed and ready for use, but it is not operational or performing active tasks while the primary data center is operational. The secondary data center is activated only in disaster / failover situations and can then replace the previously active primary data center.
[0016] Embodiments of this subject matter can further improve disaster recovery approaches by further reducing recovery time. For example, this approach can shorten the time between the shutdown of the source database in the primary center and the availability of the secondary accelerator in the secondary center for query acceleration. Embodiments of this subject matter can enable seamless handover of accelerated workloads (e.g., SQL workloads) and continuous replication of accelerator shadow tables in disaster recovery situations. This can minimize accelerator downtime to near-zero RTO and completely eliminate manual intervention to restore replication functionality in disaster recovery situations. For example, embodiments of this approach can use predefined stored procedures to prevent data loading from Db2 in the secondary data center to the secondary accelerator. This is advantageous because procedures can be limited to run within specific timeframes to ensure the normal recovery of the equipment and can become complex enough that errors could directly affect subsequent procedures and extend the overall machine downtime.
[0017] According to one embodiment, the approach further includes setting the primary target database system to read-only mode before executing parsing queries on the primary target database system. The primary target database system may be set to read-only mode at least to provide query execution services. The primary target database system may be in read-only mode while copying data from the primary target database system to the secondary target database system in order to enable query acceleration on the secondary target database system. For example, the primary accelerator may be used as a temporary measure for query acceleration until the secondary accelerator is fully caught up and reloaded from a newly started Db2 at the secondary center. Since the secondary accelerator may need to be resynchronized with all changes from the primary accelerator, using read-only mode can significantly reduce recovery time for query acceleration.
[0018] According to one embodiment, the approach further includes receiving one or more log locations in a secondary target database system up to the time when the data was replicated from a primary source database system to a primary target database system, and using the log locations to replicate any further changes that occurred in the secondary source database system after a failure.
[0019] According to one embodiment, the failure is a failure of the disk subsystem of the primary source database system. The failure may cause a partial or complete shutdown of at least a portion of the primary source database system (e.g., storage).
[0020] According to one embodiment, the primary source database system is configured to actively mirror the disk subsystem of the primary source database system to the disk subsystem of the secondary source database system so that the secondary source database system has a copy of the up-to-date source data.
[0021] According to one embodiment, copying is performed in parallel with or simultaneously with executing an analytic query. Thereby, the time to prepare a second accelerator for query acceleration can be further reduced compared to the case where the copy is performed after the analytic query is executed.
[0022] FIG. 1 is a block diagram of a suitable data center 100 according to an example of the present subject matter. The data center 100 may include, for example, an IBM Db2 Analytic Accelerator for z / OS (IDAA). The data center 100 includes a source database system 101 connected to a target database system 121. The source database system 101 may include, for example, an IBM Db2 for z / OS. The target database system 121 may include, for example, an IBM Db2 Warehouse (Db2 LUW).
[0023] The source database system 101 includes a processor 102, a memory 103, I / O circuitry 104, and a network interface 105 coupled by a bus 106.
[0024] The processor 102 may represent one or more processors (e.g., microprocessors). The memory 103 may include one or a combination of volatile memory elements (e.g., random access memory (RAM such as DRAM, SRAM, SDRAM, etc.)) and non-volatile memory elements (e.g., ROM, erasable programmable read-only memory (EPROM), electronically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM)). Note that the memory 103 may have a distributed architecture in which various components are located far apart from each other but are accessible by the processor 102.
[0025] Memory 103 may be used in combination with a persistent storage device 107 for storing local data and instructions. The storage device 107 includes one or more persistent storage devices and media controlled by an I / O circuit device 104. The storage device 107 may include, for example, magnetic, optical, magneto-optical, or solid-state devices for storing digital data, having fixed or removable media. Sample devices include hard disk drives, optical disk drives, and floppy(R) disk drives. Sample media include hard disk platters, CD-ROMs, DVD-ROMs, BD-ROMs, floppy(R) disks, and similar media. Storage 107 may include a source database 112. The source database 112 may include, for example, a source table 190. The source table 190 may include att1, ...att n It may have a set of attributes (columns) named as follows.
[0026] Memory 103 may include one or more separate programs, such as database management system DBMS1 109, each of which comprises an ordered list of executable instructions for performing logical functions, such as those particularly included in embodiments of the present invention. The software within memory 103 also typically includes a suitable operating system (OS) 108. OS 108 essentially controls the execution of other computer programs for performing at least part of the method as described herein. DBMS1 109 includes a log reader 111 and a query optimizer 110. Log reader 111 may read log records 180 (not shown) of the transaction recovery log of source database system 101 and provide the modified records to target database system 121. Log reader 111 may read log records from the recovery log and extract relevant modification or change information (insert / update / delete target tables during replication). The extracted information may be transmitted to target database system 121 (e.g., as a request to apply the changes). Query optimizer 110 may be configured to generate or define a query plan for executing a query, e.g., on source database 112.
[0027] Target database system 121 includes a processor 122, a memory 123, an I / O circuitry 124, and a network interface 125 coupled by a bus 126.
[0028] The processor 122 may represent one or more processors (e.g., microprocessors). The memory 123 may include one or a combination of volatile memory elements (e.g., random access memory (RAM such as DRAM, SRAM, SDRAM, etc.)) and non-volatile memory elements (e.g., ROM, erasable programmable read-only memory (EPROM), electronically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM)). Note that the memory 123 may have a distributed architecture in which various components are located far apart from each other but are accessible by the processor 122.
[0029] Memory 123 may be used in combination with a persistent storage device 127 for storing local data and instructions. The storage device 127 includes one or more persistent storage devices and media controlled by the I / O circuit equipment 124. The storage device 127 may include, for example, magnetic, optical, magneto-optical, or solid-state devices for storing digital data, having fixed or removable media. Sample devices include hard disk drives, optical disk drives, and floppy(R) disk drives. Sample media include hard disk platters, CD-ROMs, DVD-ROMs, BD-ROMs, floppy(R) disks, and similar media.
[0030] Memory 123 may contain one or more separate programs, such as a database management system DBMS2 129 and an application component 155, each comprising an ordered list of executable instructions for performing logical functions, particularly those included in embodiments of the present invention. The software in memory 123 also typically includes a suitable OS 128. The OS 128 essentially controls the execution of other computer programs for performing at least a portion of the methods described herein. DBMS2 129 comprises a DB application 131 and a query optimizer 130. The DB application 131 may be configured to process data stored in a storage device 127. The query optimizer 130 may be configured to generate or define a query plan for executing queries on, for example, a target database 132. The application component 155 can buffer log records sent from the log reader 111 and consolidate changes into batches to improve efficiency when applying modifications to the target database 132 via a bulk load interface. This enables replication. To maintain stable latency, replication can be advantageous if it can keep up with the amount of modifications. If modifications exceed the replication speed, latency can accumulate and become excessively high. For this reason, the source database system 101 may be configured to perform bulk loads. A bulk load can load the entire table data at a specific point in time, or load a set of partitions of a table. The data on the target database system 121 will also reflect the state of the source database system at the time the load is performed.
[0031] The source database system 101 and the target database system 121 may be independent computer hardware platforms communicating via a high-speed connection 142 or a network 141 via network interfaces 105, 125. The network 141 may include, for example, a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof. Each of the source and target database systems 101 and 121 may be responsible for managing its own copy of the data.
[0032] Although shown as separate systems in Figure 1, the source database system and the target database system may belong to a single system, for example, sharing the same memory and processor hardware, while each of the source and target database systems may be associated with its respective DBMS and dataset, for example, the two DBMSs may be stored in shared memory. In another example, two database management systems, DBMS1 and DBMS2, may form part of a single DBMS that enables the communication and methods performed by DBMS1 and DBMS2 as described herein. The first and second datasets may be stored in the same storage or in separate storage.
[0033] Figure 2 is a diagram of computer system 200, an example of the subject of this paper. Computer system 200 provides a passive-standby disaster recovery architecture. Computer system 200 comprises a primary data center 200A and a secondary data center 200B. Each of the primary center 200A and secondary data center 200B may be a data center as described with reference to Figure 1. Primary data center 200A comprises a primary source database system 201A and a primary target database system 221A (sometimes called a primary accelerator). Secondary data center 200B comprises a secondary source database system 201B and a secondary target database system 221B (sometimes called a disaster recovery (DR) accelerator).
[0034] The secondary data center 200B may be connected to the primary data center 200A through one or more connections. For example, the connections may include Fibre Channel Protocol (FCP) links that can link pairs of disk subsystems, such as disk subsystems 207A and 207B, and disk subsystems 227A and 227B. The FCP connection can be made directly, through a switch, or through other supported distance solutions (e.g., a high-density wavelength division multiplexer (DWDM) or a channel extender).
[0035] The primary source database system 201A includes storage 207A, such as a disk subsystem for storing Db2. The secondary source database system 201B (sometimes called the DR Db2 z / OS system) also includes storage 207B, such as a disk subsystem. As shown in Figure 2, the primary source database system 201A and the secondary source database system 201B are configured to actively mirror the storage 207A of the primary source database system 201A to the storage 207B of the secondary source database system 201B, so that the data in storage 207B matches the data in storage 207A.
[0036] The primary target database system 221A is equipped with storage 227A, and the secondary target database system 221B is equipped with storage 227B.
[0037] Primary data center 200A is active, while secondary data center 200B is inactive. The activity of primary data center 200A may mean that data replication occurs between primary source database system 201A and primary target database system 221A, and that primary target database system 221A can execute complex queries received in primary source database system 201A.
[0038] Since the primary data center 200A is active and mirroring is also active, the data in the three storage devices 207A, 227A, and 207B can be in a synchronized state.
[0039] In the event of a disaster, the secondary data center 200B may be activated, the secondary source database system 201B located in the secondary data center 200B may be started, recover from its current state on the disk subsystem 207B, and the changes may be propagated to the connected DR accelerator 221B. For example, a query distribution unit (not shown) may activate the secondary data center in the enterprise data network and send all query workloads to the Db2 / z logical partition (LPAR) of the secondary data center 200B, or the query workloads may be redistributed by activating specific network settings, for example. The query distribution unit may be configured to receive queries for processing data in the primary data center 200A and forward the queries to the primary source database system 201A. In the event of a failure of the primary source database system 201A (for example, in the event of a disaster), the query distribution unit may be configured to favorably transfer received queries to the secondary source database system 201B in accordance with this subject.
[0040] Figure 3 is a flowchart of an approach to operating a data center (referred to as a primary data center) using an example from this subject. For illustrative purposes, the approach described in Figure 3 may be implemented in the system illustrated in Figure 1 or Figure 2, but is not limited to such implementations. The primary data center may be, for example, data center 100 or 200A, as described with reference to Figures 1 and 2, respectively.
[0041] In step 301, a secondary data center 200B may be provided, for example, as described with reference to Figure 2. The secondary data center can enable a disaster recovery center for the primary data center 200A. The secondary data center 200B may be connected to the primary data center 200A via one or more connections. The connections may include, for example, FCP links or fiber optic connection (FICON) links or both.
[0042] The process may determine in step 303 whether a failure has occurred in the primary source database system 201A. The failure may be, for example, a failure in the disk subsystem 207A of the primary source database system 201A.
[0043] In response to detecting a failure in the primary source database system 201A, the process may, in step 305, offload queries targeting the primary source database system 201A to the secondary source database system 201B. At time t0 when the failure occurred, storages 207A, 227A, and 207B may be in a matching state and may have the same data named DS0, which may be the most recent data from the primary source database system 201A. Data matching on storage may be ensured, for example, by the owner / administrator monitoring and checking the operation of the primary and secondary data centers. Data DS0 may be the most recent data from the primary source database system 201A immediately before the failure was detected in the primary source database system 201A.
[0044] The process may determine in step 307 whether the primary target database system 221A is available. If the system is available, it means that the system can receive and execute queries.
[0045] If the primary target database system 221A is available and receives queries offloaded by the secondary source database system 201B, the secondary source database system 201B may parse the offloaded queries to identify the queries (complex queries) and, in step 309, transfer the parsed queries to the primary target database system 221A. In step 311, the primary target database system 221A may execute the parsed queries and copy the data DS0 from the primary target database system 221A to the secondary target database system 221B. The copying of data DS0 and the execution of the parsed queries in step 311 may be performed in parallel or simultaneously. Copying the data DS0 allows the DR accelerator 221B to prepare for query acceleration. This may allow the primary target database system 221A to be operational for query acceleration until the DR accelerator 221B is ready for query acceleration. Copying data from the primary target database system 221A can be advantageous. For example, when resynchronizing the DR accelerator, the data may be copied directly from the primary accelerator instead of the database system 201B under recovery, since the data is still available and its replicated recovery metadata (e.g., bookmark tables) is synchronized with the database's recovery state (only committed transactions are replicated).
[0046] If the primary target database system 221A is unavailable, in step 313, the secondary source database system 201B may copy the data DS0 to the secondary target database system 221B.
[0047] After the copying of data DS0 is complete, in step 315, the replication acceleration function may be activated in the secondary data center 200B (for example, at time t1). This means that after time t1, replication may be performed from the secondary source database system 201B to the secondary target database system 221B, and parsing queries may be executed in the secondary target database system 221B. The replication may be performed so that changes to data DS0 that occurred between t0 and t1 can be applied in the secondary target database system 221B. For example, after the secondary data center 200B is activated, these changes that occurred between t0 and t1 may be propagated to the secondary accelerator 221B. For example, there may be a time T0 between t0 and t1 that marks the availability of the Db2 / z mainframe in the secondary data center 200B. The time from t0 to T0 can be very short, possibly just a few minutes or even seconds if everything is automated, and neither the mainframe Db2 at the primary data center nor the mainframe Db2 at the secondary data center may be available. There may be no data written to the secondary Db2 z / OS system between t0 and T0. There may be data written to the secondary Db2 z / OS system between T0 and t1. The changes will be written to the secondary Db2 / z, and this data will be replicated to the secondary accelerator after t1.
[0048] Copying data from the primary target database system 221A to the secondary target database system 221B may involve the steps of: identifying the tables to be copied on the primary target database system 221A and initiating the copy process to the secondary target database system 221B; copying replication metadata that identifies one or more log locations until the data is accurately replicated from the primary source database system 201A; changing the DR db2 z / OS system and DR accelerator to read / write mode after recovery is complete; and restarting replication on the DR db2 z / OS system and DR accelerator to continue replicating any new changes after failover and recovery are complete. Since the recovery procedure is complete, the DR db2 z / OS system and DR accelerator may be switched to read / write mode. The DR accelerator has been protected from write activity until now, but since the data on the DR accelerator is currently up-to-date, the mode can be set to read / write to continue the remaining work (duplicating any newly modified data from db2 and query workloads). Switching to read / write mode completes the entire data recovery procedure.
[0049] Figure 4 is a flowchart of a recovery approach using an example from this subject. For illustrative purposes, the approach described in Figure 4 may be implemented in the system illustrated in Figure 2, but is not limited to this implementation.
[0050] Since Db2 in the primary data center 200A is down, disaster recovery may be initiated in step 401. The process may determine in step 403 whether the DR accelerator 221B is functioning. If the DR accelerator 221B is functioning, the process may execute steps 405-415; otherwise, the process may execute steps 407-411, 413-415, and 416-417. The process may determine in step 405 whether the primary accelerator 221A is functioning. If the process determines that the primary accelerator is functioning, the process may execute steps 407, 409-411, 413, and 415; otherwise, the process may execute steps 412-413, and 415. As used herein, “not functioning” means that the DR accelerator does not yet have the data necessary for it to run the workload. A DR accelerator can function, for example, after preparing and copying data from a primary accelerator to a secondary accelerator.
[0051] In step 407, the process may set the primary accelerator to read-only mode. In step 409, the process may run the workload on the primary accelerator. In step 410, the process may identify the tables to be copied, and in step 411, the primary accelerator 221A may start the copy process to the DR accelerator 221B. In step 411, the process may identify the tables and copy their contents. The operating system image, configuration, and settings may not be copied, as the operating system, network settings, and default settings may already be available on the secondary accelerator. In step 411, the primary accelerator 221A may further copy replication metadata indicating the log location to the DR accelerator 221B. After the copy is complete, in step 413, the process may start replication on the DR accelerator, and in step 415, the process may start acceleration to run the workload on the DR accelerator 221B.
[0052] In step 412, the process begins copying data from the DR Db2 z / OS system 201A to the DR accelerator 221B, and then the process executes steps 413 and 415.
[0053] In step 416, the process may determine whether the primary accelerator 221A is functioning. If the process determines that the primary accelerator 221A is functioning, the process performs steps 407, 409-411, 413, and 415. If the process determines that the primary is not functioning, then in step 417, the workload may fail because none of the accelerators are functioning.
[0054] Thus, Figure 4 shows the flow of events triggered to recover from a primary Db2 failure on z / OS system 201A. As described above, the flow starts when it is detected that Db2 is unresponsive in the primary data center. This causes Db2 to fail over to DR z / OS system 201B, and once system 201B has completed startup and recovered to a matching state from the primary transaction log, it will begin searching for an accelerator to use. If both the DR accelerator and the primary accelerator are still available, the primary accelerator can be used directly for acceleration, and the copy process can also start directly from the primary accelerator in parallel. This works because only committed changes are replicated, and the primary accelerator is in the state of the changes in DR Db2 z / OS system 201B or a previous matching state. Once the copy is complete, the replication can select from the most recent commit copied from DR Db2 z / OS system 201B. If DR accelerator 221B is not functioning but primary accelerator 221A is functioning, primary accelerator 221A will be put into read-only mode, and query acceleration from DR Db2 z / OS system 201B may occur immediately. In parallel, the copy process from primary accelerator 221A to DR accelerator 221B will begin, and once complete, replication and query acceleration will be able to resume. Finally, if primary accelerator 221A is not functioning but DR accelerator 221B is functioning, data may be copied from DR Db2 z / OS system 201B to DR accelerator 221B, and query acceleration may not occur until this copy process is complete.
[0055] This topic includes the following sections:
[0056] Section 1. A computer execution method, comprising providing a primary data center, wherein the primary data center comprises a primary source database system and a primary target database system, has a copy of the data of the primary source database system, receives parsing queries from the primary source database system, and causes the primary target database system to execute parsing queries on the data, the primary data center is activated to provide, and in response to detecting a failure in the primary source database system, offloads queries targeting the primary source database system to a secondary source database system, the secondary data A computer execution method comprising: a secondary source database system of a data center further comprising a secondary target database system, wherein the secondary source database system has a copy of the data and offloads the data, deactivating its functionality in the secondary data center; receiving parsed queries processed by the secondary source database system in response to the primary target database system being available; copying the data to the secondary target database system; and activating its functionality in the secondary data center.
[0057] The computer execution method of Section 1, further comprising setting the primary target database system to read-only mode before executing the parsing query in Section 2.1.
[0058] Section 3. A computer execution method of Section 1 or 2, further comprising transmitting log positions from the primary target database system to the secondary target database system until data is replicated from the primary source database system to the primary target database system, wherein the log positions are used to replicate further changes that occur in the secondary source database system after a failure.
[0059] Section 4. The computer execution method of Section 3, in which the log location is transmitted within the replicated metadata of the primary target database system.
[0060] Section 5. A computer execution method according to any of Sections 1 through 4, wherein the failure is a failure of the disk subsystem of the primary source database system.
[0061] A computer execution method according to any of Sections 1 to 5, wherein the primary source database system actively mirrors the storage of the primary source database system to the storage of the secondary source database system, such that the secondary source database system has a copy of the data.
[0062] Section 7. A computer execution method according to any of Sections 1 through 6, wherein copying data to a secondary target database system is performed in parallel with executing the parsing query.
[0063] A computer execution method according to any of Sections 1 through 7, wherein activating a function in a secondary data center includes changing the primary source database system to read / write mode.
[0064] A computer execution method according to any of Sections 1 through 8, wherein the primary source database system is an online transaction processing (OLTP) system and the primary target database system is an online analysis processing (OLAP) system.
[0065] A computer execution method according to any of Sections 1 through 8, wherein the secondary source database system is an online transaction processing (OLTP) system and the secondary target database system is an online analysis processing (OLAP) system.
[0066] A computer execution method according to any of Sections 1 through 10, wherein a primary data center is connected to a secondary data center through one or more links, the links comprising a selection from a group consisting of Fibre Channel Protocol (FCP) links and Fibre Connectivity (FICON) links.
[0067] The present invention may be a system, method, or computer program product, or a combination thereof, at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium (or more mediums) having computer-readable program instructions for causing a processor to execute an aspect of the present invention.
[0068] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction-executing device. A computer-readable storage medium may, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. A less-than-exclusive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random-access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random-access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or grooved-reinforced structures on which instructions are recorded, and any suitable combination thereof. Computer-readable storage media as used herein should not be interpreted as, in essence, radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or transient signals such as electrical signals transmitted through wires.
[0069] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to each computing / processing device, or to an external computer or external storage device via a network such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. The network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on the computer-readable storage media within each computing / processing device.
[0070] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine language instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuit equipment, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk(R), C++, or similar, and procedural programming languages such as the C programming language or similar programming languages. The computer-readable program instructions may be executed as a standalone software package, either entirely or partially on the user's computer, or partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, electronic circuit equipment, including, for example, a programmable logic circuit device, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer-readable program instructions by using state information of computer-readable program instructions to personalize the electronic circuit equipment in order to carry out aspects of the present invention.
[0071] Aspects of the present invention will be described herein with reference to flowcharts or block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block in the flowchart or block diagram, or both, and combinations of blocks in the flowchart or block diagram, or both, are executable by computer-readable program instructions.
[0072] These computer-readable program instructions may be provided to a computer or other programmable data processing device processor for generating a machine, so as to create means for instructions to be executed via the processor of the computer or other programmable data processing device to perform functions / actions specified in one or more blocks of a flowchart or block diagram, or both. These computer-readable program instructions may be further stored on a computer-readable storage medium so as to provide a product containing instructions that perform a manner of functions / actions specified in one or more blocks of a flowchart or block diagram, or both, and can instruct a computer, a programmable data processing device, or other device, or a combination thereof, to function in a particular manner.
[0073] Computer-readable program instructions may also be loaded onto a computer, another programmable device, or another device in order to cause a series of operational steps to be performed on the computer, another programmable device, or another device in order to generate computer execution processing to perform a function / action specified in one or more blocks of a flowchart or block diagram, or both.
[0074] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions containing one or more executable instructions for performing a specified logical function. In some alternative implementations, the functions described in a block may be performed independently of the order shown in the diagram. For example, two consecutively shown blocks may actually be performed as a single step in a manner that overlaps in time, substantially simultaneously, partially or fully, or blocks may be executed in reverse order depending on the functions they sometimes contain. It should also be noted that each block in a block diagram or flowchart, or both, and combinations of blocks in a block diagram or flowchart, or both, can be performed by a special-purpose hardware-based system that performs a specified function or action, or executes a combination of special-purpose hardware and computer instructions.
[0075] While various embodiments of the present invention have been presented for illustrative purposes, they are not intended to be exhaustive or to limit oneself to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The terminology used herein has been chosen to best describe the principles, practical applications, or technical improvements to the technologies available on the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method executed by one or more processors, To provide a primary data center comprising a primary source database system and a primary target database system, wherein the primary target database system includes Having a copy of the data from the aforementioned primary source database system, Receiving parsing queries from the aforementioned primary source database system, and Execute the aforementioned analysis query on the aforementioned data. The function to enable this is activated in the primary data center, and is provided as follows: In response to detecting a failure in the primary source database system, Offloading queries targeting the primary source database system to a secondary source database system, wherein the secondary source database system in the secondary data center further comprises a secondary target database system, the secondary source database system has a copy of the data, and the function is deactivated in the secondary data center. In response to the availability of the aforementioned primary target database system, The parsing queries of the offloaded queries, processed by the secondary source database system, are received by the primary target database system, and The primary target database system copies the data to the secondary target database system, and To enable the aforementioned function in the secondary data center. A computer execution method, including...
2. The computer execution method according to claim 1, further comprising setting the primary target database system to read-only mode before executing the parsing query in the primary target database system.
3. A computer execution method according to claim 1, further comprising transmitting from the primary target database system to the secondary target database system the log positions up to the time when the data is replicated from the primary source database system to the primary target database system, wherein the log positions are used to replicate any further changes that occurred in the secondary source database system after the failure.
4. The computer execution method according to claim 3, wherein the log location is transmitted within the replicated metadata of the primary target database system.
5. The computer execution method according to claim 1, wherein the failure is a failure in the disk subsystem of the primary source database system.
6. The computer execution method according to claim 1, wherein the primary source database system actively mirrors the storage of the primary source database system to the storage of the secondary source database system so that the secondary source database system has the copy of the data.
7. The computer execution method according to claim 1, wherein copying the data to the secondary target database system is performed in parallel with executing the analysis query.
8. The computer execution method according to claim 1, wherein activating the function in the secondary data center includes changing the primary source database system to read / write mode.
9. The computer execution method according to claim 1, wherein the primary source database system is an online transaction processing (OLTP) system, and the primary target database system is an online analysis processing (OLAP) system.
10. The computer execution method according to claim 1, wherein the secondary source database system is an online transaction processing (OLTP) system, and the secondary target database system is an online analysis processing (OLAP) system.
11. The primary data center is connected to the secondary data center through one or more links. The link includes a selection from the group consisting of Fibre Channel Protocol (FCP) links and Fibre Connectivity (FICON) links. The computer execution method according to claim 1.
12. A computer program that causes one or more processors to perform the method according to any one of claims 1 to 11.
13. A computer-readable storage medium recording a computer program for causing one or more processors to perform the method according to any one of claims 1 to 11.
14. A computer system, The system comprises one or more computer processors, one or more computer-readable storage media, and program instructions stored collectively in the one or more computer-readable storage media for execution by at least one of the one or more computer processors, wherein the program instructions are A program instruction for providing a primary data center comprising a primary source database system and a primary target database system, wherein the primary target database system includes Having a copy of the data from the aforementioned primary source database system, Receiving parsing queries from the aforementioned primary source database system, and Execute the aforementioned analysis query on the aforementioned data. A program instruction for providing a primary data center, in which a function to perform the above is activated in the primary data center, and In response to detecting a failure in the primary source database system, Offloading queries targeting the primary source database system to a secondary source database system, wherein the secondary source database system in the secondary data center further comprises a secondary target database system, the secondary source database system has a copy of the data, and the function is deactivated in the secondary data center. In response to the availability of the aforementioned primary target database system, The parsing queries of the offloaded queries, processed by the secondary source database system, are received by the primary target database system, and The primary target database system copies the data to the secondary target database system, and To enable the aforementioned function in the secondary data center. Program instructions for performing A computer system that includes [a specific feature / function].
Citation Information
Patent Citations
Control method for making data dual between computer systems
JP2004348701A
Data center system and its control method
JP2005018510A
Information processor and information processing method
JP2014120123A
Data management system
JP2015165357A
High availability and disaster recovery in large-scale data warehouse
US20160132576A1