Low-latency Demultiplexer for Propagating Sequential Data to Multiple Sinks
The database client buffers data in memory and uses order indicators from commit acknowledgments to efficiently propagate ordered data to multiple sinks with varying rates, addressing latency and scalability issues in large-scale distributed databases.
Patent Information
- Application Number
- JP2024574753
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-20
- Filing Date
- 2023-06-19
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2043-06-19
AI Technical Summary
The challenge of propagating ordered data from a large-scale distributed database to multiple sinks with varying data consumption rates, where existing solutions either require separate queries for each sink or limit all sinks to the rate of the slowest, leading to increased latency or scalability issues.
A database client that buffers data in transient memory, uses commit acknowledgments with order indicators to manage data propagation to multiple sinks, ensuring data is sent in the order of commitment to the database without relying on additional storage or database reads, and handles varying sink consumption rates using watermarks and fallback mechanisms.
Enables efficient, low-latency propagation of ordered data to multiple sinks with varying consumption rates, minimizing latency and resource usage by intercepting data commits and using in-memory buffers with order tracking.
Smart Images

Figure 2025520601000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a low-latency demultiplexer for propagating ordered data.
Background Art
[0002] As the popularity of large-scale distributed databases (i.e., "cloud" databases) continues to grow, the need to propagate data from a database to data sinks (e.g., applications) also continues to increase. However, when each sink consumes data at a different rate, it can be a problem to continuously propagate ordered data from the database to multiple sinks. A simple solution to this problem involves individually addressing data propagation to each sink. However, this requires a separate and independent database query for each stream of data to a given sink. Alternatively, each sink may be limited to the rate of the slowest sink, in which case the latency of all sinks except the slowest increases.
Summary of the Invention
[0003] One aspect of the present disclosure provides a method for a low-latency demultiplexer that propagates data to a plurality of sinks. The method, when executed by a computer, causes the data processing hardware to perform operations. The operations include receiving a series of writes. Each write in the series of writes requests that the respective data be stored in a database that communicates with the data processing hardware. For each write in the series of writes, the operations include storing the respective data in a first buffer associated with a first data streaming application, storing the respective data in a second buffer associated with a second data streaming application, and sending the respective data to the database. The operations also include receiving from the database an acknowledgement that the respective data of each write has been committed to the database. The acknowledgement includes an order indicator that indicates the order in which the respective data of each write has been committed to the database relative to other writes in the series of writes. In response to receiving the acknowledgement that the respective data of each write has been committed to the database, the operations include sending the respective data of each write from the first buffer to the first data streaming application based on the order indicator that indicates that the respective data has been committed to the database relative to other writes in the series of writes, and sending the respective data of each write from the second buffer to the second data streaming application based on the order indicator that indicates that the respective data has been committed to the database relative to other writes in the series of writes.
[0004] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, sending each data from a first buffer to a first data streaming application based on the order in which each data was committed to the database relative to other writes in a series of writes includes using an order indicator to determine that an acknowledgement of a previous write has been received. Each data of the previous write was committed to the database immediately before each data of each write. In some of these embodiments, determining that an acknowledgement of a previous write has been received includes determining a level of a watermark of the first buffer. Optionally, sending each data from a second buffer to a first data streaming application based on the order in which each data was committed to the database relative to other writes in a series of writes includes using an order indicator to determine that an acknowledgement of a previous write has not been received, where each data of the previous write was committed to the database immediately before each data of each write, and after determining that an acknowledgement of a previous write has not been received, receiving an acknowledgement that each data of the previous write has been committed to the database. The operation may also include, in response to receiving an acknowledgement that each data of a previous write has been committed to the database, sending each data of the previous write from the second buffer to a second data streaming application, and after sending each data of the previous write, sending each data of each write from the second buffer to a second data streaming application.
[0005] In some examples, the sequence indicator includes an increment identifier. The sequence indicator includes a timestamp in some embodiments. In some of these embodiments, the operation further includes generating a linked list that orders a series of writes in the order in which each respective data of each received write was committed to the database, using the timestamp of each received acknowledgement. Generating the linked list may include generating a hash map.
[0006] Optionally, receiving an acknowledgement from the database that each respective data of each write was committed to the database includes determining that a threshold period has elapsed without receiving an acknowledgement and, in response to determining that the threshold period has elapsed without receiving an acknowledgement, obtaining a change log from the database. The operation may further include determining from the change log that each respective data of each write was committed to the database. The database may include a Structured Query Language (SQL) database.
[0007] Other aspects of the present disclosure provide a system for a low-latency dense multiplexer that propagates data to multiple sinks. The system includes data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving a series of writes. Each write within the series of writes requests that respective data be stored in a database that communicates with the data processing hardware. For each write within the series of writes, the operations include storing the respective data in a first buffer associated with a first data streaming application, storing the respective data in a second buffer associated with a second data streaming application, and transmitting the respective data to the database. The operations also include receiving, from the database, an acknowledgement that the respective data of each write has been committed to the database. The acknowledgement includes an order indicator that indicates the order in which the respective data of each write has been committed to the database relative to other writes within the series of writes. In response to receiving the acknowledgement that the respective data of each write has been committed to the database, the operations include transmitting the respective data of each write from the first buffer to the first data streaming application based on the order indicator that indicates that the respective data has been committed to the database relative to other writes within the series of writes, and transmitting the respective data of each write from the second buffer to the second data streaming application based on the order indicator that indicates that the respective data has been committed to the database relative to other writes within the series of writes.
[0008] This aspect may include one or more of the following optional features. In some embodiments, transmitting each data from a first buffer to a first data streaming application based on the order in which each data was committed to the database relative to other writes in a series of writes includes using an order indicator to determine that an acknowledgement of a preceding write has been received. Each data of the preceding write was committed to the database immediately before each data of each write. In some of these embodiments, determining that an acknowledgement of a preceding write has been received includes determining a level of a watermark of the first buffer. Optionally, transmitting each data from a second buffer to a first data streaming application based on the order in which each data was committed to the database relative to other writes in a series of writes includes using an order indicator to determine that no acknowledgement of a preceding write has been received, where each data of the preceding write was committed to the database immediately before each data of each write, and after determining that no acknowledgement of a preceding write has been received, receiving an acknowledgement that each data of the preceding write has been committed to the database. The operation may also include, in response to receiving an acknowledgement that each data of a preceding write has been committed to the database, transmitting each data of the preceding write from the second buffer to a second data streaming application, and after transmitting each data of the preceding write, transmitting each data of each write from the second buffer to a second data streaming application.
[0009] In some examples, the sequence indicator includes an increment identifier. The sequence indicator includes, in some embodiments, a timestamp. In some of these embodiments, the operation further includes generating a linked list that orders a series of writes in the order in which each respective data of each received write was committed to the database, using the timestamp of each received acknowledgement. Generating the linked list may include generating a hash map.
[0010] Optionally, receiving an acknowledgement from the database that each respective data of each write was committed to the database includes determining that a threshold period has elapsed without receiving an acknowledgement, and in response to determining that the threshold period has elapsed without receiving an acknowledgement, obtaining a change log from the database. The operation may further include determining from the change log that each respective data of each write was committed to the database. The database may include a Structured Query Language (SQL) database.
[0011] Details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5
DETAILED DESCRIPTION OF THE INVENTION
[0013] Like reference symbols in the various drawings refer to like elements. Large-scale distributed databases (i.e., “cloud” databases) are in increasing demand. With this increasing demand, it is common for a given database to have hundreds or thousands of data sinks that require any updates to the database in real time or near real time. For example, these data sinks require that any data written to the database be immediately streamed to the data sink with minimal latency. Generally speaking, these data sinks need to receive the data in the same order as the order in which the data was written or committed to the database. However, when each sink consumes data at a different rate, it can be a problem to continuously propagate ordered data from the database to multiple sinks.
[0014] A simple solution to this problem involves dealing with data propagation to each sink individually. However, this requires separate and independent database queries for each data stream to a given sink. Alternatively, each sink may be limited to the rate of the slowest sink, in which case the latency of all sinks except the slowest increases. In yet another alternative, the system may obtain database logs (e.g., Write-Ahead Log (WAL)) directly from the database to determine updates to the database, but this technique has several drawbacks. First, the number of data streams is limited by the number of connections the database can maintain, which is typically on the order of a few hundred or less. This lacks the scalability required to support modern distributed computing systems. Further, this technique also requires database polling, which consumes significant database resources when the polling frequency is high and consumes data latency when the polling frequency is low.
[0015] Typically, a database includes a database client that provides an interface between the database and a user. The database client receives data (e.g., from one or more users) for writing to the database. The database client provides the data to the database. When the data is successfully written (i.e., committed) to the database, the database provides a commit confirmation to the database client, which can then notify the user(s).
[0016] Embodiments of the present specification include a database client that runs in the application layer and mediates between a user and a database. The database client enables continuous propagation of ordered data in an efficient manner with minimized latency without the need for an additional storage system. By running along the write path of data to the database rather than along the read path of the database log, the database client can obtain the written data without having to execute a database read. This enables the database client to avoid serializing data objects between the binary data formats required by many databases (e.g., Structured Query Language (SQL) databases). The database client includes one or more buffer data structures that buffer data written to the database in transient or volatile memory. When the database client receives confirmation that the data has been committed to the database, the database client propagates the data from the buffer to each waiting data sink. Each data sink can receive data from its respective buffer at any rate, provided the rate is one that the data sink can process. The database client may maintain separate "watermarks" or pointers to appropriately order the data individually for each data sink.
[0017] Referring now to FIG. 1, in some embodiments, an exemplary data dissemination system 100 includes a remote system 140 that communicates with one or more user devices 10 via a network 112. The remote system 140 may be a scalable / elastic resource 142 that includes a single computer, multiple computers, or computing resources 144 (e.g., data processing hardware) and / or storage resources 146 (e.g., memory hardware) in a distributed system (e.g., a cloud environment). A data store 150 (i.e., a remote storage device) may be overlaid on the storage resource 146, enabling scalable use of the storage resource 146 by one or more of a client (e.g., user device 10) or computing resources 144. The data store 150 is configured to store a set of data blocks 154, 154a - n (also simply referred to herein as data 154) in one or more tables or databases 156 (i.e., cloud databases) that each include, for example, multiple rows and columns.
[0018] The remote system 140 is configured to receive a series of writes 152, 152a - n from user devices 10 associated with respective users 12, for example, via the network 112. The user devices 10 may correspond to any computing device such as a desktop workstation, a laptop workstation, or a mobile device (i.e., a smartphone). The user devices 10 include computing resources 18 (e.g., data processing hardware) and / or storage resources 16 (e.g., memory hardware). The user 12 may construct a query 20 using an SQL interface. Each write 152 in the series of writes 152 includes respective data 154 for writing to the database 156. The database 156 (e.g., an SQL database) may be stored in the data store 150.
[0019] The database client 160 mediates communication between the user 12 and the database 156 (e.g., on the application layer). In addition to interfacing with the database 156, the database client 160 propagates the ordered data 154 to one or more applications 170, 170a - n (or any other data sink), such as one or more data streaming applications 170. Each application 170 receives the ordered data 154 with minimal latency when the data 154 is committed to the database 156. That is, the application 170 receives the data 154 in the same order in which the data 154 is written to the database 156.
[0020] For each write 152 within the series of writes 152 received by the database client 160, the database client 160 stores the data 154 of the write 152 in buffers 210, 210a - n associated with each application 170 to which the database client 160 streams the data 154. For example, if the database client 160 streams the data 154 to three applications 170, the database client 160 stores the data 154 from each write 152 in a first buffer 210 associated with the first application 170, a second buffer 210 associated with the second application 170, and a third buffer 210 associated with the third application. In other examples, the database client 160 stores the data 154 from each write 152 in a single buffer 210 associated with all three applications 170. Each buffer 210 represents a data structure that uses transient storage (e.g., volatile memory) to temporarily store the data 154.
[0021] Database client 160 sends each piece of data 154 to database 156 for each write 152. For each write 152, database 156 stores, or commits, each piece of data 154 to database 156 (i.e., persistent or non-volatile storage), and then sends acknowledgments 220, 220a - n to database client 160 confirming that data 154 has been committed to database 156. Each acknowledgment 220 includes an order indicator 222 that indicates the order in which the data 154 of each write 152 was committed to database 156 relative to the data 154 of other writes 152 within a series of writes 152. As discussed in more detail below, order indicator 222 may include an incrementing identifier or a timestamp.
[0022] In response to receiving an acknowledgment 220 that the data 154 of a write 154 has been committed to database 156, database client 160 sends each piece of data 154 from each buffer 210 to the corresponding application 170. For example, the first buffer 210 sends data 154 to the first application when it is confirmed that the data will be committed to database 156. In this example, database client 160 uses the second buffer 210 to send data 154 to the second application 170. In this way, database client 160 propagates data 154 to application 170 with minimal latency by not having to fetch, or read, data 154 from database 156 or from a database log. Instead, database client 160 "intercepts" data 154 on its way to database 156 and locally stores data 154 in buffer 210 until data 154 is committed to database 156. When database client 160 determines that data 154 has been committed to database 156, database client 160 may use buffer 210 to propagate data 154 to application 170.
[0023] Referring now to FIG. 2, in some examples, application 170 (or any other data sink) consumes data 154 at different rates. For example, one application 170 consumes data at an unlimited rate (i.e., the same rate at which database client 160 transmits data 154), and another application 170 consumes data at a rate that is at least temporarily slower than the rate at which database client 160 can transmit data 154 (i.e., slower than the rate at which database 156 receives data 154). In some embodiments, database client 160 implements watermarks 212, 212a - n for each buffer 210, which enables database client 160 to track which data 154 has been sent to each application 170 and which data 154 has not yet been sent. When database client 160 sends new data 154 to application 170, database client 160 updates the corresponding watermark 212 to reflect the new position of the watermark 212 (i.e., to reflect that new data 154 has been sent).
[0024] In some examples, the database client 160 expires (e.g., removes or deletes) data 154 from the buffer 210 based on the watermark 212. That is, if the watermark 212 indicates that data 154 has been sent to the corresponding application 170, the database client 160 may recover space in the buffer 210 by deleting or overwriting the sent data 154. In some embodiments, multiple applications 170 share a single buffer 210. In these embodiments, each application 170 that shares the buffer 210 has an independent watermark 212 to track which data 154 has been sent to which application 170. In this case, the database client 160 may ensure that data 154 is expired only if data 154 has been sent to all applications 170 that use the buffer 210 (i.e., based on each of the independent watermarks 212).
[0025] In the example shown in the schematic diagram 200, the database client 160 receives four writes 152a - d with data 154a - d to be written to the database 156. The database client 160 also streams data 154 to two applications 170a - b. Each application 170 is associated with a corresponding buffer 210a - b. When the database client 160 receives the writes 154a - d, it queues the data 154a - d in each of the buffers 210a - b. The database client 160 also sends the data 154a - d to the database 156. In this example, the database 156 responds with four acknowledgments 220a - d (each including a sequence indicator 222) that confirm that the data 154a - d has been committed. Here, the sequence indicator 222 is an incrementing identifier (i.e., an incrementing integer). That is, the database 156 increments the sequence indicator 222 each time a commit to the database 156 is successful.
[0026] However, the first application 170a can only consume data 154a - b (and cannot yet consume data 154c - d), and the second application 170b can only consume data 154a - c (and cannot yet consume data 154d). Thus, the database client 160 updates the first watermark 212a to reflect that the database client 160 has sent data 154a - b to the first application 170a, and similarly updates the second watermark 212b to reflect that the database client 160 has sent data 154a - c to the second application 170b. The unsent data 154 remains in the buffer 210 until the application 170 can consume or receive the data 154 (i.e., data 154c - d remains in the first buffer 210a and data 154d remains in the second buffer 210b). Thus, the database client 160 propagates the data 154 to different sinks (i.e., applications 170) without relying on the connection to the database 156, despite the sinks having different data consumption rates.
[0027] Referring now to FIG. 3A, in some embodiments, when the database client 160 sends each data 154 of each write 152 from the buffer 210 to the application 170 based on the order in which the respective data 154 of other writes 152 sent by the database client 160 to the database 156 are committed to the database 156, the database client 160 uses the order indicator 222 to determine that, immediately prior to each data 154 of each write 152, an acknowledgement 220 of a previous write 152 has been received, where the respective data 154 of the previous write 152 has been committed to the database 156. For example, the database client 160 determines that an acknowledgement 220 of a previous write has been received by determining the level of the watermark 212 of the buffer 210.
[0028] For example, as shown in schematic diagram 300a, database client 160 receives four writes 152a - d and queues the corresponding data 154a - d into two buffers 210a - b of two applications 170a - b. The database client 160 transmits the data 154a - d to the database 156. In this example, the database client 160 receives confirmations 220a, b, d that the data 154a, b, d have been committed, but for unknown reasons (e.g., network congestion or malfunction), does not receive a confirmation that the data 154c has been committed. In this scenario, the database client 160 transmits the data 154a - b from the first buffer 210a to the first application 170a and the data 154a - b from the second buffer 210b to the second application 170b. This is to reflect the order in which the data 154 was committed to the database 156. However, the database client 160 does not transmit the data 154d to the applications 170a - b even though it has received a confirmation 220d that the data 154d has been committed to the database 156. This is because the database client 160 determines (e.g., based on the order indicator 222) that there may be data 154 (i.e., data 154c) that was committed before the data 154d for which the database client 160 has not received a confirmation.
[0029] Referring now to FIG. 3B, in some embodiments, database client 160 waits for an acknowledgement of data 154c so that data 154 is propagated to application 170 in the appropriate order (i.e., the order in which database 156 committed data 154). For example, database client 160 uses order indicator 222 to determine that an acknowledgement 220 of a previous write 154 has not been received, and after determining that an acknowledgement 220 of a previous write 154 has not been received, receives an acknowledgement 220 that each respective data 154 of the previous write 152 has been committed to database 156. In response to receiving an acknowledgement 220 that each respective data 154 of the previous write 152 has been committed to database 156, database client 160 may transmit each respective data 154 of the previous write 152 from buffer 210 to the corresponding application 170. After transmitting each respective data 154 of the previous write 152, database client 160 may transmit each respective data 154 of subsequent writes 152 from buffer 210 to application 170.
[0030] For example, continuing with the example of FIG. 3A, after transmitting data 154a - b to applications 170a - b, database client 160 receives confirmation 220c that data 154c has been committed to database 156. Thus, in this example, data 154a - d is committed to database 156 in an order different from the order in which confirmations 220a - d are received by database client 160. After receiving confirmation 220c, database client 160 transmits the remaining buffered data 154c - d from each buffer 210a - b to the corresponding applications 170a - b. In this example, applications 170a - b consume data at the same rate as discussed with respect to FIG. 2, although applications 170 may consume data at different rates, and database client 160 depends on buffer 210 to buffer data 154 until all of data 154 has been consumed by applications 170.
[0031] Referring now to FIG. 3C and continuing again with the example of FIG. 3A, in some embodiments, database client 160 may not receive missing confirmation 220c within a threshold period. In these embodiments, the database client may include a fallback controller 310 that, after determining that the threshold period has been met (e.g., at least 5 ms has elapsed), retrieves change log 330 from database 156. For example, fallback controller 310 may transmit a change log request 320 requesting change log 330 in response to the threshold period being met. Fallback controller 310 may determine from change log 330 that each respective data 154 has been committed to database 156. In this example, change log 330 reveals that data 154c has been committed to database 156, and database client 160 transmits the remaining buffered data 154c - d from each buffer 210a - b to the corresponding applications 170a - b.
[0032] In the example shown, the sequence indicator 222 is an increment identifier such as a monotonically increasing integer. Some databases 156 (e.g., SQL databases) use such a scheme with commit acknowledgments. With such a scheme, the database client 160 can easily determine which data 154 was committed to the database 156 immediately prior to each write 152, so that the data 154 can be properly ordered for the application 170. For example, if the database client 160 receives an acknowledgment 220 for a first write 152 that includes a sequence identifier "4", the database client 160 determines that the data 154 associated with the acknowledgment 220 for a second write 152 that includes a sequence identifier "3" was committed to the database 156 immediately prior to the first write 152, and thus the data 154 of the second write 152 needs to be propagated to the application 170 before the data 154 of the first write 152. The database 156 does not need to commit data in the order received from the database client 160. The database 156 may provide an indication of which data 154 is committed by the acknowledgment 220 (e.g., using a description, identifier, hash, etc.) in addition to the sequence indicator 222 to enable the database client 160 to match the acknowledgment 220 with specific data 154.
[0033] Some databases 156 may not include an increment identifier and instead may include a timestamp as the sequence indicator 222. In this scenario, the database client 160 needs to determine the order in which to propagate the data 154 based on the timestamp received in the confirmation 220. In some embodiments, the database client 160 uses the timestamp to generate a linked list that orders a series of writes 152 in the order in which each respective data 154 of each write 152 was committed to the database 156. For example, the database client 160 generates a hash map of the data 154 that can be a link to an ordered linked list. The database client 160 may insert the data 154 and / or the write 152 into the hash map using a pointer to the immediately preceding write timestamp. By maintaining a pointer to the most recent data 154 propagated to the application 170, the database client 160 buffers the data 154 in the hash map until the data 154 is ready to be propagated. In some examples, the database client 160 modifies the schema of the database 156 to assist in determining the order for timestamps. For example, with the modified schema, each write may update fields in different rows of the database using the exact same timestamp as the timestamp of the confirmation 220. In this example, the database client 160 may perform a "stale read" of the modified row (e.g., at a point immediately prior to the timestamp) to determine the previous state of the row. Such a stale read tends to be very low in computational cost (i.e., much lower cost than a direct pull of data from a change log table).
[0034] The examples in this specification show buffer 210 storing data 154 until a commit confirmation 220 is received, although other buffers or storage locations may be used as well. For example, data 154 may first be stored in a first storage location (e.g., a first buffer) until a confirmation 220 that the data 154 has been committed is received. Upon receiving confirmation 220 that the data 154 has been committed, the data 154 may be moved to buffer 210, thereby alleviating the difficulty in purging data 154 that has not been successfully committed. In this scenario, the data 154 is provided from buffer 210 to application 170 at the same speed that the application 170 can consume the data 154 because all of the data 154 within buffer 210 has already been committed.
[0035] FIG. 4 is a flowchart of an exemplary operational arrangement for a method 400 that provides a low-latency demultiplexer for propagating data to a plurality of sinks. The computer-implemented method 400, when executed by data processing hardware 144, causes the data processing hardware 144 to perform operations. The method 400, in operation 402, includes receiving a series of writes 152 each requesting that respective data 154 be stored in a database 156 that communicates with the data processing hardware 144. For each respective write 152 within the series of writes 152, the method 400, in operation 404, includes storing each respective data 154 in a first buffer 210 associated with a first data streaming application 170 and storing each respective data 154 in a second buffer 210 associated with a second data streaming application 170. In operation 406, the method 400 includes transmitting each respective data 154 to the database 156. The method 400, in operation 408, includes receiving from the database 156 an acknowledgement 220 that the respective data 154 of each respective write 152 has been committed to the database 156. The acknowledgement 220 includes an order indicator 222 that indicates the order in which the respective data 154 of each respective write 152 has been committed to the database 156 relative to other writes 152 within the series of writes 152.In response to receiving a confirmation 220 that each piece of data 154 of each write 152 has been committed to the database 156, at operation 410, method 400, based on a sequence indicator 222 that indicates that each piece of data 154 has been committed to the database 156 relative to other writes 152 within a series of writes 152, transmits each piece of data 154 of each write 152 from the first buffer 210 to the first data streaming application 170, and based on a sequence indicator 222 that indicates that each piece of data 154 has been committed to the database 156 relative to other writes 152 within a series of writes 152, transmits each piece of data 154 of each write 152 from the second buffer 210 to the second data streaming application 170.
[0036] FIG. 5 is a schematic diagram of an exemplary computing device 500 that can be used to implement the systems and methods described in this document. Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the embodiments of the invention described and / or claimed in this document.
[0037] The computing device 500 includes a processor 510, a memory 520, a storage device 530, a high-speed interface / controller 540 connected to the memory 520 and the high-speed expansion port 550, and a low-speed interface / controller 560 connected to the low-speed bus 570 and the storage device 530. Each component 510, 520, 530, 540, 550, and 560 is interconnected using various buses and can be mounted on a common motherboard or exist in other ways as required. The processor 510 processes instructions for execution within the computing device 500, including instructions stored in the memory 520 or the storage device 530, and can display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 580 connected to the high-speed interface 540. In other embodiments, multiple memories and memory types, as well as multiple processors and / or multiple buses, may be used as required. Also, multiple computing devices 500 may be connected, and each device may provide a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).
[0038] Memory 520 stores information non - transiently within computing device 500. Memory 520 may be a computer - readable medium, a volatile memory unit(s), or a non - volatile memory unit(s). The non - transient memory 520 may be a physical device used to store programs (e.g., instruction sequences) or data (e.g., program state information) temporarily or permanently for use by computing device 500. Examples of non - volatile memory include, but are not limited to, flash memory and read - only memory (ROM) / programmable read - only memory (PROM) / erasable programmable read - only memory (EPROM) / electrically erasable programmable read - only memory (EEPROM) (e.g., used for firmware such as a boot program typically). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase - change memory (PCM), and disks or tapes.
[0039] Storage device 530 can provide large - capacity storage to computing device 500. In some embodiments, storage device 530 is a computer - readable medium. In various different embodiments, storage device 530 may be an array of devices including a floppy (registered trademark) disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid - state memory device, or a storage area network or other configured device. In additional embodiments, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer - readable or machine - readable medium such as memory 520, storage device 530, or memory on processor 510.
[0040] The high-speed controller 540 manages the bandwidth-intensive operations of the computing device 500, and the low-speed controller 560 manages the low-bandwidth-intensive operations. Such role assignments are merely examples. In some embodiments, the high-speed controller 540 is coupled to the memory 520, the display 580, and a high-speed expansion port 550 that can receive various expansion cards (not shown), for example, via a graphics processor or an accelerator. In some embodiments, the low-speed controller 560 is coupled to the storage device 530 and a low-speed expansion port 590. The low-speed expansion port 590, which may include various communication ports (e.g., USB, Bluetooth®, Ethernet®, wireless Ethernet), can be connected to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a network device such as a switch or a router, for example, via a network adapter.
[0041] As shown in the figure, the computing device 500 can be implemented in many different forms. For example, it may be implemented as a standard server 500a, or multiple times within a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.
[0042] Various embodiments of the systems and techniques described herein can be realized in digital and / or optical circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include an implementation in one or more computer programs executable and / or interpretable in a programmable system including at least one programmable processor, which may be special or general purpose, at least one input device, and at least one output device, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system.
[0043] A software application (i.e., a software resource) can refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an “application,” an “app,” or a “program.” Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and game applications.
[0044] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor that receives the machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0045] The processes and logical flows described in this specification can be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to act on input data and generate output. The processes and logical flows can also be implemented by special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory, a random access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer also includes, or is operatively coupled to receive from, or transfer data to, or both, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0046] To provide interaction with a user, one or more aspects of this specification can be implemented on a computer having display devices such as a CRT (cathode ray tube), an LCD (liquid crystal display) monitor, a touch screen, etc. for displaying information to the user, and optionally, a keyboard and a pointing device such as a mouse or a trackball by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input received from the user can be in any form, including acoustic, voice, or tactile input. Further, the computer can interact with the user by sending and receiving documents to and from the devices used by the user, such as by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0047] Multiple embodiments have been described. Nevertheless, it is understood that various modifications can be made without departing from the spirit and scope of this disclosure. Accordingly, other embodiments are within the scope of the following claims.
Claims
**Claim 1** A computer-executable method (400) executed by data processing hardware (144) to cause the data processing hardware (144) to perform operations, the operations comprising: The operations include: Receiving a series of writes (152), each write (152) in the series of writes (152) requiring that respective data (154) be stored in a database (156) that communicates with the data processing hardware (144), and the operations further comprising: For each respective write (152) in the series of writes (152): Storing the respective data (154) in a first buffer (210) associated with a first data streaming application (170); Storing the respective data (154) in a second buffer (210) associated with a second data streaming application (170); Transmitting the respective data (154) to the database (156); Receiving from the database (156) an acknowledgement (220) that the respective data (154) of the respective write (152) has been committed to the database (156), the acknowledgement (220) including an order indicator (222) indicating the order in which the respective data (154) of the respective write (152) has been committed to the database (156) relative to other writes (152) in the series of writes (152), and the operations further comprising: For each respective write (152) in the series of writes (152): In response to receiving the acknowledgement (220) that the respective data (154) of the respective write (152) has been committed to the database (156), Transmitting the respective data (154) of the respective write (152) from the first buffer (210) to the first data streaming application (170) based on the order indicator (222) indicating that the respective data (154) has been committed to the database (156) relative to other writes (152) in the series of writes (152); Based on the order indicator (222) indicating that each of the data (154) has been committed to the database (156) with respect to other writes (152) within the series of writes (152), transmitting each of the data (154) of each of the writes (152) from the second buffer (210) to the second data streaming application (170); A method comprising. **Claim 2** Transmitting each of the data (154) from the first buffer (210) to the first data streaming application (170) based on the order in which each of the data (154) has been committed to the database (156) with respect to other writes (152) within the series of writes (152) includes using the order indicator (222) to determine that an acknowledgement (220) of a preceding write (152) has been received, and immediately before each of the data (154) of each of the writes (152), each of the data (154) of the preceding write (152) is committed to the database (156). The method (400) according to claim 1. **Claim 3** Determining that the acknowledgement (220) of the preceding write (152) has been received includes determining the level of the watermark (212) of the first buffer (210). The method (400) according to claim 2. **Claim 4** Transmitting each of the data (154) from the second buffer (210) to the first data streaming application (170) based on the order in which each of the data (154) has been committed to the database (156) with respect to other writes (152) within the series of writes (152) includes using the order indicator (222) to determine that an acknowledgement (220) of a preceding write (152) has not been received, and immediately before each of the data (154) of each of the writes (152), each of the data (154) of the preceding write (152) is committed to the database (156), and the transmitting further includes After determining that the confirmation (220) of the preceding write (152) has not been received, receiving the confirmation (220) that each respective data (154) of the preceding write (152) has been committed to the database (156); In response to receiving the confirmation (220) that each respective data (154) of the preceding write (152) has been committed to the database (156), Transmitting each respective data (154) of the preceding write (152) from the second buffer (210) to the second data streaming application (170); After transmitting each respective data (154) of the preceding write (152), transmitting each respective data (154) of each respective write (152) from the second buffer (210) to the second data streaming application (170); The method (400) according to claim 2 or claim 3, comprising: **Claim 5** The method (400) according to any one of claims 1 to 4, wherein the sequence indicator (222) includes an increment identifier. **Claim 6** The method (400) according to any one of claims 1 to 5, wherein the sequence indicator (222) includes a timestamp. **Claim 7** The method (400) according to claim 6, wherein the operation further includes generating a linked list that orders the series of writes (152) in the order in which each respective data (154) of each respective write (152) of the series of writes (152) was committed to the database (156) using the timestamp of each received confirmation (220). **Claim 8** The method (400) according to claim 7, wherein generating the linked list includes generating a hash map. **Claim 9** Receiving, from the database (156), the confirmation (220) that each respective data (154) of each respective write (152) has been committed to the database (156) comprises: Determining that a threshold period has elapsed without receiving the confirmation (220); In response to determining that the threshold period has elapsed without receiving the confirmation (220), obtaining a change log (330) from the database (156); Determining from the change log (330) that the respective data (154) of the respective writes (152) have been committed to the database (156); The method (400) according to any one of claims 1 to 8, including.
10. The method (400) according to any one of claims 1 to 9, wherein the database (156) includes a structured query language (SQL) database.
11. Data processing hardware (144), and Memory hardware (146) communicating with the data processing hardware (144), A system (100) including: The memory hardware (146) stores instructions which, when executed on the data processing hardware (144), cause the data processing hardware (144) to perform operations, The operations include: Receiving a series of writes (152), each write (152) in the series of writes (152) requiring that respective data (154) be stored in a database (156) communicating with the data processing hardware (144), and the operations further include, For each respective write (152) in the series of writes (152), Storing the respective data (154) in a first buffer (210) associated with a first data streaming application (170), Storing the respective data (154) in a second buffer (210) associated with a second data streaming application (170), Transmitting the respective data (154) to the database (156), Receiving from the database (156) an acknowledgement (220) that the respective data (154) of the respective writes (152) have been committed to the database (156), the acknowledgement (220) including an order indicator (222) indicating the order in which the respective data (154) of the respective writes (152) have been committed to the database (156) relative to other writes (152) in the series of writes (152), and the operations include, For each respective write (152) in the series of writes (152), In response to receiving the confirmation (220) that each of the respective data (154) of the respective writes (152) has been committed to the database (156), based on the order indicator (222) indicating that each of the respective data (154) has been committed to the database (156) with respect to other writes (152) within the series of writes (152), transmitting each of the respective data (154) of the respective writes (152) from the first buffer (210) to the first data streaming application (170), based on the order indicator (222) indicating that each of the respective data (154) has been committed to the database (156) with respect to other writes (152) within the series of writes (152), transmitting each of the respective data (154) of the respective writes (152) from the second buffer (210) to the second data streaming application (170), and the system comprising.
12. Transmitting each of the respective data (154) from the first buffer (210) to the first data streaming application (170) based on the order in which each of the respective data (154) has been committed to the database (156) with respect to other writes (152) within the series of writes (152) includes using the order indicator (222) to determine that a confirmation (220) of a preceding write (152) has been received, and immediately before each of the respective data (154) of the respective writes (152), each of the respective data (154) of the preceding write (152) has been committed to the database (156). The system (100) according to claim 11.
13. Determining that the confirmation (220) of the preceding write (152) has been received includes determining a level of a watermark (212) of the first buffer (210). The system (100) according to claim 12.
14. Transmitting each of the data (154) from the second buffer (210) to the first data streaming application (170) based on the order in which each of the data (154) was committed to the database (156) relative to other writes (152) within the series of writes (152) comprises: using the order indicator (222) to determine that an acknowledgement (220) of a preceding write (152) has not been received, and further comprising, for each of the data (154) of each of the writes (152), before transmitting the data (154) of each of the writes (152), after determining that an acknowledgement (220) of the preceding write (152) has not been received, receiving the acknowledgement (220) that the data (154) of the preceding write (152) has been committed to the database (156); in response to receiving the acknowledgement (220) that the data (154) of the preceding write (152) has been committed to the database (156), transmitting the data (154) of the preceding write (152) from the second buffer (210) to the second data streaming application (170); after transmitting the data (154) of the preceding write (152), transmitting the data (154) of each of the writes (152) from the second buffer (210) to the second data streaming application (170); The system (100) according to claim 12 or claim 13, comprising: **Claim 15** The system (100) according to any one of claims 11 to 14, wherein the order indicator (222) includes an increment identifier. **Claim 16** The system (100) according to any one of claims 11 to 15, wherein the order indicator (222) includes a timestamp. **Claim 17** The system (100) of claim 16, wherein the operation further includes generating a linked list that orders the series of writes (152) in the order in which the respective data (154) of each of the series of writes (152) was committed to the database (156) using the timestamps of each received confirmation (220). [
18. ] The system (100) of claim 17, wherein generating the linked list includes generating a hash map. [
19. ] Receiving, from the database (156), the confirmation (220) that the respective data (154) of the respective write (152) was committed to the database (156) comprises: determining that a threshold period has elapsed without receiving the confirmation (220); in response to determining that the threshold period has elapsed without receiving the confirmation (220), obtaining a change log (330) from the database (156); determining, from the change log (330), that the respective data (154) of the respective write (152) was committed to the database (156); The system (100) according to any one of claims 11 to 18, comprising: [
20. ] The system (100) according to any one of claims 11 to 19, wherein the database (156) includes a Structured Query Language (SQL) database.
Citation Information
Patent Citations
Remote transfer method for magnetic disk controller
JP1999305950A
Storage system
JP2002049517A
Storage device system and data replication method
JP2004145855A
Storage system, control program for information processing unit, and control method for storage system
JP2014199581A
Database replication
US8838539B1