Cache Update via Distributed Message Queue
The distributed message queue system updates the local cache of client nodes in a multi-partition database, which solves the problem of outdated client node information, improves system performance and scalability, and reduces database load.
Patent Information
- Application Number
- CN202080104292.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-03
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-08-03
AI Technical Summary
In a multi-partition database, the local cache of client nodes may cause information outdated due to dynamic changes in database routing or other events, causing uneven timing of system resources and slow response time.
The distributed message queue system is adopted to update the local cache on the client node through cache update messages, and to periodically invalidate the local cache using message broker protocols such as AMQP, and to generate cache update messages immediately when the database data changes, and to update the local cache asynchronously.
Reduces client nodes' access to databases, improves system performance and scalability, ensures data timeliness of local caches, and reduces database load for read and write-intensive workloads.
Smart Images

Figure CN116134435B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the technical field of data storage. Background Art
[0002] A multi-partition database can provide horizontal scalability by partitioning data services among multiple computing devices (also referred to as "nodes"). For example, high availability and fault tolerance of data in the database can be achieved by replicating the database across multiple nodes and increasing the number of nodes as needed to handle increased amounts of data and / or workload. A client service can access database nodes to read or write data corresponding to the database. In some cases, a client node can maintain a local cache of a subset of data from the database so that the client can route read and write requests to the correct database node. However, database routing may change dynamically, or other events may occur that can cause the client to have stale information in its local cache, which can lead to uneven timing issues among system resources, slow system response times, etc. Summary of the Invention
[0003] Some embodiments include a first computing device that can receive a first request from a second computing device. Additionally, the first computing device can be capable of communicating with multiple database nodes, each of which maintains a portion of a database distributed across the multiple database nodes. Further, the first computing device can maintain a local cache of a subset of the information maintained in the database. The first computing device can send a second request to a first database node based on the first request to cause the first database node to change data in the database. Additionally, the first computing device receives a cache update message from a message queue among multiple distributed message queues based on the change to the data in the database. The first computing device can update the local cache based on the cache update message. Brief Description of the Drawings
[0004] The detailed description is set forth with reference to the accompanying drawings. In the drawings, the leftmost (one or more) digits of the reference numeral identify the figure in which the reference numeral first appears. The same reference numerals are used in different drawings to indicate similar or identical items or features.
[0005] Figure 1 Illustrates an example architecture of a system that uses message passing for local cache updates according to some embodiments.
[0006] Figure 2 Is a block diagram illustrating an example logical configuration of a system according to some embodiments.
[0007] Figure 3It is a block diagram illustrating an example of creating a new bucket according to some embodiments.
[0008] Figure 4 It is a block diagram illustrating an example of invalidating a local cache using a distributed messaging queue according to some embodiments.
[0009] Figure 5 It is a block diagram illustrating an example of updating a local cache according to some embodiments.
[0010] Figure 6 It is a flowchart of an example process of updating a local cache using a distributed messaging queue according to some embodiments.
[0011] Figure 7 Illustrated are selected example components of a (one or more) service computing device that can be used to implement at least some of the functions described herein. Detailed Description
[0012] Some embodiments herein relate to techniques and arrangements for a distributed computing system, where a distributed messaging queue system is used to aggregate and deliver cache invalidation messages to relevant targets. For example, the system can include a distributed database that can tolerate the client services using slightly stale data for some operations to benefit performance and improve scaling characteristics. This allows the system to greatly reduce the database load for both read-intensive workloads and write-intensive workloads, especially in the case of including one or more additional optimizations, as discussed further below.
[0013] Some examples include innovative distributed caches capable of operating in heterogeneous cloud (and / or multi-cloud) environments. For example, in a heterogeneous cloud environment, various distributed nodes with different resource characteristics (e.g., computing, memory, network, etc.) can work together. However, standard synchronization techniques such as chat publisher / subscriber protocols for synchronizing caches do not operate well in such an environment. Thus, some embodiments herein employ multiple in-memory local caches on the respective client nodes of the client services implementing the system. For example, the local cache can mirror certain database values used by the client services. Additionally, the systems herein can employ a message broker queue, such as by using the Advanced Message Queuing Protocol (AMQP) to periodically invalidate and / or synchronize the local cache.
[0014] In some cases, each cache data element can be configured to expire after a configurable time so that the data element does not become stale. When a new value is written to the database, each client can be notified via an invalidation message that the corresponding data item has been invalidated. For additional optimization, the invalidation message can contain information about the new data value. Thus, in some cases herein, the client node only performs a database read when the client's local cache does not have a record of the specified data item or if the data item has become invalid without any updated value.
[0015] Some examples herein use a message broker protocol to invalidate local caches and to achieve local cache synchronization across a distributed set of local caches. Additionally, some implementations employ delayed publication of messages to improve performance and scalability in a distributed system based on message broker queuing. For example, message queuing is inherently less costly than some other techniques due to the ability to hold messages for a longer period of time before delivery.
[0016] In some examples herein, the client node routes its corresponding read and write requests to the appropriate subset of database nodes for each request. Additionally, database routing can change dynamically, so the client device can maintain current routing information based on the implementations described herein, even though the computing resources, network resources, and storage resources on the database nodes and / or client nodes may vary, which may, for example, cause uneven timing issues among the participating entities in the system. Thus, some examples herein connect multiple heterogeneous systems, which can include connecting a public cloud storage device to a local or private system.
[0017] The implementations herein address cache issues encountered in scalable cloud storage configurations having multiple distributed database nodes storing and serving information and multiple client nodes locally storing a subset of the information stored in the database for efficient access. Additionally, some examples herein can include a distributed system that includes a set of database nodes (in some examples, metadata gateway devices) and a set of client services executed by client nodes that are clients of the distributed database provided by the database nodes. For example, the database nodes can store and serve information, and the client nodes can access or mirror the information in the database.
[0018] For discussion purposes, some example embodiments are described in the context of one or more service computing devices that communicate with a cloud storage system for using a distributed metadata database to manage storage of and access to data. However, the embodiments herein are not limited to the specific examples provided and can be extended to other types of computing system architectures, other types of databases, other types of storage environments, other types of client configurations, other types of data, etc., as will be apparent to those skilled in the art in view of the disclosure herein.
[0019] Figure 1 FIG. illustrates an example architecture of a system 100 that employs messaging for local cache updates according to some embodiments. System 100 includes a plurality of service computing devices 102 that are capable of communicating with or otherwise coupled to at least one network storage system 104, such as via one or more networks 106. Additionally, service computing devices 102 are capable of communicating with one or more user devices 108 and one or more administrator devices 110 via (one or more) networks 106, and the one or more user devices 108 and one or more administrator devices 110 can be any of a variety of types of computing devices, as discussed further below.
[0020] In some examples, service computing devices 102 can include one or more servers that can be embodied in any number of ways. For example, at least a portion of the programs, other functional components, and data storage of service computing devices 102 can be implemented on at least one server (such as in a server cluster, server farm, data center, cloud-hosted computing service, etc.), but other computer architectures can be used additionally or alternatively. Additional details regarding Figure 7 service computing devices 102 are discussed below.
[0021] Service computing devices 102 can be configured to provide storage and data management services to a user 112. As several non-limiting examples, user 112 can include users who perform functions of a business, enterprise, organization, government entity, academic entity, etc., and in some examples, can include those who store very large amounts of data. However, the embodiments herein are not limited to any particular use or application of system 100 and the other systems and arrangements described herein.
[0022] In some examples, the network storage system(s) 104 may be referred to as a "cloud storage device" or "cloud-based storage device", and in some cases, may implement a lower cost per megabyte / gigabyte storage solution than the local storage devices available at the service computing device 102. Additionally, in some examples, the network storage system(s) 104 may include commercially available cloud storage devices known in the art, while in other examples, the network storage system(s) 104 may include private or enterprise storage systems that are only accessible by entities associated with the service computing device 102, or a combination thereof.
[0023] The network(s) 106 may include any suitable network, including a wide area network such as the Internet; a local area network (LAN) such as an intranet; a wireless network such as a cellular network, a local wireless network (such as Wi-Fi), and / or short-range wireless communication (such as ); a wired network including fiber channel, fiber optic, Ethernet or any other such network, a direct wired connection, or any combination thereof. Thus, the network(s) 106 may include wired and / or wireless communication technologies. The components for such communication may depend at least in part on the type of network, the selected environment, or both. The protocols for communicating over such networks are well known and will not be discussed in detail herein. Thus, the service computing device 102, the network storage system(s) 104, the user device 108, and the management device 110 are capable of communicating over the network(s) 106 using wired or wireless connections and combinations thereof.
[0024] Additionally, the service computing devices 102 may be able to communicate with each other over the network(s) 107. In some cases, the network(s) 107 may be a LAN, a private network, etc., while in other cases, the network(s) 107 may include any of the networks 106 discussed above.
[0025] Each user device 108 may be any suitable type of computing device, such as a desktop computer, a laptop computer, a tablet computing device, a mobile device, a smart phone, a wearable device, a terminal, and / or any other type of computing device capable of sending data over a network. The user 112 may be associated with the user device 108, such as through a corresponding user account, user login credentials, etc. Additionally, the user device 108 may be able to communicate with the service computing device(s) 102 over the network(s) 106, over a separate network, or through any other suitable type of communication connection. Many other variations will be apparent to those skilled in the art who benefit from the disclosure herein.
[0026] In addition, each user device 108 may include a corresponding example of a user application 114 that may be executed on the user device 108, such as for communicating with a user web application 116 that may be executed as a service on one or more service computing devices 102, such as for sending user data for storage on the network storage system 104 and / or for receiving stored data from the network storage system 104 via a data request 118, etc. In some cases, the application 114 may include a browser or may be operable via a browser, while in other cases, the application 114 may include any other type of application having communication capabilities enabling communication with the user web application 116 via one or more networks 106.
[0027] In system 100, user 112 may store data to and receive data from the (one or more) service computing devices 102 with which their respective user devices 108 communicate. Thus, the service computing devices 102 may provide a storage service for user 112 and the corresponding user devices 108. During steady-state operation, there may be periodic communication between user 108 and the service computing devices 102, such as to read or write data.
[0028] Additionally, the administrator device 110 may be any suitable type of computing device, such as a desktop computer, laptop computer, tablet computing device, mobile device, smart phone, wearable device, terminal, and / or any other type of computing device capable of sending data over a network. Administrator 120 may be associated with the administrator device 110, such as via a corresponding administrator account, administrator login credentials, etc. Further, the administrator device 110 may be capable of communicating with the (one or more) service computing devices 102 via one or more networks 106, via a separate network, or via any other suitable type of communication connection.
[0029] In addition, each administrator device 110 may include a corresponding instance of an administrator application 122 that may be executed on the administrator device 110, such as for communicating with an administrative web application 124 that may be executed as a service on one or more service computing devices 102. For example, administrator 120 may use the administrator application to send administrative instructions for managing system 100, as well as for sending administrative data for storage on the (one or more) network storage systems 104 and / or for retrieving stored administrative data from the (one or more) network storage systems 104, such as via an administrative request 126, etc. In some cases, the administrator application 122 may include a browser or may be operable via a browser, while in other cases, the administrator application 122 may include any other type of application having communication capabilities enabling communication with the administrative web application 124 via one or more networks 106.
[0030] The service computing device 102 may execute a storage program 130 that may provide a gateway to a (one or more) network storage system 104, such as for sending data to be stored to the (one or more) network storage system 104 and for retrieving requested data from the (one or more) network storage system 104. Additionally, the storage program 130 may manage data stored by the system 100, such as for managing data retention periods, data protection levels, data replication, etc.
[0031] The service computing device 102 may further include a metadata database (DB) 132 that may be divided into a plurality of DB partitions 134(1)-134(N), and the metadata database (DB) 132 may be distributed across multiple service computing devices 102. For example, the metadata DB 132 may be used to manage object data 136 stored at the (one or more) network storage system 104. The metadata DB 132 may include a lot of metadata about the object data 136, such as information about individual objects, how to access individual objects, the storage protection level of the objects, the storage retention period, object owner information, object size, object type, etc. In addition, a DB management program 138 may manage and maintain the metadata DB 132, such as for updating the metadata DB 132 when storing new objects, deleting old objects, migrating objects, etc. The service computing device 102 including the database partition 134 may be referred to as a database node 140, and each may maintain a portion of the database 132 corresponding to one or more of the partitions 134.
[0032] Additionally, services executed thereon ( Figure 1Examples of the services shown include the user web application 116 and the service computing device 102 of the management web application 124, which may be referred to as a client node 142. Each client node 142 may maintain a corresponding local cache 146 (also referred to as a "near cache" or "local view" in some cases), such as the first local cache 146(1) and the second local cache 146(2) in the illustrated example. The client node 142 may operate as a client with respect to the metadata database 132. In some cases, the client node 142 may update the local cache 146 that may be maintained on the client node 142. For example, the local cache 146 may be periodically updated based on updates to the database 132 and / or by other techniques discussed further below. Thus, as an example, when the user web application 116 receives a data request 118 from the user device 108, the user web application 116 may access the local cache 146(1) to determine the database node 140 with which to communicate to execute the data request 118. By using the local cache 146(1), the user web application 116 is able to reduce the number of queries for obtaining the desired information from the metadata DB 132 to execute the data request 118.
[0033] In addition, some or all of the service computing devices 102 may include corresponding instances of a node hypervisor 148, which is executed by the corresponding service computing device 102 to manage the corresponding service computing device 102 as part of the system 100 and perform other functions attributed to the service computing device 102 herein. In the case where the service computing device 102 is a database node 140, the node hypervisor may also manage the configuration of the database node 140 to perform actions such as configuring the database node 140 into a partition group and controlling the operations of the partition group.
[0034] As a non-limiting example, database node 140 can be configured in a Raft group according to the Raft consensus algorithm to ensure data redundancy and consistency of database partition 134 of the distributed metadata database. According to the RAFT algorithm, one database node 140 of each partition group can be elected as the leader and can be responsible for serving all read and write operations of that database partition 134. Thus, the leader node can be used as a metadata gateway for client nodes 142. The other database nodes 140 are follower nodes that receive copies of all transactions so that they can update their own metadata database information. If the leader node fails or times out, one of the follower nodes among the follower nodes can be elected as the leader and can take over serving read and write transactions. The client nodes of the metadata system in this document are able to discover (e.g., by accessing the corresponding local cache 146 or sending a query) which database node 140 is the leader of each partition 134 and direct requests to that database node 140.
[0035] Thus, the examples in this document include systems capable of routing requests to a highly available and scalable distributed metadata database 132. The metadata database 132 in this document can provide high availability by maintaining strongly consistent copies of the metadata on separate metadata nodes 140. In addition, the distributed metadata database 132 provides scalability by partitioning the metadata and distributing the metadata across different metadata nodes 140. In addition, the solutions in this document optimize the ability of client applications to find the partition leader for a given request.
[0036] To be able to update the local cache in an efficient manner, at least some of the service computing devices 102 can execute a messaging program 150. For example, the messaging program 150 can enable the creation of cache update messages 152 for updating the local queue 146 after changes to database data, database configuration, etc. In some examples, the messaging program employed in this document can include a message broker that implements one or more of the Advanced Message Queuing Protocol (AMQP), the Streaming Text Oriented Messaging Protocol (STOMP), the Message Queuing Telemetry Transport (MQTT), and / or other suitable messaging protocols. Several non-limiting examples of software that can be used in some embodiments include APACHE QPID, JORAM, APACHE ACTIVEMQ, and RABBITMQ. For example, AMQP is a standard protocol capable of connecting applications on different platforms.
[0037] In some cases, items in the local cache 146 may become invalid by timing out or upon receipt of a message indicating that the value has been updated. The invalid items can be effectively removed from the local cache 146 by the client nodes 142 (e.g., by the respective service(s) executing on the respective client nodes 142). For example, each service executing on a client node may maintain its own local cache 146, which can be used by the respective service and can be updated by the service based on received cache update messages 152.
[0038] In some examples, when the value of a data item in the database 132 changes, a cache invalidation message 152 that provides a notification of the change using the AMQP messaging protocol can be immediately published, such as based on an instruction from the metadata node 140 that has information about the change. The update to the database 132 and the generation of the cache update message 152 can be performed online. For example, the cache update message 152 can be generated immediately after the database is updated, but the update to the database and the generation of the cache update message 152 can be performed asynchronously relative to each other.
[0039] The client nodes 142 can be configured to listen for cache update messages 152 that indicate data change events related to the (one or more) cache types of their respective local caches 146. For example, there may be different types of local caches 146 for different data types, such as for different types of services executed by the client nodes 142. The cache update messages 152 can be individually routed for each different cache type, such that client nodes 142 having a cache type different from the cache type to which a particular cache update message 152 belongs do not need to process data that they will not use. Additionally, each local cache instance of the appropriate type receives the cache update message 152 of that data type indicating the invalidation of the data item.
[0040] In some examples, a node such as database node 140, client node 142, or other computing nodes herein can be a single physical machine or virtual machine that can maintain one or more of the programs, services, or data described herein. All logical components (e.g., metadata gateway) as well as client services can be executed on any physical service computing device 102 within system 100. The distributed metadata database 132 can use dynamic partitioning, where the data stored by corresponding metadata nodes 140 can be divided into a set of manageable chunks (partitions 134) to distribute the data of database 132 across multiple metadata nodes 140. As a partition 134 grows, the system can dynamically split the partition 134 to form two or more new partitions, and can migrate the new partitions to metadata nodes 140 that have sufficient storage capacity to receive them and / or to newly added metadata nodes 140.
[0041] When a new metadata node 140 is added to system 100, communication within system 100 may slow down, such as to accommodate data growth. For example, at least a portion of the information included in local cache 146 may become invalid. Similarly, the client service nodes can also be scaled by adding new client nodes 142 to match the incoming workload. As described above, each client node 142 can maintain one or more local caches 146, which are in-memory caches that mirror a subset of the information stored in the metadata database nodes. The local cache 146 can increase system efficiency by greatly reducing the need for and frequency of database queries. For example, in a highly distributed system, constant access from a client node that directly results in a database query (e.g., where the requested data is stored in a permanent medium such as a hard disk) can be expensive and can increase system latency.
[0042] In some cases, the data in the distributed database 132 can be updated by a user request, and the update can occur on only a specific database node 140. Thus, when a mirror of the updated data exists in the local cache 146 of one or more client nodes 142, then the data becomes stale or invalid. Accordingly, some examples herein can use a distributed invalidation scheme to keep the data in the local cache 146 asynchronously refreshed. As an example, the service can periodically mark the data in the local cache 146 as invalid. Upon arrival of a new request for the data, the service can update the local cache 146 by querying the database 132 for the latest value in database 132.
[0043] Example algorithms for invalidating and updating local cache 146 include the following. (1) Each client node 142 maintains a local cache 146 of data that the client node 142 has previously retrieved from database 132 as needed. (2) The client node 142 may have multiple local caches 146, each local cache 146 storing a different type of data and configured with different optimization parameters, such as for use by one or more services executing on the client node 142. (3) Each data item in the local cache 146 may expire after a configurable amount of time. For example, the expiration time may be chosen to minimize database access while also preventing data in the local cache 146 from becoming stale. (4) When the value of a data item in database 132 changes, a cache update message 152 announcing the change may be generated and published immediately using a messaging protocol such as AMQP. Updates in database 132 and generation of the message may be performed online. (5) Each service listens for cache update messages 152, which indicate a data change event for the corresponding type of cache used by the service. In some examples, cache update messages 152 may be routed separately for each cache type, such that the client node 142 and the service managing the local cache do not need to process data that they will not use. (6) Each local cache instance (of a specified type) receives a cache update message 152 of that data type indicating invalidation of a data item. (7) An item may become invalid by either timing out or upon receipt of a cache update message 152 indicating that the item's value has been updated. The invalid item is effectively removed from the local cache, such as by marking the item as deleted or otherwise allowing the storage location of the item to be overwritten in due course. (8) When a client needs to access a data item, the cache immediately returns any value that it has stored. If no value is stored, or if the value has been invalidated, the program managing the local cache 146 may request the current value from database 132. (9) Cache update messages 152 may be configured with a "time-to-live" value so that they cannot survive beyond their useful life. For example, items in the local cache 146 may expire automatically after a certain time. (10) As an additional optimization, in some examples, cache update messages 152 may contain part or all of the value of the updated data. This increases the likelihood that multiple client nodes may update the value simultaneously. In such cases, the program managing the corresponding local cache 146 may use a distributed beat counter to identify which value is the most up-to-date. For write-intensive workloads, this optimization may further reduce the database load.
[0044] Using the architectures and algorithms discussed above, for read-intensive workloads, the amount of access by client node 142 to database 132 can be greatly reduced. Additionally, system 100 can be configured with an expiration threshold that prevents local cache 146 from becoming stale beyond that threshold. Separate expiration thresholds can be configured for each different type of local cache 146 and / or for individual local caches of the same type on different client nodes 142. For example, local cache 146 can be configured to expire least recently used data to enforce memory usage limits.
[0045] In addition, local cache 146 can be updated due to internal events. For example, local cache 146 can be configured to store all system metadata except user object metadata, which can grow to trillions of bytes of data. In such cases, the system metadata mirrored in local cache 146 can be user-driven or internal system metadata, such as a metadata partition map. A metadata partition map is a table or other data structure that includes partition identifiers (IDs) and the IDs of the database nodes 140 on which the corresponding partitions reside. All user requests related to object management, such as putObject and getObject requests, can cause a service (e.g., user web app 116) to look up at least four different types of metadata, such as user information, bucket information, partition information, and object information. Thus, in some cases, it is possible that all four types of metadata are refreshed due to a single user request. To avoid this, some examples in this document can employ a dynamic metadata partitioning technique that involves partitioning all metadata types and tables into partitions 134. Partitions 134 are distributed across database nodes 140 to provide unified load management. When a metadata partition is split into two or more partitions, partition map invalidation can occur. Although partition splits may not be driven by user requests, the invalidation and further refresh processes can be similar to the refreshes caused by user operations such as the putObject and getObject requests discussed above.
[0046] In addition, to achieve better response times for end users, cache updates in this document can be performed asynchronously relative to user write requests. As a result, there may be a small delay before the local cache 146 of services distributed across the cross-system 100 on the client node is invalidated or otherwise updated. For example, the actual time to invalidate or otherwise update the local cache 146 can vary based on network and system activity. The AMQP messaging protocol that can be used in this document is inherently robust because cache update messages 152 can be queued before delivery. Thus, aggregated cache update messages 152 can be updated across multiple databases. For example, assuming that bucket, user, and partition map updates all occur and are enqueued simultaneously, only one cache update message 152 may actually be sent to the corresponding service on the client node 142. However, in some cases, such as due to temporary network failures, etc., cache update messages 152 via AMQP may still be lost. Thus, the embodiments in this document can include a mechanism to retry delivering messages up to a specific attempt threshold.
[0047] In addition, in the case of a delivery failure of a queued cache update message 152, the local cache 146 can perform invalidation or other cache updates based on exceeding a time threshold. For example, when the last update for a given cache stored value exceeds a certain timeout value, the local cache memory 146 can be configured to automatically invalidate the entry. The timeout threshold for invalidation in this document can be configurable such that the timeout threshold can be adjusted based on system workload dynamics, etc.
[0048] In some cases, the service computing devices 102 can be arranged in one or more groups, clusters, systems, etc. at the site 154. In some cases, multiple sites 154 can be geographically dispersed from each other, such as for providing data replication, disaster recovery protection, etc. In addition, in some scenarios, the service computing devices 102 at multiple different sites 154 can be configured to communicate securely with each other, such as for providing a federation of multiple sites 154.
[0049] Figure 2 is a block diagram illustrating an example logical configuration of a system 200 according to some embodiments. In some examples, the system 200 can correspond to the system 100 discussed above or any of various other possible computing system architectures, which will be apparent to those skilled in the art who benefit from the disclosure in this document. The system 200 can implement distributed object storage and can include a front-end service using a web application for users and administrators. In some cases, the system 200 can connect network storage devices ( Figure 2Objects on the (not shown in the figure) are stored in buckets that can be created by end users 112, 120. System 200 can use resources distributed across local deployments and cloud systems to achieve complex management and storage of data. In system 200, scalability can be provided by logically partitioning the stored metadata in distributed database 132.
[0050] Figure 2 System 200 can include a distributed system of client nodes 142 and database nodes 140. Database nodes 140 can tolerate client nodes 142 using stale cache data under certain operations, which can improve system performance and scalability characteristics. This provides a reduced database load for read-intensive workloads and, in some examples, also for write-intensive workloads. For example, when a new value is written to distributed database 132, each client node 142 with a corresponding local cache 146 can receive a cache update message 152 indicating that the corresponding data item has been invalidated. For additional optimization, cache update message 152 can contain information about the new data value. Additionally, a database read can be performed when local cache 146 does not have a record of the data item or the data item has become invalid without any updated values.
[0051] In this example, system 200 includes a message passing queue grid 202 for queuing and routing cache update messages 152 to services executed on client nodes 142. For example, message passing queue grid 202 can be provided by message passing program 150 discussed above with respect to Figure 1 and can be hosted on one or more message passing nodes 204. Message passing nodes 204 can correspond to one or more of the service computing devices 102 that can execute message passing program 150 discussed above. In some cases, message passing program 150 can be executed on the same service computing device 102 that serves as database node 140, and / or on the service computing device 102 that serves as client node 142, and / or on other service computing devices 102 in system 100 discussed above with respect to Figure 1 Message passing queue grid 202 can include multiple message queues, such as message queue 208(1), message queue 208(2), and message queue 208(3), each of which can be maintained in a separate virtual container, such as a DOCKER container, etc. In some examples, message queues 208(1)-208(3) can be maintained on separate physical machines or virtual machines.
[0052] In this example, the distributed database 132 includes a plurality of metadata gateways 210 that can correspond to database nodes 140. For example, as described above, in some examples, each database partition 134 can be maintained by a partition group of two or more database nodes 140. Each partition group can have a leader node that responds to user read and write requests for the corresponding partition 134. Thus, the partition leader of each partition group can serve as the metadata gateway 210 for that partition 134. In this example, for purposes of explanation, four metadata gateways 210(1)-210(4) are illustrated; however, in an actual implementation, some examples of the systems herein can include a greater number of metadata gateways 210, depending on the number of database partitions 134.
[0053] The message queue 208 is configured to deliver cache update messages 152 to services executed on the client nodes 142. In this example, the first service program 212 can correspond to the user web app 116 that can provide user device 108 data access services as discussed above. For example, the first service program 212 can maintain a local cache 146(1) that contains information that can be used to access the metadata gateway 210. For example, the first service program 212 can provide client functionality that enables the client node 142(1) to interact with the metadata gateway 210 to retrieve metadata. In addition, the first service program 212 can provide functionality for receiving cache update messages 152 for updating the associated local cache 146(1). In addition, the first service program 212 can interact with the storage program 130( Figure 2 (not shown in the figure), such as for retrieving object data 136 based on the retrieved metadata and / or metadata maintained in the local cache 146(1), such as discussed above with respect to Figure 1 . In addition, the first service program 212 can exchange communications 214 with the user device 108, such as for sending or receiving user data.
[0054] In this example, the second client node 142(2) can also execute an instance of the first service program 212, can maintain a local cache 146(2), and can exchange communications 216 with another user device 108. Additionally, in this example, the third client node 142(3) executes two services, namely the second service program 218 and the third service program 220. For example, the second service program can correspond to the one discussed above with respect to Figure 1The management web application 124 that provides management services to an administrator is discussed. For example, the second service program 218 may include client functions for interacting with other nodes in the system 200, including client nodes 142, messaging nodes 204, and / or database nodes 140. The second service program 218 may exchange communications 222 with the administrator device 110, such as for receiving management instructions, providing status updates, and the like. The second service program 218 may maintain a local cache 146(3), which in some examples may contain one or more data types different from the data types maintained in local caches 146(1) and 146(2), or vice versa.
[0055] In addition, the third service program 220 may provide another type of service different from the services provided by the first service program 212 and the second service program 218. As several non-limiting examples, the third service may include garbage collection, object data management, and the like. The third service program 220 may exchange communications 224 with the administrator device 110, such as for receiving management instructions, providing status updates, and the like. The third service program 220 may maintain a local cache 146(4), which in some examples may include one or more data types different from the data types maintained by local caches 146(1), 146(2), and 146(3), or vice versa.
[0056] When the metadata gateway 210 changes the value of a data item in the distributed database 132 or otherwise makes a change to the database 132, the metadata gateway 210 may send an enqueue instruction 230, which may include information about the changed value for one of the message queues 208. In some examples, a message queue 208 may be randomly selected, but other selection techniques may alternatively be used.
[0057] Receipt of the enqueue instruction may cause the messaging program 150 at the corresponding messaging node 204 to generate a cache update message 152 and add the cache update message 152 to the corresponding message queue 208. For example, the cache update message 152 may be generated, queued, and distributed according to the AMQP messaging protocol. As described above, the cache update message 152 may be advertised or otherwise routed to the corresponding service programs 212, 218, 220 executing on the client nodes 142.
[0058] As an example, cache update messages 152 can be routed separately for each different type of local cache 146 based on the data types included therein and the data types affected by an update to database 132. For example, if local caches 146(1) and 146(2) have one or more data types corresponding to the update and local caches 146(3) and 146(4) do not include these one or more data types, then based on the identification of the data types affected by the change, cache update messages 152 for the first service program 112 are not routed to the second service program 218 or the third service program 220. An indication of the data types affected by the change can be provided, for example, by the metadata gateway 210 that makes the change to database 132. Thus, service programs whose caches are not affected by changes to database 132 may not receive or process cache update messages 152 that are not relevant to their respective local caches 146.
[0059] Figures 3 to 5 Illustrates an example of creating a new bucket in database 132; invalidates existing local caches 146 of several services that include bucket information in their local caches due to a change in the database; and subsequently updates several local caches to include the updated database information. Figures 3 to 5 The example can correspond in part to the Figure 2 example system 200 discussed above; however, for clarity of illustration, client node 142(2) is omitted. Other non-participating components are also excluded.
[0060] Figure 3 is a block diagram illustrating an example 300 of creating a new bucket according to some embodiments. In this example, it is assumed that user 112 of user device 108 sends a write request 302 to cause a new bucket to be created at (one or more) network storage systems 104 ( Figure 3 not shown in the figure). The first service program 212 can receive the write request 302 and, in response, can send a write request 303 to the metadata gateway 210(3), such as based on routing information currently included in local cache 146(1). In response, the metadata gateway 210(3) can update the metadata in database 132, such as by creating a new record 304 for the new bucket. In this example, it is assumed that record 304 includes the name 306 of the bucket (i.e., "bucket 1") and the settings 308 of the bucket (which include a 3-day retention period and synchronization settings). A bucket can also be created at network storage system 104 ( Figure 3 not shown in the figure). The metadata gateway 210(3) returns a write response 310 to the first service program 212. In turn, the first service program 212 returns a write response 312 to the user device 108 indicating that the bucket has been created.
[0061] Figure 4 is a block diagram illustrating an example 400 of invalidating local caches using a distributed messaging queue according to some embodiments. In this example, associated with the creation of a new bucket, as discussed above with respect to Figure 3 the metadata gateway 210(3) can send an enqueue instruction 402 to one of the messaging nodes 204 that includes one of the message queues (i.e., message queue 208(2) in this example). As described above, in some examples, the message queue 208(2) and / or the messaging node 204 can be randomly selected by the metadata gateway 210(3). In other examples, the metadata gateway 210(3) can employ any other suitable technique for selecting one of the message queue 208(2) / messaging node 204 to receive the enqueue instruction 402. In some examples, the metadata gateway 210(3) can include a record 304 created in the database using the enqueue instruction 402; however, in this example, it is assumed that the metadata gateway 210(3) only identifies the data types of the metadata records affected by the changes made to the database (i.e., the bucket).
[0062] Based on receiving the enqueue instruction 402, the messaging node 204 can create a cache update message 152 to send to the services that maintain local caches with bucket information. In this example, it is assumed that all of the first service program 212, the second service program 218, and the third service program 220 include bucket information in their respective local caches 146(1), 146(3), and 146(4). The messaging node 204 can add the cache update message 152 to the message queue 208 and distribute the cache update message 152 to the services. For example, the messaging node can determine the type of local cache maintained by each service for correctly routing the cache update message 152. Thus, the messaging node 204 can use the message queue 208-2 to deliver the cache update message 152 to the first service program 212, the second service program 218, and the third service program 220. In response, the first service program 212 can invalidate or otherwise update the bucket portion of the local cache 146(1); the second service program 218 can invalidate or otherwise update the bucket portion of the local cache 146(3); and the third service program 220 can invalidate or otherwise update the bucket portion of the local cache 146(4).
[0063] Figure 5FIG. is a block diagram illustrating an example 500 of updating a local cache according to some embodiments. In this example, it is assumed that user 112 submits a get bucket request 502 to the first service program 212 using a user device. In response, the first service program 212 determines that the bucket portion of the local cache 146 has been invalidated and sends a get bucket request 504 for bucket information. In response, the metadata gateway 210(3) may provide a get bucket response 506 that includes a copy of the record 304 from the database 132, which can be added by the first service program 212 to the local cache 146(1).
[0064] Similarly, it is assumed that the administrator 120 sends a get object request 508 to the third service program 220 using the administrator device 110. The third service program 220 may send a get bucket request 510 to query the metadata gateway 210(3) to request information related to the bucket (bucket 1) that contains the requested object. In response, the metadata gateway 210(3) may send a get bucket response 512 to the third service program 220, which may include a copy of the metadata record 304 that the third service program 220 can add to the local cache 146(4) to refresh the bucket portion of the local cache 146(4).
[0065] In addition, although in this example, the service is queried by the metadata gateway to obtain updated information for a new bucket from the database 132, in other examples, as described above, the record 304 may have been included in a cache update message 152 previously sent to the service to invalidate the associated local cache 146. Thus, in this alternative example, the first service 212 and the third service 220 do not have to query the metadata gateway 210 for the bucket record 304 because the information was already included in the respective local caches 146(1) and 146(4).
[0066] Figure 6is a flowchart illustrating an example process for updating a local cache using a distributed messaging queue according to some embodiments. The process is illustrated as a collection of blocks in a logical flowchart that represents a sequence of operations, some or all of which can be implemented in hardware, software, or a combination thereof. In the context of software, the blocks can represent computer-executable instructions stored on one or more computer-readable media that, when executed by one or more processors, program the processors to perform the operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the blocks are described should not be construed as limiting. Any number of the described blocks can be combined in any order and / or in parallel to implement the process or an alternative process, and not all blocks need to be executed. For purposes of discussion, the process is described with reference to the environments, frameworks, and systems described in the examples herein, but the process can be implemented in a variety of other environments, frameworks, and systems. In Figure 6 , process 600 can be performed, at least in part, by a client node that executes one or more of service programs 212, 218, or 220.
[0067] At 602, a service computing device can partition a database across multiple database nodes to provide multiple partitions that are distributed across the multiple database nodes.
[0068] At 604, a client node can execute a service that maintains a local cache of a subset of the information maintained in the database.
[0069] At 606, the client node can receive a first request from a user computing device that affects data in the database. For example, the client node can receive a write request or other request that will change the data in the database.
[0070] At 608, the client node can send a second request to a first database node among the multiple database nodes based on the first request, the second request causing the first database node to change the data in the database.
[0071] At 610, the client node can receive a cache update message based on a change to the data in the database from one of the multiple distributed message queues.
[0072] At 612, the client node can determine whether the received cache update message includes updated data. If so, the process proceeds to 614. If not, the process proceeds to 616.
[0073] At 614, the client node can update the local cache to include the updated data included in the cache update message.
[0074] At 616, the client node can invalidate at least a portion of the local cache in response to a cache update message.
[0075] At 618, the client node can receive a third request from the user computing device to access data corresponding to data in the database.
[0076] At 620, the client node can send a query to at least one of the plurality of database nodes to determine information related to the third request from the database.
[0077] At 622, the client node can update the local cache at least in part based on the response to the query received from the database node.
[0078] The example processes described herein are merely examples of processes provided for discussion purposes. Given the disclosure herein, many other changes will be apparent to those skilled in the art. Additionally, while the disclosure herein sets forth several examples of suitable frameworks, architectures, and environments for performing processes, the embodiments herein are not limited to the specific examples shown and discussed. Additionally, the present disclosure provides various example embodiments as described and as shown in the figures. However, the present disclosure is not limited to the embodiments described and shown herein, but may extend to other embodiments as would be known or become known to those skilled in the art.
[0079] Figure 7 Selected example components of the (one or more) service computing devices 102 that can be used to implement at least some of the functions described herein are illustrated. The (one or more) service computing devices 102 can include one or more servers or other types of computing devices that can be embodied in any number of ways. For example, in the case of a server, programs, other functional components, and data can be implemented on a single server, a server cluster, a server farm, or a data center, a cloud-hosted computing service, etc., but other computer architectures can be additionally or alternatively used. The plurality of service computing devices 102 can be located together or separately and can be organized, for example, as virtual servers, server libraries, and / or server farms. The functions described can be provided by the servers of a single entity or enterprise, or can be provided by the servers and / or services of multiple different entities or enterprises.
[0080] In the illustrated example, the service computing device(s) 102 includes one or more processors 702, one or more computer-readable media 704, and one or more communication interfaces 706, or may be associated with one or more processors 702, one or more computer-readable media 704, and one or more communication interfaces 706. Each processor 702 may be a single processing unit or multiple processing units, and may include a single or multiple computing units or multiple processing cores. The processor(s) 702 may be implemented as one or more central processing units, microprocessors, microcomputers, microcontrollers, digital signal processors, state machines, logic circuits, and / or any device that manipulates signals based on operational instructions. As an example, the processor(s) 702 may include any suitable type of one or more hardware processors and / or logic circuits that are specifically programmed or configured to execute the algorithms and processes described herein. The processor(s) 702 may be configured to obtain and execute computer-readable instructions stored in the computer-readable media 704, which may program the processor(s) 702 to perform the functions described herein.
[0081] The computer-readable media 704 may include volatile and non-volatile memory and / or removable and non-removable media implemented in any type of technology for storing information such as computer-readable instructions, data structures, program modules, or other data. For example, the computer-readable media 704 may include, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, optical storage devices, solid-state storage devices, magnetic tape, magnetic disk storage devices, storage arrays, network-attached storage devices, storage area networks, cloud storage devices, or any other medium that can be used to store the desired information and can be accessed by a computing device. Depending on the configuration of the service computing device(s) 102, the computer-readable media 704 may be a tangible non-transitory medium, provided that non-transitory computer-readable media excludes media such as energy, carrier signals, electromagnetic waves, and / or signals themselves when mentioned. In some cases, the computer-readable media 704 may be in the same location as the service computing device 102, while in other examples, the computer-readable media 704 may be partially remote from the service computing device 102. For example, in some cases, the computer-readable media 704 may include a portion of the storage in the network storage system(s) 104 discussed above Figure 1 as part of the storage discussed above.
[0082] The computer-readable medium 704 can be used to store any number of functional components executable by the processor(s) 702. In many embodiments, these functional components include instructions or programs executable by the processor(s) 702 and, when executed, specifically programming the processor(s) 702 to perform the actions attributed herein to the service computing device 102. The functional components stored in the computer-readable medium 704 can include the user web application 116, the management web application 124, the storage program 130, the database management program 138, the node management program 148, and the messaging program 150, each of which can include one or more computer programs, applications, executable code, or portions thereof. Additionally, although these programs are illustrated together in this example, some or all of these programs can be executed on separate service computing devices 102 during use.
[0083] Additionally, the computer-readable medium 704 can store data, data structures, and other information for performing the functions and services described herein. For example, the computer-readable medium 704 can store the metadata database 132 including the database partitions 134. Additionally, the computer-readable medium can store the local cache(s) 146. Additionally, although these data structures are illustrated together in this example, some or all of these data structures can be stored on separate service computing devices 102 during use. The service computing device 102 can also include or maintain other functional components and data, which can include programs, drivers, etc., and data used or generated by these functional components. Additionally, the service computing device 102 can include many other logical, programming, and physical components, and those described above are merely examples relevant to the discussion herein.
[0084] One or more communication interfaces 706 can include one or more software and hardware components for enabling communication with various other devices, such as via one or more networks 106. For example, the communication interface(s) 706 can enable communication through one or more of a LAN, the Internet, a cable network, a cellular network, a wireless network (e.g., Wi-Fi), and a wired network (e.g., Fibre Channel, fiber optic, Ethernet), a direct connection, and proximity communication such as etc., as otherwise enumerated elsewhere herein.
[0085] The various instructions, methods, and techniques described herein can be considered in the general context of computer-executable instructions, such as computer programs and applications stored on a computer-readable medium and executed by one or more processors herein. Generally, the terms program and application can be used interchangeably and can include instructions, routines, modules, objects, components, data structures, executable code, etc. for performing a particular task or implementing a particular data type. These programs, applications, etc. can be executed as native code or can be downloaded and executed, such as in a virtual machine or other just-in-time compilation execution environment. Generally, the functionality of the programs and applications can be combined or distributed as needed in various embodiments. Embodiments of these programs, applications, and techniques can be stored on a computer storage medium or transmitted via some form of communication medium.
[0086] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.
Claims
1. A system, comprising: A first computing device capable of communicating with a plurality of database nodes, each database node maintaining a portion of the database based on a partition of the database to distribute the database across the plurality of database nodes, the first computing device maintaining at least one local cache of a subset of the information maintained in the database, the first computing device being configured by executable instructions to perform operations including the following: Receiving, by the first computing device, a first request from a second computing device, the first request affecting data in the database; Sending, by the first computing device, a second request to a first database node among the plurality of database nodes based on the first request, the second request causing the first database node to change data in the database; Receiving, by the first computing device, a cache update message from a message queue among a plurality of distributed message queues based on the change to the data in the database; And Updating, by the first computing device, the local cache based on the cache update message, The at least one local cache being configured to maintain a predetermined data type for different types of services performed by the first computing device, The first database node being configured to include the data type of the change to the data in the database, and The message queue among the plurality of distributed message queues being configured to route the cache update message to the first computing device at least based on determining that the local cache includes the data type of the change to the data in the database.
2. The system according to claim 1, wherein: Receiving the cache update message includes receiving updated metadata added to the database as the change to the data in the database; and Updating the local cache based on the cache update message includes updating the local cache to include the updated metadata.
3. The system according to claim 1, wherein updating the local cache based on the cache update message includes invalidating at least a portion of the local cache.
4. The system according to claim 1, the operations further including receiving the cache update message according to the Advanced Message Queuing Protocol.
5. The system according to claim 1, wherein the system includes a plurality of messaging nodes, and the plurality of distributed message queues are respectively provided by the plurality of messaging nodes.
6. The system according to claim 1, wherein the first computing device performs a first service on the first computing device, the first service maintaining the local cache, wherein the first service enables a second computing device to access data stored in a storage system corresponding to the database.
7. The system according to claim 6, wherein the first service is one of the following: A user web application; or A management web application.
8. The system according to claim 1, wherein the first request is a data write request for storing data at a storage system associated with the database.
9. The system according to claim 1, wherein updating the local cache based on the cache update message includes invalidating at least a portion of the local cache, and the operations further include: Receiving a third request from the second computing device; Determining that at least the portion of the local cache is invalidated; And Sending a query to at least one database node among the plurality of database nodes to determine information related to the third request from the database.
10. The system according to claim 9, further comprising sending metadata information to another computing device to obtain data from a remote storage system via a network based on a response received from the at least one database node.
11. A method, comprising: Receiving, by a first computing device, a first request from a second computing device, wherein the first computing device is capable of communicating with a plurality of database nodes, each database node maintaining a portion of the database based on a partition of the database to distribute the database across the plurality of database nodes, and the first computing device maintains at least one local cache of a subset of information maintained in the database; Sending, by the first computing device, a second request to a first database node among the plurality of database nodes based on the first request, the second request causing the first database node to change data in the database; Receiving, by the first computing device, a cache update message from a message queue among a plurality of distributed message queues based on the change to the data in the database; And Updating, by the first computing device, the local cache based on the cache update message, wherein the at least one local cache is configured to maintain a predetermined data type for different types of services performed by the first computing device, the first database node is configured to include the data type of the change to the data in the database, and the message queue among the plurality of distributed message queues is configured to route the cache update message to the first computing device at least based on determining that the local cache includes the data type of the change to the data in the database.
12. The method according to claim 11, wherein: Receiving the cache update message includes receiving updated metadata added to the database as the change to the data in the database; and Updating the local cache based on the cache update message includes updating the local cache to include the updated metadata.
13. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, configure the one or more processors to perform operations including the following: A first computing device receives a first request from a second computing device, where the first computing device is capable of communicating with a plurality of database nodes, and each database node maintains a portion of the database based on a partition of the database to distribute the database across the plurality of database nodes. The first computing device maintains at least one local cache of a subset of the information maintained in the database. The first computing device sends a second request to a first database node among the plurality of database nodes based on the first request, and the second request causes the first database node to change data in the database. The first computing device receives a cache update message from a message queue among a plurality of distributed message queues based on the change to the data in the database. And The first computing device updates the local cache based on the cache update message. The at least one local cache is configured to maintain a predetermined data type for different types of services performed by the first computing device. The first database node is configured to include the data type of the change to the data in the database. And The message queue among the plurality of distributed message queues is configured to route the cache update message to the first computing device at least based on determining that the local cache includes the data type of the change to the data in the database.
14. The one or more non-transitory computer-readable media according to claim 13, wherein: Receiving the cache update message includes receiving updated metadata added to the database as the change to the data in the database. And Updating the local cache based on the cache update message includes updating the local cache to include the updated metadata.
Citation Information
Patent Citations
Updating cache data
CN110347707A