Domain Name System-based global server load balancing service
The system optimizes DNS-based load balancing by monitoring server health and location to efficiently distribute traffic, addressing server overload and underutilization issues.
Patent Information
- Application Number
- JP2025501835
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-31
- Filing Date
- 2023-07-14
- Publication Date
- 2025-07-17
AI Technical Summary
Existing DNS-based load balancing systems fail to efficiently distribute traffic across servers, leading to server overload or underutilization, which degrades performance and availability.
A system that monitors server health periodically, maintains metadata on server health and location, and uses this information to process DNS queries, selecting servers based on health and location to optimize load balancing.
Enhances server utilization by ensuring that traffic is directed to healthy servers in proximity to the client, improving performance and availability.
Smart Images

Figure 2025523115000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Application No. 63 / 389,791, filed Jul. 15, 2022, and U.S. Provisional Application No. 63 / 470,140, filed May 31, 2023, each of which is hereby incorporated by reference in its entirety.
[0002] The subject matter described generally relates to load balancing of servers, and more specifically to domain name system (DNS) - based global server load balancing services.
Background Art
[0003] Domain Name System (DNS) load balancing distributes traffic across multiple servers to improve performance and availability. Organizations use various forms of load balancing to speed up websites and private networks. Without load balancing, web applications and websites cannot effectively handle traffic. DNS translates a website domain (e.g., www.xyz.com) into a network address, e.g., an IP (Internet Protocol) address. A DNS server receives DNS queries requesting the IP address of a domain. DNS load balancing configures the domain within the Domain Name System (DNS) so that client requests to the domain are distributed across a group of server machines. A domain can correspond to a website, a mail system, a print server, or another service. If DNS load balancing is insufficient, the performance of requests degrades due to overloading a particular server, not enough requests are sent to some of the servers, and the servers cannot be effectively utilized by keeping them busy, resulting in a decrease in server utilization.
Summary of the Invention
[0004] The system performs efficient Domain Name System (DNS)-based global server load balancing. The system periodically monitors the server health of the servers that process requests directed to the virtual servers. The updated server health information is used to process DNS queries that request an assignment of servers for processing requests directed to the virtual servers.
[0005] According to one embodiment, the system maintains metadata that describes the servers based on user requests associated with the virtual servers. The virtual servers are identified by a Uniform Resource Locator (URL), and requests to the virtual servers are processed by one or more of a plurality of servers. The system receives requests associated with the virtual servers, such as requests to generate, update, or delete information that describes the virtual servers. The system updates the information stored in the database based on the requests. The database stores records that map the virtual servers to the servers. The system propagates the updated information to a plurality of data plane clusters. Each data plane cluster includes a database that stores metadata that describes the plurality of servers.
[0006] For each of a plurality of servers, the system periodically updates one or more measurements of server health. The measurements of server health are updated for each of a plurality of data plane clusters. The measurements of server health are relative to a location within the data plane cluster and are associated with the communication protocol for reaching the server from computing devices within the location. For example, the communication protocol associated with the measurement of server health can be one of TCP (transmission control protocol), HTTP (hypertext transfer protocol), or HTTPS (hypertext transfer protocol secure), or ICMP (internet control message protocol). The system uses information describing server health to process DNS queries. Accordingly, the system receives a DNS query from a client device requesting a server for processing a request directed to a URL of a particular virtual server. The system identifies one or more candidate servers for processing a request directed to the URL of the particular virtual server based on information stored in the DNS cache. The system selects a candidate server from the one or more candidate servers based on factors including the measurement of the server health of the server and the location associated with the client device. The system sends a response to the DNS query to the client device. The response identifies the candidate server for processing the request directed to the URL of the virtual server.
[0007] According to one embodiment, the system monitors the server health of servers used to process requests directed to virtual servers and uses the monitored health information to respond to DNS queries. The system identifies a plurality of servers, and each server is associated with a virtual server. For example, the system can receive a request associated with a virtual server identified by a URL (uniform resource locator) and update information stored in a database based on the request, where the database can include records that map virtual servers to servers. The request can be to generate, update, or delete metadata information of the virtual server. The system generates a plurality of server health check tasks. Each server health check task is for determining a measurement of the server health of a server with respect to a location. Each measurement of server health is associated with a communication protocol for reaching the server from a computing device within the location. The communication protocol associated with the measurement of server health can be one of tcp, http, https, or icmp. The system divides the plurality of server health check tasks into a sequence of a plurality of buckets of tasks.
[0008] The system monitors the health of multiple servers by periodically repeating a series of steps. The system processes multiple server health check tasks within a time interval. The time interval includes multiple sub-time intervals. Multiple buckets of tasks are processed in sequence, and each bucket of tasks is processed during the sub-time interval assigned to the bucket. Processing of a bucket of tasks includes the following steps. The system sends each server health check task to a worker process associated with a communication protocol. The worker process determines a measurement of the server health of the server by communicating with the server using the communication protocol of the worker process. The system receives the result of the server health check from the worker process and propagates the result of the server health check to multiple data plane clusters. The system processes DNS queries, and each DNS query requests a server for processing a request directed to a virtual server. The system processes DNS queries by selecting a server based on the result of the server health check.
[0009] The disclosed technology can be implemented as a computer-implemented method, as a non-transitory computer-readable storage medium including instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the steps of the method disclosed herein, or as a computer system comprising a non-transitory computer-readable storage medium and one or more computer processors that, when executed by the one or more computer processors, cause the one or more computer processors to perform the steps of the method disclosed herein.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 4A
Figure 4B
Figure 4C
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11A
Figure 11B
Figure 11C
Figure 11D
Figure 11E
Figure 11F
Figure 11G
Figure 12
Best Mode for Carrying Out the Invention
[0011] The drawings and the following description illustrate specific embodiments for the purpose of illustration only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structure and method can be adopted without departing from the principles described. Whenever possible, similar or like reference numbers are used in the figures to indicate similar or like functions. If elements share a common number followed by a different letter, this indicates that the elements are similar or like. Unless the context indicates otherwise, references to numbers alone generally refer to any one or any combination of such elements.
[0012] FIG. 1 shows the overall system environment of a DNS-based global server load balancing service according to an embodiment. The load balancing system 100 enables the efficient generation of virtual servers for the backend where users are running services on physical servers. The load balancing system performs load balancing of traffic to virtual servers based on regions, operating states, weights, or any combination thereof. The load balancing system includes a control plane implemented as a control plane cluster 110 and data planes implemented as a plurality of data plane clusters 120a, 120b, 120c, 120d, 120e, 120f, etc. References to the control plane in this specification correspond to the control plane cluster 110. References to the data plane in this specification correspond to the data plane cluster 120.
[0013] Different data plane clusters can be implemented in different geographical regions. For example, data plane clusters 120a and 120b can be in one continent, while data plane clusters 120c and 120d can be in different continents. The control plane processes change requests entering the system via API (application programming interface) requests or UI (user interface) requests. The data plane is responsible for various functions including (1) providing records to clients, (2) performing health checks on target backends, (3) synchronizing data from the control plane, (4) providing changes in health check states, etc. Each data plane cluster 120 is a fully functional version of the data plane that operates independently. The system implements various types of load balancing mechanisms, such as weight-based load balancing or other types of load balancing. The system performs health checks and if a server is down, the system does not return that server as a result of DNS queries, so requests are not sent to that server. The system can be configured to enable responses from specific regions.
[0014] The system performs DNS-based load balancing aimed at providing load balancing through DNS resolution. The system receives a DNS query from an application, processes it, and maps the received domain name or alias to an IP address provided as a result of the DNS query and used by the application to connect to a server. This is different from conventional load balancers that intercept traffic, such as network load balancers. The system receives requests from applications such as browsers, for example, Chrome (registered trademark). The request can be for a connection to a target domain or alias, for example, www.xyz.com. The application attempts to execute a DNS query to resolve the alias. The DNS query can be forwarded to a caching server. If there is no response to the DNS query in the cache, the system forwards the request to an authority server. The system executes the query to determine the network address (for example, an IP address) provided as a result of the query. The application directly connects to the target server via this network address.
[0015] The system can perform different types of load balancing, such as round-robin load balancing, set-based load balancing, or other types of load balancing. The system implements a DNS round-robin load balancing strategy by using a DNS resource record set with multiple IP addresses that are cycled for each DNS request returned by the DNS caching infrastructure. This provides a uniform near-time distribution across many IP addresses because the cache layer can hold all the information from all the servers in its set and provide a load-balanced response for each query.
[0016] Figure 1 uses like reference numerals to identify like elements. The letters after a reference numeral, such as "120A", indicate that the text specifically refers to the element having that particular reference numeral. A reference numeral in text without a trailing letter, such as "120", refers to any or all of the elements in the drawing having that reference numeral. For example, "120" in the text refers to reference numerals "120A", "120B", and / or "120N" in the figure.
[0017] The various systems shown in FIG. 1 communicate via a network. The network provides a communication channel through which other elements of the networked computing environment can communicate. The network can include any combination of local area and wide area networks using wired or wireless communication systems. In one embodiment, the network uses standard communication technology protocols. For example, the network can include communication links using technologies such as Ethernet, 802.11, WiMAX (worldwide interoperability for microwave access), 3G, 4G, 5G, CDMA (code division multiple access), digital subscriber line (DSL), etc. Examples of network protocols used to communicate via the network include MPLS (multiprotocol label switching), TCP / IP (transmission control protocol / Internet protocol), HTTP (hypertext transport protocol), SMTP (simple mail transfer protocol), and FTP (file transfer protocol). Data exchanged via the network can be represented using any suitable format such as HTML (hypertext markup language) or XML (extensible markup language). In some embodiments, some or all of the communication links of the network may be encrypted using any suitable technology or technologies.
[0018] FIG. 2 shows an example of a traffic flow with set - based load balancing according to one embodiment. A control plane cluster 110 (also referred to as the control plane) receives a request 210 from a service provider 205 that associates a virtual server alias with real servers, e.g., server 1 and server 2. Real servers are also referred to as servers herein. The request 210 may be received from a user interface (UI) or an API. The request 210 is recorded in a database 220 for auditing purposes and added to a queue 225, e.g., RabbitMQ (registered trademark). The request is processed by a data plane cluster 120 and processed by an authorization server 230. The authorization server 230 interacts with a health checker 235 that polls various servers including server 1 and server 2. The authorization server 230 responds (245) based on various factors including the user's location, the health of the target server, and the weight associated with each server. When a virtual server alias is associated with a real server, e.g., server 2, the mapping is stored in one or more DNS caches 260. A client 265 can execute a DNS query and search for the virtual server alias (250). In response to the DNS cache 260 that processes the request, the client 265 receives the server (e.g., server 2) associated with the virtual server alias. According to one embodiment, the weight of a server is configurable and can be assigned by a user. However, in other embodiments, the system can automatically determine the weight based on various factors such as the type of machine on which the server is running, the speed of the network connection to the server, etc.
[0019] FIG. 3A shows an example of a screenshot of a user interface showing a virtual server mapped to a physical server identified using an IP address, according to one embodiment. To generate virtual servers using round-robin load balancing, the system generates physical servers with the same set name. The drawback of this approach is that the system can only use IP addresses within the set and cannot use host names or combine host names and IPs within the set. As shown in FIG. 3A, virtual server 310 is mapped to physical servers 320a, 320b, and 320c of the tree.
[0020] The system also supports set-based load balancing that allows the use of host names within the set assigned to the virtual server. The system generates multiple sets within a location. According to one embodiment, the client makes a DNS query to the DNS cache infrastructure. The system may configure a record with a default time-to-live (TTL) value of T seconds (e.g., 5 seconds). Every T seconds, the cache forwards the request to an authoritative server that has the authority to respond with one of the sets associated with the virtual server. This response is cached for T seconds. During those T seconds, a client querying the cache gets the same set as the response.
[0021] Figure 3B shows an example of a screenshot of a user interface showing virtual servers mapped to physical servers identified using hostnames, according to one embodiment. As shown in Figure 3B, virtual server 310 is mapped to two physical servers 320a, 320b identified using hostnames. However, unlike round-robin DNS load balancing, there are two sets, S1 and S2. Each set contains a single hostname-based physical server. The virtual host is global but is assigned to a region 315 which can be, for example, America, EMEA, Asia, etc. Each set has weights shown as 325A, 325B. Each circle shown in the user interface corresponding to entities such as virtual hosts, physical hosts, regions, etc. indicates the health of the entity. For example, a green circle may indicate good health and a red circle may indicate poor health.
[0022] Figure 3C shows an example of another screenshot of a user interface showing virtual servers mapped to physical servers identified using hostnames, according to one embodiment. Figure 3C shows that the health of physical servers 320a, 320b is relative to each data plane cluster shown in the set of data plane clusters 330. The health of a physical server can be relative to its location within each data plane cluster, such as a building, for example. Thus, the health of the same physical server may be different for two different data plane clusters. Similarly, the health of a physical server may be different for two different locations within the same data plane cluster.
[0023] Figure 4A is a flowchart showing an overall process for processing DNS queries according to one embodiment. The system maintains metadata describing virtual servers (402). The control plane 110 of the system receives requests associated with virtual servers, such as requests for creation, update, and deletion, and updates the metadata describing the virtual servers in the database 220. The metadata of the control plane is propagated to the data plane. The health checker 235 of the data plane 120 of the system monitors and maintains the health of various physical servers of the system (405). The data plane 120 of the system receives and processes DNS queries based on the metadata and the health of the physical servers.
[0024] Details of step 402 are further shown in Figure 4B and are described in connection with Figure 4B. Details of step 405 are shown in Figure 8 and are described in connection with Figure 8. Details of step 408 are shown in Figure 4C and are described in connection with Figure 4C.
[0025] Figure 4B is a flowchart showing a process for maintaining metadata describing virtual servers according to one embodiment. The process shown in Figure 4B provides details of step 402 of Figure 4A.
[0026] The control plane 110 receives requests via a UI or API for updating metadata describing virtual servers, such as requests for creating a new virtual server, requests for modifying an existing virtual server, or requests for deleting an existing virtual server (410). The system updates the data stored in the database 220 based on the request. If the received request is a request for creating a new virtual server, the system adds a new record associated with the new virtual server to the database 220. If the received request is a request for modifying an existing virtual server, the system modifies the record associated with the virtual server in the database 220. If the received request is a request for deleting a virtual server, the system deletes the record associated with the virtual server from the database 220.
[0027] The record describes metadata related to a physical server associated with a virtual server, such as a physical server ID; a record type (whether the record is a physical server is represented using an IP address or a host name); record data (the actual IP address or host name corresponding to the physical server); a set name; a set weight; the region where the physical server exists (global, America, EMEA, Asia); a location, such as the name of the data center where the physical server is maintained, the zone within a cloud platform, the name of the building where the physical server is maintained (enumerated values for identifying the location); a health check type representing the type of network communication or communication protocol used for the health check of the physical server, such as none, icmp (Internet Control Message Protocol supporting ping operations), tcp (transmission control protocol), http (hypertext transfer protocol), https (hypertext transfer protocol secure); health check data representing the value received from the physical server as a response to a health check request; the time when the health was last checked; the current state of the physical server (UP, DOWN, DISABLED), etc.
[0028] The control plane pushes updates to the record into queue 225. Changes to the record are propagated to various data planes (418). As a result, a subset of the data stored in database 220 related to responding to DNS queries is sent to the data plane and stored in the data plane's database 228. The data plane updates the DNS cache based on the database updates to the record (420). The data stored in the DNS cache is used for processing DNS queries.
[0029] When a DNS authority server receives a request to resolve an alias, the DNS authority server provides a response based on several factors, namely, the weight of the real server, the health status of the real server (if health checks are enabled), and the location of the originating client. For DNS usage, traffic can also be directed to real servers (even if they have a down health check or a zero (0) weight). DNS records are retrieved via the cache, and all DNS records have a TTL, for example, 5 seconds. Thus, depending on when the client last searched for the record and which DNS cache hit, since the record is cached, the client may continue to obtain a down (or zero-weighted) record.
[0030] When a user attempts to resolve a virtual server alias, there are several factors that affect which target backend server is provided in the DNS query response. This is due to the DNS cache layer. For example, assume that there is a virtual server with a 2:1 weight for two real servers and that health checks are enabled for those real servers. The health of the servers observed by the system may be different from the actual health. For example, in the presence of network latency, even if a server is healthy, due to the latency in reaching the server, the system may observe the server as unhealthy. According to the system's observation, assuming that the server is healthy regardless of its actual health, the first and second queries to the DNS authority server result in the first real server being returned as a response. Assume that as a result of the third query to the DNS authority server, the second real server is returned as a response. That response is cached by the DNS cache, and a new client attempting to use that alias (a client passing through the same cache server) will obtain that response for the next 5 seconds (since this is the TTL of the Nimbus record).
[0031] Furthermore, there are multiple (e.g., 12) authority servers that can respond to user queries, each of which may be in various iterative stages to balance the weights of the actual servers, and there are hundreds of caching servers, all of which may be in various stages of TTL for their records. As a result, the system may return different results depending on the context.
[0032] The system enables regional routing to map virtual server aliases to actual servers. According to one embodiment, the system uses the location of the user / client device that sent the request to determine which server to provide in response to the DNS query. The system determines the location of the user (or the client device where the request was received) by detecting the location of the DNS cache where the request was forwarded. The system also has a database of all subnets deployed across the enterprise and their locations. When a request enters the system, the system fetches the source IP of the DNS request (usually the DNS cache IP address), calculates the subnet of that IP, and determines the location of the client device based on the location associated with the subnet.
[0033] The system relies on the DNS cache infrastructure to determine the user's location. If the local or regional set of DNS caches is down, the system may misinterpret the user's location and provide a response for a different region than expected.
[0034] FIG. 4C is a flowchart showing a process for processing a DNS query according to an embodiment. The system receives a DNS query and validates the DNS query (430). The DNS query may specify a virtual server and may request a real server to process a request directed to the virtual server. The system detects all real servers for the virtual server (432). If the system determines that no real servers were found (435), the system returns an error indicating that no real servers were found for the specified virtual server (438).
[0035] If the server finds one or more real servers, the system identifies the real servers by performing the following steps. The system excludes real servers known to be unhealthy based on the health checks performed (440). The system also excludes a server if the set to which the server belongs is disabled or has a weight of zero (or if the weight value indicates that the set should not be used) (440). If there are multiple regions associated with the set, the system filters the servers based on the region to identify only the servers that match the region of the client sending the request. If multiple real servers or sets of real servers remain, the system selects a server having a region within the proximity of the client's region (445). According to one embodiment, the system stores a routing plan for the regions. If all servers for a given region are down or unavailable and the system is configured such that the service provider uses regional routing with health checks, the system uses another set of regions to determine which real servers can be provided as a response to the DNS query. Examples of regions include, but are not limited to, the United States, Europe, Asia, etc. If regions A, B, C, D represent various regions, the routing plan for region A may be an ordered list of regions such as region A → region C → region B → region D, and the routing plan for region C may be region C → region D → region A → region B.
[0036] According to one embodiment, the system stores a routing plan for determining an alternative area when there is no availability of servers in an area. The routing plan identifies a sequence of areas in order of priority. The system identifies areas across the sequence of areas specified in the order of the routing plan. For example, assume that the routing plan specifies the sequence of areas Area A → Area B → Area C → Area D. If there are no available servers in Area A, the system searches for a server corresponding to Area B; if there are no available servers in Area B, the system searches for a server corresponding to Area C; if there are no available servers in Area C, the system searches for a server corresponding to Area D. The system identifies a server from the areas identified based on the routing plan. If a server cannot be identified based on the routing plan, the system can return a healthy server.
[0037] If there are multiple servers or sets of servers, the system selects a server or set of servers based on weights (450), similar to the previous servers returned in response to a DNS query received for this particular virtual server. Thus, the system returns the remaining physical servers such that, over a given time interval, the number of times a set of servers (or a server) is returned is proportional to the weight of the set of servers (or the weight of the set associated with the server). For example, if the weight of set of servers S1 is w1 and the weight of set of servers S2 is w2 and w1 is greater than w2, then set S1 (or the servers of set S1) is more likely to be returned in response to a DNS query for the virtual server than set S2. Further, the system ensures that the ratio of the probability p1 of returning set S1 and the probability p2 of returning set S2 matches the ratio of the weights w1 and w2. The system stores the selected server and returns the server in response to the query (455). The system saves information indicating the number of times a server or set of servers has been returned in response to a DNS query identifying a virtual server and can use that historical information to determine its response when the DNS query is received again.
[0038] Figures 5A - 5B show the system architecture of the control plane and data plane and their interactions according to one embodiment. A user can use the system to efficiently (in seconds) generate aliases (virtual servers) of the back - end that are running services (real servers or physical servers), and load - balance traffic to those servers based on regions, operating states, weights, or any combination thereof. There can be multiple data - plane clusters that operate independently of each other. Thus, if one of the data - plane clusters is not operating or is unreachable, another data - plane cluster can be used. As a result, the data plane can be upgraded independently of other data planes. For example, during the upgrade of one data - plane cluster, the remaining data - plane clusters can provide the necessary services. The control - plane cluster supports high - rate requests, such as thousands of requests per minute. Similarly, the system can execute very high - rate, e.g., thousands of health checks per second.
[0039] The control center node 510 executes a control center process 512 that receives requests via a user interface or API. Data describing the requests can be stored in a database 517, such as a document database like MongoDB (registered trademark). When a change request is received via the control plane, database transactions are used to ensure that the changes are executed and sent to the data plane. When a change request is received via the control plane, database transactions are used to ensure that the changes are executed and sent to the data plane. The control plane mode interacts with the data plane via a queue 522 (515). When the system needs to write information based on a change request received via the API, the system writes the change to the data and the audit information regarding the change to a transaction in the database 517. By auditing the changes, experts can review the changes made and the users who executed the changes, thus enhancing the security of changes to the system. Examples of change events to be processed include the generation of a new virtual server alias, the deletion of a virtual server alias, the update of a server weight, and the change of a virtual server alias definition. According to one embodiment, the database 517 supports a change stream and enables the process to subscribe to changes in the change stream. As a result, every time there is a change, the change is pushed down to the queue 522. The changes are consumed by any of the data plane clusters.
[0040] The control plane node executes various tasks such as request logging, authorization checks, and audit backups. The control plane node is stateless and can scale horizontally. The data plane cluster can continue to execute even if the control plane node stops functioning.
[0041] The data plane performs various functions. The data plane nodes provide records to clients via the DNS protocol. The data plane nodes calculate the responses provided to clients based on various factors such as the location of the clients, record configurations including weighting and health checks. The data plane performs health checks on the target backend. The data plane provides a mechanism to synchronize data from the control plane. The data plane provides a mechanism to return changes in the health check status to the control plane.
[0042] Changes received by the data plane are processed by a data synchronization service that propagates the changes to the data store of the data plane or to sites within the data store. Even if the same change is propagated to multiple data planes, two different data planes may respond to the same DNS query in different ways. For example, a first data plane cluster may be able to respond to a DNS query specifying virtual host V1 by returning server S1, while a second data plane cluster may be able to respond to the same DNS query specifying virtual host V1 by returning server S2. There are various reasons for this, for example, different data planes may determine different values regarding the health of the server, and the period during which data is stored in the cache of each data plane may also be different. The data plane nodes determine the appropriate results of DNS queries received from clients, also considering factors such as the location of the clients.
[0043] The system associates each change event with a serial ID to ensure that change events are not lost. The serial ID increases monotonically. The nodes within the data plane cluster include a cache that stores information about the data store of the data plane. The data plane receives and processes client requests. The data plane nodes determine the appropriate real servers for virtual server aliases.
[0044] Figure 6 shows the processing of change events using database transactions according to one embodiment. The database is a document database that stores documents as collections. The received change event 605 is processed using an appropriate service 610 to effect the change. The service 610 generates a database transaction 612. The change event is stored as a document 620 in an appropriate collection of the database. Data regarding the change event is passed to an audit service 615. The audit service generates an audit document 625. The audit document 625 has historical data and cannot be changed. The system uses a database 630 to generate a monotonically increasing timestamp. A serial ID for each transaction is also generated. Some serial IDs may need to be discarded (in the case of rollback). However, the system does not track discarded serial IDs (to avoid situations where the data plane tries to read a serial ID that has not been committed). The system uses a wait timeout that is longer than the automatic rollback timeout from the database. The system determined that a missing serial ID was the result of a rollback.
[0045] The database 630 includes an audit collection. The system stores the audit document (within the same transaction as above) in the audit collection. Any kind of failure within the transaction boundary does not result in a change to the database because the failure is properly rolled back (or via a timeout). There may be multiple threads / processes that read the change stream 635 for high availability. These publish the same message to a queue 645 (e.g., RabbitMQ). The consumer (data plane) manages duplicate messages.
[0046] The provisioning service 640 listens for changes to the audit collection via a database change stream. A single audit message may pass multiple messages to a queue exchange (or vice versa).
[0047] Figure 7 shows the interaction of various components of the data plane and control plane according to one embodiment. Interaction A-1 indicates that when changes pass through the control plane, they are sent to the queue (as an event log). These changes are picked up and processed in order (by serial ID). The loss of a serial ID on one virtual server does not affect other virtual servers. Skipping of serial IDs is supported. Interaction A-2 indicates that when an out-of-order update is received (or when the queue mechanism breaks), the API is used as a fallback mechanism and the serial ID is fetched via HTTP. Interaction A-2 indicates that the data is stored in the database as a fully flattened record. Since consistency is prioritized in this write, a slow write speed is not a problem (for example, 1000 writes per second).
[0048] The data synchronization service supports both incremental (event log) updates and full data synchronization. Full data synchronization functions as a backup mechanism to ensure that the data on both sides is correct. Full data synchronization is executed periodically. This service also synchronizes the subnet information required for regional routing. Interaction B-1 represents that the client resolves DNS queries via the caching infrastructure. Interaction B-2 represents that the query is transferred to the name server. This component is responsible for DNS protocol conversion. When decoded, the request is sent to the DNS response service. Interaction B-3 represents that the DNS response service receives the decoded structured request. The DNS response service requests the data attached to the requested record from the database (possibly via the cache). The DNS response service executes the necessary calculation logic (health, region, weight, etc.), returns the final structured response to the name server, and converts the DNS protocol into the response. Interaction B-4 represents that the data is retrieved by indexing the data store. The data store needs to operate properly in a large-scale environment and its reads need to be optimized. The data model needs to be lightweight and simple for fast data acquisition. Interaction B-5 represents that all queries received by the name server are logged. Ideally, the log aggregation of specific metrics should be performed for long-term storage. The validity period of the complete request-response may be short.
[0049] Data store D-1 is the central component of the system. Since the workload has a large number of read profiles, the reads of the data store are fast. The data is mastered and managed via the control plane and thus not saved or backed up. The data store contains the subnet data and metadata of the site.
[0050] Log Store D-2 is used to store a large amount of data. Log Store D-2 is optimized for writing and has a policy to automatically delete data according to the amount of DNS and health check logs. The log store is shared across all data planes.
[0051] Site Health Monitoring Service D-3 is responsible for verifying that all components of the site are operating optimally. If a component is not operating optimally, Site Health Monitoring Service D-3 issues an alert, shuts down the site, and is responsible for protecting the entire data plane from a site operating maliciously.
[0052] Before shutting down the site when a component fails, Site Health Monitoring Service D-3 reports to the control plane and also makes adjustments (via the control plane) to ensure that at least one site is operational. Site Health Monitoring Service D-3 also functions as a web server for site management functions.
[0053] The Health Check Scheduler Service attempts to detect all target endpoints that need to be checked periodically (every second) and schedule the work (C-1). The scheduler is multi-instance and distributed for reliability.
[0054] A worker is a simple execution engine for directly performing checks on the target backend. Workers (all workers within a site) process a minimum of 15,000 checks per second. When the work is completed, the results are reported to the result processor (C-3).
[0055] The processor also stores the final state and state change information for records for simple lookups (C-4). The change trigger is used to push health audit events to the state synchronization service (C-5). The message is transformed and returned to the control plane via the state change queue (C-6). The system supports full (and current) state synchronization of all records. All operations are reported to the log store for querying and analysis (C-7).
[0056] Figure 8 shows the overall flow of a server health check system according to one embodiment. The health check system consists of two parts: (1) a master scheduler process (MSP) responsible for managing and scheduling health checks and managing the results, and (2) a worker manager responsible for starting and managing health check workers that perform the tasks of health checking the servers. The tasks can also be referred to as jobs herein.
[0057] This system is highly scalable and can perform a very large number of health check operations. For example, using state-of-the-art processing hardware, the system was able to process approximately 30,000 checks per second across the cluster with half the CPU core capacity. The master scheduler process consists of the following four components: (1) a service engine 802 responsible for managing inter-process communication, resilience, and failover; (2) a multiplexer 810 (or work multiplexer) responsible for distributing jobs to workers and managing responses from workers; (3) a writer 830 responsible for propagating responses to the database, name server, and control plane; and (4) a scheduler 805 responsible for scheduling jobs. Each component can be implemented as a process, such as a daemon process. Each component is configurable and reads a set of configuration parameters to start processing.
[0058] The service engine 802 can check the heartbeats of all major components such as the multiplexer 810, the writer 830, and the scheduler 805, and switch their on / off based on the system's soundness. For example, if the service engine 802 determines that a component, such as the scheduler 805, is down, the service engine will attempt to restart the scheduler 805. If the service engine 802 determines that a component such as the scheduler 805 cannot be restarted, since the data plane cluster cannot monitor the soundness of the server and cannot process DNS queries accurately, the service engine 802 can take the entire data plane cluster offline. As a result, DNS queries are redirected to another data plane cluster. However, if the system performs a health check on the site and determines that this data plane cluster is the only functioning data plane cluster, the system will keep the data plane cluster running instead of shutting down the entire system so that the system can process DNS queries even if the results are inaccurate.
[0059] Figure 9 shows a flowchart depicting a scheduler process according to an embodiment. The scheduler 805 identifies various servers (910) to check their soundness. The scheduler 805 opens a connection to a queue and a connection to a database to monitor ongoing changes, and reads all records from the database into memory. The scheduler 805 reads all initial health check status data from the database into memory and also reads any DNS namespace information. The scheduler 805 opens a connection to the multiplexer 810, for example, a TCP socket connection. The records loaded by the scheduler 805 store information describing various servers including their mappings to virtual servers, and the current known health information of the servers indicating whether the servers are running or stopped.
[0060] The scheduler 805 repeatedly performs the following steps. The scheduler 805 reads (912) the queue of ongoing changes to the records and verifies that the records in the database match the in-memory copies of the records. The scheduler 805 can filter the records, for example, by checking the metadata of the records to determine whether the health check is disabled for any record.
[0061] The scheduler 805 generates a plurality of server health check tasks. Each server health check task is configured to determine a measurement of the server health for a location, for example, a data plane cluster or a building within the data plane cluster. Each measurement of the server health is associated with a communication protocol for reaching the server from a computing device within a particular location. Examples of communication protocols that may be associated with the worker include tcp (Transmission Control Protocol), http (Hypertext Transfer Protocol), or https (Hypertext Transfer Protocol Secure), or icmp (Internet Control Message Protocol). The worker accesses the server using the corresponding communication protocol. If the worker can access the server within a threshold time, the worker determines that the health of the server is operational; otherwise, the worker determines that the health of the server is down.
[0062] The scheduler divides the plurality of server health check tasks into a sequence of task buckets (915). Thus, the tasks are divided into a plurality of task buckets, and each bucket is assigned a sequence number indicating the order in which the task buckets are to be processed within a time interval.
[0063] The system monitors the health of multiple servers by periodically repeating a process where, for example, all buckets are processed at a time interval (e.g., 10 seconds), and then the processing is repeated at the next time interval. The system processes multiple server health check tasks within the time interval. The scheduler divides that time interval into multiple sub-time intervals, with one sub-time interval corresponding to each bucket. For example, if there are 10 buckets and the time interval is 10 seconds, each second of the 10 seconds is assigned to a bucket. The multiple buckets of the task are processed in the order of the sequence of the buckets of the task, and each bucket of the task is processed during the sub-time interval assigned to the bucket.
[0064] Accordingly, the scheduler repeats steps 918, 920, 923, 925 for each sub-time interval. The system selects the bucket corresponding to the sub-time interval based on the sequence number of the bucket (918). If the scheduler determines that there are unfinished tasks in the previous sub-time interval, the scheduler adds the unfinished tasks to the current bucket so that the unfinished tasks are carried over (920). The scheduler sends the server health check task of the bucket of the task to the worker process (923). Each worker process determines the measured value of the server health of the server by communicating with the server using the communication protocol of the worker process. The scheduler determines statistical information such as the number of carried-over unfinished tasks and the number of functioning workers, and sends the statistical information for display via the user interface of the control plane (925). The scheduler also sends a heartbeat signal indicating the health of the scheduler itself to the service engine.
[0065] Figure 10 shows a flowchart illustrating a multiplexer process according to an embodiment. The multiplexer process receives jobs from a scheduler and ensures that the jobs are processed by worker processes. The multiplexer process opens a connection (e.g., a TCP socket) to the scheduler (1010). The multiplexer process opens a connection (e.g., a TCP socket) to the worker processes (1012). The multiplexer process periodically repeats steps 1015, 1018, 1020, 1023, 1025. The multiplexer process receives a set of jobs, e.g., a job bucket, from the scheduler (1015). The multiplexer process sends the jobs to the worker processes (1018). The multiplexer process receives the results of the job executions indicating the health of the server from the worker processes (1020). The multiplexer process also receives a heartbeat signal from the worker processes indicating the health of the worker processes. The multiplexer process writes the job results to a writer process (1023), and the writer process stores the information in a database and propagates the information to various data planes via a queue. The multiplexer process adjusts the worker process to be assigned to the next job based on the received status (health or heartbeat signal) of the worker processes (1025).
[0066] The writer process opens connections to the queue and the database. The writer process also opens a connection to the multiplexer to listen for information provided by the multiplexer process. The writer process reads the results of the server health check from the multiplexer and determines whether the health state of any server has changed, i.e., whether it has changed from running to stopped or from stopped to running. The writer process writes data describing the state change information of the server whose state has changed to the database and the queue. The writer process also provides statistical information, such as the number of servers whose state has changed, to the control plane. The writer process also sends a heartbeat describing the health of the writer process to the service engine.
[0067] The worker manager process starts the worker processes. Each worker process opens a connection to the multiplexer process, listens for the multiplexer process for tasks to be executed by the worker, and provides the result of the server health check to the multiplexer. The worker process receives information identifying the server for performing the health check and attempts to connect to the server using the communication protocol associated with the worker to perform the health check. The worker process can receive an indication that the communication was successful or an indication that the communication failed. The indication of failure can be a timeout or an error message received as a result of the communication protocol. The worker process writes the result to the multiplexer process. The worker process also sends a heartbeat signal indicating the health of the worker process itself to the multiplexer process.
[0068] The result of the health check is provided to the writer process by the multiplexer process, and the writer process provides the result to the database and the queue. The health check information is propagated from the database and the queue to each data plane cluster. The data plane cluster stores the health check information in the DNS cache and uses the information to respond to DNS queries to provide the latest results.
[0069] Screenshots of the user interface displayed via the control plane are shown in FIGS. 11A-11G described below.
[0070] FIG. 11A shows an aspect of the user interface of the management health check view according to one embodiment. The user interface enables a user, for example, a system administrator, to request a health check. The user interface 1105 shows various types of health checks. The user interface 1108 shows statistical information describing the health check information.
[0071] Figure 11B shows an aspect of the home page of the user interface according to an embodiment. The user interface enables the user to perform various types of operations including generating a namespace (1112), generating a virtual server (1114), and updating the weights of a set of physical servers (1118).
[0072] Figure 11C shows an aspect of the user interface of the namespace view according to an embodiment. The user interface shows a list of all namespaces (1132). The user can filter namespaces using search criteria (1134).
[0073] Figure 11D shows an aspect of the user interface of the virtual server view according to an embodiment. The user interface shows a list of all virtual servers (1142). The user can filter virtual servers using search criteria (1144).
[0074] Figure 11E shows an aspect of the user interface of the virtual server audit view according to an embodiment. The audit view shows various actions (1152) and the users (1154) who performed the actions.
[0075] Figure 11F shows an aspect of the user interface of the virtual server details view according to an embodiment. The user interface shows the mapping of the virtual server to different physical servers (1162) and the weights of the physical servers from different locations (1164) corresponding to different data plane clusters.
[0076] Figure 11G shows an aspect of the user interface of the health check status view of a virtual server according to an embodiment. The user interface shows various types of metadata including various physical servers and their health (1172), as well as the timestamps at which the health was checked, the corresponding virtual servers, etc.
[0077] Computing System Architecture FIG. 12 is a block diagram of an exemplary computer 1200 suitable for use as a server or client device. The exemplary computer 1200 includes at least one processor 1202 coupled to a chipset 1204. The chipset 1204 includes a memory controller hub 1220 and an input / output (I / O) controller hub 1222. A memory 1206 and a graphics adapter 1212 are coupled to the memory controller hub 1220, and a display 1218 is coupled to the graphics adapter 1212. A storage device 1208, a keyboard 1210, a pointing device 1214, and a network adapter 1216 are coupled to the I / O controller hub 1222. Other embodiments of the computer 1200 have different architectures.
[0078] In the embodiment shown in FIG. 12, the storage device 1208 is a non-transitory computer-readable storage medium such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 1206 holds instructions and data used by the processor 1202. The pointing device 1214 is a mouse, trackball, touch screen, or other type of pointing device and is used in combination with the keyboard 1210 (which may be an on-screen keyboard) to input data into the computer system 1200. The graphics adapter 1212 displays images and other information on the display 1218. The network adapter 1216 couples the computer system 1200 to one or more computer networks such as network 170.
[0079] The type of computer used by the entities of FIGS. 1 and 2 may vary depending on the embodiment and the processing capabilities required by the entity. For example, a computer may lack some of the components described above such as the keyboard 1210, the graphics adapter 1212, and the display 1218.
[0080] Additional Considerations Some parts of the above description explain the embodiments from the perspective of algorithm processes or operations. These descriptions and expressions of algorithms are commonly used by those skilled in the computing art to effectively convey the content of their work to other skilled persons. While these operations are described functionally, computationally, or logically, they are understood to be executed by a computer program including instructions for execution by a processor or equivalent electrical circuit, microcode, etc. Further, without loss of generality, sometimes it may be convenient to refer to the arrangement of these functional operations as modules.
[0081] As used herein, a reference to "an embodiment" or "embodiments" means that the particular element, feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" in various places in this specification are not necessarily all referring to the same embodiment. Similarly, the use of "a" or "an" before an element or component is done merely for convenience. This description should be understood to mean that one or more of the elements or components are present unless the contrary meaning is apparent from the context.
[0082] When a value is described as "about" or "substantially" (or their derivatives), such a value should be interpreted with an accuracy of + / - 10% unless another meaning is apparent from the context. For example, "about 10" should be understood to mean the range of 9 to 11.
[0083] As used herein, the terms "comprise," "comprising," "include," "including," "have," "having" or any other variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements, but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to an inclusive "or" and not to an exclusive "or". For example, the condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).
[0084] Upon reading this disclosure, those skilled in the art will appreciate additional alternative structural and functional designs. Accordingly, while specific embodiments and applications have been illustrated and described, it should be understood that the described subject matter is not limited to the exact structures and components disclosed. The scope of protection should be limited only by the following claims.
Claims
1. A method for global server load balancing based on the Domain Name System (DNS), comprising: Receiving a request associated with a virtual server identified by a URL (uniform resource locator), wherein the request to the virtual server is processed by one or more of a plurality of servers; Updating information stored in a database based on the request, wherein the database stores records mapping virtual servers to servers; Propagating the updated information to a plurality of data plane clusters, each data plane cluster comprising a database storing metadata describing the plurality of servers; For each of the plurality of data plane clusters, periodically updating one or more measurements of server health for each of the plurality of servers, wherein the measurements of server health are relative to a location within the data plane cluster and are associated with a communication protocol for reaching the server from a computing device at the location; Receiving a DNS query from a client device requesting a server to process a request directed to the URL of a particular virtual server; Identifying one or more candidate servers for processing a request directed to the URL of the particular virtual server based on information stored in the DNS cache; Selecting a candidate server from the one or more candidate servers based on the measurements of server health of the servers and factors including the location associated with the client device; Sending a response to the DNS query to the client device, the response identifying the candidate server for processing a request directed to the URL of the virtual server; A method comprising the above steps.
2. The method of claim 1, wherein the request creates a new virtual server, associates the new virtual server with a set of one or more servers, and each of the set of servers has a weight indicating the likelihood that the server of the set is assigned to the virtual server.
3. The method of claim 2, wherein the factors include the weights of the sets of servers to which a particular server belongs.
4. The method of claim 1, wherein the factor includes a server previously returned in response to a DNS query requested from the server for the virtual server.
5. The method of claim 1, wherein the communication protocol associated with the measurement value of the server health of the server is one of tcp (transmission control protocol), http (hypertext transfer protocol), https (hypertext transfer protocol secure), and icmp (internet control message protocol).
6. The method of claim 1, wherein identifying one or more servers for the virtual server includes selecting a server from a region, the region being determined based on a routing plan associated with the virtual server, the routing plan including a series of regions.
7. The method of claim 1, wherein the position of the client device is determined based on the DNS cache that received the request from the client device.
8. A method for domain name system (DNS)-based global server load balancing, comprising: identifying a plurality of servers, each server being associated with a virtual server; generating a plurality of server health check tasks, each server health check task determining a measurement value of the server health of the server with respect to a position, and each measurement value of the server health being associated with a communication protocol for reaching the server from a computing device at that position; dividing the plurality of server health check tasks into a sequence of buckets of tasks; monitoring the health of the plurality of servers by periodically repeating processing the plurality of server health check tasks within a time interval, the time interval including a plurality of sub-time intervals, the plurality of buckets of tasks being processed in the order of the sequence of the plurality of buckets of tasks, each bucket of tasks being processed during a sub-time interval assigned to the bucket, and the processing including: To send each server health check task to a worker process associated with a communication protocol, wherein the worker process determines a measured value of the server health of the server by communicating with the server using the communication protocol of the worker process, Receiving the result of the server health check from the worker process, Propagating the result of the server health check to a plurality of data plane clusters, Including, Processing a DNS query that requests a server to process a request directed to a virtual server, wherein the processing includes selecting a server based on the result of the server health check, A method comprising.
9. Processing the bucket of tasks comprises Determining that one or more tasks in the bucket of tasks were incomplete at the end of the sub-time interval corresponding to the bucket of tasks, Adding one or more tasks of the bucket to the next bucket in the sequence of buckets of tasks, The method of claim 8, including.
10. Receiving a request associated with a virtual server identified by a URL (uniform resource locator), wherein the request to the virtual server is processed by one or more of a plurality of servers, Updating information stored in a database based on the request, wherein the database includes a record that maps a virtual server to a server, The method of claim 8, further comprising.
11. Processing a DNS query received from a client device comprises Identifying one or more candidate servers for the virtual server based on information stored in a DNS cache, Selecting a particular server from the one or more candidate servers based on factors including the measured value of the health of the server and the location associated with the client device, The method of claim 8, including.
12. The method of claim 8, wherein the communication protocol associated with the measured value of the server health is one of TCP (Transmission Control Protocol), HTTP (Hypertext Transfer Protocol), HTTPS (Hypertext Transfer Protocol Secure), and ICMP (Internet Control Message Protocol).
13. Propagating the result of the server health check to a plurality of data plane clusters is writing the result of the server health check to a queue, wherein each of the plurality of data plane clusters listens for changes written to the queue and updates data stored in the DNS cache, the method of claim 8, comprising.
14. The method of claim 8, wherein the result of the server health indicates whether the server is operating or stopped.
15. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more computer processors, cause one or more computers to perform the method of any one of claims 1 to 14.
16. A processor, and a non-transitory computer-readable storage medium storing instructions that, when executed by one or more computer processors, cause one or more computers to perform the method of any one of claims 1 to 14, a system comprising.