Resilient computer systems and methods
Patent Information
- Application Number
- PCT/GB2025/050613
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-18
- Filing Date
- 2025-03-21
- Publication Date
- 2026-08-27
Smart Images

Figure GB2025050613_27082026_PF_FP_ABST
Abstract
Description
[0001] RESILIENT COMPUTER SYSTEMS AND METHODS
[0002] BACKGROUND OF THE INVENTION
[0003] 1. Field of the Invention
[0004] The field of the invention relates to computer systems including computer data centres in communication with each other, in which the computer data centres are configured to route requests between the computer data centres, and / or to route requests within the same computer data centre, and to related computer-implemented methods.
[0005] 2. Technical Background
[0006] Computer systems including computer data centres in communication with each other exist, but the maintenance and operation of such computer systems may be defined by a network service provider, with limited scope for a user of the network services to control the operation of the network services, such as to be able to control the criteria under which one computer data centre should be used instead of another computer data centre, for providing a particular service on the computer network. This can impact the ability to control the resilience of a particular service that is provided on the computer network.
[0007] 3. Discussion of Related Art
[0008] The publication “Resilience on AWS Using the shared responsibility model,” Amazon Web Services, Inc. (2023) discloses AWS Regions and Availability Zones (AZs), in which AWS regions are physical locations around the world in which data centers are clustered, in which each AWS Region has multiple AZs, in which a region is a physical location in the world, in which each AZ includes one or more discrete data centers, in which data centers, each with redundant power, networking, and connectivity, are housed in separate facilities.SUMMARY OF THE INVENTION
[0009] According to a first aspect of the invention, there is provided a computer system including a first data centre and a second data centre, the first data centre configured to communicate with the second data centre via a network, wherein the first data centre and the second data centre are separated by between 2 km and 500 km, the first data centre including a first service unit configured to provide a first service, and a plurality of second service units each configured to provide a second service, the first service unit including a first router proxy, and the first service unit including software executable to receive first requests and to send second requests relating to the first requests to the first router proxy;
[0010] the plurality of second service units each including a respective second router proxy and respective software executable to receive at least some of the second requests via the respective second router proxy, to process the at least some of the second requests to provide respective responses, and to respond to the at least some of the second requests using the respective responses via the respective second router proxy; the first data centre including a first default gateway ingress service, and a first default gateway egress service configured to communicate with a first default gateway ingress service of the second data centre via the network;
[0011] the second data centre including the first default gateway ingress service of the second data centre, a first default gateway egress service, and a plurality of third service units each configured to provide the second service, the plurality of third service units each including a respective third router proxy and respective software executable to receive at least some of the second requests via the respective third router proxy, to process the at least some of the second requests to provide respective responses, and to respond to the at least some of the second requests using the respective responses via the respective third router proxy;
[0012] wherein the first router proxy is configured to route a first set of at least some of the second requests to the first default gateway ingress service of the first data centre, and the first default gateway ingress service of the first data centre is configured to send the first set of at least some of the second requests to the first default gateway ingress service of the first data centre to a second service unit of the plurality of second service units, and to receive the respective responses from the second service unit, andto send the respective responses to the first service unit of the first data centre; wherein the first router proxy is configured to route a second set of at least some of the second requests to the first default gateway egress service of the first data centre, and the first default gateway egress service of the first data centre is configured to send the second set of at least some of the second requests to the first default gateway ingress service of the second data centre, wherein the first default gateway ingress service of the second data centre is configured to send the second set of at least some of the second requests to a service unit of the plurality of third service units of the second data centre, wherein the service unit of the plurality of third service units of the second data centre is configured to respond to the second set of at least some of the second requests using the respective responses via the respective third router proxy, wherein the respective third router proxy is configured to route the respective responses to the first default gateway egress service of the second data centre, and the first default gateway egress service of the second data centre is configured to send the respective responses to the first default gateway ingress service of the first data centre, wherein the first default gateway ingress service of the first data centre is configured to send the respective responses to the first service unit of the first data centre. Optionally, the first service unit of the first data centre may store the responses it receives.
[0013] An advantage is that the computer system is resilient to faults, because if one or more, or all, second service units of the first data centre is determined to have become faulty, it is possible in response to route second requests to the third service units of the second data centre, which means that a computer terminal in connection with the system making first requests to the first service of the first data centre will experience no, or negligible, reduction in performance of the first service by the first data centre.
[0014] The computer system may be one wherein the first router proxy is configured to route the second requests, the first default gateway ingress service of the first data centre is configured to route the first set of at least some of the second requests routed via the first router proxy, and the first default gateway ingress service of the second data centre is configured to route the second set of at least some of the second requests routed via the first router proxy.The computer system may be one wherein the first service unit is configured to use the software executable to receive the first requests to use the responses received to the second requests to respond to the first requests.
[0015] The computer system may be one wherein the first service unit includes a structured database, wherein the first service unit software is executable to receive first requests which relate to contents of the structured database, and to send the second requests relating to the first requests which relate to contents of the structured database to the first router proxy, wherein the second requests include data extracted from the structured database of the first service unit.
[0016] The computer system may be one wherein the second data centre further includes a first service unit configured to provide the first service, the first service unit of the second data centre including a fourth router proxy and software executable to receive first requests and to send second requests relating to the first requests to the fourth router proxy;
[0017] the second data centre including the first default gateway egress service configured to communicate with the first default gateway ingress service of the first data centre via the network;
[0018] wherein the fourth router proxy is configured to route a third set of at least some of the second requests to the first default gateway ingress service of the second data centre, and the first default gateway ingress service of the second data centre is configured to send the third set of at least some of the second requests to the first default gateway ingress service of the second data centre to a third service unit of the plurality of third service units, and to receive the respective responses from the third service unit, and to send the respective responses to the first service unit of the second data centre;
[0019] wherein the fourth router proxy is configured to route a fourth set of at least some of the second requests to the first default gateway egress service of the second data centre, and the first default gateway egress service of the second data centre is configured to send the fourth set of at least some of the second requests to the first default gateway ingress service of the first data centre, wherein the first defaultgateway ingress service of the first data centre is configured to send the fourth set of at least some of the second requests to a service unit of the plurality of second service units of the first data centre,
[0020] wherein the service unit of the plurality of second service units of the first data centre is configured to respond to the fourth set of at least some of the second requests using the respective responses via the respective second router proxy, wherein the respective second router proxy is configured to route the respective responses to the first default gateway egress service of the first data centre, and the first default gateway egress service of the first data centre is configured to send the respective responses to the first default gateway ingress service of the second data centre, wherein the first default gateway ingress service of the second data centre is configured to send the respective responses to the first service unit of the second data centre. Optionally, the first service unit of the second data centre may store the responses it receives.
[0021] An advantage is that the computer system is more resilient to faults, because if one or more, or all, third service units of the second data centre is determined to have become faulty, it is possible in response to route second requests to the second service units of the first data centre, which means that a computer terminal in connection with the system making first requests to the first service of the second data centre will experience no, or negligible, reduction in performance of the first service by the second data centre.
[0022] The computer system may be one wherein the fourth router proxy is configured to route the second requests, the first default gateway ingress service of the second data centre is configured to route the third set of at least some of the second requests routed via the fourth router proxy, and the first default gateway ingress service of the first data centre is configured to route the fourth set of at least some of the second requests routed via the fourth router proxy.
[0023] The computer system may be one wherein the first service unit of the second data centre is configured to use the software executable to receive the first requests to use the responses received to the second requests to respond to the first requests.The computer system may be one wherein the system does not include a network load balancer, between the first data centre and the second data centre, configured to load balance the second requests. An advantage is that the computer system is more resilient to faults, because a network load balancer is a possible point of failure in a networked computer system. An advantage is reduced latency, because a network load balancer typically introduces latency. An advantage is reduced cost, because a network load balancer consumes computing resources.
[0024] The computer system may be one wherein the first data centre is located in a first zone, and wherein the second data centre is located in a second zone, and wherein the first zone and the second zone are separated by between 2 km and 500 km.
[0025] The computer system may be one including a third data centre, the third data centre configured to communicate with the first data centre and with the second data centre via the network, wherein the third data centre and the second data centre are separated by between 2 km and 500 km, wherein the first data centre and the third data centre are separated by between 2 km and 500 km, the third data centre including a first service unit configured to provide a first service, and a plurality of fourth service units each configured to provide the second service, the first service unit including a fifth router proxy, and the first service unit including software executable to receive first requests and to send second requests relating to the first requests to the fifth router proxy;
[0026] the plurality of fourth service units each including a respective sixth router proxy and respective software executable to receive at least some of the second requests via the respective sixth router proxy, to process the at least some of the second requests to provide respective responses, and to respond to the at least some of the second requests using the respective responses via the respective sixth router proxy;
[0027] the third data centre including a first default gateway ingress service, and a first default gateway egress service configured to communicate with a first default gateway ingress service of the first data centre via the network; the third data centre including a second default gateway egress service configured to communicate with a first default gateway ingress service of the second data centre via the network;the first data centre including a second default gateway egress service, and the second data centre including a second default gateway egress service;
[0028] wherein the fifth router proxy is configured to route a first set of at least some of the second requests to the first default gateway ingress service of the third data centre, and the first default gateway ingress service of the third data centre is configured to send the first set of at least some of the second requests to the first default gateway ingress service of the third data centre to a fourth service unit of the plurality of fourth service units, and to receive the respective responses from the fourth service unit, and to send the respective responses to the first service unit of the third data centre; wherein the fifth router proxy is configured to route a second set of at least some of the second requests to the first default gateway egress service of the third data centre, and the first default gateway egress service of the third data centre is configured to send the second set of at least some of the second requests to the first default gateway ingress service of the first data centre, wherein the first default gateway ingress service of the first data centre is configured to send the second set of at least some of the second requests to a service unit of the plurality of second service units of the first data centre, wherein the service unit of the plurality of second service units of the first data centre is configured to respond to the second set of at least some of the second requests using the respective responses via the respective second router proxy, wherein the respective second router proxy is configured to route the respective responses to the second default gateway egress service of the first data centre, and the second default gateway egress service of the first data centre is configured to send the respective responses to the first default gateway ingress service of the third data centre, wherein the first default gateway ingress service of the third data centre is configured to send the respective responses to the first service unit of the third data centre;
[0029] wherein the fifth router proxy is configured to route a third set of at least some of the second requests to the second default gateway egress service of the third data centre, and the second default gateway egress service of the third data centre is configured to send the third set of at least some of the second requests to the first default gateway ingress service of the second data centre, wherein the first default gateway ingress service of the second data centre is configured to send the third set of at least some of the second requests to a service unit of the plurality of third service units of thesecond data centre, wherein the service unit of the plurality of third service units of the second data centre is configured to respond to the third set of at least some of the second requests using the respective responses via the respective third router proxy, wherein the respective third router proxy is configured to route the respective responses to the second default gateway egress service of the second data centre, and the second default gateway egress service of the second data centre is configured to send the respective responses to the first default gateway ingress service of the third data centre, wherein the first default gateway ingress service of the third data centre is configured to send the respective responses to the first service unit of the third data centre. Optionally, the first service unit of the third data centre may store the responses it receives.
[0030] An advantage is that the computer system is more resilient to faults, because if one or more, or all, fourth service units of the third data centre is determined to have become faulty, it is possible in response to route second requests to the second service units of the first data centre, or to the third service units of the second data centre, which means that a computer terminal in connection with the system making first requests to the first service of the third data centre will experience no, or negligible, reduction in performance of the first service by the third data centre.
[0031] The computer system may be one wherein the system does not include a network load balancer, between the first data centre and the second data centre and the third data centre, configured to load balance the second requests. An advantage is that the computer system is more resilient to faults, because a network load balancer is a possible point of failure in a networked computer system. An advantage is reduced latency, because a network load balancer typically introduces latency. An advantage is reduced cost, because a network load balancer consumes computing resources.
[0032] The computer system may be one wherein spacing the data centres apart reduces the risk of the data centres all going down if a disaster (e.g. explosion, fire, power outage) occurs locally in a geographic region. An advantage is that the computer system is more resilient to faults.The computer system may be one wherein the system includes a control plane and a data plane, wherein the data plane includes all the proxies deployed as sidecars or as gateways, or as other configurable proxies, and wherein the control plane manages and configures the proxies to route traffic. An advantage is that the computer system is more resilient to faults.
[0033] The computer system may be one wherein the data plane is monitored by measuring at least traffic, errors, latency and saturation.
[0034] The computer system may be one wherein the control plane includes four components: Componentl which provides service discovery for the gateway sidecars; component2 which is the router’s configuration validation, ingestion, processing and distribution component; components which enables strong service-to-service and enduser authentication with built-in identity and credential management; component4 which enforces access control and usage policies across the service mesh, and collects telemetry data from the gateway proxy and other services. An advantage is that the computer system is more resilient to faults.
[0035] The computer system may be one wherein each data centre includes at least one server cluster, wherein each server cluster runs its own control plane, and / or wherein the router control plane is deployed independently into each server cluster. An advantage is that the computer system is more resilient to faults.
[0036] The computer system may be one wherein service units each run a router-compatible sidecar proxy.
[0037] The computer system may be one wherein sidecar proxy injection into a new service unit happens automatically every time a new service unit is created.
[0038] The computer system may be one wherein traffic is directed from application services to and from the sidecars.
[0039] The computer system may be one wherein when router software is updated, sidecarsin existing service units remain the same until those service units are deleted.
[0040] The computer system may be one wherein when those service units are deleted, replacement service units are created.
[0041] The computer system may be one wherein http 503 error response codes lead to one or more retries by a sidecar proxy.
[0042] The computer system may be one wherein when a service unit or any intermediate proxies returns a http 503 error, the client gateway sidecar proxy automatically retries for another service unit.
[0043] The computer system may be one wherein a fully jittered exponential back-off algorithm is used for retries.
[0044] The computer system may be one in which router circuit breakers are configured to automatically failover traffic to the closest working server cluster or service unit. An advantage is that the computer system is more resilient to faults.
[0045] The computer system may be one in which there are configured two sets of circuit breaking rules:
[0046] one rule for communication within a server cluster, which applies to requests coming from the ingress gateway going to the destination service units;
[0047] one rule for communication across server clusters.
[0048] The computer system may be one in which the system includes a mesh which provides service uni t-to- service unit communication.
[0049] The computer system may be one in which all service uni t-to- service unit communication within the mesh is end-to-end encrypted, and the encryption is handled by the router proxy e.g. router sidecar proxy.The computer system may be one in which each service includes its own security certificate, which is included into the router proxy e.g. router sidecar proxy.
[0050] The computer system may be one in which certificates are automatically rotated and dynamically reloaded by the router proxy e.g. router sidecar proxy, without any disruption to the service unit.
[0051] The computer system may be one in which a service unit (e.g. its service owner) must declare the list of services that the router proxy, e.g. router sidecar proxy, can see or call.
[0052] The computer system may be one in which these services are the only services within the mesh.
[0053] The computer system may be one wherein all sidecar proxies in the mesh are programmed with the necessary configuration required to reach every workload instance in the mesh.
[0054] The computer system may be one wherein the amount of traffic each router is handling is monitored, to ensure the volume of traffic handled by different routers does not differ by more than a threshold e.g. to ensure the volume of traffic handled by different routers does not differ by more than a factor of ten, or e.g. to ensure the volume of traffic handled by different routers does not differ by more than a factor of three. An advantage is more balanced use of the computer system.
[0055] The computer system may be one in which in response to detecting that the volume of traffic handled by different routers does differ by more than a threshold, the router proxies are reconfigured to ensure that the volume of traffic handled by different reconfigured router proxies does not differ by more than a threshold, e.g. does not differ by more than a factor of ten, or e.g does not differ by more than a factor of three. An advantage is more balanced use of the computer system.The computer system may be one wherein a default gateway in the first or second data centre connects to between one hundred and ten thousand service units in the same data centre.
[0056] The computer system may be one wherein a default gateway in the first or second data centre connects to between one thousand and five thousand service units in the same data centre.
[0057] The computer system may be one wherein services are not deployed to a single server cluster.
[0058] The computer system may be one wherein each server cluster has its own service mesh control plane.
[0059] The computer system may be one wherein each service unit has its own sidecar which all inbound and outbound traffic goes through.
[0060] The computer system may be one wherein for traffic within each subnet, it is preferred to use the local subnet; the next preference is other subnets of the data centre of the local subnet, and the next preference is subnets of other data centres. An advantage is that the computer system is more resilient to faults
[0061] The computer system may be one wherein server clusters’ traffic is drained and shifted to other working server clusters in other Data Centres.
[0062] The computer system may be one wherein only server clusters within a single Data Centre are drained at one time. An advantage is that the computer system is more resilient to faults.
[0063] The computer system may be one wherein at least three server clusters are present in each data centre, e.g. to ensure at least one server cluster is working. An advantage is that the computer system is more resilient to faults.The computer system may be one wherein each route to a service goes via a default gateway; each default gateway is configured to route to any service units providing a particular service, whereas the mesh is configured to route to, or to receive traffic from, any of the default gateways.
[0064] The computer system may be one wherein the default gateways are configured in the routers as a Service Entry in every server cluster, not just in the server clusters the service is deployed in.
[0065] The computer system may be one wherein when updating routers, routers are updated in at most one Data Centre at a time. An advantage is that the computer system is more resilient to faults.
[0066] The computer system may be one wherein updates to routers are done after draining the server clusters in the Data Centre in which the routers are being updated. An advantage is that the computer system is more resilient to faults.
[0067] The computer system may be one wherein every server cluster provides an isolated failure zone. An advantage is that the computer system is more resilient to faults.
[0068] The computer system may be one wherein there is no connectivity between different Data Centres in different geographic regions of the world. An advantage is that the computer system is more resilient to faults.
[0069] The computer system may be one wherein the path for traffic is specified with routing rules, and destination rules are used to configure a set of policies that gateway proxies apply to a request at a specific destination.
[0070] The computer system may be one wherein a Router Management Service in a local server cluster uses a Discovery tool to discover endpoints within other server clusters, to do the following:
[0071] (i) Read a secret which contains credentials and server URL of another server cluster; (ii) Adds this other server cluster to a list of remote server clusters to communicatewith;
[0072] (iii) Communicates with a network API server in that remote server cluster and collects the Services and their endpoint IPs;
[0073] (iv) Adds these Services, with their endpoints, to router configurations in Sidecars and Gateways in the local server cluster. An advantage is that the computer system is more resilient to faults.
[0074] The computer system may be one wherein requests are sent to a single IP address from the first service unit of the first data centre to keep full-server cluster outlier detection working as normal, so requests are not sent straight to the remote service name, as the router management service would add all Default Gateway service units in the second Data Centre as endpoints, and will no longer outlier detect entire server clusters, but instead would just outlier detect single service units. An advantage is that the computer system is more resilient to faults.
[0075] According to a second aspect of the invention, there is provided a computer implemented method of routing second requests, and of routing responses to the second requests, in a computer system including a first data centre and a second data centre, the first data centre communicating with the second data centre via a network, wherein the first data centre and the second data centre are separated by between 2 km and 500 km, the first data centre including a first service unit providing a first service, and a plurality of second service units each providing a second service, the first service unit including a first router proxy, and the first service unit including software executing to receive first requests and to send second requests relating to the first requests to the first router proxy;
[0076] the plurality of second service units each including a respective second router proxy and respective software executing to receive at least some of the second requests via the respective second router proxy, to process the at least some of the second requests to provide respective responses, and to respond to the at least some of the second requests using the respective responses via the respective second router proxy; the first data centre including a first default gateway ingress service, and a first default gateway egress service communicating with a first default gateway ingress service of the second data centre via the network;the second data centre including the first default gateway ingress service of the second data centre, a first default gateway egress service, and a plurality of third service units each providing the second service, the plurality of third service units each including a respective third router proxy and respective software executing to receive at least some of the second requests via the respective third router proxy, to process the at least some of the second requests to provide respective responses, and to respond to the at least some of the second requests using the respective responses via the respective third router proxy;
[0077] wherein the first router proxy routes a first set of at least some of the second requests to the first default gateway ingress service of the first data centre, and the first default gateway ingress service of the first data centre sends the first set of at least some of the second requests to the first default gateway ingress service of the first data centre to a second service unit of the plurality of second service units, and receives the respective responses from the second service unit, and sends the respective responses to the first service unit of the first data centre;
[0078] wherein the first router proxy routes a second set of at least some of the second requests to the first default gateway egress service of the first data centre, and the first default gateway egress service of the first data centre sends the second set of at least some of the second requests to the first default gateway ingress service of the second data centre, wherein the first default gateway ingress service of the second data centre sends the second set of at least some of the second requests to a service unit of the plurality of third service units of the second data centre, wherein the service unit of the plurality of third service units of the second data centre responds to the second set of at least some of the second requests using the respective responses via the respective third router proxy, wherein the respective third router proxy routes the respective responses to the first default gateway egress service of the second data centre, and the first default gateway egress service of the second data centre sends the respective responses to the first default gateway ingress service of the first data centre, wherein the first default gateway ingress service of the first data centre sends the respective responses to the first service unit of the first data centre. Optionally, the first service unit of the first data centre may store the responses it receives.
[0079] An advantage is that the computer system is resilient to faults, because if one or more,or all, second service units of the first data centre is determined to have become faulty, it is possible in response to route second requests to the third service units of the second data centre, which means that a computer terminal in connection with the system making first requests to the first service of the first data centre will experience no, or negligible, reduction in performance of the first service by the first data centre.
[0080] The method may be one including use of a system of any aspect of the first aspect of the invention.
[0081] According to a third aspect of the invention, there is provided a computer system including a first data centre, a second data centre and a third data centre, the first data centre configured to communicate with the second data centre and with the third data centre via a network, wherein the first data centre and the second data centre are separated by between 2 km and 500 km, wherein the third data centre and the second data centre are separated by between 2 km and 500 km, and wherein the first data centre and the third data centre are separated by between 2 km and 500 km, wherein each data centre includes a respective default gateway;
[0082] the first data centre including a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a first service, in which the first respective server cluster is configured to receive a query in respect of the first service, and in response to query a second service, and to receive a response from the second service, and to use the response received from the second service to provide a response from the first service,
[0083] the first data centre including a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the second service, in which the second respective server cluster is configured to receive a query in respect of the second service and to provide a response;
[0084] the second data centre including a respective subnet which includes a respective server cluster configured to provide the second service, in which the respective server cluster is configured to receive a query in respect of the second service and to provide a response;the third data centre including a respective subnet which includes a respective server cluster configured to provide the second service, in which the respective server cluster is configured to receive a query in respect of the second service and to provide a response;
[0085] wherein in the first data centre the first respective server cluster is configured to provide the first service, the first respective server cluster including a first service client unit, the first service client unit including a first service application executable in the first service client unit, and wherein the first service client unit includes a first router;
[0086] wherein in the first data centre the second respective server cluster includes a second service unit, the second service unit including a second service application executable in the second service unit, and the second service unit includes a second router configured to receive a request from the first service via the default gateway of the first data centre, to route the request to the second service application for processing by the second service application;
[0087] wherein in a first configuration of the first router, the first router is configured to route greater than 90% (e.g. greater than 97%, e.g. 99%) of the requests from the first service to the second service via the default gateway of the first data centre to the second respective subnet in the first data centre, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the second service;
[0088] wherein in the first configuration of the first router, the first router is configured to route between 0.1% and 5% (e.g. 1%) of the requests from the first service to the second service via the default gateway of the second data centre to the respective subnet in the second data centre, the respective subnet including the respective server cluster configured to provide the second service;
[0089] wherein the respective server cluster of the second data centre includes a second service unit, the second service unit including a second service application executable in the second service unit, and the second service unit includes a third router configured to receive a request from the first service via the default gateway of the second data centre, to route the request to the second service application for processing by the second service application;wherein in the first configuration of the first router, the first router is configured to route between 0.1% and 5% (e.g. 1%) of the requests from the first service to the second service via the default gateway of the third data centre to the respective subnet in the third data centre, the respective subnet including the respective server cluster configured to provide the second service;
[0090] wherein the respective server cluster of the third data centre includes a second service unit, the second service unit including a second service application executable in the second service unit, and the second service unit includes a fourth router configured to receive a request from the first service via the default gateway of the third data centre, to route the request to the second service application for processing by the second service application;
[0091] wherein the computer system does not include a network load balancer between the first data centre and the second data centre, and wherein the computer system does not include a network load balancer between the first data centre and the third data centre, and wherein the first data centre does not include a network load balancer between first respective subnet and the second respective subnet, wherein the network load balancers are configured to load balance the second requests. Optionally, the first router includes configuration settings, in which the first router may store its configuration settings.
[0092] An advantage is a more resilient architecture, because any network load balancer is a potential point of failure. An advantage is reduced latency, because a network load balancer typically introduces latency. An advantage is reduced cost, because a network load balancer consumes computing resources. An advantage is that the first router may be configured to monitor responses from the second respective server clusters, to identify if a particular second respective server cluster of the second respective server clusters appears to be faulty and therefore to be unsuitable to be used to receive an increased amount of traffic from the first service client unit of the first respective server cluster of the first data centre, in the event that the first router detects that the second respective server cluster of the first data centre including the respective second service unit becomes faulty, and traffic should be sent not to the second service unit of the second respective server cluster of the first data centre, but instead to the respective second service units of the non-faulty second respectiveserver clusters.
[0093] The computer system may be one wherein the percentages of traffic to the default gateways add up to one hundred percent.
[0094] The computer system may be one wherein in a second configuration of the first router, the first router is configured to route zero percent of the requests from the first service to the second service via the default gateway of the first data centre to the second respective subnet in the first data centre, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the second service, and wherein in the second configuration of the first router, the first router is configured to route a non-zero fraction of the requests from the first service to the second service via the default gateway of the second data centre to the respective subnet in the second data centre, the respective subnet including the respective server cluster, wherein the respective server cluster is configured to provide the second service; and wherein in the second configuration of the first router, the first router is configured to route a non-zero fraction of the requests from the first service to the second service via the default gateway of the third data centre to the respective subnet in the third data centre, the respective subnet including the respective server cluster, wherein the respective server cluster is configured to provide the second service.
[0095] An advantage is improved resilience of the computer system.
[0096] The computer system may be one wherein the percentages of traffic to the default gateways add up to one hundred percent.
[0097] The computer system may be one wherein the first router is configured to change from the first configuration of the first router to the second configuration of the first router, in response to the first router detecting that the second respective server cluster of the first data centre including the respective second service unit becomes faulty. An advantage is that system routing is reconfigured quickly, in response to detecting thatthe second respective server cluster of the first data centre including the respective second service unit has become faulty. This means for example that there is no need to wait for a notification from a network service provider that a fault has been detected within the network in relation to the second respective server cluster of the first data centre including the respective second service unit to make the reconfiguration. This also means for example that the reconfiguration can be performed in relation to faults that a network service provider would not be expected to detect, such as the respective second service unit providing output that has traffic characteristics which are acceptable to the network service provider, but which are not acceptable to the provider of the second service: for example the second service unit operating at a reduced response time which is acceptable to the network service provider, but which is not acceptable to the provider of the second service unit.
[0098] The computer system may be one wherein in a third configuration of the first router, the first router is configured to route greater than 90% (e.g. greater than 97%, e.g.
[0099] 99%) of the requests from the first service to the second service via the default gateway of the first data centre to the second respective subnet in the first data centre, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the second service, and wherein in the third configuration of the first router, the first router is configured to route zero percent of the requests from the first service to the second service via the default gateway of the second data centre to the respective subnet in the second data centre, the respective subnet including the respective server cluster, wherein the respective server cluster is configured to provide the second service, and wherein in the third configuration of the first router, the first router is configured to route a fraction between 0.1% and 5% (e.g. 1%) of the requests from the first service to the second service via the default gateway of the third data centre to the respective subnet in the third data centre, the respective subnet including the respective server cluster, wherein the respective server cluster is configured to provide the second service.The computer system may be one wherein the percentages of traffic to the default gateways add up to one hundred percent.
[0100] The computer system may be one wherein the first router is configured to change from the first configuration of the first router to the third configuration of the first router, in response to the first router detecting that the respective server cluster of the second data centre including the respective second service unit has become faulty. An advantage is that the system is resilient to faulty upgrades.
[0101] The computer system may be one wherein all traffic between the server clusters is routed by the routers via the default gateways.
[0102] The computer system may be one wherein spacing the data centres apart reduces the risk of the data centres all going down if a disaster (e.g. explosion, fire, power outage) occurs locally in a geographic region. An advantage is that the computer system is more resilient to faults.
[0103] The computer system may be one wherein the system includes a control plane and a data plane, wherein the data plane includes all the proxies deployed as sidecars or as gateways, or as other configurable proxies, and wherein the control plane manages and configures the proxies to route traffic. An advantage is that the computer system is more resilient to faults.
[0104] The computer system may be one wherein the data plane is monitored by measuring at least traffic, errors, latency and saturation.
[0105] The computer system may be one wherein the control plane includes four components: Componentl which provides service discovery for the gateway sidecars; component2 which is the router’s configuration validation, ingestion, processing and distribution component; components which enables strong service-to-service and enduser authentication with built-in identity and credential management; component4 which enforces access control and usage policies across the service mesh, and collectstelemetry data from the gateway proxy and other services. An advantage is that the computer system is more resilient to faults.
[0106] The computer system may be one wherein each data centre includes at least one server cluster, wherein each server cluster runs its own control plane, and / or wherein the router control plane is deployed independently into each server cluster. An advantage is that the computer system is more resilient to faults.
[0107] The computer system may be one wherein service units each run a router-compatible sidecar proxy.
[0108] The computer system may be one wherein sidecar proxy injection into a new service unit happens automatically every time a new service unit is created.
[0109] The computer system may be one wherein traffic is directed from application services to and from the sidecars.
[0110] The computer system may be one wherein when router software is updated, sidecars in existing service units remain the same until those service units are deleted.
[0111] The computer system may be one wherein when those service units are deleted, replacement service units are created.
[0112] The computer system may be one wherein http 503 error response codes lead to one or more retries by a sidecar proxy.
[0113] The computer system may be one wherein when a service unit or any intermediate proxies returns a http 503 error, the client gateway sidecar proxy automatically retries for another service unit.
[0114] The computer system may be one wherein a fully jittered exponential back-off algorithm is used for retries.The computer system may be one in which router circuit breakers are configured to automatically failover traffic to the closest working server cluster or service unit. An advantage is that the computer system is more resilient to faults.
[0115] The computer system may be one in which there are configured two sets of circuit breaking rules:
[0116] one rule for communication within a server cluster, which applies to requests coming from the ingress gateway going to the destination service units;
[0117] one rule for communication across server clusters.
[0118] The computer system may be one in which the system includes a mesh which provides service uni t-to- service unit communication.
[0119] The computer system may be one in which all service uni t-to- service unit communication within the mesh is end-to-end encrypted, and the encryption is handled by the router proxy e.g. router sidecar proxy.
[0120] The computer system may be one in which each service includes its own security certificate, which is included into the router proxy e.g. router sidecar proxy.
[0121] The computer system may be one in which certificates are automatically rotated and dynamically reloaded by the router proxy e.g. router sidecar proxy, without any disruption to the service unit.
[0122] The computer system may be one in which a service unit (e.g. its service owner) must declare the list of services that the router proxy, e.g. router sidecar proxy, can see or call.
[0123] The computer system may be one in which these services are the only services within the mesh.The computer system may be one wherein all sidecar proxies in the mesh are programmed with the necessary configuration required to reach every workload instance in the mesh.
[0124] The computer system may be one wherein a default gateway in the first or second data centre connects to between one hundred and ten thousand service units in the same data centre.
[0125] The computer system may be one wherein a default gateway in the first or second data centre connects to between one thousand and five thousand service units in the same data centre.
[0126] The computer system may be one wherein services are not deployed to a single server cluster.
[0127] The computer system may be one wherein each server cluster has its own service mesh control plane.
[0128] The computer system may be one wherein each service unit has its own sidecar which all inbound and outbound traffic goes through.
[0129] The computer system may be one wherein for traffic within each subnet, it is preferred to use the local subnet; the next preference is other subnets of the data centre of the local subnet, and the next preference is subnets of other data centres. An advantage is that the computer system is more resilient to faults.
[0130] The computer system may be one wherein server clusters’ traffic is drained and shifted to other working server clusters in other Data Centres.
[0131] The computer system may be one wherein only server clusters within a single Data Centre are drained at one time. An advantage is that the computer system is more resilient to faults.The computer system may be one wherein at least three server clusters are present in each data centre, e.g. to ensure at least one server cluster is working. An advantage is that the computer system is more resilient to faults.
[0132] The computer system may be one wherein each route to a service goes via a default gateway; each default gateway is configured to route to any service units providing a particular service, whereas the mesh is configured to route to, or to receive traffic from, any of the default gateways.
[0133] The computer system may be one wherein the default gateways are configured in the routers as a Service Entry in every server cluster, not just in the server clusters the service is deployed in.
[0134] The computer system may be one wherein when updating routers, routers are updated in at most one Data Centre at a time. An advantage is that the computer system is more resilient to faults.
[0135] The computer system may be one wherein updates to routers are done after draining the server clusters in the Data Centre in which the routers are being updated. An advantage is that the computer system is more resilient to faults.
[0136] The computer system may be one wherein every server cluster provides an isolated failure zone. An advantage is that the computer system is more resilient to faults.
[0137] The computer system may be one wherein there is no connectivity between different Data Centres in different geographic regions of the world. An advantage is that the computer system is more resilient to faults.
[0138] The computer system may be one wherein the path for traffic is specified with routing rules, and destination rules are used to configure a set of policies that gateway proxies apply to a request at a specific destination.
[0139] The computer system may be one wherein a Router Management Service in a localserver cluster uses a Discovery tool to discover endpoints within other server clusters, to do the following:
[0140] (i) Read a secret which contains credentials and server URL of another server cluster; (ii) Adds this other server cluster to a list of remote server clusters to communicate with;
[0141] (iii) Communicates with a network API server in that remote server cluster and collects the Services and their endpoint IPs;
[0142] (iv) Adds these Services, with their endpoints, to router configurations in Sidecars and Gateways in the local server cluster. An advantage is that the computer system is more resilient to faults.
[0143] The computer system may be one wherein requests are sent to a single IP address from the first service unit of the first data centre to keep full-server cluster outlier detection working as normal, so requests are not sent straight to the remote service name, as the router management service would add all Default Gateway service units in the second Data Centre as endpoints, and will no longer outlier detect entire server clusters, but instead would just outlier detect single service units. An advantage is that the computer system is more resilient to faults.
[0144] According to a fourth aspect of the invention, there is provided a computer-implemented method of reconfiguring a first router in a computer system from a first configuration to a second configuration, wherein the computer system includes a first data centre, a second data centre and a third data centre, the first data centre communicating with the second data centre and with the third data centre via a network, wherein the first data centre and the second data centre are separated by between 2 km and 500 km, wherein the third data centre and the second data centre are separated by between 2 km and 500 km, and wherein the first data centre and the third data centre are separated by between 2 km and 500 km, wherein each data centre includes a respective default gateway;
[0145] the first data centre including a first respective subnet which includes a first respective server cluster, the first respective server cluster providing a first service, in which the first respective server cluster receives a query in respect of the first service, and in response queries a second service, and receives a response from the second service,and uses the response received from the second service to provide a response from the first service,
[0146] the first data centre including a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster provides the second service, in which the second respective server cluster receives a query in respect of the second service and provides a response;
[0147] the second data centre including a respective subnet which includes a respective server cluster providing the second service, in which the respective server cluster receives a query in respect of the second service and provides a response;
[0148] the third data centre including a respective subnet which includes a respective server cluster providing the second service, in which the respective server cluster receives a query in respect of the second service and provides a response;
[0149] wherein in the first data centre the first respective server cluster provides the first service, the first respective server cluster including a first service client unit, the first service client unit including a first service application executing in the first service client unit, and wherein the first service client unit includes the first router; wherein in the first data centre the second respective server cluster includes a second service unit, the second service unit including a second service application executing in the second service unit, and the second service unit includes a second router which receives a request from the first service via the default gateway of the first data centre, and routes the request to the second service application for processing by the second service application;
[0150] wherein in a first configuration of the first router, the first router routes greater than 90% (e.g. greater than 97%, e.g. 99%) of the requests from the first service to the second service via the default gateway of the first data centre to the second respective subnet in the first data centre, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster provides the second service;
[0151] wherein in the first configuration of the first router, the first router routes between 0.1% and 5% (e.g. 1%) of the requests from the first service to the second service via the default gateway of the second data centre to the respective subnet in the seconddata centre, the respective subnet including the respective server cluster providing the second service;
[0152] wherein the respective server cluster of the second data centre includes a second service unit, the second service unit including a second service application executing in the second service unit, and the second service unit includes a third router which receives a request from the first service via the default gateway of the second data centre, and routes the request to the second service application for processing by the second service application;
[0153] wherein in the first configuration of the first router, the first router routes between 0.1% and 5% (e.g. 1%) of the requests from the first service to the second service via the default gateway of the third data centre to the respective subnet in the third data centre, the respective subnet including the respective server cluster providing the second service;
[0154] wherein the respective server cluster of the third data centre includes a second service unit, the second service unit including a second service application executing in the second service unit, and the second service unit includes a fourth router which receives a request from the first service via the default gateway of the third data centre, to route the request to the second service application for processing by the second service application;
[0155] wherein the computer system does not include a network load balancer between the first data centre and the second data centre, and wherein the computer system does not include a network load balancer between the first data centre and the third data centre, and wherein the first data centre does not include a network load balancer between first respective subnet and the second respective subnet, wherein the network load balancers are configured to load balance the second requests;
[0156] wherein in a second configuration of the first router, the first router routes zero percent of the requests from the first service to the second service via the default gateway of the first data centre to the second respective subnet in the first data centre, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster provides the second service, and wherein in the second configuration of the first router, the first router routes a non-zero fraction of the requests from the first service to the second service via the default gateway of thesecond data centre to the respective subnet in the second data centre, the respective subnet including the respective server cluster, wherein the respective server cluster provides the second service; and wherein in the second configuration of the first router, the first router routes a non-zero fraction of the requests from the first service to the second service via the default gateway of the third data centre to the respective subnet in the third data centre, the respective subnet including the respective server cluster, wherein the respective server cluster provides the second service;
[0157] the method including the step of the first router changing from the first configuration of the first router to the second configuration of the first router, in response to the first router detecting that the second respective server cluster of the first data centre including the respective second service unit becomes faulty.
[0158] An advantage is that system routing is reconfigured quickly, in response to detecting that the second respective server cluster of the first data centre including the respective second service unit has become faulty. This means for example that there is no need to wait for a notification from a network service provider that a fault has been detected within the network in relation to the second respective server cluster of the first data centre including the respective second service unit to make the reconfiguration. This also means for example that the reconfiguration can be performed in relation to faults that a network service provider would not be expected to detect, such as the respective second service unit providing output that has traffic characteristics which are acceptable to the network service provider, but which are not acceptable to the provider of the second service: for example the second service unit operating at a reduced response time which is acceptable to the network service provider, but which is not acceptable to the provider of the second service unit.
[0159] The method may be one including use of a system of any aspect of the third aspect of the invention.
[0160] Aspects of the invention may be combined.BRIEF DESCRIPTION OF THE FIGURES
[0161] Aspects of the invention will now be described, by way of example(s), with reference to the following Figures, in which:
[0162] Figure 1 shows an example of a resilient server architecture.
[0163] Figure 2 shows an example of a resilient server architecture in a virtual cloud system.
[0164] Figure 3 shows an example of a resilient server architecture.
[0165] Figure 4 shows an example of a resilient server architecture in a virtual cloud system.
[0166] Figure 5 shows an example of a configuration of a resilient server architecture.
[0167] Figure 6 shows an example of a configuration of a resilient server architecture.
[0168] Figure 7 shows an example of a configuration of a resilient server architecture.
[0169] Figure 8 shows an example of a service C making requests to a service D using connected Zone Data Centres.
[0170] Figure 9 shows an example of routing a request from a service A to a service B. Figure 10 shows an example of connected Zone Data Centres configured to route requests from service A to service B.DETAILED DESCRIPTION
[0171] A data centre in a zone is a single Data Center or a group of Data Centers. A data centre in a zone is located many miles (e.g. between one mile and one hundred miles, or between 3 km and 100 km, or between 2 km and 500 km) from any data centre in a different zone. Spacing the zones apart, or spacing the data centres apart, reduces the risk of the data centres all going down if a disaster (e.g. explosion, fire, power outage) occurs locally in a geographic region. Examples of geographic regions are Western North America, Western Europe, Japan.
[0172] In an example, there is provided a resilient server architecture. In the resilient server architecture, in a data centre in zone a there is provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone a there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. In the resilient server architecture, in a data centre in zone b there is provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone b there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. In the resilient server architecture, in a data centre inzone c there is provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone c there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. The resilient server architecture is present with respect to data centres in zones a, b and c. If a first respective server cluster in a first data centre in a zone becomes faulty, a different first respective server cluster in a different zone may be used to provide service A, which means that the server architecture has the advantage that it is resilient to faults. If a second respective server cluster in a first data centre in a zone becomes faulty, a different second respective server cluster in a different zone may be used to provide service B, which means that the server architecture has the advantage that it is resilient to faults. An example is shown in Figure 1.
[0173] In an example, there is provided a virtual cloud system, the virtual cloud system including a resilient server architecture. In the resilient server architecture, in a data centre in zone a there is provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone a there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. In the resilient server architecture, in a data centre in zone b there is provided a first respective subnet which includes a first respectiveserver cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone b there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. In the resilient server architecture, in a data centre in zone c there is provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone c there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. The resilient server architecture is present with respect to data centres in zones a, b and c. If a first respective server cluster in a first data centre in a zone becomes faulty, a different first respective server cluster in a different zone may be used to provide service A, which means that the server architecture has the advantage that it is resilient to faults. If a second respective server cluster in a first data centre in a zone becomes faulty, a different second respective server cluster in a different zone may be used to provide service B, which means that the server architecture has the advantage that it is resilient to faults. An example is shown in Figure 2.
[0174] In an example, there is provided a resilient server architecture. In the resilient server architecture, in a data centre in zone a there is provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster isconfigured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone a there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. In the resilient server architecture, in a data centre in zone b there is provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone b there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. In the resilient server architecture, in a data centre in zone c there is provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone c there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. The resilient server architecture is present with respect to data centres in zones a, b and c. A first load balancer (LB1) is configured to manage traffic to the first respective server clusters. A second load balancer (LB2) is configured to manage traffic to the second respective server clusters. Load balancers comprise the first loadbalancer (LB1) and the second load balancer (LB2). If a first respective server cluster in a first data centre in a zone becomes faulty, a different first respective server cluster in a different zone may be used to provide service A, which means that the server architecture has the advantage that it is resilient to faults. If a second respective server cluster in a first data centre in a zone becomes faulty, a different second respective server cluster in a different zone may be used to provide service B, which means that the server architecture has the advantage that it is resilient to faults. An example is shown in Figure 3.
[0175] In an example, there is provided a virtual cloud system, the virtual cloud system including a resilient server architecture. In the resilient server architecture, in a data centre in zone a there is provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone a there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. In the resilient server architecture, in a data centre in zone b there is provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone b there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. In the resilient server architecture, in a data centre in zone c there isprovided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone c there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. The resilient server architecture is present with respect to data centres in zones a, b and c. A first load balancer (LB1) is configured to manage traffic to the first respective server clusters. A second load balancer (LB2) is configured to manage traffic to the second respective server clusters. Load balancers comprise the first load balancer (LB1) and the second load balancer (LB2). If a first respective server cluster in a first data centre in a zone becomes faulty, a different first respective server cluster in a different zone may be used to provide service A, which means that the server architecture has the advantage that it is resilient to faults. If a second respective server cluster in a first data centre in a zone becomes faulty, a different second respective server cluster in a different zone may be used to provide service B, which means that the server architecture has the advantage that it is resilient to faults. An example is shown in Figure 4.
[0176] In an example, there is provided a resilient server architecture. In the resilient server architecture, in a data centre in zone a there is provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone a there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a serviceB, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. In the resilient server architecture, in a data centre in zone b there may be provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone b there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. In the resilient server architecture, in a data centre in zone c there may be provided a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a service A, in which the first respective server cluster is configured to receive a query in respect of service A and in response to query a service B and to receive a response from service B, and to provide a response from service A. In the resilient server architecture, in the data centre in zone c there is provided a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide a service B, in which the second respective server cluster is configured to receive a query in respect of service B and to provide a response. The resilient server architecture is present with respect to data centres in zones a, b and c.
[0177] In the data centre in zone a which includes the first respective subnet which includes the first respective server cluster, wherein the first respective server cluster is configured to provide a service A, the first respective server cluster includes a Service A Client Unit, the Service A client unit including a Service A application executable in the Service A Client Unit, and the first respective server cluster includes a first router. In a first configuration of the first router, the first router is configured to route a large fraction (e.g. >90%, e.g. >97%, e.g. 99%) of the requests from Service A toService B via a default gateway of the zone a Data Centre to the second respective subnet in the data centre in zone a, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the service B. The second respective server cluster includes a Service B Unit, the Service B unit including a Service B application executable in the Service B Unit, and the Service B Unit includes a second router configured to receive a request from service A via the default gateway of the zone a Data Centre, to route the request to the Service B application for processing by the Service B application.
[0178] In the data centre in zone a, in the first configuration, the first router is configured to route a small fraction (e.g. between 0.1% and 5%, e.g. 1%) of the requests from Service A to Service B via a default gateway of the zone b Data Centre to the second respective subnet in the data centre in zone b, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the service B. The second respective server cluster includes a Service B Unit, the Service B unit including a Service B application executable in the Service B Unit, and the Service B Unit includes a third router configured to receive a request from service A via the default gateway of the zone b Data Centre, to route the request to the Service B application for processing by the Service B application.
[0179] In the data centre in zone a, in the first configuration, the first router is configured to route a small fraction (e.g. between 0.1% and 5%, e.g. 1%) of the requests from Service A to Service B via a default gateway of the zone c Data Centre to the second respective subnet in the data centre in zone c, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the service B. The second respective server cluster includes a Service B Unit, the Service B unit including a Service B application executable in the Service B Unit, and the Service B Unit includes a fourth router configured to receive a request from service A via the default gateway of the zone c Data Centre, to route the request to the Service B application for processing by the Service B application. Anexample is shown in Figure 5. The percentages of traffic to the default gateways add up to one hundred percent.
[0180] An advantage is that the first router may be configured to monitor responses from the second respective server clusters, to identify if a particular second respective server cluster of the second respective server clusters appears to be faulty and therefore to be unsuitable to be used to receive an increased amount of traffic from the Service A Client Unit of the first respective server cluster of the zone a data centre, in the event that the first router detects that the second respective server cluster of the zone a data centre including the respective Service B Unit becomes faulty, and traffic should be sent not to the Service B Unit of the second respective server cluster of the zone a data centre, but instead to the respective Service B Units of the non-faulty second respective server clusters.
[0181] In a second configuration of the first router, the first router is configured to route zero percent of the requests from Service A to Service B via a default gateway of the zone a Data Centre to the second respective subnet in the data centre in zone a, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the service B. In the second configuration of the first router, the first router is configured to route a non-zero fraction of the requests from Service A to Service B via a default gateway of the zone b Data Centre to the second respective subnet in the data centre in zone b, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the service B. In the second configuration of the first router, the first router is configured to route a non-zero fraction of the requests from Service A to Service B via a default gateway of the zone c Data Centre to the second respective subnet in the data centre in zone c, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the service B. An example is shown in Figure 6. The percentages of traffic to the default gateways add up to onehundred percent.
[0182] In an example, the first router is configured to change from the first configuration of the first router to the second configuration of the first router, in response to the first router detecting that the second respective server cluster of the zone a data centre including the respective Service B Unit becomes faulty. An advantage is that system routing is reconfigured quickly, in response to detecting that the second respective server cluster of the zone a data centre including the respective Service B Unit has become faulty. This means for example that there is no need to wait for a notification from a network service provider that a fault has been detected within the network in relation to the second respective server cluster of the zone a data centre including the respective Service B Unit to make the reconfiguration. This also means for example that the reconfiguration can be performed in relation to faults that a network service provider would not be expected to detect, such as the respective Service B Unit providing output that has traffic characteristics which are acceptable to the network service provider, but which are not acceptable to the provider of Service B: for example Service B operating at a reduced response time which is acceptable to the network service provider, but which is not acceptable to the provider of Service B.
[0183] In a third configuration of the first router, the first router is configured to route a large fraction (e.g. >90%, e.g. >97%, e.g. 99%) of the requests from Service A to Service B via a default gateway of the zone a Data Centre to the second respective subnet in the data centre in zone a, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the service B. In the third configuration of the first router, the first router is configured to route zero percent of the requests from Service A to Service B via a default gateway of the zone b Data Centre to the second respective subnet in the data centre in zone b, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the service B. In the third configuration of the first router, the first router is configured to route a small fraction (e.g. between 0.1% and 5%, e.g. 1%) of the requests fromService A to Service B via a default gateway of the zone c Data Centre to the second respective subnet in the data centre in zone c, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the service B. An example is shown in Figure 7. The percentages of traffic to the default gateways add up to one hundred percent.
[0184] In an example, in the resilient server architecture, in the data centre in zone b which includes the second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, in which the second respective server cluster is configured to provide the service B, the service B is upgraded, but the respective Service B Unit becomes faulty as a result of the upgrade. In this example, the first router is configured to change from the first configuration of the first router to the third configuration of the first router, in response to the first router detecting that the second respective server cluster of the zone b data centre including the respective Service B Unit has become faulty. An advantage is that the system is resilient to faulty upgrades.
[0185] In an example, the first router is configured to change from the first configuration of the first router to the third configuration of the first router, in response to the first router detecting that the second respective server cluster of the zone b data centre including the respective Service B Unit becomes faulty.
[0186] In an example, all traffic between server clusters is routed by the routers via default gateways. In an example, all traffic between server clusters is routed by the routers via default gateways without passing through any network load balancers. An advantage is a more resilient architecture, because any network load balancer is a potential point of failure. An advantage is reduced latency, because a network load balancer typically introduces latency.
[0187] In an example, a default gateway in a zone data centre may connect to between one hundred and ten thousand service units (e.g. service B units) in the zone data centre. In an example, a default gateway in a zone data centre may connect to between onethousand and five thousand service units (e.g. service B units) in the zone data centre.
[0188] In an example, service A includes using a structured database, and providing a response to a query received relating to contents of the structured database, the response process including sending data from the structured database to service B for analysis by service B, and receiving the analysis results from service B, and then responding to the query received relating to contents of the structured database using the analysis results from service B.
[0189] In an example, there is provided a resilient server architecture. In the resilient server architecture, multiple single server clusters are provided from which services (e.g. user services) are deployed. This differs from an alternative less resilient server architecture including server clusters, where services are only deployed to a single cluster.
[0190] The resilient server architecture poses a challenge to route requests from client services, to their dependent services. In the alternative server architecture, clients know if their dependent services are in the same cluster, in a remote cluster, or not running in the alternative server architecture at all. This information known by the clients is fairly static. In the resilient server architecture, dependent services might be located in the same cluster as the client, in a different cluster in the same zone data centre, or if there is a failure of a cluster in the same zone data centre, they might be in a cluster in another zone data centre.
[0191] During preliminary design considerations, we saw two main ways to tackle this issue:
[0192] • Mutli-cluster ingress - i.e. a single (e.g. Layer 7) Load balancer above all clusters which all traffic goes through. Layer 7 load balancing describes traffic distribution at the application layer of the Open Systems Interconnection (OSI) model, which is where human-application interaction occurs and where applications access network services.
[0193] • Client Side Routing - i.e. a router which routes traffic to configured locations, preferably with the nearest locality preferred unless in a failover state.
[0194] In an example, the second option is preferred.In an example, no instance of a service is deployed across two or more zone data centres.
[0195] In an example, services are spread across clusters, not confined to one cluster. Each server cluster may have its own service mesh control plane. Each service unit may have its own sidecar which all inbound and outbound traffic goes through. In an example, for traffic within each subnet, it is preferred to use the local subnet; the next preference is other subnets of the zone data centre of the local subnet, and the next preference is subnets of other zone data centres. This may be implemented with a Service Entry in each cluster, performance outlier detection from each client, and load balancing across service units by a gateway associated with each subnet. A service entry describes the properties of a service (e.g. domain name server (DNS) name, virtual IP addresses (VIPs), ports, protocols, endpoints).
[0196] For example, if a service is running on n>l clusters in a plurality of zone data centres, each cluster should receive a 1 / n of the traffic for that service from external services, but local clients will prefer their closest service in the same zone data centre.
[0197] A service might become unavailable inside one or multiple server clusters. This could be due to a bad deployment, a bug, or a problem on the respective one or multiple server clusters. In this scenario, traffic should stop being sent to the faulty one or faulty multiple server clusters and only be routed to the non-faulty server clusters.
[0198] There is provided the ability to shift all or most of the traffic away from a server cluster in order to perform critical maintenance or upgrades; this may be called “draining the server cluster”. There is provided the ability to shift all or most of the traffic away from a zone data centre in order to perform critical maintenance or upgrades; this may be called “draining the zone data centre”.
[0199] Communication between services may be secure by default.
[0200] An advantage is providing a resilient platform that can provide 99.99% globalavailability for all services through the creation of server clusters (e.g. isolated units of resilience, scale and capacity) that ensure that a smallest fault isolation zone in our system is a server cluster within a Zone Data Centre, not an entire region (e.g. North America, Western Europe) of the world. In particular this design is to ensure resilient traffic routing.
[0201] An advantage is providing a secure network. This design ensures that service to service communication is secure by default and end-to-end encrypted.
[0202] Example Scenario
[0203] We have a service C making requests to a service D using connected Zone Data Centres. The service C is deployed in three service C server clusters, one in each of three Zone Data Centres: a, b, c. The service D is deployed on two service D server clusters across two Zone Data Centres, one in each of Zone Data Centres a and b. In the service C server cluster (in Zone Data Centre a), half of the requests for service D in Zone Data Centre a return an http503 error. An example is shown in Figure 8.
[0204] What is the availability perceived by the upstream service C? In order to answer this question, we need to explain how requests are routed inside the mesh:
[0205] each gateway will retry a request up to three times, each time hitting a different service D unit in that service D server cluster until it finds one that does not return a http 502 / 503 / 504 error. If all retries fail it returns a http 503 error to the service C caller.
[0206] Each router prefers its local gateway. In our example, service C in Zone Data Centre a sends all its requests to service D via the gateway in Zone Data Centre a; service C in Zone Data Centre b sends all its requests to service D via the gateway in Zone Data Centre b; service C in Zone Data Centre c sends half of its requests to service D via the gateway in Zone Data Centre a, and half of its requests to service D via the gateway in Zone Data Centre b.
[0207] Requests will only go to another Zone Data Centre when a circuitbreaker function kicks-in as per its circuit breaker settings. In our example, we consider the failures within each gateway to be too low to trigger the circuit breaker.Let's now compute the availability perceived by service C in each availability zone. In server C service cluster in Zone Data Centre a:
[0208] service C communicates with service D via the local gateway (traffic tends to stay in the
[0209] local Zone Data Centre) by default; retry is up to 3-1=2 times, covering most of the errors. The probability of service C seeing an error in this server C service cluster is 0.5A3= 12.5%.
[0210] Hence availability in this server C service cluster is 87.5%.
[0211] in the server C service cluster cell in Zone Data Centre b:
[0212] service C communicates with service D via the local gateway
[0213] availability is then 100%
[0214] in server C service cluster cell in Zone Data Centre c:
[0215] service C communicates with service D across availability zones (50% to Zone Data Centre a, 50% to Zone Data Centre b)
[0216] as mentioned above, service C sees 12.5% of the request to service D in Zone Data Centre a failing. However all those requests will be covered up by retries to service D in Zone Data Centre b.
[0217] Availability is then 100%
[0218] In summary:
[0219] The actual availability of service D is 75% (50% failure in Zone Data Centre a, 0% failure in Zone Data Centre b).
[0220] The perceived availability seen by service C is 95% (12.5% failure in Zone Data Centre a, 0% failure in Zone Data Centre b, 0% failure in Zone Data Centre c).
[0221] Example in which service units are injected with a sidecar proxy
[0222] In an example, service units run a router-compatible sidecar proxy. The sidecar proxy injection into a new service unit may happen automatically every time a new service unit is created. Traffic may be directed from application services to and from these sidecars. Service applications are connected to the router service mesh.
[0223] When router software is updated, sidecars in existing service units may remain the same until service units are deleted; replacement service units may then be created.There might be times when a given service cluster includes different service units which include different versions of the sidecar proxy. Most times replacement of older sidecar proxies is left to naturally occur in the life cycle of the server cluster, but it is possible to update the sidecar proxies in a server cluster one after the other, in a sequence. The version of the sidecar proxy and all other sidecar proxy templates may be stored in a configuration map.
[0224] All the incoming traffic the port application is listening to may be redirected to the gateway proxy. The same holds true for all outgoing traffic.
[0225] Benefits of this setup are:
[0226] • gives the service unit access to all the features the mesh has to offer.
[0227] • additional visibility on all the egress traffic, e.g. inside and outside of the mesh.
[0228] Example in which Router Circuit breaker enabled by default
[0229] We may use router circuit breakers in order to automatically failover traffic to the closest working server cluster or service unit.
[0230] The router may use outlier detection as the circuit breaking mechanism. A Circuit breaker implementation may be one that tracks the status of each individual host in the upstream service. For HTTP services, hosts that continually return 502, 503 and / or 504 errors (aka gateway failures) may be ejected from the pool for a pre-defined period of time.
[0231] In an example, we configure two sets of circuit breaking rules:
[0232] one rule for communication within a server cluster: applies to requests coming from the ingress gateway going to the destination service units;
[0233] one rule for communication across server clusters.
[0234] We can set sensible defaults for both rules, but let service owners tune those parameters when needed.Let's go through an example rule. This rule configures upstream hosts to be scanned every minute, such that any host that fails 7 consecutive times with http 5XX error code will be ejected for 5 minutes. Outlier detection is enabled as long as the associated load balancing pool has at least 10 percent of the hosts in working mode. At most 50% of the hosts will be ejected.
[0235] Automatic retries e.g. on HTTP 503 errors
[0236] In an example, http 503 error response codes lead to one or more retries by a sidecar proxy. When a service unit or any intermediate proxies returns a http 503 error, the client gateway sidecar proxy may automatically retry for another service unit. A gateway automatically retrying up to 3-1=2 times (total of 3 requests) may occur when such an http error is returned. A fully jittered exponential back-off algorithm may be used for retries e.g. with a default base interval of 25ms. For example, given the default interval, the first retry will be delayed randomly by 0-24ms, the 2nd retry by 0-74ms, a 3rd retry by 0-174ms, and so on.
[0237] The behaviour of default retries can be tuned by service owners at deployment time by tuning a virtual service object. The benefits for this setup include: Increased resiliency to network failures.
[0238] Example: Automatic End to end encryption and service identity
[0239] In an example:
[0240] 1. All service unit-to-service unit communication within the mesh may be end-to-end encrypted. The encryption may be handled by the router sidecar proxy. Each service gets its own security certificate, which may be included into the router-sidecar proxy. All the generated certificates may be derived from the same root Certificate Authority. Certificates may be automatically rotated and dynamically reloaded by the sidecar proxy without any disruption to the service unit.
[0241] 2. The identity of the client may be propagated to the server.
[0242] 3. HTTP protocol may be used in applications. For example, when I make an HTTPrequest within my application, it will first reach the sidecar proxy. The connection to the proxy is secure, because both the service unit and the proxy are part of the same (e.g. linux) network namespace, invisible from other service units running on the same host (e.g. it requires privileged kernel capabilities). The gateway proxy may then establish a secure communication to the next hop over HTTPS. This process may be fully transparent from the application point of view.
[0243] Example: Declaration of downstream dependencies.
[0244] In an example, a service unit (e.g. its service owner) must declare the list of services that the sidecar proxy can see / call. In an example, these are just the services within the mesh. External services may be passed through the mesh if it is in a mode to allow any egress.
[0245] By default, regarding the routers, all sidecar proxies in the mesh may be programmed with the necessary configuration required to reach every workload instance in the mesh. But this may not be a desired configuration. As the number services and workload in the mesh grows, the size of all sidecar proxies configuration increases. This can quickly become an issue at scale. In an example, the routers include a Sidecar resource to restrict the set of services that the proxy can reach when forwarding outbound traffic from workload instances. This may provide the ability to enforce boundaries of control.
[0246] Example Draining of server clusters
[0247] Server clusters may be drained (e.g. frequently) to mitigate issues or as part of routine maintenance. There is only so much resiliency we can guarantee within a single server cluster. Core components such as the DNS service, may have bugs and might fail. Also, there is no such thing as risk free maintenance operations. Changing core components might require complex orchestration and it might be safer to perform without traffic.
[0248] In an example:
[0249] Server clusters may be drained and traffic may be shifted to other working server clusters in other Zone Data Centres.In an example, we drain only server clusters within a single Zone Data Centre at a time.
[0250] In an example, we use the n+2 redundancy model (n>l) by deploying applications to multiple server clusters across different Zone Data Centres. At any point in time, one server cluster might be down for maintenance and one server cluster may down due to failure, so we should have at least three server clusters available in general, to ensure at least one server cluster is working in such a scenario.
[0251] Note that the mesh may be configured to prefer to keep traffic within the local Zone Data Centre. But at any point in time, cluster operators can decide to reconfigure the mesh and change this behaviour.
[0252] Example Configuration of services in all server clusters
[0253] The route to a service goes via a gateway. The default gateway is aware of all of the service units providing the service, whereas the mesh is aware of the default gateways. The default gateways are configured in the routers as a Service Entry in every server cluster, not just in the server clusters the service is deployed in.
[0254] In an example, there is a server cluster in each of Zone Data Centres a, b, and c. The service is only deployed in Zone Data Centres a, and b. However a service entry needs to be deployed in Zone Data Centre c also, so services can route to a server cluster in Zone Data Centres a and b via the mesh. These service entries may be critical to the server cluster routing and may be deployed when installing a new service.
[0255] Typical latency costs of communication between Zone Data Centres
[0256] We find that typically the latency costs of communication between services in different Zone Data Centres, when compared to communication between services within a single Zone Data Centre, are 2ms to 7ms. Typically it is advantageous to tradeoff the added latency in order to gain the added resiliency provided by the multiple Zone Data Centres architecture.Typical effect of using the routers in the multiple Zone Data Centres architecture
[0257] Possible advantages of using the routers in the multiple Zone Data Centres architecture are identity creation, automatic retries, a fair loadbalancing algorithm, and traffic shifting. We find that typically the latency costs of using the routers are 2ms to 5 ms. Typically it is advantageous to tradeoff the added latency in order to gain the added functionality provided by the routers.
[0258] The latency added by enabling mutual Transport Layer Security (mTLS) inside the mesh is about 0.5ms. This cost being minimal, we may enforce mTLS by default inside the mesh.
[0259] Monitoring and Alerting
[0260] The key question to answer is "how do we know the mesh is working?". In this section we sketch an example approach to monitoring a large scale mesh.
[0261] The problem can be split in two, monitoring the control plane and monitoring the data plane.
[0262] The data plane may include all the intelligent proxies deployed as sidecars or as gateways, or as other configurable proxies. These proxies mediate and control all network communication between microservices.
[0263] The control plane may manage and configure the proxies to route traffic.
[0264] Monitoring the data plane
[0265] In an example, every single request transmitted between services in the mesh transits via the sidecar proxies. To understand the scope of this task, we are talking about many thousands of proxies distributed throughout the whole infrastructure.
[0266] Our example approach is to focus on the golden metrics:
[0267] traffic: how many requests (rps / concurrency) are flowing through each proxy. We need to give particular attention to ingress gateways due to the huge amount of trafficthey need to handle.
[0268] errors: we may differentiate between the service failures and failures of the proxies, latency: again, we may differentiate between the latency added by the proxy and the latency added by the target service.
[0269] saturation: proxies in general are CPU bound. We need to give particular attention to cpu throttling that can have a huge impact on tail latency.
[0270] We aim to monitor real traffic first, and we may rely on synthetic traffic on services when no metrics are available.
[0271] Monitoring the control plane
[0272] The control plane may include four core components. The mesh can tolerate those components to be down for a short period of time. The four core components may be packaged into a single component.
[0273] Componentl: componentl provides service discovery for the gateway sidecars. It converts high level routing rules that control traffic behavior into gateway-specific configurations. When componentl goes down, sidecar proxies won't be able discover new services and new service units in the local zone data centre. New service units will fail to start.
[0274] Component2: component2 is the router’s configuration validation, ingestion, processing and distribution component. It is responsible for insulating the rest of the router components from the details of obtaining user configuration from the underlying platform.
[0275] When component2 goes down, sidecar proxies won't be able to discover new services and new service units in the local cell. New service units will fail to start.
[0276] Components: components enables strong service-to-service and end-user authentication with built-in identity and credential management.
[0277] When components goes down, certificates within the cell won't be rotated (certificates may be valid for 90 days by default).
[0278] Component4: component4 enforces access control and usage policies across the service mesh, and collects telemetry data from the gateway proxy and other services. When component2 goes down, telemetry will be impacted.Multi-cluster mesh including Multiple control planes
[0279] This is the preferred model because:
[0280] it's symmetric. Each server cluster runs it’s own control plane.
[0281] it's simpler: the only requirement for cross-cell communication is a shared root certificate authority (CA).
[0282] Example Updating in the system
[0283] The routers are updated in at most one Zone Data Centre at a time, to minimize the odds of disruption e.g. to users.
[0284] Significant updates to routers are done after draining the server clusters in the Zone Data Centre in which the routers are being updated.
[0285] Example Key design aspects of the system
[0286] Every server cluster provides an isolated failure zone.
[0287] The router control plane is deployed independently into each server cluster.
[0288] The amount of traffic each router is handling is monitored, to ensure the volume of traffic handled by different routers does not differ by more than a threshold e.g. to ensure the volume of traffic handled by different routers does not differ by more than a factor of ten, or e.g. to ensure the volume of traffic handled by different routers does not differ by more than a factor of three. In response to detecting that the volume of traffic handled by different routers does differ by more than a threshold, the routers may be reconfigured to ensure that the volume of traffic handled by different reconfigured routers does not differ by more than a threshold, e.g. does not differ by more than a factor of ten, or e.g does not differ by more than a factor of three.
[0289] Geographic Connectivity
[0290] In an example, there is no connectivity between different Zone Data Centres in different geographic regions of the world, e.g. between Western Europe and North America. Providing failover coverage for different Zone Data Centres in differentgeographic regions of the world, e.g. between Western Europe and North America would add too much latency. Instead, service resiliency is provided within each of a plurality of geographic regions of the world e.g. for Western Europe, North America, Japan.
[0291] Router features which may be used
[0292] traffic management: including one or more or all of: request routing, circuit breaking, mirroring, traffic shifting.
[0293] security: including one or more or all of: mutual TLS (mTLS), identity and authorisation.
[0294] policies: including one or more or all of: policy enforcement, rate limiting, white list, black list.
[0295] telemetry: including one or more or all of: metrics, logs, tracing.
[0296] Router and Routing Key Concepts
[0297] Gateway
[0298] A gateway may be used to manage inbound and outbound traffic in a mesh. One can manage multiple types of traffic with a gateway.
[0299] Gateway configurations apply to gateway proxies that are running at the edge of the mesh, which means that the gateway proxies are not running as service sidecars. To configure a gateway means configuring a gateway proxy to allow or block certain traffic from entering or leaving the mesh.
[0300] A mesh can have any number of gateway configurations, and multiple gateway workload implementations can co-exist within your mesh. You might use multiple gateways to have one gateway for private traffic and another for public traffic, so you can keep all private traffic inside a firewall, for example.
[0301] Service Entry
[0302] A service entry is used to add an entry to the system of routers’ abstract model, or service registry, that the system of routers maintain internally. After service entry isadded, the gateway proxies can send traffic to the service as if it was a service in your mesh. Configuring service entries allows you to manage traffic for services running outside of the mesh:
[0303] • Redirect and forward traffic for external destinations, such as APIs consumed from the web, or traffic to services in legacy infrastructure.
[0304] • Define retry, timeout, and fault injection policies for external destinations.
[0305] • Add a service running in a Virtual Machine (VM) to the mesh to expand the mesh.
[0306] • Logically add services from a different cluster to the mesh to configure a multicluster router mesh.
[0307] You don’t need to add a service entry for every external service that you want your mesh services to use. By default, the router system configures the gateway proxies to pass through requests to unknown services, although you can’t use system of routers’ features to control the traffic to destinations that are not registered in the mesh.
[0308] You can use service entries to perform the following configurations:
[0309] • Access secure external services over plain text ports, to configure the gateway to perform TLS Origination.
[0310] • Ensure, together with an egress gateway, that all external services are accessed through a single exit point.
[0311] Service Entry objects may be used to define server cluster failover. The list of endpoints may include the local server cluster, and the other server clusters where the service is provided.
[0312] Locality can be added to an endpoint to define where the service is.
[0313] Locality can then be used to Prefer local server cluster, and in combination with a Destination Rule, failover to remote server clusters.
[0314] Destination Rule
[0315] One can specify the path for traffic with routing rules, and then one may use destination rules to configure the set of policies that gateway proxies apply to a request at a specific destination. Destination rules are applied after the routing rules are evaluated.Configurations set in destination rules apply to traffic that is routed through a platform’s basic connectivity. You can use wildcard prefixes in a destination rule to specify a single rule for multiple services. You can use destination rules to specify service subsets, that is, to group all the instances of your service with a particular version together. You then may configure routing rules that route traffic to your subsets to send certain traffic to particular service versions.
[0316] One can specify explicit routing rules to service subsets. This model allows to:
[0317] • Refer to a specific service version across different virtual services.
[0318] • Simplify the stats that the router proxies emit.
[0319] • Encode subsets in Server Name Indication (SNI) headers.
[0320] Destination rules allow service owners to define different load balancing strategies for their services. Service owners are also required to configure the TLS mode, for example, to tell the gateway to upgrade the HTTP connection to mutual TLS.
[0321] Outlier detection is used to define when an endpoint should be ejected; this may allow cell failover to work transparently.
[0322] In an example, there is not a concept of health checking, only outlier detection.
[0323] Removing Network Load Balancers
[0324] In an example, a gateway is not associated with a related network load balancer. A network load balancer is a possible point of failure in a network. Therefore a network load balancer not being present removes a possible point of failure in a network. An advantage is reduced latency, because a network load balancer typically introduces latency. An advantage is reduced cost, because a network load balancer consumes computing resources.
[0325] In an example, each gateway is not associated with a related network load balancer or with a respective network load balancer. A network load balancer is a possible point of failure in a network. Therefore a network load balancer not being present for each gateway removes a plurality of possible points of failure in a network. An advantage is reduced latency, because a network load balancer typically introduces latency. Anadvantage is reduced cost, because a network load balancer consumes computing resources.
[0326] In an example, each gateway is associated with a respective network load balancer. This is an improvement over a system in which a network load balancer is associated with a plurality of gateways, because when a network load balancer which is associated with a plurality of gateways fails, this impacts the plurality of gateways, whereas when a network load balancer which is associated with a respective gateway of the network load balancer fails, this impacts only the respective gateway, which provides a more resilient system.
[0327] In an example, each server cluster in a Zone Data Centre is associated with a respective network load balancer. This is an improvement over a system in which a network load balancer is associated with a plurality of server clusters in a Zone Data Centre, because when a network load balancer which is associated with a plurality of server clusters in a Zone Data Centre fails, this impacts the plurality of server clusters, whereas when a network load balancer which is associated with a respective server cluster in a Zone Data Centre of the network load balancer fails, this impacts only the respective server cluster in a Zone Data Centre, which provides a more resilient system.
[0328] Default Gateway Design
[0329] Default Gateway relates to routing architecture and is used to route traffic between server clusters, which enables some core platform features such as:
[0330] • Draining of clusters
[0331] • Routing requests to a service that exists in a different server cluster from its origin • Outlier detecting a whole server cluster
[0332] Service entries to a service typically offer at least three routes to a client service: a heavily weighted route to stay in the same Zone Data Centre, and at least two lower-weighted routes to other Zone Data Centres in the same geographic region, where example geographic regions are North America, Western Europe, Japan. In anexample, same-cluster routes go directly to Default Gateway service units via networking rather than needlessly leave the server cluster, and then re-enter the server cluster.
[0333] Networking
[0334] In an example, service units in server clusters may be given internal IP addresses via a virtual private cloud (VPC) - these are real elastic network interface IP addresses that can be called directly from other resources within the same VPC. In an example, each server cluster uses a different subnet, but service units in a first server cluster can technically connect directly to service units in a second, different server cluster via their IP address, but do not do so because of security issues and mTLS issues.
[0335] In an example, a network service has a virtual IP address. Connections to this service (either via the DNS name or directly to the cluster IP) then route using a network proxy to one of the available endpoints. An endpoint in this context is a service unit, so if Default Gateway has 10 service units, there are 10 endpoints available, these endpoints being the VPC network interface IP addresses allocated as described above. When a service unit is created and becomes healthy, or is deleted, a network API updates the list of endpoints associated with this service dynamically, and this list of endpoints can then be read by other services (e.g. by a router management service).
[0336] Router Management Service
[0337] While service units across server clusters can directly communicate with each other via their VPC IPs, in practice this is difficult. Service units often rotate so service units IP addresses are variable. What we need is a way to dynamically discover the endpoints (service units) in remote server clusters, and make these available via a fixed route within the given server cluster, so we can make a connection and it routes to service units in another server cluster.
[0338] The Router Management Service solves this for us. In an example, it uses a Discovery tool to discover endpoints within server clusters, to do the following:
[0339] 1. Read a secret which contains credentials and server URL of another server cluster2. Adds this server cluster to a list of remote server clusters to communicate with 3. Communicates with a network API server in that remote server cluster and collects the Services and their endpoint IPs
[0340] 4. Adds these Services, with their endpoints, to router configurations in Sidecars and Gateways in the local server cluster.
[0341] In an example, the Router Management Service is configured such that we can select all services by default to be server cluster local, so their endpoints are not shared with other server clusters and we don’t combine services, and only share server clusterspecific Default Gateway services to expose those service unit IPs to other server clusters.
[0342] In an example, for traffic originating in a first server cluster (e.g. in Zone Data Centre a) destined for a second server cluster (e.g. in Zone Data Centre b), instead of routing it to a network Load Balancer, we instead route the request to an egress gateway within Zone Data Centre a, which is associated with that destination server cluster in Zone Data Centre b. This is a one-to-one mapping. The egress gateway’s purpose includes to replicate that of a network Load Balancer: to receive a request though a single IP route and to load balance it to Default Gateway pods in the remote destination cluster.
[0343] In an example, a request from Service A is routed by a first router proxy on to a default gateway egress service unit. This is via a default gateway egress service cluster IP address. Here, the service name of the egress gateway not important, we just need server cluster’s IP address. The request is sent from the default gateway egress service unit to a default gateway ingress service unit. Here the service name is important, and is unique per server cluster. There are no retries. Failures are propagated back to the egress gateway to trigger client-side outlier detection for a whole server cluster. The default gateway ingress service unit sends the request to a second router proxy using a virtual service. The second router proxy sends the request to service B. An example is shown in Figure 9.In an example, Zone a Data Centre is configured such that when Service A operating in Zone a Data Centre makes a request to Service B, the request may be routed via a default gateway ingress service in Zone a Data Centre via a default gateway service unit inside Zone a Data Centre to a Service B unit inside Zone a Data Centre. Or the request may be routed via a default gateway egress service in Zone a Data Centre, via a default gateway ingress service in Zone b Data Centre, then via a default gateway service unit inside Zone b Data Centre to a Service B unit inside Zone b Data Centre. Or the request may be routed via a default gateway egress service in Zone a Data Centre, via a default gateway ingress service in Zone c Data Centre, then via a default gateway service unit inside Zone c Data Centre to a Service B unit inside Zone c Data Centre. In the same example, Zone b Data Centre is configured such that when Service A operating in Zone b Data Centre makes a request to Service B, the request may be routed via a default gateway ingress service in Zone b Data Centre via a default gateway service unit in Zone b Data Centre to a Service B unit inside Zone b Data Centre. Or the request may be routed via a default gateway egress service in Zone b Data Centre, via a default gateway ingress service in Zone a Data Centre, then via a default gateway service unit in Zone a Data Centre to a Service B unit inside Zone a Data Centre. Or the request may be routed via a default gateway egress service in Zone b Data Centre, via a default gateway ingress service in Zone c Data Centre, then via a default gateway service unit inside Zone c Data Centre to a Service B unit inside Zone c Data Centre. In the same example, Zone c Data Centre is configured such that when Service A operating in Zone c Data Centre makes a request to Service B, the request may be routed via a default gateway ingress service in Zone c Data Centre via a default gateway service unit in Zone c Data Centre to a Service B unit inside Zone c Data Centre. Or the request may be routed via a default gateway egress service in Zone c Data Centre, via a default gateway ingress service in Zone a Data Centre, then via a default gateway service unit in Zone a Data Centre to a Service B unit inside Zone a Data Centre. Or the request may be routed via a default gateway egress service in Zone c Data Centre, via a default gateway ingress service in Zone b Data Centre, then via a default gateway service unit inside Zone b Data Centre to a Service B unit inside Zone b Data Centre. An example is shown in Figure 10.
[0344] In an example, it is necessary to continue sending requests to a single IP address fromthe client service to keep full-server cluster outlier detection working as normal, so we cannot send requests straight to the remote service name, as the router management service would add all Default Gateway service units in Zone Data Centre b as endpoints, and will no longer outlier detect entire server clusters, but instead just outlier detect single service units. This is a reason behind sending requests to an egress gateway tied to another server cluster via a single Cluster IP, as it retains the whole-server cluster outlier detection features.
[0345] Example Gateways
[0346] As an example, consider a system with six server clusters, and we’re trying to route traffic to server cluster A from any other server cluster (B, C, D, E or F).
[0347] This design includes, instead of a single network Load Balancer used by Server Clusters B, C, D, E, and F, pointing at Default Gateway service units in Cluster A, one-to-one egress-type gateways in each of the clusters B, C, D, E and F, which forward to Default Gateway service units in Cluster A. This means that in every server cluster, there is one Default Gateway, and five egress-type gateways (one for each other server cluster’s Default Gateway). The Default Gateway and Egress Gateway services have fixed and well-known Cluster IP addresses which are constant in every server cluster.
[0348] The egress gateways do not terminate TLS (so full mTLS passthrough) and have simple routing - any request they receive gets forwarded to a Default Gateway service unit of the server cluster that egress gateway is assigned to. There is a Destination Rule for configuring traffic from the Default Egress Gateway in Server Cluster A to the Default Gateway pods in Server Cluster B.
[0349] Service Discovery time-to-live (TTL) - Router Management Service vs. Network Load Balancer Controller
[0350] The Router Management Service may be used to dynamically discover service unit IP addresses within the same server cluster, and update the list of available endpoints (each service unit) in Default Gateway, using the network’s Application ProgrammingInterface (API) on endpoint resources. This is almost instantaneous and hasn’t presented us with any known issues. This concept may also be applied to multiple server clusters - Cluster A will use the network’s watch API request on remote Cluster B endpoint resources to discover that server cluster’s Default Gateway service unit IP addresses, and then update the local server cluster gateway configurations so requests know where to go when attempting to route to another server cluster.
[0351] A network Load Balancer (NLB) controller may employ a similar method: it may watch endpoint resources and keep target groups updated to match the list of IP addresses. The difference however is the delay in registration and de-regi strati on of targets - testing in our development environment showed remote Default Gateway service unit IP addresses being added and removed almost instantly, e.g. in less than 5 seconds, e.g. in less than 1.0 seconds, whereas it can take over a minute for a new service unit IP address to register and actually become available on an NLB even after the service unit becomes ready, and a few seconds to deregister the service unit when it becomes unhealthy.
[0352] Basic connectivity testing has also showed very slightly lower latency using the multiserver cluster solution compared to NLBs.
[0353] Access to other Server Cluster API servers from the Router Management Service
[0354] We need to let service units from server cluster A access API control planes of server clusters B, C..., and vice-versa. We can create this access in line with security requirements.
[0355] The Router Management Service needs to be able to communicate with the API control plane of other server clusters so it can determine the endpoints (service unit IP addresses) of the desired services to be shared. This essentially means the node that the Router Management Service runs on needs to be able to access other server clusters in line with security requirements.
[0356] NoteIt is to be understood that the above-referenced arrangements are only illustrative of the application for the principles of the present invention. Numerous modifications and alternative arrangements can be devised without departing from the spirit and scope of the present invention. While the present invention has been shown in the drawings and fully described above with particularity and detail in connection with what is presently deemed to be the most practical and preferred example(s) of the invention, it will be apparent to those of ordinary skill in the art that numerous modifications can be made without departing from the principles and concepts of the invention as set forth herein.
Claims
CLAIMS1. A computer system including a first data centre and a second data centre, the first data centre configured to communicate with the second data centre via a network, wherein the first data centre and the second data centre are separated by between 2 km and 500 km, the first data centre including a first service unit configured to provide a first service, and a plurality of second service units each configured to provide a second service, the first service unit including a first router proxy, and the first service unit including software executable to receive first requests and to send second requests relating to the first requests to the first router proxy;the plurality of second service units each including a respective second router proxy and respective software executable to receive at least some of the second requests via the respective second router proxy, to process the at least some of the second requests to provide respective responses, and to respond to the at least some of the second requests using the respective responses via the respective second router proxy; the first data centre including a first default gateway ingress service, and a first default gateway egress service configured to communicate with a first default gateway ingress service of the second data centre via the network;the second data centre including the first default gateway ingress service of the second data centre, a first default gateway egress service, and a plurality of third service units each configured to provide the second service, the plurality of third service units each including a respective third router proxy and respective software executable to receive at least some of the second requests via the respective third router proxy, to process the at least some of the second requests to provide respective responses, and to respond to the at least some of the second requests using the respective responses via the respective third router proxy;wherein the first router proxy is configured to route a first set of at least some of the second requests to the first default gateway ingress service of the first data centre, and the first default gateway ingress service of the first data centre is configured to send the first set of at least some of the second requests to the first default gateway ingress service of the first data centre to a second service unit of the plurality of second service units, and to receive the respective responses from the second service unit, and to send the respective responses to the first service unit of the first data centre;wherein the first router proxy is configured to route a second set of at least some of the second requests to the first default gateway egress service of the first data centre, and the first default gateway egress service of the first data centre is configured to send the second set of at least some of the second requests to the first default gateway ingress service of the second data centre, wherein the first default gateway ingress service of the second data centre is configured to send the second set of at least some of the second requests to a service unit of the plurality of third service units of the second data centre, wherein the service unit of the plurality of third service units of the second data centre is configured to respond to the second set of at least some of the second requests using the respective responses via the respective third router proxy, wherein the respective third router proxy is configured to route the respective responses to the first default gateway egress service of the second data centre, and the first default gateway egress service of the second data centre is configured to send the respective responses to the first default gateway ingress service of the first data centre, wherein the first default gateway ingress service of the first data centre is configured to send the respective responses to the first service unit of the first data centre.
2. The computer system of Claim 1, wherein the first router proxy is configured to route the second requests, the first default gateway ingress service of the first data centre is configured to route the first set of at least some of the second requests routed via the first router proxy, and the first default gateway ingress service of the second data centre is configured to route the second set of at least some of the second requests routed via the first router proxy.
3. The computer system of Claims 1 or 2, wherein the first service unit is configured to use the software executable to receive the first requests to use the responses received to the second requests to respond to the first requests.
4. The computer system of any previous Claim, wherein the first service unit includes a structured database, wherein the first service unit software is executable to receive first requests which relate to contents of the structured database, and to send the second requests relating to the first requests which relate to contents of the structured database to the first router proxy, wherein the second requests include dataextracted from the structured database of the first service unit.
5. The computer system of any previous Claim, wherein the second data centre further includes a first service unit configured to provide the first service, the first service unit of the second data centre including a fourth router proxy and software executable to receive first requests and to send second requests relating to the first requests to the fourth router proxy;the second data centre including the first default gateway egress service configured to communicate with the first default gateway ingress service of the first data centre via the network;wherein the fourth router proxy is configured to route a third set of at least some of the second requests to the first default gateway ingress service of the second data centre, and the first default gateway ingress service of the second data centre is configured to send the third set of at least some of the second requests to the first default gateway ingress service of the second data centre to a third service unit of the plurality of third service units, and to receive the respective responses from the third service unit, and to send the respective responses to the first service unit of the second data centre;wherein the fourth router proxy is configured to route a fourth set of at least some of the second requests to the first default gateway egress service of the second data centre, and the first default gateway egress service of the second data centre is configured to send the fourth set of at least some of the second requests to the first default gateway ingress service of the first data centre, wherein the first default gateway ingress service of the first data centre is configured to send the fourth set of at least some of the second requests to a service unit of the plurality of second service units of the first data centre,wherein the service unit of the plurality of second service units of the first data centre is configured to respond to the fourth set of at least some of the second requests using the respective responses via the respective second router proxy, wherein the respective second router proxy is configured to route the respective responses to the first default gateway egress service of the first data centre, and the first default gateway egress service of the first data centre is configured to send the respective responses to the first default gateway ingress service of the second data centre,wherein the first default gateway ingress service of the second data centre is configured to send the respective responses to the first service unit of the second data centre.
6. The computer system of Claim 5, wherein the fourth router proxy is configured to route the second requests, the first default gateway ingress service of the second data centre is configured to route the third set of at least some of the second requests routed via the fourth router proxy, and the first default gateway ingress service of the first data centre is configured to route the fourth set of at least some of the second requests routed via the fourth router proxy.
7. The computer system of Claims 5 or 6, wherein the first service unit of the second data centre is configured to use the software executable to receive the first requests to use the responses received to the second requests to respond to the first requests.
8. The computer system of any previous Claim, wherein the system does not include a network load balancer, between the first data centre and the second data centre, configured to load balance the second requests.
9. The computer system of any previous Claim, wherein the first data centre is located in a first zone, and wherein the second data centre is located in a second zone, and wherein the first zone and the second zone are separated by between 2 km and 500 km.
10. The computer system of any previous Claim, the computer system including a third data centre, the third data centre configured to communicate with the first data centre and with the second data centre via the network, wherein the third data centre and the second data centre are separated by between 2 km and 500 km, wherein the first data centre and the third data centre are separated by between 2 km and 500 km, the third data centre including a first service unit configured to provide a first service, and a plurality of fourth service units each configured to provide the second service, the first service unit including a fifth router proxy, and the first service unit includingsoftware executable to receive first requests and to send second requests relating to the first requests to the fifth router proxy;the plurality of fourth service units each including a respective sixth router proxy and respective software executable to receive at least some of the second requests via the respective sixth router proxy, to process the at least some of the second requests to provide respective responses, and to respond to the at least some of the second requests using the respective responses via the respective sixth router proxy;the third data centre including a first default gateway ingress service, and a first default gateway egress service configured to communicate with a first default gateway ingress service of the first data centre via the network; the third data centre including a second default gateway egress service configured to communicate with a first default gateway ingress service of the second data centre via the network; the first data centre including a second default gateway egress service, and the second data centre including a second default gateway egress service;wherein the fifth router proxy is configured to route a first set of at least some of the second requests to the first default gateway ingress service of the third data centre, and the first default gateway ingress service of the third data centre is configured to send the first set of at least some of the second requests to the first default gateway ingress service of the third data centre to a fourth service unit of the plurality of fourth service units, and to receive the respective responses from the fourth service unit, and to send the respective responses to the first service unit of the third data centre; wherein the fifth router proxy is configured to route a second set of at least some of the second requests to the first default gateway egress service of the third data centre, and the first default gateway egress service of the third data centre is configured to send the second set of at least some of the second requests to the first default gateway ingress service of the first data centre, wherein the first default gateway ingress service of the first data centre is configured to send the second set of at least some of the second requests to a service unit of the plurality of second service units of the first data centre, wherein the service unit of the plurality of second service units of the first data centre is configured to respond to the second set of at least some of the second requests using the respective responses via the respective second router proxy, wherein the respective second router proxy is configured to route the respective responses to the second default gateway egress service of the first data centre, and thesecond default gateway egress service of the first data centre is configured to send the respective responses to the first default gateway ingress service of the third data centre, wherein the first default gateway ingress service of the third data centre is configured to send the respective responses to the first service unit of the third data centre;wherein the fifth router proxy is configured to route a third set of at least some of the second requests to the second default gateway egress service of the third data centre, and the second default gateway egress service of the third data centre is configured to send the third set of at least some of the second requests to the first default gateway ingress service of the second data centre, wherein the first default gateway ingress service of the second data centre is configured to send the third set of at least some of the second requests to a service unit of the plurality of third service units of the second data centre, wherein the service unit of the plurality of third service units of the second data centre is configured to respond to the third set of at least some of the second requests using the respective responses via the respective third router proxy, wherein the respective third router proxy is configured to route the respective responses to the second default gateway egress service of the second data centre, and the second default gateway egress service of the second data centre is configured to send the respective responses to the first default gateway ingress service of the third data centre, wherein the first default gateway ingress service of the third data centre is configured to send the respective responses to the first service unit of the third data centre.
11. The computer system of Claim 10, wherein the system does not include a network load balancer, between the first data centre and the second data centre and the third data centre, configured to load balance the second requests.
12. The computer system of any previous Claim, wherein spacing the data centres apart reduces the risk of the data centres all going down if a disaster (e.g. explosion, fire, power outage) occurs locally in a geographic region.
13. The computer system of any previous Claim, wherein the system includes a control plane and a data plane, wherein the data plane includes all the proxiesdeployed as sidecars or as gateways, or as other configurable proxies, and wherein the control plane manages and configures the proxies to route traffic.
14. The computer system of Claim 13, wherein the data plane is monitored by measuring at least traffic, errors, latency and saturation.
15. The computer system of Claims 13 or 14, wherein the control plane includes four components: Componentl which provides service discovery for the gateway sidecars; component2 which is the router’s configuration validation, ingestion, processing and distribution component; components which enables strong service-to-service and end-user authentication with built-in identity and credential management; component4 which enforces access control and usage policies across the service mesh, and collects telemetry data from the gateway proxy and other services.
16. The computer system of any of Claims 13 to 15, wherein each data centre includes at least one server cluster, wherein each server cluster runs its own control plane, and / or wherein the router control plane is deployed independently into each server cluster.
17. The computer system of any previous Claim, wherein service units each run a router-compatible sidecar proxy.
18. The computer system of Claim 17, wherein sidecar proxy injection into a new service unit happens automatically every time a new service unit is created.
19. The computer system of Claims 17 or 18, wherein traffic is directed from application services to and from the sidecars.
20. The computer system of any of Claims 17 to 19, wherein when router software is updated, sidecars in existing service units remain the same until those service units are deleted.
21. The computer system of Claim 20, wherein when those service units aredeleted, replacement service units are created.
22. The computer system of any of Claims 17 to 21, wherein http 503 error response codes lead to one or more retries by a sidecar proxy.
23. The computer system of any of Claims 17 to 22, wherein when a service unit or any intermediate proxies returns a http 503 error, the client gateway sidecar proxy automatically retries for another service unit.
24. The computer system of Claims 22 or 23, wherein a fully jittered exponential back-off algorithm is used for retries.
25. The computer system of any previous Claim, in which router circuit breakers are configured to automatically failover traffic to the closest working server cluster or service unit.
26. The computer system of Claim 25, in which there are configured two sets of circuit breaking rules:one rule for communication within a server cluster, which applies to requests coming from the ingress gateway going to the destination service units;one rule for communication across server clusters.
27. The computer system of any previous Claim, in which the system includes a mesh which provides service unit-to-service unit communication.
28. The computer system of Claim 27, in which all service unit-to-service unit communication within the mesh is end-to-end encrypted, and the encryption is handled by the router proxy e.g. router sidecar proxy.
29. The computer system of Claims 27 or 28, in which each service includes its own security certificate, which is included into the router proxy e.g. router sidecar proxy.
30. The computer system of Claim 29, in which certificates are automatically rotated and dynamically reloaded by the router proxy e.g. router sidecar proxy, without any disruption to the service unit.
31. The computer system of any of Claims 27 to 30, in which a service unit (e.g. its service owner) must declare the list of services that the router proxy, e.g. router sidecar proxy, can see or call.
32. The computer system of Claim 31, in which these services are the only services within the mesh.
33. The computer system of any of Claims 27 to 32, wherein all sidecar proxies in the mesh are programmed with the necessary configuration required to reach every workload instance in the mesh.
34. The computer system of any of Claims 27 to 33, wherein the amount of traffic each router is handling is monitored, to ensure the volume of traffic handled by different routers does not differ by more than a threshold e.g. to ensure the volume of traffic handled by different routers does not differ by more than a factor of ten, or e.g. to ensure the volume of traffic handled by different routers does not differ by more than a factor of three.
35. The computer system of Claim 34, in which in response to detecting that the volume of traffic handled by different routers does differ by more than a threshold, the router proxies are reconfigured to ensure that the volume of traffic handled by different reconfigured router proxies does not differ by more than a threshold, e.g. does not differ by more than a factor of ten, or e.g does not differ by more than a factor of three.
36. The computer system of any previous Claim, wherein a default gateway in the first or second data centre connects to between one hundred and ten thousand service units in the same data centre.
37. The computer system of any previous Claim, wherein a default gateway in the first or second data centre connects to between one thousand and five thousand service units in the same data centre.
38. The computer system of any previous Claim, wherein services are not deployed to a single server cluster.
39. The computer system of any previous Claim, wherein each server cluster has its own service mesh control plane.
40. The computer system of any previous Claim, wherein each service unit has its own sidecar which all inbound and outbound traffic goes through.
41. The computer system of any previous Claim, wherein for traffic within each subnet, it is preferred to use the local subnet; the next preference is other subnets of the data centre of the local subnet, and the next preference is subnets of other data centres.
42. The computer system of any previous Claim, wherein server clusters’ traffic is drained and shifted to other working server clusters in other Data Centres.
43. The computer system of Claim 42, wherein only server clusters within a single Data Centre are drained at one time.
44. The computer system of any previous Claim, wherein at least three server clusters are present in each data centre, e.g. to ensure at least one server cluster is working.
45. The computer system of any previous Claim, wherein each route to a service goes via a default gateway; each default gateway is configured to route to any service units providing a particular service, whereas the mesh is configured to route to, or to receive traffic from, any of the default gateways.
46. The computer system of any previous Claim, wherein the default gateways are configured in the routers as a Service Entry in every server cluster, not just in the server clusters the service is deployed in.
47. The computer system of any previous Claim, wherein when updating routers, routers are updated in at most one Data Centre at a time.
48. The computer system of any previous Claim, wherein updates to routers are done after draining the server clusters in the Data Centre in which the routers are being updated.
49. The computer system of any previous Claim, wherein every server cluster provides an isolated failure zone.
50. The computer system of any previous Claim, wherein there is no connectivity between different Data Centres in different geographic regions of the world.
51. The computer system of any previous Claim, wherein the path for traffic is specified with routing rules, and destination rules are used to configure a set of policies that gateway proxies apply to a request at a specific destination.
52. The computer system of any previous Claim, wherein a Router Management Service in a local server cluster uses a Discovery tool to discover endpoints within other server clusters, to do the following:(i) Read a secret which contains credentials and server URL of another server cluster; (ii) Adds this other server cluster to a list of remote server clusters to communicate with;(iii) Communicates with a network API server in that remote server cluster and collects the Services and their endpoint IPs;(iv) Adds these Services, with their endpoints, to router configurations in Sidecars and Gateways in the local server cluster.
53. The computer system of any previous Claim, wherein requests are sent to a single IP address from the first service unit of the first data centre to keep full-server cluster outlier detection working as normal, so requests are not sent straight to the remote service name, as the router management service would add all Default Gateway service units in the second Data Centre as endpoints, and will no longer outlier detect entire server clusters, but instead would just outlier detect single service units.
54. A computer implemented method of routing second requests, and of routing responses to the second requests, in a computer system including a first data centre and a second data centre, the first data centre communicating with the second data centre via a network, wherein the first data centre and the second data centre are separated by between 2 km and 500 km, the first data centre including a first service unit providing a first service, and a plurality of second service units each providing a second service, the first service unit including a first router proxy, and the first service unit including software executing to receive first requests and to send second requests relating to the first requests to the first router proxy;the plurality of second service units each including a respective second router proxy and respective software executing to receive at least some of the second requests via the respective second router proxy, to process the at least some of the second requests to provide respective responses, and to respond to the at least some of the second requests using the respective responses via the respective second router proxy; the first data centre including a first default gateway ingress service, and a first default gateway egress service communicating via the network with a first default gateway ingress service of the second data centre;the second data centre including the first default gateway ingress service of the second data centre, a first default gateway egress service, and a plurality of third service units each providing the second service, the plurality of third service units each including a respective third router proxy and respective software executing to receive at least some of the second requests via the respective third router proxy, to process the at least some of the second requests to provide respective responses, and to respond to the at least some of the second requests using the respective responses via therespective third router proxy;wherein the first router proxy routes a first set of at least some of the second requests to the first default gateway ingress service of the first data centre, and the first default gateway ingress service of the first data centre sends the first set of at least some of the second requests to the first default gateway ingress service of the first data centre to a second service unit of the plurality of second service units, and receives the respective responses from the second service unit, and sends the respective responses to the first service unit of the first data centre;wherein the first router proxy routes a second set of at least some of the second requests to the first default gateway egress service of the first data centre, and the first default gateway egress service of the first data centre sends the second set of at least some of the second requests to the first default gateway ingress service of the second data centre, wherein the first default gateway ingress service of the second data centre sends the second set of at least some of the second requests to a service unit of the plurality of third service units of the second data centre, wherein the service unit of the plurality of third service units of the second data centre responds to the second set of at least some of the second requests using the respective responses via the respective third router proxy, wherein the respective third router proxy routes the respective responses to the first default gateway egress service of the second data centre, and the first default gateway egress service of the second data centre sends the respective responses to the first default gateway ingress service of the first data centre, wherein the first default gateway ingress service of the first data centre sends the respective responses to the first service unit of the first data centre.
55. The method of Claim 54, including use of a system of any of Claims 1 to 53.
56. A computer system including a first data centre, a second data centre and a third data centre, the first data centre configured to communicate via a network with the second data centre and with the third data centre, wherein the first data centre and the second data centre are separated by between 2 km and 500 km, wherein the third data centre and the second data centre are separated by between 2 km and 500 km, and wherein the first data centre and the third data centre are separated by between 2 km and 500 km, wherein each data centre includes a respective default gateway;the first data centre including a first respective subnet which includes a first respective server cluster, wherein the first respective server cluster is configured to provide a first service, in which the first respective server cluster is configured to receive a query in respect of the first service, and in response to query a second service, and to receive a response from the second service, and to use the response received from the second service to provide a response from the first service,the first data centre including a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the second service, in which the second respective server cluster is configured to receive a query in respect of the second service and to provide a response;the second data centre including a respective subnet which includes a respective server cluster configured to provide the second service, in which the respective server cluster is configured to receive a query in respect of the second service and to provide a response;the third data centre including a respective subnet which includes a respective server cluster configured to provide the second service, in which the respective server cluster is configured to receive a query in respect of the second service and to provide a response;wherein in the first data centre the first respective server cluster is configured to provide the first service, the first respective server cluster including a first service client unit, the first service client unit including a first service application executable in the first service client unit, and wherein the first service client unit includes a first router;wherein in the first data centre the second respective server cluster includes a second service unit, the second service unit including a second service application executable in the second service unit, and the second service unit includes a second router configured to receive a request from the first service via the default gateway of the first data centre, to route the request to the second service application for processing by the second service application;wherein in a first configuration of the first router, the first router is configured to route greater than 90% (e.g. greater than 97%, e.g. 99%) of the requests from the firstservice to the second service via the default gateway of the first data centre to the second respective subnet in the first data centre, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the second service;wherein in the first configuration of the first router, the first router is configured to route between 0.1% and 5% (e.g. 1%) of the requests from the first service to the second service via the default gateway of the second data centre to the respective subnet in the second data centre, the respective subnet including the respective server cluster configured to provide the second service;wherein the respective server cluster of the second data centre includes a second service unit, the second service unit including a second service application executable in the second service unit, and the second service unit includes a third router configured to receive a request from the first service via the default gateway of the second data centre, to route the request to the second service application for processing by the second service application;wherein in the first configuration of the first router, the first router is configured to route between 0.1% and 5% (e.g. 1%) of the requests from the first service to the second service via the default gateway of the third data centre to the respective subnet in the third data centre, the respective subnet including the respective server cluster configured to provide the second service;wherein the respective server cluster of the third data centre includes a second service unit, the second service unit including a second service application executable in the second service unit, and the second service unit includes a fourth router configured to receive a request from the first service via the default gateway of the third data centre, to route the request to the second service application for processing by the second service application;wherein the computer system does not include a network load balancer between the first data centre and the second data centre, and wherein the computer system does not include a network load balancer between the first data centre and the third data centre, and wherein the first data centre does not include a network load balancer between first respective subnet and the second respective subnet, wherein the network load balancers are configured to load balance the second requests.
57. The computer system of Claim 56, wherein the percentages of traffic to the default gateways add up to one hundred percent.
58. The computer system of Claim 56, wherein in a second configuration of the first router, the first router is configured to route zero percent of the requests from the first service to the second service via the default gateway of the first data centre to the second respective subnet in the first data centre, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the second service, and wherein in the second configuration of the first router, the first router is configured to route a non-zero fraction of the requests from the first service to the second service via the default gateway of the second data centre to the respective subnet in the second data centre, the respective subnet including the respective server cluster, wherein the respective server cluster is configured to provide the second service; and wherein in the second configuration of the first router, the first router is configured to route a non-zero fraction of the requests from the first service to the second service via the default gateway of the third data centre to the respective subnet in the third data centre, the respective subnet including the respective server cluster, wherein the respective server cluster is configured to provide the second service.
59. The computer system of Claim 58, wherein the percentages of traffic to the default gateways add up to one hundred percent.
60. The computer system of Claims 58 or 59, wherein the first router is configured to change from the first configuration of the first router to the second configuration of the first router, in response to the first router detecting that the second respective server cluster of the first data centre including the respective second service unit becomes faulty.
61. The computer system of Claims 56, 58, 60 or 61, wherein in a third configuration of the first router, the first router is configured to route greater than 90%(e.g. greater than 97%, e.g. 99%) of the requests from the first service to the second service via the default gateway of the first data centre to the second respective subnet in the first data centre, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster is configured to provide the second service, and wherein in the third configuration of the first router, the first router is configured to route zero percent of the requests from the first service to the second service via the default gateway of the second data centre to the respective subnet in the second data centre, the respective subnet including the respective server cluster, wherein the respective server cluster is configured to provide the second service, and wherein in the third configuration of the first router, the first router is configured to route a fraction between 0.1% and 5% (e.g. 1%) of the requests from the first service to the second service via the default gateway of the third data centre to the respective subnet in the third data centre, the respective subnet including the respective server cluster, wherein the respective server cluster is configured to provide the second service.
62. The computer system of Claim 61, wherein the percentages of traffic to the default gateways add up to one hundred percent.
63. The computer system of Claims 61 or 62, wherein the first router is configured to change from the first configuration of the first router to the third configuration of the first router, in response to the first router detecting that the respective server cluster of the second data centre including the respective second service unit has become faulty.
64. The computer system of any of Claims 56 to 63, wherein all traffic between the server clusters is routed by the routers via the default gateways.
65. The computer system of any of Claims 56 to 64, wherein spacing the data centres apart reduces the risk of the data centres all going down if a disaster (e.g. explosion, fire, power outage) occurs locally in a geographic region.
66. The computer system of any of Claims 56 to 65, wherein the system includes a control plane and a data plane, wherein the data plane includes all the proxies deployed as sidecars or as gateways, or as other configurable proxies, and wherein the control plane manages and configures the proxies to route traffic.
67. The computer system of Claim 66, wherein the data plane is monitored by measuring at least traffic, errors, latency and saturation.
68. The computer system of Claims 66 or 67, wherein the control plane includes four components: Componentl which provides service discovery for the gateway sidecars; component2 which is the router’s configuration validation, ingestion, processing and distribution component; components which enables strong service-to-service and end-user authentication with built-in identity and credential management; component4 which enforces access control and usage policies across the service mesh, and collects telemetry data from the gateway proxy and other services.
69. The computer system of any of Claims 66 to 68, wherein each data centre includes at least one server cluster, wherein each server cluster runs its own control plane, and / or wherein the router control plane is deployed independently into each server cluster.
70. The computer system of any of Claims 56 to 69, wherein service units each run a router-compatible sidecar proxy.
71. The computer system of Claim 70, wherein sidecar proxy injection into a new service unit happens automatically every time a new service unit is created.
72. The computer system of Claims 70 or 71, wherein traffic is directed from application services to and from the sidecars.
73. The computer system of any of Claims 70 to 72, wherein when router software is updated, sidecars in existing service units remain the same until those service units are deleted.
74. The computer system of Claim 73, wherein when those service units are deleted, replacement service units are created.
75. The computer system of any of Claims 70 to 74, wherein http 503 error response codes lead to one or more retries by a sidecar proxy.
76. The computer system of any of Claims 70 to 75, wherein when a service unit or any intermediate proxies returns a http 503 error, the client gateway sidecar proxy automatically retries for another service unit.
77. The computer system of Claims 75 or 76, wherein a fully jittered exponential back-off algorithm is used for retries.
78. The computer system of any of Claims 56 to 77, in which router circuit breakers are configured to automatically failover traffic to the closest working server cluster or service unit.
79. The computer system of Claim 78, in which there are configured two sets of circuit breaking rules:one rule for communication within a server cluster, which applies to requests coming from the ingress gateway going to the destination service units;one rule for communication across server clusters.
80. The computer system of any of Claims 56 to 79, in which the system includes a mesh which provides service unit-to-service unit communication.
81. The computer system of Claim 80, in which all service unit-to-service unit communication within the mesh is end-to-end encrypted, and the encryption is handled by the router proxy e.g. router sidecar proxy.
82. The computer system of Claims 80 or 81, in which each service includes its own security certificate, which is included into the router proxy e.g. router sidecar proxy.
83. The computer system of Claim 82, in which certificates are automatically rotated and dynamically reloaded by the router proxy e.g. router sidecar proxy, without any disruption to the service unit.
84. The computer system of any of Claims 80 to 83, in which a service unit (e.g. its service owner) must declare the list of services that the router proxy, e.g. router sidecar proxy, can see or call.
85. The computer system of Claim 84, in which these services are the only services within the mesh.
86. The computer system of any of Claims 80 to 85, wherein all sidecar proxies in the mesh are programmed with the necessary configuration required to reach every workload instance in the mesh.
87. The computer system of any of Claims 80 to 86, wherein a default gateway in the first or second data centre connects to between one hundred and ten thousand service units in the same data centre.
88. The computer system of any of Claims 80 to 86, wherein a default gateway in the first or second data centre connects to between one thousand and five thousand service units in the same data centre.
89. The computer system of any of Claims 56 to 88, wherein services are not deployed to a single server cluster.
90. The computer system of any of Claims 56 to 89, wherein each server cluster has its own service mesh control plane.
91. The computer system of any of Claims 56 to 90, wherein each service unit has its own sidecar which all inbound and outbound traffic goes through.
92. The computer system of any of Claims 56 to 91, wherein for traffic within each subnet, it is preferred to use the local subnet; the next preference is other subnets of the data centre of the local subnet, and the next preference is subnets of other data centres.
93. The computer system of any of Claims 56 to 92, wherein server clusters’ traffic is drained and shifted to other working server clusters in other Data Centres.
94. The computer system of Claim 93, wherein only server clusters within a single Data Centre are drained at one time.
95. The computer system of any of Claims 56 to 94, wherein at least three server clusters are present in each data centre, e.g. to ensure at least one server cluster is working.
96. The computer system of any of Claims 56 to 95, wherein each route to a service goes via a default gateway; each default gateway is configured to route to any service units providing a particular service, whereas the mesh is configured to route to, or to receive traffic from, any of the default gateways.
97. The computer system of any of Claims 56 to 96, wherein the default gateways are configured in the routers as a Service Entry in every server cluster, not just in the server clusters the service is deployed in.
98. The computer system of any of Claims 56 to 97, wherein when updating routers, routers are updated in at most one Data Centre at a time.
99. The computer system of any of Claims 56 to 98, wherein updates to routers are done after draining the server clusters in the Data Centre in which the routers are being updated.
100. The computer system of any of Claims 56 to 99, wherein every server cluster provides an isolated failure zone.
101. The computer system of any of Claims 56 to 100, wherein there is no connectivity between different Data Centres in different geographic regions of the world.
102. The computer system of any of Claims 56 to 101, wherein the path for traffic is specified with routing rules, and destination rules are used to configure a set of policies that gateway proxies apply to a request at a specific destination.
103. The computer system of any of Claims 56 to 102, wherein a Router Management Service in a local server cluster uses a Discovery tool to discover endpoints within other server clusters, to do the following:(i) Read a secret which contains credentials and server URL of another server cluster; (ii) Adds this other server cluster to a list of remote server clusters to communicate with;(iii) Communicates with a network API server in that remote server cluster and collects the Services and their endpoint IPs;(iv) Adds these Services, with their endpoints, to router configurations in Sidecars and Gateways in the local server cluster.
104. The computer system of any of Claims 56 to 103, wherein requests are sent to a single IP address from the first service unit of the first data centre to keep full-server cluster outlier detection working as normal, so requests are not sent straight to the remote service name, as the router management service would add all Default Gateway service units in the second Data Centre as endpoints, and will no longer outlier detect entire server clusters, but instead would just outlier detect single service units.
105. A computer-implemented method of reconfiguring a first router in a computer system from a first configuration to a second configuration, wherein the computersystem includes a first data centre, a second data centre and a third data centre, the first data centre communicating with the second data centre and with the third data centre via a network, wherein the first data centre and the second data centre are separated by between 2 km and 500 km, wherein the third data centre and the second data centre are separated by between 2 km and 500 km, and wherein the first data centre and the third data centre are separated by between 2 km and 500 km, wherein each data centre includes a respective default gateway;the first data centre including a first respective subnet which includes a first respective server cluster, the first respective server cluster providing a first service, in which the first respective server cluster receives a query in respect of the first service, and in response queries a second service, and receives a response from the second service, and uses the response received from the second service to provide a response from the first service,the first data centre including a second respective subnet different to the first respective subnet which includes a second respective server cluster different to the first respective server cluster, wherein the second respective server cluster provides the second service, in which the second respective server cluster receives a query in respect of the second service and provides a response;the second data centre including a respective subnet which includes a respective server cluster providing the second service, in which the respective server cluster receives a query in respect of the second service and provides a response;the third data centre including a respective subnet which includes a respective server cluster providing the second service, in which the respective server cluster receives a query in respect of the second service and provides a response;wherein in the first data centre the first respective server cluster provides the first service, the first respective server cluster including a first service client unit, the first service client unit including a first service application executing in the first service client unit, and wherein the first service client unit includes the first router; wherein in the first data centre the second respective server cluster includes a second service unit, the second service unit including a second service application executing in the second service unit, and the second service unit includes a second router which receives a request from the first service via the default gateway of the first data centre, and routes the request to the second service application for processing by the secondservice application;wherein in a first configuration of the first router, the first router routes greater than 90% (e.g. greater than 97%, e.g. 99%) of the requests from the first service to the second service via the default gateway of the first data centre to the second respective subnet in the first data centre, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster provides the second service;wherein in the first configuration of the first router, the first router routes between 0.1% and 5% (e.g. 1%) of the requests from the first service to the second service via the default gateway of the second data centre to the respective subnet in the second data centre, the respective subnet including the respective server cluster providing the second service;wherein the respective server cluster of the second data centre includes a second service unit, the second service unit including a second service application executing in the second service unit, and the second service unit includes a third router which receives a request from the first service via the default gateway of the second data centre, and routes the request to the second service application for processing by the second service application;wherein in the first configuration of the first router, the first router routes between 0.1% and 5% (e.g. 1%) of the requests from the first service to the second service via the default gateway of the third data centre to the respective subnet in the third data centre, the respective subnet including the respective server cluster providing the second service;wherein the respective server cluster of the third data centre includes a second service unit, the second service unit including a second service application executing in the second service unit, and the second service unit includes a fourth router which receives a request from the first service via the default gateway of the third data centre, to route the request to the second service application for processing by the second service application;wherein the computer system does not include a network load balancer between the first data centre and the second data centre, and wherein the computer system does not include a network load balancer between the first data centre and the third data centre,and wherein the first data centre does not include a network load balancer between first respective subnet and the second respective subnet, wherein the network load balancers are configured to load balance the second requests;wherein in a second configuration of the first router, the first router routes zero percent of the requests from the first service to the second service via the default gateway of the first data centre to the second respective subnet in the first data centre, different to the first respective subnet, the second respective subnet including the second respective server cluster different to the first respective server cluster, wherein the second respective server cluster provides the second service, and wherein in the second configuration of the first router, the first router routes a non-zero fraction of the requests from the first service to the second service via the default gateway of the second data centre to the respective subnet in the second data centre, the respective subnet including the respective server cluster, wherein the respective server cluster provides the second service; and wherein in the second configuration of the first router, the first router routes a non-zero fraction of the requests from the first service to the second service via the default gateway of the third data centre to the respective subnet in the third data centre, the respective subnet including the respective server cluster, wherein the respective server cluster provides the second service;the method including the step of the first router changing from the first configuration of the first router to the second configuration of the first router, in response to the first router detecting that the second respective server cluster of the first data centre including the respective second service unit becomes faulty.
106. The method of Claim 105, including use of a system of any of Claims 56 to 104.