Cross-cluster disaster recovery method based on service grid and electronic equipment

By adding disaster recovery identifiers to service requests and combining Envoy and Istio configurations, the flexibility and accuracy of the service mesh cross-cluster disaster recovery method is achieved, and the flexibility and fine-grained problems of cross-cluster disaster recovery strategies in the existing technology are solved, ensuring the continuity and timely response of service requests.

CN120455453APending Publication Date: 2025-08-08中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510584367.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the prior art, service mesh technology is used to implement cross-cluster disaster recovery strategies without flexibility and fine-grainedness, so it is impossible to accurately process specific services or requests, and it is inefficient when network isolation between clusters.

Method used

By attaching disaster recovery identifiers to service requests, using local cluster identification and redirecting to off-site cluster processing, combining Envoy source code modification and Istio virtual service configuration, intelligent redirection and dynamic routing of service requests are realized, and selective disaster recovery switching for specific services or traffic is supported.

Benefits of technology

It realizes service request continuity and timely response in the case of local service failure or network isolation, improves the flexibility and adaptability of fault handling, and ensures dynamic routing adjustment capabilities when service status and network conditions change.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455453A_ABST
    Figure CN120455453A_ABST
Patent Text Reader

Abstract

The invention provides a cross-cluster disaster recovery method based on a service grid and electronic equipment, and the method comprises the steps that a local cluster receives a service request sent by a client, and the local cluster represents a Kubernetes cluster where an environment initially initiated by the service request is located; the local cluster determines whether a disaster recovery backup identifier exists in header information of the service request, and redirects the service request to a remote cluster under the condition that the disaster recovery backup identifier exists in the header information, so that the remote cluster processes and responds to the service request, and the disaster recovery backup identifier is used for indicating that the service request needs to be redirected to the remote cluster; the remote cluster represents a Kubernetes cluster which is different from the local cluster and is located at other geographic positions or in a network environment. According to the invention, the problems of lack of flexibility and relatively coarse fine granularity when a cross-cluster disaster recovery strategy is implemented by using a service grid technology in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a cross-cluster disaster recovery method based on a service grid, a computer-readable storage medium, a computer program product, and an electronic device. Background Art

[0002] Currently, there are several common technical challenges and limitations when implementing cross-cluster disaster recovery strategies using service mesh technologies (such as Istio). First, most existing methods rely on switching all ingress traffic to a backup system. This full traffic switching strategy not only lacks flexibility but also fails to implement fine-grained disaster recovery for specific services or specific types of requests. Second, traditional solutions rely on service registries (such as Nacos or Pilot), which are ineffective when two Kubernetes clusters are network-isolated, limiting their availability and efficiency in specific scenarios. Finally, currently popular disaster recovery solutions (such as the official Istio method) require synchronizing service mesh metadata between clusters, which not only increases the resource burden on control plane components (such as causing a surge in Pilot's memory usage) but also complicates the deployment and configuration process. This is especially true when configuring a dual-active mode across networks, which increases the difficulty and complexity of implementation. Summary of the Invention

[0003] The main purpose of this application is to provide a cross-cluster disaster recovery method based on a service grid, a computer-readable storage medium, a computer program product, and an electronic device, so as to at least solve the problem in the prior art of lack of flexibility and coarse granularity when using service grid technology to implement cross-cluster disaster recovery strategies.

[0004] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a cross-cluster disaster recovery method based on a service grid is provided, including: a local cluster receives a service request sent by a client, and the local cluster represents the Kubernetes cluster where the environment in which the service request is initially initiated is located; the local cluster determines whether there is a disaster recovery identifier in the header information of the service request, and if the disaster recovery identifier exists in the header information, redirects the service request to an off-site cluster so that the off-site cluster processes and responds to the service request, and the disaster recovery identifier is used to indicate that the service request needs to be redirected to the off-site cluster, and the off-site cluster represents a Kubernetes cluster that is different from the local cluster and is located in other geographical locations or network environments.

[0005] Optionally, the method further includes: when the disaster recovery identifier does not exist in the header information, directing the service request to the service production end of the local cluster, so that the service instance in the local cluster processes and responds to the service request.

[0006] Optionally, there are multiple off-site clusters, and redirecting the service request to the off-site cluster includes: determining that the off-site cluster corresponding to the identification value of the disaster recovery identifier is the target off-site cluster, and the identification value corresponds one-to-one to the off-site cluster; redirecting the service request to the service instance in the target off-site cluster.

[0007] Optionally, determine whether there is a disaster recovery identifier in the header information of the service request, and redirect the service request to an off-site cluster if the disaster recovery identifier exists in the header information, including: based on the virtual service pre-configured in the local cluster, determine whether there is the disaster recovery identifier in the header information by checking the field of the header information in the service request, and if the disaster recovery identifier exists, determine that the off-site cluster corresponding to the identifier value is the target off-site cluster; based on the service export pre-configured in the local cluster, redirect the service request to the target off-site cluster, and the service export defines the entry gateway of the target off-site cluster.

[0008] Optionally, determine whether there is a disaster recovery identifier in the header information of the service request, and redirect the service request to an off-site cluster if the disaster recovery identifier exists in the header information, including: based on the pre-modified Envoy source code in the local cluster, determine whether there is a disaster recovery identifier in the header information, and if the disaster recovery identifier exists, determine that the off-site cluster corresponding to the identification value of the disaster recovery identifier is the target off-site cluster; based on the service export in the local cluster, redirect the service request to the target off-site cluster, and the service export defines the entry gateway of the target off-site cluster.

[0009] Optionally, based on the pre-modified Envoy source code in the local cluster, determining whether the disaster recovery identifier exists in the header information includes: based on the pre-modified decodeHeaders function in the Envoy source code, determining whether the disaster recovery identifier exists by checking the fields of the header information in the service request.

[0010] Optionally, receiving a service request sent by a client includes: receiving the service request, where the service request is generated after the disaster recovery identifier is added to the initial service request when the client sends the initial service request to the local cluster within a preset number of times and has never received a request response from the local cluster; or, where the service request is generated without modifying the initial service request when the client sends the initial service request to the local cluster within the preset number of times and receives the request response from the local cluster, and the disaster recovery identifier does not exist in the initial service request.

[0011] According to another aspect of the present application, a computer-readable storage medium is provided, which includes a stored program, wherein when the program is run, the device where the computer-readable storage medium is located is controlled to execute any one of the cross-cluster disaster recovery methods based on the service grid.

[0012] According to another aspect of the present application, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement any one of the service grid-based cross-cluster disaster recovery methods.

[0013] According to another aspect of the present application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a method for executing any one of the service grid-based cross-cluster disaster recovery methods.

[0014] Applying the technical solution of the present application, first, the local cluster receives the service request sent by the client, and then the local cluster determines whether there is a disaster recovery identifier in the header information of the service request, and redirects the service request to the remote cluster when the disaster recovery identifier exists, so that the remote cluster processes and responds to the service request. The disaster recovery identifier is used to indicate that the service request needs to be redirected to the remote cluster. Compared with the problem of lack of flexibility and coarse granularity when using service grid technology to implement cross-cluster disaster recovery strategies in the prior art, the present application attaches a disaster recovery identifier to the service request, so that the local cluster can identify the disaster recovery identifier in the header information of the service request, and can intelligently and efficiently redirect the service request to the remote cluster for processing, thereby ensuring the continuity of the service request and the timeliness of the response in the event of a local service failure or network isolation, realizing selective disaster recovery switching for specific services or traffic, supporting the directional transfer of some service traffic, more accurately matching the needs of disaster recovery scenarios, improving the flexibility and adaptability of fault handling, and ensuring the dynamic routing adjustment capability when the service status and network conditions change. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings that constitute part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation on this application. In the drawings:

[0016] Figure 1 A hardware structure block diagram of a mobile terminal for executing a cross-cluster disaster recovery method based on a service grid provided in an embodiment of the present application is shown;

[0017] Figure 2 A schematic diagram of a process flow of a cross-cluster disaster recovery method based on a service grid provided according to an embodiment of the present application is shown;

[0018] Figure 3 A technical architecture diagram corresponding to a cross-cluster disaster recovery method based on a service grid provided according to an embodiment of the present application is shown;

[0019] Figure 4 A call flow chart of a specific service grid-based cross-cluster disaster recovery method provided according to an embodiment of the present application is shown.

[0020] The above drawings include the following reference numerals:

[0021] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. DETAILED DESCRIPTION

[0022] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0023] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] For ease of description, some nouns or terms involved in the embodiments of the present application are explained below:

[0026] Service Mesh: In a microservices architecture, a service mesh typically refers to an infrastructure layer dedicated to handling inter-service communication. It provides load balancing, service discovery, traffic management, and inter-service encryption.

[0027] Istio: An open-source service mesh that provides a unified way to connect, secure, control, and observe microservices. Istio manages network communication between microservices by deploying a lightweight sidecar proxy on each service in an application.

[0028] Envoy: An open-source edge and service proxy for cloud-native applications. Envoy is the sidecar proxy used in the Istio service mesh, responsible for dynamic routing, service discovery, load balancing, TLS termination, and HTTP / 2 & gRPC proxying.

[0029] Envoy Filter: Envoy's configuration can be modified dynamically through EnvoyFilter, which includes Lua scripts that can insert custom logic to handle traffic in and out of the Envoy proxy;

[0030] Pilot: An Istio component responsible for propagating service routing information to Istio's Envoy instances. Pilot allows service mesh users to define advanced routing rules, which are then implemented through Envoy.

[0031] ServiceEntry: A resource type in the Istio configuration model that allows external service entries to be added to Istio's internal service registry, allowing internal services to communicate with external services;

[0032] Virtual Service (VS): A resource defined in Istio that controls the distribution of traffic within the mesh through a set of routing rules. Virtual Services can be used to configure the behavior of ingress and egress traffic.

[0033] psbc-ingress-gateway: This is assumed to be a proper noun, referring to an ingress gateway designed specifically for a service mesh. It is responsible for forwarding external requests to the appropriate service in the service mesh.

[0034] Metadata synchronization: In a multi-cluster configuration, service status and configuration information across different clusters need to be shared and updated in real time. Metadata synchronization refers to this real-time update process of cross-cluster status and configuration information.

[0035] Failover: When a system failure occurs, the process of automatically transferring requests from the failed node to the normal node to ensure the continuous availability and stability of the service;

[0036] Dynamic DNS resolution: The process of dynamically resolving domain names to IP addresses, allowing services to dynamically find the location of target services without knowing the IP address in advance.

[0037] As introduced in the background technology, the existing technology lacks flexibility and has coarse granularity when implementing cross-cluster disaster recovery strategies using service grid technology. To solve the above problems, the embodiments of the present application provide a cross-cluster disaster recovery method based on service grid, a computer-readable storage medium, a computer program product and an electronic device.

[0038] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.

[0039] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure diagram of a mobile terminal of a cross-cluster disaster recovery method based on a service grid according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0040] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the cross-cluster disaster recovery method based on the service grid in the embodiment of the present invention. The processor 102 executes the computer program stored in the memory 104 to perform various functional applications and data processing, thereby implementing the above-mentioned method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or transmit data via a network. Specific examples of such networks may include a wireless network provided by the mobile terminal's telecommunications provider. In one example, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0041] In this embodiment, a cross-cluster disaster recovery method based on a service grid is provided, which runs on a mobile terminal, a computer terminal or a similar computing device. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0042] Figure 2 This is a flow chart of a cross-cluster disaster recovery method based on a service grid according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps:

[0043] Step S201: A local cluster receives a service request sent by a client. The local cluster represents the Kubernetes cluster where the environment where the service request is initially initiated is located.

[0044] Specifically, a "service request" refers to a request sent by a client to a service within the service mesh, aiming to obtain certain data or perform a specific task. This request usually follows a network communication protocol such as HTTP or gRPC, and carries necessary parameters, identity authentication information, and possible business logic identifiers so that the target service can understand and process the content of the request.

[0045] In step S202, the local cluster determines whether there is a disaster recovery identifier in the header information of the service request, and redirects the service request to the remote cluster if the disaster recovery identifier exists in the header information, so that the remote cluster processes and responds to the service request. The disaster recovery identifier is used to indicate that the service request needs to be redirected to the remote cluster. The remote cluster represents a Kubernetes cluster that is different from the local cluster and is located in other geographical locations or network environments.

[0046] Specifically, directing the request service to a remote cluster or a local cluster means routing the traffic carried by the request service to the remote cluster or the local cluster.

[0047] Specifically, other geographical locations or network environments refer to geographical locations or network environments that are different from those where the local cluster is located.

[0048] Specifically, the header information also contains the metadata of the request, which is used to describe key information such as the request attributes, source, destination, and how to process the request.

[0049] Through the above embodiment, first, the local cluster receives the service request sent by the client, and then the local cluster determines whether there is a disaster recovery identifier in the header information of the service request, and when there is a disaster recovery identifier, redirects the service request to the remote cluster so that the remote cluster processes and responds to the service request. The disaster recovery identifier is used to indicate that the service request needs to be redirected to the remote cluster. Compared with the problem of lack of flexibility and coarse granularity when using service grid technology to implement cross-cluster disaster recovery strategies in the prior art, the present application attaches a disaster recovery identifier to the service request, so that the local cluster can identify the disaster recovery identifier in the header information of the service request, and can intelligently and efficiently redirect the service request to the remote cluster for processing, thereby ensuring the continuity of the service request and the timeliness of the response in the event of a local service failure or network isolation, realizing selective disaster recovery switching for specific services or traffic, supporting the directional transfer of some service traffic, more accurately matching the disaster recovery scenario requirements, improving the flexibility and adaptability of fault handling, and ensuring the dynamic routing adjustment capability when the service status and network conditions change.

[0050] In an optional solution, the method further includes: when the disaster recovery identifier does not exist in the header information, directing the service request to the service production end of the local cluster, so that the service instance in the local cluster processes and responds to the service request. In this embodiment, when the disaster recovery identifier does not exist in the header information of the service request, it can be ensured that the service request is processed normally within the local cluster, and the service production end of the local cluster directly parses the request, executes business logic, and generates a response, thereby giving priority to using local resources to quickly respond to requests, improving service response speed and user experience, while reducing delays in cross-cluster communication and potential network bottlenecks, and optimizing the resource utilization efficiency and performance of the service grid.

[0051] Specifically, before a service request carrying a disaster recovery identifier is transmitted across clusters, the disaster recovery identifier will be encrypted. The disaster recovery identifier can be encrypted using a public key infrastructure (PKI) or a symmetric encryption scheme to avoid interception and tampering during transmission, thereby ensuring the integrity and authenticity of the disaster recovery request.

[0052] According to some exemplary embodiments of the present application, there are multiple off-site clusters, and redirecting the service request to the off-site cluster includes: determining that the off-site cluster corresponding to the identification value of the disaster recovery identifier is the target off-site cluster, and the identification value corresponds one-to-one to the off-site cluster; redirecting the service request to the service instance in the target off-site cluster. In this embodiment, by assigning a specific identification value to the disaster recovery identifier, it is possible to achieve refined control of multiple off-site clusters, ensuring that service requests can be directed to the most appropriate disaster recovery cluster. For example, if there are multiple off-site clusters, each cluster may have different loads, geographical locations or network conditions, then through different identification values, service requests can be directed to a disaster recovery cluster with lower load, smaller network latency or closer geographical location. This method solves the problems of uneven resource allocation, high network latency or inappropriate geographical location that may exist in a single disaster recovery cluster. Through flexible configuration of disaster recovery identification values, optimal utilization of disaster recovery resources is achieved, and service availability and user experience are improved.

[0053] In other embodiments, determining whether there is a disaster recovery identifier in the header information of the above-mentioned service request, and redirecting the above-mentioned service request to an off-site cluster if the above-mentioned disaster recovery identifier exists in the above-mentioned header information includes: based on the virtual service pre-configured in the above-mentioned local cluster, determining whether there is the above-mentioned disaster recovery identifier in the above-mentioned header information by checking the field of the above-mentioned header information in the above-mentioned service request, and if the above-mentioned disaster recovery identifier exists, determining that the above-mentioned off-site cluster corresponding to the above-mentioned identifier value is the above-mentioned target off-site cluster; based on the service export pre-configured in the above-mentioned local cluster, redirecting the above-mentioned service request to the above-mentioned target off-site cluster, and the above-mentioned service export defines the entry gateway of the above-mentioned target off-site cluster. In this embodiment, by pre-configuring the Istio virtual service (VirtualService) in the local cluster, this method can quickly check and parse the specific disaster recovery identification field in the request header information in the early stage of service request processing. Once the existence of the disaster recovery identification is identified, the target remote cluster can be determined based on the identification value, and intelligent routing decisions for the service request can be made. This method fully utilizes the routing capabilities and policy flexibility of the Istio service grid, making the disaster recovery process both efficient and controllable; in addition, the pre-configuration of the service exit (ServiceEntry) clarifies the entry gateway of the target remote cluster, ensuring that even in the case of network isolation or limited inter-cluster communication, the service request can be accurately directed to the correct disaster recovery resource. This method not only provides immediate response capabilities to disaster recovery events, but also enhances the disaster tolerance and service continuity of the service grid in a multi-cluster environment through clear routing logic and configuration policies, while reducing the complexity of implementation and maintenance and improving the overall stability of the system.

[0054] According to other exemplary embodiments of the present application, determining whether there is a disaster recovery identifier in the header information of the above-mentioned service request, and redirecting the above-mentioned service request to a remote cluster if the above-mentioned disaster recovery identifier exists in the above-mentioned header information, including: based on the Envoy source code pre-modified in the above-mentioned local cluster, determining whether there is the above-mentioned disaster recovery identifier in the above-mentioned header information, and if the above-mentioned disaster recovery identifier exists, determining that the above-mentioned remote cluster corresponding to the identifier value of the above-mentioned disaster recovery identifier is the target remote cluster; based on the service export in the above-mentioned local cluster, redirecting the above-mentioned service request to the above-mentioned target remote cluster, the above-mentioned service export defines the entry gateway of the above-mentioned target remote cluster. In this embodiment, by pre-modifying the Envoy proxy source code in the local cluster, the method can directly identify and parse the disaster recovery identifier in the service request header information in the underlying logic of HTTP request processing, thereby achieving immediate response and intelligent routing to disaster recovery requests without affecting the daily operation of the service grid. This method significantly improves the efficiency of disaster recovery switching, because Envoy can directly determine the target remote cluster based on the identifier value, without the need to perform complex queries and decisions step by step through Istio's Pilot or other service registration centers, greatly reducing the delay in disaster recovery processing.

[0055] In some other optional solutions of the present application, based on the pre-modified Envoy source code in the local cluster, determining whether the above-mentioned disaster recovery identifier exists in the above-mentioned header information includes: based on the pre-modified decodeHeaders function in the Envoy source code, by checking the fields of the above-mentioned header information in the above-mentioned service request, determining whether the above-mentioned disaster recovery identifier exists. In this embodiment, since the decodeHeaders function is the key entry point for Envoy to process HTTP requests, modifying it allows Envoy to immediately execute routing logic based on the disaster recovery identifier in the header information without waiting for the decision results of other components (such as Pilot), thereby significantly reducing the response time and resource consumption of disaster recovery switching.

[0056] In some further optional schemes of the present application, receiving a service request sent by a client includes: receiving the above-mentioned service request, wherein the above-mentioned service request is generated after adding the above-mentioned disaster recovery identifier to the above-mentioned initial service request when the above-mentioned client sends the initial service request to the above-mentioned local cluster within a preset number of times and has never received a request response from the above-mentioned local cluster, or the above-mentioned service request is generated when the above-mentioned client sends the above-mentioned initial service request to the above-mentioned local cluster within the above-mentioned preset number of times and receives the above-mentioned request response from the above-mentioned local cluster without modifying the above-mentioned initial service request, and the above-mentioned disaster recovery identifier does not exist in the above-mentioned initial service request. In this embodiment, the client request strategy is dynamically adjusted to realize intelligent control of the internal traffic of the service grid and efficient use of disaster recovery resources, which not only ensures that traffic can be redirected quickly and accurately when the service is abnormal, but also maintains the performance advantage under normal service conditions to the greatest extent, providing solid technical support for building a flexible and highly available service grid architecture.

[0057] Specifically, this application ensures that cross-cluster traffic switching and failover can be flexibly implemented based on real-time business needs and system status without the need to synchronize metadata between clusters. This aims to optimize problems existing in traditional disaster recovery solutions, such as the lack of flexibility in full traffic switching, service unreachability due to network isolation, and performance bottlenecks caused by metadata synchronization. Specifically, it includes the implementation of the following key functions: 1. Environment configuration and initialization: In the initial stage, the system settings of the local cluster are optimized by executing Kubernetes commands, cleaning up the old configuration of the local cluster and applying the new configuration files (envoy_bootstrap_tmpl.json and gcp_envoy_bootstrap_tmpl.json).json), then deploy the psbc-ingress-gateway of the local cluster, reference the newly created configuration, and ensure that the gateway is started and running correctly; 2. Consumer-side retry strategy: On the service requester (i.e., the client), by setting the number of retries and timeout mechanism, priority is given to trying to access internal services (i.e., services in the local cluster). If consecutive attempts fail, the system will start the failover mechanism by adding a specific identifier (i.e., disaster recovery identifier) (such as dcdus:f) in the request header (i.e., the header information of the request service) to indicate that the traffic carried by the request service needs to be redirected to the backup cluster (i.e., the remote cluster); 3. Dynamic traffic routing strategy: Relying on Istio's VirtualService (i.e. virtual service), dynamically decides traffic routing based on the dcdus field in the service request. If it is marked as "f", the traffic will be directed to the ingress gateway psbc-ingress-gateway of the backup cluster. Otherwise, the traffic will be directed to the default service scserver-service of the local cluster. 4. Service exit point configuration: ServiceEntry (i.e. service exit) configuration defines the entry point of the external cluster (i.e. remote cluster), and specifies in detail how to route service requests to external addresses and ports to ensure that service requests can reach their destination accurately. 5. Traffic routing and failover execution: Under normal circumstances, service requests are routed to the Spring For the Boot service (i.e., the local consumer service), when a service request carries the specific identifier dcdus:f, it is redirected to the backup cluster according to pre-set logic. Dynamic DNS resolution is used to adjust routing policies in real time to address service anomalies or disaster recovery needs. 6. Dynamic forwarding mechanism on the production side: By configuring Envoy's dynamic_forward_proxy filter, dynamic DNS resolution and routing adjustments are implemented for traffic. This mechanism enables the system to automatically locate the appropriate backup cluster and forward traffic to it, ensuring fast and accurate traffic processing even in changing network environments or disaster recovery scenarios. 7. Testing and verification process: To verify the correctness of the system configuration and the effectiveness of the traffic routing logic, a series of tests were conducted. First, curl requests were executed using the kubectl tool to verify that service instances within the local cluster could be successfully accessed under standard conditions without special header information. Next, the kubectl tool was used again, but this time the dcdus disaster recovery identifier header was added to the request to test and verify that the system could correctly redirect traffic to services in the external cluster according to pre-set conditions.

[0058] Specifically, multiple backup clusters can be set up. When there are more backup clusters, if a service request fails to access the first backup cluster, the identification value of the disaster recovery identifier can be modified (for example, the disaster recovery identifier corresponding to the first backup cluster can be changed from dcdus=f to the disaster recovery identifier corresponding to the second backup cluster, dcdus=h). Then, the service request is redirected to the second backup cluster. In this way, the service request can be redirected by simply modifying the identification value of the disaster recovery identifier, ensuring that the service request is ultimately successfully accessed.

[0059] Figure 3 This is the technical architecture diagram of this application, such as Figure 3As shown, S1. Client request (initiating request): The client initiates a request to Kubernetes cluster A (i.e., service request); S2. Spring Boot application (local consumer service): The application deployed in Kubernetes cluster A (i.e., local cluster) is responsible for processing client requests; S3. Envoy Sidecar Proxy: The Envoy proxy is deployed in Kubernetes cluster A and works with the Spring Boot application to manage network communications as part of the service mesh. It is not only responsible for intelligent routing of traffic, but also can implement policy decisions and routing for traffic with disaster recovery identification through configured VirtualService rules or directly modified Envoy routing logic; S4. VirtualService or Envoy decodeHeaders (virtual service or Envoy parses the header, i.e. routing rules - transit control): determines whether it is necessary to change the request route based on the request header information (such as the presence of a disaster recovery identifier dcdus). The VS component can dynamically adjust the routing direction of the traffic based on the disaster recovery identifier in the request header information (such as dcdus=f). In addition, the decodeHeaders function of the Envoy proxy is used to detect and parse the disaster recovery identifier in the request header, and accordingly decide whether to route the request to the service instance of Kubernetes cluster A or the backup Kubernetes cluster B or cluster C; S5. Spring Boot application (local production service): If disaster recovery redirection is not required, the request is handled by the Spring application in Kubernetes cluster A. Boot application processing; S6. ServiceEntry (remote service egress configuration): When the disaster recovery flag dcdus=f is detected, the request is routed to the backup Kubernetes cluster B according to the ServiceEntry configuration. ServiceEntry defines the entry point for external services in the Istio service mesh, ensuring that traffic is correctly directed from Kubernetes cluster A to Kubernetes cluster B when the disaster recovery mechanism is triggered. It specifies the route mapping of service requests from the source cluster to the destination cluster, including the virtual service hostname and port configuration, which facilitates flexible cross-cluster traffic management; S7.psbc-ingress-gateway (ingress gateway, also known as a namespace-level gateway): Serving as the ingress gateway for Kubernetes cluster B, it receives requests from cluster A based on routing rules. As the front-end entry point for Kubernetes cluster B, it receives traffic redirected from cluster A based on the routing configuration in ServiceEntry. This gateway handles dynamic DNS resolution and load balancing, adjusting routing policies based on real-time resolution results to provide traffic support for the Spring Boot application in the backup cluster. S8. Spring Boot application (remote production service): Deployed in Kubernetes cluster B, it serves as a disaster recovery system to handle redirected requests. S9. psbc-ingress-gateway (another backup ingress gateway): The following is an optional extension step: When there are more backup clusters (such as cluster C), this process can further route requests to other backup clusters. For example, if a service request fails to access cluster B, the disaster recovery identifier value is changed from f to h (that is, the disaster recovery identifier is changed from dcdus=f to dcdus=h), and the request is redirected to cluster C.

[0060] Specifically, the production-side service refers to the service production side.

[0061] Figure 4 This is a specific call flow chart of the cross-cluster disaster recovery method based on service grid in this application. The system can flexibly and automatically decide traffic routing based on real-time disaster recovery requirements and service status, and achieve efficient cross-cluster failover and disaster recovery operations. Figure 4The call flow chart in the description details each step: S1. Client request processing: The client initiates a service request and first tries to access the service instance in the local cluster. If it is still inaccessible after the set number of attempts, the system determines that the service may need disaster recovery processing and adds a disaster recovery identifier dcdus=f to the request header so that subsequent steps can identify and make corresponding routing decisions; S2-1. Sidecar proxy reception and processing: The Sidecar proxy intercepts the client request, checks the request header information, and determines the processing path of the request based on whether the request contains a disaster recovery identifier. The proxy is responsible for inter-service communication and request routing decisions; S3. Local service instance processing: If there is no disaster recovery identifier in the request or the identifier does not indicate that disaster recovery is required, the request is processed by the service instance in the production service of the local cluster and the result is returned to the client; S2-2. Sidecar disaster recovery routing decision: As part of the same Sidecar service in the second step, if the request contains a disaster recovery identifier, the Sidecar proxy will use Istio's VirtualService rule to route the request and call port 30819 to redirect it to the Ingress of the backup cluster specified in the configuration. Gateway; S4. ServiceEntry configuration guidance: ServiceEntry requests through a specific gateway IP port according to the configuration guidance to ensure that the request can reach the correct entrance of the backup cluster; S5. Remote cluster ingress gateway reception: The ingress gateway of the backup cluster receives the request redirected by ServiceEntry and determines which service instance to route the request to based on its own service discovery and load balancing mechanism. Among them, the Istio ingress gateway can be published in any of the following ways: nodeport, hostport, etc.; S6. Remote cluster service instance processing: The request forwarded by the Ingress Gateway is ultimately processed by the service instance in the remote backup cluster, and the processing result will be returned to the client.

[0062] Specifically, environment configuration and initialization configure and initialize an environment that supports cross-cluster disaster recovery, ensure that traffic can be flexibly and accurately redirected in the event of service anomalies or disaster recovery, and create a configuration environment for the cross-cluster disaster recovery process that does not require metadata synchronization. This includes extracting existing configurations from the cluster, updating configuration files, creating configuration mappings, and redeploying the gateway with updated configurations. The key components and steps of the process are as follows: 1. Extraction of configuration files: First, extract the envoy_bootstrap_tmpl.json and gcp_envoy_bootstrap_tmpl.json files from the Pod of the running Kubernetes cluster. These files contain the basic configuration of the Envoy proxy and provide a template for disaster recovery switching, requiring envoy_bootstrap _tmpl.json, gcp_envoy_bootstrap_tmpl.json versions and configurations are consistent with the production environment; 2. Configuration update: The extracted configuration files need to be updated to support cross-cluster disaster recovery solutions. Edit the envoy_bootstrap_tmpl.json file and add additional listener configurations. These listeners are specifically used to handle inbound requests marked as disaster recovery transfers. At the same time, update the "clusters" section and introduce a new dynamic_forward_proxy_cluster to ensure that service instances can be dynamically resolved and load balancing can be achieved; 3. Configuration map creation ConfigMap update: In order to apply the updated Envoy configuration to the service grid, a new Kubernetes needs to be created ConfigMap, the newly created ConfigMap is named ggateway-configmap, which is used to store the modified envoy_bootstrap_tmpl.json and gcp_envoy_bootstrap_tmpl.json files. Before updating the ConfigMap, if the old ConfigMap exists, it must be deleted first; 4. Mount ConfigMap: Mount the updated ggateway-configmap to psbc-ingress-gateway so that the gateway loads the latest configuration and handles cross-cluster traffic; 5. Gateway deployment: After creating a new ConfigMap, you need to deploy or update psbc-ingress-gateway. This operation synchronizes the existing deployment of the gateway with the latest ggateway-configmap to ensure that the gateway's Envoy agent starts with the latest configuration file.

[0063] Specifically, the retry mechanism on the consumer side is designed to ensure a smooth transition to the backup cluster when the local cluster service is unreachable in the construction of a cross-cluster disaster recovery system, thereby ensuring the continuity and stability of the service: 1. Retry strategy definition: The retry strategy adopted by the consumer side aims to prioritize the service availability of the local cluster. This strategy is implemented by the application-level client (such as Spring Boot, Dubbo, Go, etc.). The specific strategies are as follows: 1) Retry mechanism configuration: turn off the default retry mechanism of Istio / Sidecar, and let the client code control the retry logic to achieve more refined fault judgment and processing; 2) Global retry mechanism: customize the retry strategy, use Ribbon's RetryRule as the strategy, and extend the RetryRule code to capture and handle exceptions including database exceptions, service unavailability exceptions, etc.; 3) Preset number of attempts: define the maximum number of attempts for a service request before triggering a failover, the default is 3 times, and the default number of requests forwarded to the gateway is 1; 4) Timeout parameter setting: set a timeout for each service request, the default is 1 second; 5) Exception judgment: the client code is responsible for capturing and judging various exceptions, including but not limited to connection timeouts, service unavailability, response timeouts, etc. 2. Failover logic: When consecutive service requests fail in the local cluster (i.e. after 3 preset attempts), the consumer end will start the failover logic. The specific process is as follows: 1) Identifier injection: Add a disaster recovery identifier (for example, dcdus=f) to the HTTP request header. This identifier serves as a trigger condition for dynamic routing, indicating that the request needs to be redirected to the backup cluster; 2) Request retransmission: The request with the disaster recovery identifier will be resent. In this step, the dynamic routing configuration of the service grid or client will automatically route the request to the corresponding gateway of the backup cluster based on the dcdus identifier.

[0064] Specifically, the key to the traffic routing rule strategy of this application is to combine the use of Istio's VirtualService and ServiceEntry resources, and modify the Envoy source code to implement dynamic routing decisions based on specific request header information, thereby realizing a flexible and efficient cross-cluster disaster recovery transfer mechanism. This strategy ensures that when the service requires disaster recovery processing, the traffic can be seamlessly redirected to the external cluster. Specifically: 1. Istio VirtualService configuration: Istio's VirtualService defines routing decisions based on HTTP header information. By checking the dcdus header field, we can determine whether to redirect the request to the disaster recovery path. When the value of the dcdus field is "f", Istio will direct the traffic to the ingress of the external cluster defined by ServiceEntry. gateway (psbc-ingress-gateway), if this field does not exist in the request header, or its value does not match the specified disaster recovery identifier, the traffic will be routed to the default internal service (scserver-service); 2. Envoy source code modification plan: In addition to using Istio configuration to implement routing decisions, this application can also directly modify the Envoy source code to control the routing logic at a lower level. In Envoy's HTTP filter, modify the decodeHeaders function in source / common / router / router.cc to check the value of the dcdus header. When the disaster recovery identifier is detected, the target cluster is controlled by modifying the return value of config_.cm_.getThreadLocalCluster. Envoy will directly direct traffic to the specified backup cluster. This process does not rely on Istio's standard routing rules, thereby greatly improving the response speed of disaster recovery switching. Through this source code-level modification, the Envoy proxy image (such as the image specified by sidecar.istio.io / proxyImage) can be replaced to apply routing policies in batches. This eliminates the need to configure VirtualService rules for each service separately, simplifying the deployment and management of disaster recovery routing policies.

[0065] Specifically, service export configuration: defines the service entry of the external cluster (i.e., remote cluster) to ensure that when disaster recovery needs are triggered, the request can smoothly pass through the mesh boundary to the predetermined destination, and the traffic can be correctly directed to the Istio ingress of the external cluster. Gateway, ServiceEntry is the key resource for declaring external services in the Istio service grid. It specifies in detail how to map requests within the service grid to specific addresses and ports outside the cluster. In order to achieve correct routing of service requests, the ServiceEntry configuration includes the following key elements: 1. Hosts definition: By defining the virtual domain name psbc-ingress-gateway.namespace.svc.cluster.local, ServiceEntry provides an access path for the external cluster that can be referenced within the service grid. Traffic using this domain name will be considered outbound traffic; 2. Location attribute: Set the location attribute to MESH_EXTERNAL to explicitly mark the target service as an external service, so that ServiceEntry can direct traffic to resources outside the service grid; 3. Port configuration: Specify the access port and protocol through port configuration. Number: 30819 defines the port where traffic enters the external service, and protocol: HTTP indicates that the port uses the HTTP protocol. If gRPC traffic needs to be supported, the corresponding port and protocol must also be specified in a similar way to ensure correct processing of gRPC traffic; 4. Endpoints mapping: The endpoints part maps the above fictitious domain name to the external Istio The specific IP address of the ingressgateway is associated with the corresponding service port, regardless of whether the nodeport, Istio ingressgateway or hostport publishing method is used.

[0066] Specifically, the traffic routing and failover mechanism of the present application is based on intelligent decision-making to ensure high availability and disaster recovery capabilities of the service. The mechanism prioritizes internal services through multi-stage request attempts and dynamic routing logic, while maintaining a sensitive response to potential failures. Specifically: 1. Fault detection and traffic redirection: On the service consumer side, the traffic will first attempt to initiate multiple attempts to the service in the local cluster through ordinary HTTP requests. This can be implemented through Java's HTTP client, setting an appropriate timeout period, and performing a limited number of retries. For example, the availability of the local service can be determined by initiating three ordinary requests. When these requests all fail, the system will determine that the local service is temporarily unable to meet the processing requirements, and then initiate a request containing the disaster recovery identifier dcdus=f. The disaster recovery identifier indicates to the service grid that the current request needs to be forwarded to the backup cluster; 2. Dynamic DNS resolution and traffic direction of the backup cluster: On the service production side, the gateway of the backup cluster is configured with envoy.filters.http.dynamic_forward_proxy, enabling it to perform dynamic DNS resolution on the traffic and direct the traffic to the final target service based on the resolution results. The dynamic_forward_proxy_cluster in the configuration uses Envoy's dynamic forwarding capability to optimize the resolution process by specifying DNS cache configuration, so that traffic can be intelligently directed to the backup service without preset routing information; 3. Implementation details: 1) Normal traffic processing: Under normal circumstances, service requests are handled by the internal Spring Boot service, and the service consumer client will look for available service instances in the local cluster for communication; 2) Failover triggering: When continuous requests fail, the service consumer client starts the failover mechanism. At this time, it will initiate a request again and add dcdus=f to the HTTP header information, indicating that the request needs to be redirected to the backup cluster; 3) Backup cluster reception: After the backup cluster's gateway receives a request with a special identifier, it uses dynamic DNS resolution to determine and select the appropriate service instance to complete traffic forwarding; 4) Final service response: After redirection, the traffic finally reaches the target service in the backup cluster, which processes the request and returns a response to the service consumer.

[0067] Specifically, when implementing cross-cluster disaster recovery strategies, the dynamic forwarding mechanism on the production side plays a vital role. This mechanism is based on Envoy's envoy.filters.http.dynamic_forward_proxy configuration, which allows dynamic DNS resolution and routing decisions for incoming traffic, ensuring that even in disaster recovery scenarios, traffic can be correctly forwarded to the target service. The key components and implementation process of dynamic forwarding on the production side are elaborated in detail below: 1. Dynamic DNS resolution, that is, by configuring the envoy.filters.http.dynamic_forward_proxy filter, the IngressGateway of the backup cluster obtains the ability to perform instant DNS queries when receiving forwarding requests. This function enables the Gateway to dynamically identify and resolve the current IP address of the target service, regardless of whether the service has changed in the network environment: 1) Resolution mechanism: When a failover request marked with dcdus=f is received, the Gateway queries DNS to obtain the real-time address of the backup service. This avoids relying on the static IP address of the service instance and improves the flexibility and reliability of the disaster recovery process; 2) Forwarding strategy: Based on the resolution results, Gateway directs traffic to the service instance of the resolved backup cluster. This process supports multiple protocols (such as HTTP and HTTPS) to ensure that different types of service requests can be handled correctly. 2. Gateway's Envoy key parameter configuration. To ensure that the dynamic forwarding logic on the production side can be executed accurately and without error, the configuration of dynamic DNS resolution and traffic routing decisions is crucial. The following is a detailed description of the Envoy key parameter configuration: 1) Listener configuration: Define the listener listener_0 to listen on port 10000 on the 0.0.0.0 network interface. Due to conflicts in the 0.0.0.0:10000 port in some versions, please change the address to "{{(index.metadata.InstanceIPs 0)}}".This listener is used to capture all incoming traffic. This setting ensures that Envoy can receive requests from any source and become a unified entry point for traffic; 2) HTTP connection manager: Use envoy.filters.network.http_connection_manager to manage HTTP connections, in which the automatic codec type (codec_type:AUTO) is configured. This allows Envoy to automatically select the HTTP / 1.x or HTTP / 2 codec based on the actual traffic, optimizing the efficiency and compatibility of traffic processing; 3) Routing configuration: Define a routing configuration named local_route, which contains a virtual host that will connect all domains (" *") and the path (" / ") are routed to the dynamic_forward_proxy_cluster cluster. This setting enables Envoy to dynamically forward traffic to the resolved service address. 4) Cluster configuration Dynamic_forward_proxy_cluster: The cluster connection timeout is set to 0.5 seconds, the ROUND_ROBIN load balancing policy is adopted, and the envoy.clusters.dynamic_forward_proxy type is configured. This cluster configuration includes DNS cache configuration and specifies the DNS lookup policy as V4_ONLY to optimize the IPv4 address resolution process and enhance the efficiency and accuracy of service scheduling.

[0068] In summary, in the face of the limitations of existing cross-cluster disaster recovery solutions, the present application specifically solves the key problems in traditional cross-cluster disaster recovery solutions. The present application not only implements refined traffic switching for abnormal services, but also realizes adaptive management of cross-cluster disaster recovery through dynamic service discovery and load balancing technology, fundamentally optimizing the disaster recovery process in the service grid. These advantages make the present application have significant technical and practical value in providing cross-cluster disaster recovery solutions, which are specifically manifested in the following points: 1. Refined request orientation and failover control: By introducing a fine-grained traffic management mechanism, the present application realizes selective disaster recovery switching for specific services or traffic compared to the traditional full traffic switching strategy. When an abnormality occurs in any link in the service chain, the present application supports the directional transfer of part of the service traffic, more accurately matching the disaster recovery scenario requirements, with the help of the consumer end's accurate judgment of the abnormal state of the production end, and the strategy of traffic redirection through specific identifiers in the request header (such as dcdus=f), the present application not only improves the flexibility and adaptability of fault handling, but also ensures the dynamic routing adjustment capability when the service status and network conditions change. 2. Optimized Disaster Recovery Policy Configuration: This application significantly reduces the complexity of cross-cluster disaster recovery configuration and the need for metadata synchronization by innovatively modifying the routing logic of the Envoy proxy. Compared to the traditional Istio dual-master model, this approach avoids the tedious process of manually configuring a large number of Virtual Services (VS) and Destination Rules (DR) for each service instance. By implementing routing decisions directly at the Envoy level without pre-synchronizing a large amount of service metadata or configuration details, the main advantages are as follows: 1) Configuration Simplification: This significantly reduces the complexity of disaster recovery policy implementation and improves configuration flexibility and manageability; 2) Intelligent Routing Decisions: This intelligently redirects disaster recovery traffic directly at the Envoy level, enhancing the adaptability and accuracy of the disaster recovery process; 3) Reduced Resource Consumption: This effectively avoids the resource waste caused by metadata synchronization due to extensive configuration, improving system performance and stability. 3. Significantly Reduced Metadata Management Burden: By eliminating the need for cross-cluster metadata synchronization, this application significantly reduces the burden on the service mesh control plane. Metadata synchronization not only consumes a large amount of network and storage resources, but can also cause control plane performance bottlenecks, especially in large-scale cluster environments.This application effectively addresses the network isolation problem by simplifying the deployment and configuration process, which not only improves the practicality and maintainability of the system, but also reduces system resource consumption and operation and maintenance complexity. In addition, this method effectively solves the configuration update problem caused by gateway changes and reduces the burden of sending full metadata due to gateway changes. In traditional solutions, any change in the gateway may trigger extensive metadata updates for all services, thereby increasing the pressure on the control plane and may affect the stability of the service. This application simplifies the implementation of disaster recovery strategies and reduces dependence on metadata synchronization, which not only improves the configuration flexibility and system maintainability, but also significantly reduces the consumption of network and storage resources, fundamentally improving the efficiency and reliability of the disaster recovery process.

[0069] Specifically, the key points of this application are: 1. Identifier-based application-level traffic redirection mechanism: This application dynamically redirects traffic to the backup cluster according to the specific identifier carried in the request (such as the disaster recovery identifier dcdus=f) by utilizing Virtual Service rules or directly modifying the Envoy routing logic, thereby achieving fine-grained control of traffic switching. Compared with traditional batch traffic switching, this method provides a more flexible and accurate disaster recovery switching strategy, which can perform disaster recovery processing for specific services or requests, thereby improving the availability and continuity of services; 2. Dynamic service call and load balancing optimization: By integrating envoy.filters.http.dynamic_forward_proxy, this application solves the service discovery and load balancing problems in cross-cluster disaster recovery scenarios. This technology resolves the target address in real time and dynamically adjusts the routing strategy according to network conditions and service availability, ensuring efficient and stable service calls, and significantly improving the system's adaptability, elasticity and scalability to dynamic environments.

[0070] An embodiment of the present application provides a computer-readable storage medium, which includes a stored program. When the program is run, the device where the computer-readable storage medium is located is controlled to execute the cross-cluster disaster recovery method based on the service grid.

[0071] Specifically, cross-cluster disaster recovery methods based on service mesh include:

[0072] Step S201: A local cluster receives a service request sent by a client. The local cluster represents the Kubernetes cluster where the environment where the service request is initially initiated is located.

[0073] In step S202, the local cluster determines whether there is a disaster recovery identifier in the header information of the service request, and redirects the service request to the remote cluster if the disaster recovery identifier exists in the header information, so that the remote cluster processes and responds to the service request. The disaster recovery identifier is used to indicate that the service request needs to be redirected to the remote cluster. The remote cluster represents a Kubernetes cluster that is different from the local cluster and is located in other geographical locations or network environments.

[0074] Optionally, the method further includes: when the disaster recovery identifier does not exist in the header information, directing the service request to the service production end of the local cluster, so that the service instance in the local cluster processes and responds to the service request.

[0075] Optionally, there are multiple off-site clusters, and redirecting the service request to the off-site cluster includes: determining that the off-site cluster corresponding to the identification value of the disaster recovery identifier is the target off-site cluster, and the identification value corresponds one-to-one to the off-site cluster; redirecting the service request to the service instance in the target off-site cluster.

[0076] Optionally, determine whether there is a disaster recovery identifier in the header information of the above-mentioned service request, and redirect the above-mentioned service request to the remote cluster if the above-mentioned disaster recovery identifier exists in the above-mentioned header information, including: based on the virtual service pre-configured in the above-mentioned local cluster, by checking the field of the above-mentioned header information in the above-mentioned service request, determine whether there is the above-mentioned disaster recovery identifier in the above-mentioned header information, and if the above-mentioned disaster recovery identifier exists, determine that the above-mentioned remote cluster corresponding to the above-mentioned identifier value is the above-mentioned target remote cluster; based on the service export pre-configured in the above-mentioned local cluster, redirect the above-mentioned service request to the above-mentioned target remote cluster, and the above-mentioned service export defines the entry gateway of the above-mentioned target remote cluster.

[0077] Optionally, determine whether there is a disaster recovery identifier in the header information of the above-mentioned service request, and redirect the above-mentioned service request to the remote cluster if the above-mentioned disaster recovery identifier exists in the above-mentioned header information, including: based on the Envoy source code pre-modified in the above-mentioned local cluster, determine whether there is the above-mentioned disaster recovery identifier in the above-mentioned header information, and if the above-mentioned disaster recovery identifier exists, determine that the above-mentioned remote cluster corresponding to the identification value of the above-mentioned disaster recovery identifier is the target remote cluster; based on the service export in the above-mentioned local cluster, redirect the above-mentioned service request to the above-mentioned target remote cluster, and the above-mentioned service export defines the entry gateway of the above-mentioned target remote cluster.

[0078] Optionally, based on the pre-modified Envoy source code in the local cluster, determine whether the above-mentioned disaster recovery identifier exists in the above-mentioned header information, including: based on the pre-modified decodeHeaders function in the Envoy source code, determine whether the above-mentioned disaster recovery identifier exists by checking the fields of the above-mentioned header information in the above-mentioned service request.

[0079] Optionally, receiving a service request sent by a client includes: receiving the above-mentioned service request, where the above-mentioned service request is generated after adding the above-mentioned disaster recovery identifier to the above-mentioned initial service request when the above-mentioned client sends the above-mentioned initial service request to the above-mentioned local cluster within a preset number of times and has never received a request response from the above-mentioned local cluster; or, the above-mentioned service request is generated without modifying the above-mentioned initial service request when the above-mentioned client sends the above-mentioned initial service request to the above-mentioned local cluster within the above-mentioned preset number of times and receives the above-mentioned request response from the above-mentioned local cluster, and the above-mentioned disaster recovery identifier does not exist in the above-mentioned initial service request.

[0080] The present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement at least the following method steps: Step S201, a local cluster receives a service request sent by a client, and the local cluster represents the Kubernetes cluster where the environment in which the service request is initially initiated is located; Step S202, the local cluster determines whether there is a disaster recovery identifier in the header information of the service request, and if the disaster recovery identifier is present in the header information, redirects the service request to an off-site cluster, so that the off-site cluster processes and responds to the service request, and the disaster recovery identifier is used to indicate that the service request needs to be redirected to the off-site cluster, and the off-site cluster represents a Kubernetes cluster that is different from the local cluster and is located in another geographical location or network environment.

[0081] Optionally, the method further includes: when the disaster recovery identifier does not exist in the header information, directing the service request to the service production end of the local cluster, so that the service instance in the local cluster processes and responds to the service request.

[0082] Optionally, there are multiple off-site clusters, and redirecting the service request to the off-site cluster includes: determining that the off-site cluster corresponding to the identification value of the disaster recovery identifier is the target off-site cluster, and the identification value corresponds one-to-one to the off-site cluster; redirecting the service request to the service instance in the target off-site cluster.

[0083] Optionally, determine whether there is a disaster recovery identifier in the header information of the above-mentioned service request, and redirect the above-mentioned service request to the remote cluster if the above-mentioned disaster recovery identifier exists in the above-mentioned header information, including: based on the virtual service pre-configured in the above-mentioned local cluster, by checking the field of the above-mentioned header information in the above-mentioned service request, determine whether there is the above-mentioned disaster recovery identifier in the above-mentioned header information, and if the above-mentioned disaster recovery identifier exists, determine that the above-mentioned remote cluster corresponding to the above-mentioned identifier value is the above-mentioned target remote cluster; based on the service export pre-configured in the above-mentioned local cluster, redirect the above-mentioned service request to the above-mentioned target remote cluster, and the above-mentioned service export defines the entry gateway of the above-mentioned target remote cluster.

[0084] Optionally, determine whether there is a disaster recovery identifier in the header information of the above-mentioned service request, and redirect the above-mentioned service request to the remote cluster if the above-mentioned disaster recovery identifier exists in the above-mentioned header information, including: based on the Envoy source code pre-modified in the above-mentioned local cluster, determine whether there is the above-mentioned disaster recovery identifier in the above-mentioned header information, and if the above-mentioned disaster recovery identifier exists, determine that the above-mentioned remote cluster corresponding to the identification value of the above-mentioned disaster recovery identifier is the target remote cluster; based on the service export in the above-mentioned local cluster, redirect the above-mentioned service request to the above-mentioned target remote cluster, and the above-mentioned service export defines the entry gateway of the above-mentioned target remote cluster.

[0085] Optionally, based on the pre-modified Envoy source code in the local cluster, determine whether the above-mentioned disaster recovery identifier exists in the above-mentioned header information, including: based on the pre-modified decodeHeaders function in the Envoy source code, determine whether the above-mentioned disaster recovery identifier exists by checking the fields of the above-mentioned header information in the above-mentioned service request.

[0086] Optionally, receiving a service request sent by a client includes: receiving the above-mentioned service request, where the above-mentioned service request is generated after adding the above-mentioned disaster recovery identifier to the above-mentioned initial service request when the above-mentioned client sends the above-mentioned initial service request to the above-mentioned local cluster within a preset number of times and has never received a request response from the above-mentioned local cluster; or, the above-mentioned service request is generated without modifying the above-mentioned initial service request when the above-mentioned client sends the above-mentioned initial service request to the above-mentioned local cluster within the above-mentioned preset number of times and receives the above-mentioned request response from the above-mentioned local cluster, and the above-mentioned disaster recovery identifier does not exist in the above-mentioned initial service request.

[0087] An embodiment of the present application also provides an electronic device, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a method for executing any one of the above-mentioned service grid-based cross-cluster disaster recovery methods.

[0088] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0089] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0090] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0091] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0092] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0093] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0094] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0095] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0096] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0097] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0098] In the cross-cluster disaster recovery method based on service grid of the present application, first, the local cluster receives the service request sent by the client, and then the local cluster determines whether there is a disaster recovery identifier in the header information of the service request, and redirects the service request to the remote cluster when the disaster recovery identifier is present, so that the remote cluster processes and responds to the service request. The disaster recovery identifier is used to indicate that the service request needs to be redirected to the remote cluster. Compared with the problem of lack of flexibility and coarse granularity when implementing cross-cluster disaster recovery strategies using service grid technology in the prior art, the present application attaches a disaster recovery identifier to the service request, so that the local cluster can identify the disaster recovery identifier in the header information of the service request, and can intelligently and efficiently redirect the service request to the remote cluster for processing, thereby ensuring the continuity of the service request and the timeliness of the response in the event of a local service failure or network isolation, realizing selective disaster recovery switching for specific services or traffic, supporting the directional transfer of some service traffic, more accurately matching the needs of disaster recovery scenarios, improving the flexibility and adaptability of fault handling, and ensuring the dynamic routing adjustment capability when the service status and network conditions change.

[0099] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A cross-cluster disaster recovery method based on service grid, characterized in that: include: The local cluster receives the service request sent by the client, where the local cluster represents the Kubernetes cluster where the environment where the service request is initially initiated is located; The local cluster determines whether there is a disaster recovery identifier in the header information of the service request, and if the disaster recovery identifier exists in the header information, redirects the service request to the remote cluster so that the remote cluster processes and responds to the service request. The disaster recovery identifier is used to indicate that the service request needs to be redirected to the remote cluster. The remote cluster represents a Kubernetes cluster that is different from the local cluster and is located in other geographical locations or network environments.

2. The cross-cluster disaster recovery method based on service grid according to claim 1, characterized in that: The method further comprises: When the disaster recovery identifier does not exist in the header information, the service request is directed to the service production end of the local cluster, so that the service instance in the local cluster processes and responds to the service request.

3. The cross-cluster disaster recovery method based on service grid according to claim 1, characterized in that: There are multiple remote clusters, and redirecting the service request to the remote cluster includes: Determine the remote cluster corresponding to the identification value of the disaster recovery identifier as the target remote cluster, where the identification value corresponds one-to-one to the remote cluster; Redirect the service request to the service instance in the target remote cluster.

4. The cross-cluster disaster recovery method based on service grid according to claim 3 is characterized in that: Determining whether a disaster recovery identifier is present in the header information of the service request, and redirecting the service request to a remote cluster if the disaster recovery identifier is present in the header information, including: Based on the virtual service pre-configured in the local cluster, by checking the field of the header information in the service request, determining whether the header information contains the disaster recovery identifier, and if the disaster recovery identifier exists, determining that the remote cluster corresponding to the identifier value is the target remote cluster; Based on the pre-configured service export in the local cluster, the service request is redirected to the target remote cluster, and the service export defines the entry gateway of the target remote cluster.

5. The cross-cluster disaster recovery method based on service grid according to claim 1, characterized in that: Determining whether a disaster recovery identifier is present in the header information of the service request, and redirecting the service request to a remote cluster if the disaster recovery identifier is present in the header information, including: Based on the pre-modified Envoy source code in the local cluster, determine whether the disaster recovery identifier exists in the header information, and if the disaster recovery identifier exists, determine that the remote cluster corresponding to the identifier value of the disaster recovery identifier is the target remote cluster; Based on the service export in the local cluster, the service request is redirected to the target remote cluster, and the service export defines the entry gateway of the target remote cluster.

6. The cross-cluster disaster recovery method based on service grid according to claim 5, characterized in that: Determining whether the disaster recovery identifier exists in the header information based on the pre-modified Envoy source code in the local cluster includes: Based on the decodeHeaders function pre-modified in the Envoy source code, it is determined whether the disaster recovery identifier exists by checking the fields of the header information in the service request.

7. The cross-cluster disaster recovery method based on service grid according to claim 1, characterized in that: Receive service requests sent by clients, including: Receive the service request, where the service request is generated after the disaster recovery identifier is added to the initial service request when the client sends the initial service request to the local cluster within a preset number of times and has never received a request response from the local cluster; or, where the service request is generated without modifying the initial service request when the client sends the initial service request to the local cluster within the preset number of times and receives the request response from the local cluster, and the disaster recovery identifier does not exist in the initial service request.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the cross-cluster disaster recovery method based on the service grid according to any one of claims 1 to 7.

9. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the cross-cluster disaster recovery method based on a service grid according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for executing the service grid-based cross-cluster disaster recovery method described in any one of claims 1 to 7.