Liveness detection method, service system, and service node

The subsystem activation is carried out through existing resources in the service system, which solves the problem of detection failure caused by detection center failure, and realizes efficient and low-dependence subsystem survival status detection.

WO2025130188A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/118105
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-27
Filing Date
2024-09-11
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the prior art, a failure of the detection center will lead to the inability to detect the active subsystem in time, affecting the performance of the service system.

Method used

Through existing resources in the service system (i.e., the service node in the first subsystem), a detection request is sent to the anchor point, the subsystem survival status is determined, and the subsystem survival status is judged through the statistics of the detection result between the service nodes, without an external detection center.

Benefits of technology

The exploration of multiple subsystems in the service system is realized, which avoids dependence on external resources, reduces deployment difficulty and resource waste, and reduces network latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024118105_26062025_PF_FP_ABST
    Figure CN2024118105_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a liveness detection method, a service system, and a service node. The service system comprises a first subsystem and a second subsystem, wherein a plurality of service nodes in the first subsystem are standby nodes of each other, a plurality of service nodes in the second subsystem are standby nodes of each other, and the first subsystem is used for: after determining that the first subsystem itself is in a liveness state, detecting, by means of at least one service node in the first subsystem, whether the second subsystem is reachable, when it is detected in the at least one service node that the number of reachable nodes in the second subsystem reaches a first threshold, or when the number of times that the at least one service node detects the second subsystem as reachable reaches a second threshold, determining that the second subsystem is in the liveness state, and when it is detected in the at least one service node that the number of reachable nodes in the second subsystem does not reach the first threshold, or when the number of times that the at least one service node detects the second subsystem as reachable does not reach the second threshold, determining that the second subsystem is in a non-liveness state.
Need to check novelty before this filing date? Find Prior Art

Description

Detection method, service system and service node

[0001] This application claims priority to the Chinese patent application with application number 202311759661.8 filed with the State Intellectual Property Office of China on December 20, 2023, and priority to the Chinese patent application with invention name “Node activation method, device and computer-readable storage medium”, as well as priority to the Chinese patent application with application number 202410217004.9 filed with the State Intellectual Property Office of China on February 27, 2024, and priority to the Chinese patent application with invention name “Activation method, service system and service node”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer network technology, and in particular to a detection method, a service system, and a service node. Background Art

[0003] In recent years, with the continuous development of Internet technology, the scale of service systems provided by enterprises to users has become larger and larger. In order to improve the reliability, stability and operational performance of the services provided to users, enterprises usually divide the entire service system into multiple subsystems responsible for providing different services. Multiple subsystems are deployed independently, and all service nodes in each subsystem (also called service devices, computing devices) have equal status and roles. They can provide and request services, synchronize data, and serve as backup nodes for each other. When some service nodes fail, the remaining service nodes can still provide services to users normally, playing a disaster recovery role.

[0004] When an enterprise has multiple subsystems, it needs to regularly detect the subsystems to see if they are alive. This means detecting whether each subsystem is alive, so that when a subsystem is determined to be non-alive, it can locate the fault in the subsystem and take timely measures to repair it.

[0005] Currently, enterprises typically perform liveness checks on multiple subsystems by setting up a dedicated detection center, which periodically sends detection requests to each subsystem. If the center receives a response from the subsystem, the subsystem is considered alive; otherwise, it is considered dead. However, if the detection center fails, it will be unable to continue performing liveness checks on the subsystems. This means that if a subsystem fails and becomes dead, it cannot be detected in time, impacting the performance of the entire service system.

[0006] Summary of the Invention

[0007] The present application provides an activity detection method, a service system, and a service node, which can detect the activity of multiple subsystems in a service system.

[0008] In a first aspect, a liveness detection method is provided, which can be applied to a service system. The service system includes a first subsystem and a second subsystem. Multiple service nodes in the first subsystem are backup nodes for each other, and multiple service nodes in the second subsystem are backup nodes for each other. The method may specifically include the following steps: the first subsystem sends a first detection request to an anchor point outside the first subsystem and the second subsystem to determine whether the first subsystem is in a live state. Thereafter, at least one service node in the first subsystem detects whether the second subsystem is reachable. If the number of nodes reachable from the second subsystem detected by the first subsystem in at least one service node reaches a first threshold, the first subsystem determines that the second subsystem is in a live state. If the number of nodes reachable from the second subsystem detected by the first subsystem in at least one service node does not reach the first threshold, the first subsystem determines that the second subsystem is in a non-live state. Alternatively, if the number of times at least one service node detects that the second subsystem is reachable reaches a second threshold, the first subsystem determines that the second subsystem is in a live state. If the number of times at least one service node detects that the second subsystem is reachable does not reach the second threshold, the first subsystem determines that the second subsystem is in a non-live state.

[0009] The anchor point is a public resource outside the first subsystem and the second subsystem, and specifically may be a node, server, or other network device with a public Internet protocol (IP) address.

[0010] In the above scheme, the remaining subsystems in the service system (such as the second subsystem mentioned above) can be detected through the existing resources in the service system (that is, the first subsystem in the above service system) without the need to introduce external resources (such as a dedicated detection center). This not only solves the problem of the existing detection method that if the detection center fails, it will be unable to continue to perform the detection task of the subsystem, but also avoids the service system's dependence on external resources, reduces the difficulty of deploying the service system, and reduces resource waste.

[0011] In addition, when the first subsystem is waiting to transmit data with the second subsystem, the first subsystem can directly know whether the second subsystem is in a survival state after executing the above scheme. When it is known that the second subsystem is in a survival state, it immediately transmits data with the second subsystem. When it is known that the second subsystem is in a non-survival state, unnecessary data transmission processes are avoided. Unlike the existing detection method, when the first subsystem is waiting to transmit data with the second subsystem, the first subsystem needs to first access the detection center, and the detection center detects whether the second subsystem is in a survival state and returns the detection result to the first subsystem. This can reduce the network delay between the first subsystem and the second subsystem. Moreover, it can avoid the situation where the detection center fails to detect whether the second subsystem is in a survival state. If the second subsystem is actually in a survival state, the first subsystem delays data transmission with the second subsystem because it cannot know whether the second subsystem is in a survival state, resulting in poor performance of the service system.

[0012] In some possible implementations, the first subsystem may specifically determine that the first subsystem is alive in the following manner:

[0013] When the first subsystem sends a first detection request to an anchor point outside the first subsystem and the second subsystem, it starts recording a first time. When the first time does not reach a first time threshold and a response is received from the anchor point based on the first detection request, it is determined that the first subsystem is in a survival state.

[0014] In some possible implementations, at least one service node in the first subsystem detects whether the second subsystem is reachable, including: each service node in the at least one service node in the first subsystem starts recording a second time when sending a second detection request to the second subsystem; when the second time does not reach a second time threshold and a response returned by the second subsystem based on the second detection request is received, it is determined that the second subsystem is reachable; when the second time reaches the second time threshold and no response returned by the second subsystem based on the second detection request is received, it is determined that the second subsystem is unreachable.

[0015] In some possible implementations, after at least one service node in the first subsystem detects whether the second subsystem is reachable, the above method also includes the following steps: each service node of the at least one service node in the first subsystem records the detection result of whether the second subsystem is reachable to the first service node, and the first service node is any one of the at least one service node; the first subsystem counts the number of nodes in the at least one service node that detect that the second subsystem is reachable, or counts the number of times at least one service node detects that the second subsystem is reachable based on the detection result on the first service node.

[0016] In some possible implementations, the above method also includes the following steps: each service node of at least one service node in the first subsystem detects whether the second service node is reachable, and the second service node is any service node in the first subsystem except at least one service node; the first subsystem determines that the second service node is in a surviving state when at least one service node detects that the number of nodes reachable by the second subsystem reaches a third threshold, and determines that the second service node is in a non-surviving state when at least one service node detects that the number of nodes reachable by the second service node does not reach the third threshold, or the first subsystem determines that the second service node is in a surviving state when at least one service node detects that the number of times the second service node is reachable reaches a fourth threshold, and determines that the second service node is in a non-surviving state when at least one service node detects that the number of times the second service node is reachable does not reach the fourth threshold.

[0017] In the above implementation method, the second service node in the first subsystem can be detected through the existing resources in the service system (i.e., at least one service node in the first subsystem in the above service system). There is no need to introduce external resources (such as a dedicated detection center) to detect the second service node in the first subsystem. This can avoid the service system's dependence on external resources, reduce the difficulty of deploying the service system, and reduce resource waste.

[0018] In some possible implementations, the first subsystem and the second subsystem serve as backup subsystems for each other, that is, the first subsystem and the second subsystem form an active-active disaster recovery system.

[0019] In some possible implementations, the first subsystem and the second subsystem are deployed in different availability zones (AZs).

[0020] In a second aspect, a service system is provided, comprising a first subsystem and a second subsystem, wherein a plurality of service nodes in the first subsystem are each other's backup nodes, and a plurality of service nodes in the second subsystem are each other's backup nodes;

[0021] The first subsystem is configured to send a first detection request to an anchor point outside the first subsystem and the second subsystem to determine whether the first subsystem is alive;

[0022] The first subsystem is configured to detect whether the second subsystem is reachable through at least one service node in the first subsystem;

[0023] the first subsystem is configured to determine that the second subsystem is in a live state if it is detected in the at least one service node that the number of nodes reachable by the second subsystem reaches a first threshold, and to determine that the second subsystem is in a non-live state if it is detected in the at least one service node that the number of nodes reachable by the second subsystem does not reach the first threshold;

[0024] Alternatively, the first subsystem is configured to determine that the second subsystem is in a live state when the at least one service node detects that the second subsystem is reachable a number of times that reaches a second threshold, and to determine that the second subsystem is in a non-live state when the at least one service node detects that the second subsystem is reachable a number of times that does not reach the second threshold.

[0025] In some possible implementations, the first subsystem is configured to start recording a first time when sending a first detection request to an anchor point outside the first subsystem and the second subsystem, and determine that the first subsystem is in a survival state when the first time does not reach a first time threshold and a response returned by the anchor point based on the first detection request is received.

[0026] In some possible implementations, each service node of at least one service node in the first subsystem is configured to start recording a second time when sending a second probe request to the second subsystem, and determine that the second subsystem is reachable when the second time does not reach a second time threshold and a response returned by the second subsystem based on the second probe request is received; and determine that the second subsystem is unreachable when the second time reaches the second time threshold and no response returned by the second subsystem based on the second probe request is received.

[0027] In some possible implementations, after at least one service node in the first subsystem detects whether the second subsystem is reachable, each service node in the at least one service node is used to record the detection result of detecting whether the second subsystem is reachable to the first service node, and the first service node is any one of the at least one service node; the first subsystem is used to count the number of nodes in the at least one service node that detect that the second subsystem is reachable based on the detection result on the first service node, or count the number of times that the at least one service node detects that the second subsystem is reachable.

[0028] In some possible implementations, each service node in at least one service node in the first subsystem is used to detect whether a second service node is reachable, and the second service node is any service node in the first subsystem except the at least one service node; the first subsystem is used to determine that the second service node is in a surviving state when the at least one service node detects that the number of nodes reachable by the second subsystem reaches a third threshold, and to determine that the second service node is in a non-surviving state when the at least one service node detects that the number of nodes reachable by the second service node does not reach the third threshold; or, the first subsystem is used to determine that the second service node is in a surviving state when the at least one service node detects that the number of times the second service node is reachable reaches a fourth threshold, and to determine that the second service node is in a non-surviving state when the at least one service node detects that the second service node is reachable does not reach the fourth threshold.

[0029] In some possible implementations, the first subsystem and the second subsystem are backup subsystems for each other.

[0030] In some possible implementations, the first subsystem and the second subsystem are deployed in different AZs.

[0031] In a third aspect, a service node is provided, which includes modules for executing the method provided by the first aspect or any possible implementation of the first aspect, which is executed by each service node in at least one service node in the first subsystem.

[0032] In a fourth aspect, a service node is provided, which includes a processor and a memory, the memory being used to store instructions. When the edge node is running, the processor executes the instructions to implement the method provided in the first aspect or any possible implementation of the first aspect as performed by each service node in the at least one service node in the first subsystem.

[0033] In a fifth aspect, a node cluster is provided, which is the first subsystem in the service system provided by the second aspect or any possible implementation of the second aspect.

[0034] In a sixth aspect, a computer-readable storage medium is provided, which stores instructions. When the computer-readable storage medium is run on a computer, the computer executes the method provided by the first aspect or any possible implementation of the first aspect, which is performed by the first edge node, the second edge node, or the third edge node. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] FIG1 is a schematic diagram of the structure of a cloud service system provided in an embodiment of the present application;

[0036] FIG2 is a schematic diagram of a deployment method of a cloud service system provided in an embodiment of the present application;

[0037] FIG3 is a schematic diagram of the structure of a cloud service system with a “multi-active peer-to-peer” architecture provided in an embodiment of the present application;

[0038] FIG4 is a flow chart of a detection method according to an embodiment of the present application;

[0039] FIG5 is a flow chart of another detection method provided in an embodiment of the present application;

[0040] FIG6 is a flow chart of another detection method provided in an embodiment of the present application;

[0041] FIG7 is a flow chart of another detection method provided in an embodiment of the present application;

[0042] FIG8 is a flow chart of another detection method provided in an embodiment of the present application;

[0043] 9 is a schematic diagram of a flow chart of a first subsystem detecting whether a second service node is alive, provided in an embodiment of the present application;

[0044] 10 is a schematic diagram of another process of the first subsystem detecting whether the second service node is alive according to an embodiment of the present application;

[0045] FIG11 is a schematic diagram of the structure of a service node provided in an embodiment of the present application;

[0046] FIG12 is a schematic diagram of the structure of another service node provided in an embodiment of the present application. DETAILED DESCRIPTION

[0047] First, the service system involved in this application is introduced in detail.

[0048] Taking the cloud service system 10 shown in Figure 1 as an example, as shown in Figure 1, the cloud service system 10 includes multiple subsystems: a registration service subsystem 101, a login service subsystem 102, an artificial intelligence service subsystem 103, a storage service subsystem 104, a computing service subsystem 105...

[0049] The multiple subsystems in the cloud service system 10 are used to provide different services to users, such as the registration service subsystem 101 for providing registration services to users, the login service subsystem 102 for providing login services to users, the artificial intelligence service subsystem 103 for providing users with artificial intelligence type services, which may include one or more of image recognition services, artificial intelligence model training services, text recognition services, natural language translation services, etc. The storage service subsystem 104 is used to provide users with storage type services, which may include but are not limited to object storage services, cloud hard disk rental services, cloud backup services, cloud hard disk backup services, etc. One or more, the computing service subsystem 105 is used to provide users with computing type services, which may include bare metal server (BMS) rental services, virtual machine (virtual One or more of the following: cloud backup service refers to a service that backs up user data to the cloud for storage; cloud hard disk backup service refers to a service that backs up data in the user's cloud hard disk; BMS refers to a general physical server, such as an ARM server or an X86 server; a virtual machine refers to a complete computer system with complete hardware system functions that runs in a completely isolated environment and simulated by software. Any work that can be done in a physical computer can be done in a virtual machine. When creating a virtual machine in a computing device, part of the hard disk and memory capacity of the physical machine needs to be used as the hard disk and memory capacity of the virtual machine. Each virtual machine has an independent basic input / output system (BIOS), hard disk and operating system, and can be operated on the virtual machine like a physical machine; a container is a portable software unit that can combine an application and all its dependencies into a software package that is not restricted by the underlying host operating system. This eliminates the need to build a complex environment, simplifying the process from application development to deployment.

[0050] Each subsystem in the cloud service system 10 may include multiple service nodes, which may also be referred to as multiple service devices or multiple computing devices, wherein the service node may be a BMS, such as an ARM server or an X86 server. All service nodes in each subsystem have equal status and role, can provide and request services, and can synchronize data, serve as backup nodes for each other, so that when some service nodes fail, the remaining service nodes can still provide services to users normally, thereby playing a role in disaster recovery. For example, all service nodes in the registration service subsystem 101 can provide registration services to users and can synchronize data, serve as backup nodes for each other. Since all service nodes in each subsystem have equal status and role, these nodes can be referred to as peer nodes, and these subsystems can be referred to as peer groups or peer-to-peer networks. For example, the registration service subsystem 101 is a peer group, and the login service subsystem 102 is a peer group.

[0051] The multiple subsystems in the cloud service system 10 can be deployed in the same region or distributed across different regions. Furthermore, the multiple subsystems can be distributed across the same availability zone (AZ) or across different AZs, with each AZ including one data center or multiple geographically close data centers. Typically, a region can include multiple AZs.

[0052] Taking the example of multiple subsystems in cloud service system 10 distributed across different AZs, as shown in Figure 2, registration service subsystem 101 is deployed in AZ1, login service subsystem 102 is deployed in AZ2, artificial intelligence service subsystem 103 is deployed in AZ3, storage service subsystem 104 is deployed in AZ4, and computing service subsystem 105 is deployed in AZ5. It can be understood that the independent deployment of multiple subsystems allows them to operate independently without affecting each other. Even if one subsystem fails, the remaining subsystems can still provide normal services to users.

[0053] There is a communication connection between the multiple subsystems in the cloud service system 10. In a specific implementation, the communication connection can be a wired network connection or a wireless network connection, which is not specifically limited in this application.

[0054] It should be understood that Figures 1 and 2 are merely examples of cloud service systems. For example, in actual applications, the number of subsystems in a cloud service system can be any number. For example, each subsystem in the cloud service system can provide multiple services to users. For another example, the cloud service system also includes network devices for forwarding communication data between multiple subsystems, such as base stations, routers, switches, etc. For another example, the cloud service system is a "multi-active peer-to-peer" architecture system, where the "multi-active peer-to-peer" architecture system refers to a system that includes multiple subsystems with the same functions. Multiple subsystems need to be in a live state (also called a normal state or an available state) at the same time, can process requests and provide services at the same time, and multiple subsystems can communicate and synchronize data with each other and serve as backups for each other. When one subsystem fails and becomes unavailable, another subsystem can take over its work to ensure the normal operation of the entire system. In addition, multiple service nodes in each subsystem have equal status and roles and can provide and request services. When some service nodes fail and become unavailable, the remaining service nodes can take over their work to ensure the normal operation of the entire subsystem. This architecture is designed to improve the availability, fault tolerance and performance of the system.

[0055] Refer to Figure 3, which is a structural diagram of a cloud service system with a "multi-active peer-to-peer" architecture provided in an embodiment of the present application. As shown in Figure 3, the cloud service system with a "multi-active peer-to-peer" architecture includes two subsystems: cloud service subsystem 310 and cloud service subsystem 320.

[0056] The cloud service subsystem 310 can provide users with one or more services including registration services, login services, artificial intelligence services, storage services, computing services, etc. Among them, artificial intelligence services may include one or more image recognition services, artificial intelligence model training services, text recognition services, natural language translation services, etc. Storage services may include but are not limited to object storage services, cloud hard disk rental services, cloud backup services, cloud hard disk backup services, etc. Computing services may include one or more BMS rental services, VM rental services, container rental services, etc.

[0057] The cloud service subsystem 310 may include multiple service nodes, which may be BMSs, such as ARM servers or X86 servers. Multiple service nodes have equal status and roles, and can provide and request services. Multiple service nodes can synchronize data and serve as backup nodes for each other. For example, when the cloud service subsystem 310 provides registration services, login services, artificial intelligence services, storage services, and computing services, each service node in the cloud service subsystem 310 can provide registration services, login services, artificial intelligence services, storage services, and computing services.

[0058] The cloud service subsystem 320 and the cloud service subsystem 310 serve as backup systems for each other, so that when one of the subsystems fails, the other subsystem can still provide services to users normally, thereby playing a disaster recovery role.

[0059] Cloud service subsystem 310 and cloud service subsystem 320 can be deployed in the same region or distributed across different regions. Furthermore, cloud service subsystem 310 and cloud service subsystem 320 can be distributed across different AZs. A failure in one AZ only affects the cloud service subsystem deployed in that AZ, without affecting other AZs or cloud service subsystems deployed in other AZs. In Figure 3, cloud service subsystem 310 is deployed in AZ001 and cloud service subsystem 320 is deployed in AZ002.

[0060] There is a communication connection between the cloud service subsystem 310 and the cloud service subsystem 320. In a specific implementation, the communication connection can be a wired network connection or a wireless network connection, which is not specifically limited in this application.

[0061] It should be understood that FIG3 is merely an example of a cloud service system with a “multi-active peer-to-peer” architecture. For example, in actual applications, the cloud service system may include more cloud service subsystems with the same functions as the cloud service subsystem 310 .

[0062] It can be understood that in addition to the above-mentioned cloud service system scenarios, there may be other service system scenarios. For example, an enterprise providing live broadcast services may divide the entire live broadcast service system into a subsystem that provides registration services for users, a subsystem that provides login services for users, a subsystem that provides video upload services for anchors, a subsystem that provides video viewing services for viewers, etc., and the multiple service nodes in each subsystem are each other's backup nodes, or the enterprise deploys the entire live broadcast service system into multiple subsystems of a "multi-active peer-to-peer" architecture. For example, an enterprise providing shopping services may divide the entire shopping service system into a subsystem that provides user registration services, a subsystem that provides login services for users, a subsystem that provides commodity trading services for users, a subsystem that provides logistics query services for users, a subsystem that provides return and exchange and commodity evaluation services for users, etc., and the multiple service nodes in each subsystem are each other's backup nodes, or the enterprise deploys the entire shopping service system into multiple subsystems of a "multi-active peer-to-peer" architecture.

[0063] When an enterprise has multiple subsystems, it needs to regularly detect the subsystems to see if they are alive. This means detecting whether each subsystem is alive, so that when a subsystem is determined to be non-alive, it can locate the fault in the subsystem and take timely measures to repair it.

[0064] Currently, enterprises typically perform liveness checks on multiple subsystems by setting up a dedicated detection center, which periodically sends detection requests to each subsystem. If the center receives a response from the subsystem, the subsystem is considered alive; otherwise, it is considered dead. However, if the detection center fails, it will be unable to continue performing liveness checks on the subsystems. This means that if a subsystem fails and becomes dead, it cannot be detected in time, impacting the performance of the entire service system.

[0065] In response to the above problems, the present application provides a detection method and a service system, etc., which can realize the detection tasks of multiple subsystems in the service system through the existing resources of the service system (i.e., the subsystems in the service system) without the introduction of external resources (such as a dedicated detection center). Not only can the above problems be solved, but also the service system can be prevented from relying on external resources, reducing the deployment difficulty of the service system, reducing resource waste, and also allowing the subsystems to directly obtain each other's survival status when they need to, without the need to obtain it indirectly through the detection center, thereby reducing the network delay between the subsystems. Next, the detection method and service system provided by the present application are described in detail in conjunction with the corresponding drawings.

[0066] Referring to FIG4 , FIG4 is a flow chart of a detection method provided in an embodiment of the present application, as shown in FIG4 , including the following steps:

[0067] S401: The first subsystem in the service system sends a message to an anchor point outside the first subsystem and the second subsystem.

[0068] A first detection request is sent to determine whether the first subsystem is alive.

[0069] The service system can be the cloud service system shown in Figures 1, 2, and 3, or it can be other service system scenarios, such as a live broadcast service system. This system can include any multiple subsystems such as a subsystem that provides registration services for users, a subsystem that provides login services for users, a subsystem that provides video upload services for anchors, a subsystem that provides video viewing services for viewers, and the like, and the multiple service nodes in each subsystem are each other's backup nodes. Alternatively, the live broadcast service system is a "multi-active peer-to-peer" architecture system. Another example is a shopping service system. This system can include any multiple subsystems such as a subsystem that provides registration services for users, a subsystem that provides login services for users, a subsystem that provides commodity trading services for users, a subsystem that provides logistics query services for users, a subsystem that provides return and exchange and commodity evaluation services for users, and the like, and the multiple service nodes in each subsystem are each other's backup nodes. Alternatively, the shopping service system is a "multi-active peer-to-peer" architecture system. For more information about the "multi-active peer-to-peer" architecture system, please refer to the relevant description in Figure 3, which will not be elaborated here.

[0070] As can be seen from the description of the service system above, the multiple subsystems it comprises are typically deployed independently and do not affect each other. Therefore, the likelihood of multiple subsystems failing simultaneously is extremely small and can be ignored. In other words, typically, some of the subsystems in a service system are alive while others are not, or all of them are alive.

[0071] The first subsystem may be any one of the multiple subsystems included in the service system, and the second subsystem may be any one of the multiple subsystems included in the service system except the first subsystem.

[0072] The anchor point is a public resource outside the first subsystem and the second subsystem, and may be a node, server, or other network device with a public Internet protocol (IP) address.

[0073] It can be understood that if the first subsystem itself is normal and the communication network between the first subsystem and the outside world is normal, then the first subsystem is in a surviving state, and the first subsystem can detect whether the second subsystem is in a surviving state. If the first subsystem itself is normal, but the communication network between the first subsystem and the outside world is abnormal (such as the cable for communication between the computer room where the first subsystem is deployed and the outside world is dug up), then the first subsystem is in a non-surviving state, and the first subsystem cannot detect whether the second subsystem is in a surviving state. If the first subsystem itself is abnormal (such as the power outage in the computer room where the first subsystem is deployed), then the first subsystem is in a non-surviving state, and the first subsystem cannot detect whether the second subsystem is in a surviving state. Therefore, it can be understood that if the first subsystem wants to be able to detect whether the second subsystem is in a surviving state, the first subsystem itself needs to be in a surviving state. In the present application, the first subsystem can determine whether it is in a surviving state by sending a first detection request to an anchor point outside the first subsystem and the second subsystem.

[0074] The specific implementation method of the first subsystem sending a first detection request to the anchor point to determine whether it is in a survival state is as follows: when the first subsystem sends the first detection request to the anchor point, it starts to record the first time. When the first time does not reach the first time threshold and the response returned by the anchor point based on the first detection request is received, it is determined that the first subsystem is in a survival state. When the first time reaches the first time threshold and the response returned by the anchor point based on the first detection request is not received, it is determined that the first subsystem is in a non-survival state.

[0075] More specifically, each service node in the first subsystem may send a first detection request to the anchor point, or take turns sending the first detection request to the anchor point, and start recording the first time when sending the first detection request. When the first time does not reach the first time threshold and receives a response returned by the anchor point based on the first detection request, it determines that it is in a survival state. When the first time reaches the first time threshold and does not receive a response returned by the anchor point based on the first detection request, it determines that it is in a non-survival state. When the first subsystem has a surviving service node in itself, it determines that the first subsystem is in a survival state. Otherwise, it determines that the first subsystem is in a non-survival state.

[0076] The above-mentioned first time threshold can be customized according to the actual scenario and is not specifically limited in this application.

[0077] S402: Each service node of the at least one service node in the first subsystem detects whether the second subsystem is reachable.

[0078] The at least one service node in the first subsystem is a service node that has determined to be alive among the multiple service nodes included in the first subsystem. The process by which the at least one service node determines whether it is alive can be described in S401, where each service node in the first subsystem sends a first probe request to an anchor point outside the first and second subsystems to determine whether it is alive. For the sake of brevity, this description is not further elaborated here.

[0079] The following describes in detail the process of each of the at least one service node in the first subsystem detecting whether the second subsystem is reachable.

[0080] Taking the example of the i-th service node in at least one service node detecting whether the second subsystem is reachable, the way for each service node in at least one service node to detect whether the second subsystem is reachable can be: the i-th service node periodically sends a second detection request to the second subsystem, which requires the second subsystem to return a response, and starts recording the second time each time the second detection request is sent. If the second time threshold is not reached in the second time and a response returned by the second subsystem based on the second detection request is received, it is determined that the second subsystem is reachable; otherwise, it is determined that the second subsystem is unreachable.

[0081] The period of the second detection request sent by the i-th service node to the second subsystem and requiring the second subsystem to return a response, as well as the second time threshold, can be set according to the actual scenario. For example, when the period of the second detection request sent by the i-th service node to the second subsystem and requiring the second subsystem to return a response is 5 seconds, the second time threshold can be 5 seconds, optionally, 6 seconds or 7 seconds, etc. When the period of the second detection request sent by the i-th service node to the second subsystem and requiring the second subsystem to return a response is 10 seconds, the second time threshold can be 10 seconds, optionally, 11 seconds or 12 seconds, etc., and this application does not make specific restrictions on this. The period of the second detection request sent by at least one service node to the second subsystem and requiring the second subsystem to return a response can be the same or different, and the second time threshold corresponding to at least one service node can be the same or different, and this application does not make specific restrictions on this.

[0082] S403: Each service node of the at least one service node in the first subsystem records the detection result to a first service node, where the first service node is any one of the at least one service node.

[0083] It can be understood that, since at least one service node in the first subsystem is in a live state, the first service node is any one of the at least one service node, and therefore, the first service node is in a live state.

[0084] Each service node in at least one service node records the detection result to the first service node. Taking the detection result recorded by the i-th service node in at least one service node to the first service node as an example, the detection result recorded by the i-th service node may include the identifier of the i-th service node (such as the ID, IP address, etc. of the i-th service node), the identifier of the second subsystem (such as the ID, IP address of the second subsystem), and information indicating whether the second subsystem is reachable. The order of the identifier of the i-th service node, the identifier of the second subsystem, and the information indicating whether the second subsystem is reachable in the detection result can be the identifier of the i-th service node, the identifier of the second subsystem, and the information indicating whether the second subsystem is reachable, or the identifier of the second subsystem, the identifier of the i-th service node, the information indicating whether the second subsystem is reachable, or the order of the identifier of the i-th service node, the information indicating whether the second subsystem is reachable, the identifier of the second subsystem, etc. This application does not make any specific restrictions on this.

[0085] For example, if the ID of the i-th service node is 0001 and the ID of the second subsystem is 200, the detection results can be in the following forms:

[0086] Form ①:

[0087] When the i-th service node detects that the second subsystem is reachable, the detection result can be "ServerNode ID: 0001, Test Server System ID: 200, value: 1", "Test ServerNode ID: 200, value: 1, ServerNode ID: 0001", "Test ServerNode ID: 200, ServerNode ID: 0001, value: 1", etc. "value: 1" means that the service node with ID 0001 determines that the second subsystem with ID 200 is reachable.

[0088] When the i-th service node detects that the second subsystem is unreachable, the detection result can be "ServerNode ID: 0001, Test Server System ID: 200, value: 0", "Test ServerNode ID: 200, value: 0, ServerNode ID: 0001", "Test ServerNode ID: 200, ServerNode ID: 0001, value: 0", etc. "value: 0" means that the service node with ID 0001 determines that the second subsystem with ID 200 is unreachable.

[0089] Form ②:

[0090] When the i-th service node detects that the second subsystem is reachable, the detection result can be "ServerNode ID: 0001, Test ServerNode ID: 200, true", "Test ServerNode ID: 200, true, ServerNode ID: 0001", "Test ServerNode ID: 200, ServerNode ID: 0001, true", etc. "true" means that the service node with ID 0001 determines that the second subsystem with ID 200 is reachable.

[0091] When the i-th service node detects that the second subsystem is unreachable, the detection result can be "ServerNode ID: 0001, Test Server System ID: 200, false", "Test ServerNode ID: 200, false, ServerNode ID: 0001", "Test ServerNode ID: 200, ServerNode ID: 0001, false", etc. "false" means that the service node with ID 0001 determines that the second subsystem with ID 200 is unreachable.

[0092] Form ③: When the i-th service node detects that the second subsystem is reachable, the detection result can be "ServerNode ID: 0001, Test ServerNode ID: 200, √", "Test ServerNode ID: 200, √, ServerNode ID: 0001", "Test ServerNode ID: 200, ServerNode ID: 0001, √", etc. "√" means that the service node with ID 0001 determines that the second subsystem with ID 200 is reachable.

[0093] When the i-th service node detects that the second subsystem is unreachable, the detection result can be "ServerNode ID: 0001, Test Server System ID: 200, ×", "Test ServerNode ID: 200, ×, ServerNode ID: 0001", "Test ServerNode ID: 200, ServerNode ID: 0001, ×", etc. "×" means that the service node with ID 0001 determines that the second subsystem with ID 200 is unreachable.

[0094] It should be understood that the identifier of the i-th service node, the identifier of the second subsystem, and the form of the detection result are merely examples and should not be considered as specific limitations. For example, the detection result may also include the time at which the i-th service node detects whether the second subsystem is reachable. This time may be the time at which the i-th service node sends a second detection request to the second subsystem, or the time at which the i-th service node determines whether the second subsystem is reachable.

[0095] S404: The first subsystem counts the number of nodes detected to be reachable by the second subsystem in at least one service node according to the detection result on the first service node.

[0096] The specific manner in which the first subsystem counts the number of nodes detected as reachable by the second subsystem in at least one service node based on the detection result on the first service node includes but is not limited to the following two methods:

[0097] Method (1): The first service node in the first subsystem counts the number of nodes detected as reachable from the second subsystem in at least one service node based on the detection result on the first service node.

[0098] Specifically, the first service node may count the number of nodes reachable by the second subsystem detected in at least one service node based on the service node identifier included in the detection result recorded on the first service node each time a service node records a detection result to the first service node.

[0099] Method (2): When each service node in at least one service node detects whether the second subsystem is reachable and records the detection result to the first service node, the number of nodes in at least one service node that detect that the second subsystem is reachable is counted based on the detection result recorded on the first service node.

[0100] Specifically, each service node may count the number of nodes detected to be reachable by the second subsystem in at least one service node according to the service node identifier included in the detection result recorded on the first service node.

[0101] S405: The first subsystem determines whether the counted number of nodes reaches a first threshold. If it is determined that the number of nodes reaches the first threshold, S406 is executed. If it is determined that the number of nodes does not reach the first threshold, S407 is executed.

[0102] The first threshold can be customized according to the actual scenario, such as being set according to the total number of service nodes in the first subsystem, such as being set to 1 / 5, 1 / 4, 1 / 3, etc. of the total number of service nodes in the first subsystem. It can be understood that the larger the first threshold is set, the higher the accuracy of the first subsystem in determining whether the second subsystem is in a survival state based on the first threshold.

[0103] Corresponding to the method (1) in S404, the method in which the first subsystem determines whether the number of nodes counted has reached the first threshold is as follows: ① The first service node determines whether the number of nodes counted has reached the first threshold.

[0104] Corresponding to the method (2) in S404, the method for the first subsystem to determine whether the number of nodes counted has reached the first threshold is ②: each service node determines whether the number of nodes counted has reached the first threshold.

[0105] S406: The first subsystem determines that the second subsystem is in a live state.

[0106] Corresponding to method ① in S405, the first subsystem determines that the second subsystem is in a survival state as follows: the first service node determines that the second subsystem is in a survival state, and then the first service node synchronizes the information that the second subsystem is in a survival state to the remaining service nodes in the first subsystem.

[0107] Corresponding to method ② in S405, the way in which the first subsystem determines that the second subsystem is in a survival state is: when there is a service node in at least one service node that determines that the second subsystem is in a survival state, the service node synchronizes the information that the second subsystem is in a survival state to the remaining service nodes in the first subsystem.

[0108] S407: The first subsystem determines that the second subsystem is in a non-survival state.

[0109] Corresponding to method ① in S405, the way the first subsystem determines that the second subsystem is in a non-survival state is: the first service node determines that the second subsystem is in a non-survival state, and then the first service node synchronizes the information that the second subsystem is in a non-survival state to the remaining service nodes in the first subsystem.

[0110] Corresponding to method ② in S405, the way in which the first subsystem determines that the second subsystem is in a non-survival state is: when there is a service node in at least one service node that determines that the second subsystem is in a non-survival state, the service node synchronizes the information that the second subsystem is in a non-survival state to the remaining service nodes in the first subsystem.

[0111] Optionally, as shown in FIG5 , S404 may be: the first subsystem counts the number of nodes detected as unreachable by the second subsystem in at least one service node based on the detection result on the first service node, and S405 may be: the first subsystem determines whether the counted number of nodes reaches a first threshold, and if it is determined that the number of nodes does not reach the first threshold, executes S406; if it is determined that the number of nodes reaches the first threshold, executes S407. The steps in FIG5 are the same as or similar to the steps in FIG4 , and reference may be made to the description of the steps in FIG4 .

[0112] Optionally, as shown in FIG6 , S403 may also be: each of the at least one service node in the first subsystem synchronizes the detection result of its own detection of whether the second subsystem is reachable to the remaining service nodes in the first subsystem; S404 may be: each of the at least one service node counts the number of nodes detected as reachable to the second subsystem in the at least one service node based on the detection results synchronized with other service nodes and its own detection result of whether the second subsystem is reachable; S405 may be: each of the at least one service node determines whether the counted number of nodes reaches a first threshold, executes S406 when it is determined that the number of nodes reaches the first threshold, and executes S407 when it is determined that the number of nodes does not reach the first threshold. The steps in FIG6 are the same as or similar to the steps in FIG4 , and reference may be made to the relevant descriptions of the steps in FIG4 .

[0113] Optionally, as shown in FIG7 , S403 may also be: each of the at least one service node in the first subsystem synchronizes the detection result of its own detection of whether the second subsystem is reachable to the remaining service nodes in the first subsystem; S404 may be: each of the at least one service node counts the number of nodes in the at least one service node that detect that the second subsystem is unreachable based on the detection results synchronized with other service nodes and its own detection result of whether the second subsystem is reachable; S405 may be: each of the at least one service node determines whether the counted number of nodes reaches a first threshold, and if it is determined that the number of nodes does not reach the first threshold, executes S406; if it is determined that the number of nodes reaches the first threshold, executes S407. The steps in FIG7 are the same as or similar to the steps in FIG4 , and reference may be made to the relevant descriptions of the steps in FIG4 .

[0114] Referring to FIG8 , FIG8 is a flow chart of another detection method provided in an embodiment of the present application, as shown in FIG8 , including the following steps:

[0115] S801: A first subsystem in a service system sends a first detection request to an anchor point other than the first subsystem and the second subsystem to determine whether the first subsystem is alive.

[0116] S801 is the same as S401. Please refer to the description of S401.

[0117] S802: Each of the at least one service node in the first subsystem detects whether the second subsystem is reachable.

[0118] S802 is the same as S402. Please refer to the description of S402.

[0119] S803: Each service node of the at least one service node in the first subsystem records the detection result to a first service node, where the first service node is any one of the at least one service node.

[0120] S803 is the same as S403. Please refer to the description of S403.

[0121] S804: The first subsystem counts the number of times that at least one service node detects that the second subsystem is reachable based on the detection result on the first service node.

[0122] The specific manner in which the first subsystem counts the number of times at least one service node detects that the second subsystem is reachable based on the detection result on the first service node includes but is not limited to the following two methods:

[0123] Mode I: The first service node in the first subsystem counts the number of times at least one service node detects that the second subsystem is reachable based on the detection result on the first service node.

[0124] Specifically, each time a service node records a detection result to the first service node, the first service node can count the number of detection results recorded on the first service node that include information indicating that the second subsystem is reachable, and use the number as the number of times at least one service node detects that the second subsystem is reachable.

[0125] Method II: When each service node in at least one service node detects whether the second subsystem is reachable and records the detection result to the first service node, the number of times at least one service node detects that the second subsystem is reachable is counted based on the detection result recorded on the first service node.

[0126] Specifically, each service node may count the number of detection results recorded on the first service node that include information indicating that the second subsystem is reachable, and use the number as the number of times at least one service node detects that the second subsystem is reachable.

[0127] S805: The first subsystem determines whether the number of counts reaches the second threshold, and executes S806 if it is determined that the number of counts reaches the second threshold; otherwise, it executes S807.

[0128] The second threshold value can be customized according to the actual scenario. It can be understood that the larger the second threshold value is set, the higher the accuracy of the first subsystem in determining whether the second subsystem is in a survival state based on the second threshold value.

[0129] Corresponding to the method I in S804, the method in which the first subsystem determines whether the number of statistics reaches the first threshold is ①: the first service node determines whether the number of statistics reaches the first threshold.

[0130] Corresponding to the method II in S804, the method in which the first subsystem determines whether the number of statistics reaches the first threshold is ②: each service node determines whether the number of statistics reaches the first threshold.

[0131] S806: The first subsystem determines that the second subsystem is in a live state.

[0132] S806 is the same as S406. Please refer to the description of S406.

[0133] S807: The first subsystem determines that the second subsystem is in a non-survival state.

[0134] S807 is the same as S407. Please refer to the description of S407.

[0135] Optionally, S804 may also be: the first subsystem counts the number of times at least one service node detects that the second subsystem is unreachable based on the detection results on the first service node, and S805 may be: the first subsystem determines whether the counted number reaches the second threshold, and executes S806 if it is determined that the number does not reach the second threshold, and executes S807 if it is determined that the number reaches the second threshold.

[0136] Optionally, S803 may also be: each service node of at least one service node in the first subsystem synchronizes the detection result of its own detection of whether the second subsystem is reachable to the remaining service nodes in the first subsystem, S804 may be: each service node of at least one service node counts the number of times at least one service node detects that the second subsystem is reachable based on the detection results synchronized by other service nodes and its own detection result of whether the second subsystem is reachable, and S805 may be: each service node of at least one service node determines whether the counted number reaches the second threshold, and executes S806 if it is determined that the number reaches the second threshold, and executes S807 if it is determined that the number does not reach the second threshold.

[0137] Optionally, S803 may also be: each service node of at least one service node in the first subsystem synchronizes the detection result of its own detection of whether the second subsystem is reachable to the remaining service nodes in the first subsystem, S804 may be: each service node of at least one service node counts the number of times at least one service node detects that the second subsystem is unreachable based on the detection results synchronized by other service nodes and its own detection result of whether the second subsystem is reachable, and S805 may be: the first subsystem determines whether the counted number reaches the second threshold, and executes S806 if it is determined that the number does not reach the second threshold, and executes S807 if it is determined that the number reaches the second threshold.

[0138] It should be understood that the liveness detection methods shown in Figures 4 to 8 are merely examples of the liveness detection methods provided in this application and should not be considered as specific limitations. For example, in a specific implementation, the liveness detection methods of Figures 4 and 8 can be combined to perform liveness detection. In the combined method, the condition for at least one service node in the first subsystem to determine whether the second subsystem is alive can be: at least one service node in the first subsystem determines whether the number of nodes detected as reachable by the second subsystem by the at least one service node in the statistics reaches a first threshold, and determines whether the number of times the at least one service node detected as reachable by the second subsystem reaches a second threshold. If it is determined that the number of nodes reaches the first threshold and the number of times reaches the second threshold, the second subsystem is determined to be alive; otherwise, the second subsystem is determined to be non-alive. Alternatively, if it is determined that the number of nodes reaches the first threshold and / or the number of times reaches the second threshold, the second subsystem is determined to be alive; otherwise, the second subsystem is determined to be non-alive.

[0139] In the embodiments of Figures 4 to 8, when the first subsystem determines that the second subsystem is in a non-survival state, if any service node in the first subsystem receives an access request for the service of the second subsystem sent by a terminal device, the service node can directly process the service access request and provide corresponding services to the user of the terminal device.

[0140] In the embodiments of Figures 4 to 8, when the first subsystem detects that the second subsystem has recovered to a survival state, if any service node in the first subsystem receives an access request for the service of the second subsystem sent by a terminal device, the service node can forward the service access request to the second subsystem, and the second subsystem will process the service access request to provide corresponding services to the user of the terminal device.

[0141] It can be seen from the embodiments of Figures 4 to 8 that the activation detection method provided in the present application can detect the remaining subsystems in the service system (such as the second subsystem mentioned above) through the existing resources in the service system (i.e., the first subsystem in the above-mentioned service system) without introducing external resources (such as a dedicated detection center). It can not only solve the problem of the existing activation detection method that if the detection center fails, it will not be able to continue to perform the activation task of the subsystem, but also avoid the service system's dependence on external resources, reduce the deployment difficulty of the service system, and reduce resource waste. In addition, when the first subsystem is waiting to transmit data with the second subsystem, the first subsystem can directly know whether the second subsystem is in a survival state after executing any of the embodiments in Figures 4 to 8. When it is known that the second subsystem is in a survival state, the first subsystem immediately transmits data with the second subsystem. When it is known that the second subsystem is in a non-survival state, unnecessary data transmission processes are avoided. Unlike the existing detection method, when the first subsystem is waiting to transmit data with the second subsystem, the first subsystem needs to first access the detection center, and the detection center detects whether the second subsystem is in a survival state and returns the detection result to the first subsystem. This can reduce the network delay between the first subsystem and the second subsystem. Moreover, it can avoid the situation where the detection center fails to detect whether the second subsystem is in a survival state. If the second subsystem is actually in a survival state, the first subsystem delays data transmission with the second subsystem because it cannot know whether the second subsystem is in a survival state, resulting in poor performance of the service system.

[0142] In the present application, the first subsystem may further detect whether any service node other than the at least one service node in the first subsystem (hereinafter referred to as the second service node) is alive. Specifically, the first subsystem detects whether the second service node is alive in the following manner, including but not limited to the two shown in Figures 9 and 10.

[0143] As shown in FIG9 , the process includes the following steps: S901 : Each service node of at least one service node in the first subsystem detects whether a second service node is reachable.

[0144] S902: Each service node of the at least one service node in the first subsystem records the detection result to a first service node, where the first service node is any one of the at least one service node.

[0145] S903: The first subsystem counts the number of nodes detected to be reachable from the second service node in at least one service node according to the detection result on the first service node.

[0146] S904: The first subsystem determines whether the counted number of nodes reaches a third threshold. If it is determined that the number of nodes reaches the third threshold, S905 is executed. If it is determined that the number of nodes does not reach the third threshold, S906 is executed.

[0147] The third threshold can be customized according to the actual scenario, such as being set according to the total number of service nodes in the first subsystem, such as being set to 1 / 5, 1 / 4, 1 / 3, etc. of the total number of service nodes in the first subsystem. It can be understood that the larger the third threshold is set, the higher the accuracy of the first subsystem in determining whether the second service node is alive based on the third threshold.

[0148] S905: The first subsystem determines that the second service node is in an alive state.

[0149] S906: The first subsystem determines that the second service node is in a non-survival state.

[0150] It can be seen that the difference between the embodiment of Figure 9 and S402 to S407 of the embodiment of Figure 4 is only that the object of the first subsystem to detect activity and the conditions for judging whether the detected object is in a survival state are different. In S402 to S407 of the embodiment of Figure 4, the object of the first subsystem to detect activity is the second subsystem, and the condition for judging whether the second subsystem is in a survival state is to judge whether the number of nodes detected to be reachable by the second subsystem in at least one service node reaches a first threshold. In the embodiment of Figure 9, the object of the first subsystem to detect activity is the second service node, and the condition for judging whether the second service node is in a survival state is to judge whether the number of nodes detected to be reachable by the second service node in at least one service node reaches a third threshold. Therefore, in the specific implementation process of the embodiment of Figure 9, the object of the first subsystem to detect activity in S402 to S407 of the embodiment of Figure 4 can be replaced by "second subsystem" with "second service node", and "first threshold" with "third threshold". For the sake of brevity of the specification, it will not be elaborated.

[0151] Optionally, S903 may also be: the first subsystem counts the number of nodes in at least one service node that detect that the second service node is unreachable based on the detection results on the first service node, and S904 may be: the first subsystem determines whether the counted number of nodes reaches a third threshold, and executes S905 if it is determined that the number of nodes does not reach the third threshold, and executes S906 if it is determined that the number of nodes reaches the third threshold.

[0152] Optionally, S902 may also be: each service node of at least one service node in the first subsystem synchronizes the detection result of its own detection of whether the second service node is reachable to the remaining service nodes in the first subsystem, S903 may be: each service node of at least one service node counts the number of nodes detected as reachable by the second service node in at least one service node based on the detection results synchronized by other service nodes and its own detection result of whether the second service node is reachable, and S904 may be: each service node of at least one service node determines whether the counted number of nodes reaches a third threshold, and executes S905 when it is determined that the number of nodes reaches the third threshold, and executes S906 when it is determined that the number of nodes does not reach the third threshold.

[0153] Optionally, S902 may also be: each service node of at least one service node in the first subsystem synchronizes the detection result of its own detection of whether the second service node is reachable to the remaining service nodes in the first subsystem, S903 may be: each service node of at least one service node counts the number of nodes in at least one service node that detect that the second service node is unreachable based on the detection results synchronized by other service nodes and its own detection result of whether the second service node is reachable, and S904 may be: each service node of at least one service node determines whether the counted number of nodes reaches a third threshold, and executes S905 when it is determined that the number of nodes does not reach the third threshold, and executes S906 when it is determined that the number of nodes reaches the third threshold.

[0154] As shown in Figure 10, the following steps are included:

[0155] S1001: Each service node of at least one service node in a first subsystem detects whether a second service node is reachable.

[0156] S1002: Each service node of the at least one service node in the first subsystem records a detection result to a first service node, where the first service node is any one of the at least one service node.

[0157] S1003: The first subsystem counts the number of times at least one service node detects that the second service node is reachable based on the detection result on the first service node.

[0158] S1004: The first subsystem determines whether the number of statistics reaches the fourth threshold, and executes S1005 if it is determined that the number of statistics reaches the fourth threshold; if it is determined that the number of statistics does not reach the fourth threshold, executes S1006.

[0159] The fourth threshold value can be customized according to actual scenarios. It is understandable that the larger the fourth threshold value is, the higher the accuracy of the first subsystem in determining whether the second service node is alive according to the fourth threshold value.

[0160] S1005: The first subsystem determines that the second service node is in an alive state.

[0161] S1006: The first subsystem determines that the second service node is in a non-survival state.

[0162] It can be seen that the difference between the embodiment of Figure 10 and S802 to S807 of the embodiment of Figure 8 is only that the object of the first subsystem's detection of activity and the conditions for judging whether the detection object is in a survival state are different. In S802 to S807 of the embodiment of Figure 8, the object of the first subsystem's detection of activity is the second subsystem, and the condition for judging whether the second subsystem is in a survival state is to judge whether the number of times at least one service node detects that the second subsystem is reachable reaches the second threshold. In the embodiment of Figure 10, the object of the first subsystem's detection of activity is the second service node, and the condition for judging whether the second service node is in a survival state is to judge whether the number of times at least one service node detects that the second service node is reachable reaches the fourth threshold. Therefore, in the specific implementation process of the embodiment of Figure 10, the object of the first subsystem's detection of activity in S802 to S807 of the embodiment of Figure 8 can be replaced by "second subsystem" with "second service node", and "second threshold" with "fourth threshold". For the sake of brevity of the specification, it will not be elaborated.

[0163] Optionally, S1003 may also be: the first subsystem counts the number of times that at least one service node detects that the second service node is unreachable based on the detection result on the first service node, and S1004 may be: the first subsystem determines whether the counted number reaches a fourth threshold, and executes S1005 if it is determined that the number does not reach the fourth threshold, and executes S1006 if it is determined that the number reaches the fourth threshold.

[0164] Optionally, S1002 may also be: each service node of at least one service node in the first subsystem synchronizes the detection result of its own detection of whether the second service node is reachable to the remaining service nodes in the first subsystem, S1003 may be: each service node of at least one service node counts the number of times at least one service node detects that the second service node is reachable based on the detection results synchronized by other service nodes and its own detection result of whether the second service node is reachable, and S1004 may be: each service node of at least one service node determines whether the counted number reaches a fourth threshold, and executes S1005 if it is determined that the number reaches the fourth threshold, and executes S1006 if it is determined that the number does not reach the fourth threshold.

[0165] Optionally, S1002 may also be: each service node of at least one service node in the first subsystem synchronizes the detection result of its own detection of whether the second service node is reachable to the remaining service nodes in the first subsystem, S1003 may be: each service node of at least one service node counts the number of times at least one service node detects that the second service node is unreachable based on the detection results synchronized by other service nodes and the detection result of its own detection of whether the second service node is reachable, and S1004 may be: the first subsystem determines whether the counted number reaches the fourth threshold, and executes S1005 if it is determined that the number does not reach the fourth threshold, and executes S1006 if it is determined that the number reaches the fourth threshold.

[0166] It should be understood that the methods for the first subsystem to detect whether the second service node is alive as shown in Figures 9 and 10 are merely examples and should not be considered as specific limitations. For example, in a specific implementation, the two detection methods of Figures 9 and 10 can be combined to perform detection. In this combination, the condition for at least one service node in the first subsystem to determine whether the second service node is alive can be: at least one service node in the first subsystem determines whether the number of nodes detected as reachable by the second service node among the at least one service node counted reaches a third threshold, and determines whether the number of times the at least one service node detected as reachable by the second service node reaches a fourth threshold. If the number of nodes reaches the third threshold and the number of times reaches the fourth threshold, the second service node is determined to be alive; otherwise, the second service node is determined to be not alive. Alternatively, if the number of nodes reaches the third threshold and / or the number of times reaches the fourth threshold, the second service node is determined to be alive; otherwise, the second service node is determined to be not alive.

[0167] In the embodiments of Figures 9 and 10, when the first subsystem determines that the second service node is in a non-survival state, if any service node in the first subsystem other than the second service node receives an access request for the service of the second service node sent by the terminal device, the service node can directly process the service access request and provide corresponding services to the user of the terminal device.

[0168] In the embodiments of Figures 9 and 10, when the first subsystem detects that the second service node has recovered its survival state, if any service node in the first subsystem other than the second service node receives an access request for the service of the second service node sent by the terminal device, the service node can forward the service access request to the second service node, and the second service node will process the service access request to provide corresponding services to the user of the terminal device.

[0169] It can be seen from the embodiments of Figures 9 and 10 that the detection method provided in the present application can detect the second service node in the first subsystem through the existing resources in the service system (i.e., at least one service node in the first subsystem of the above-mentioned service system), without introducing external resources (such as a dedicated detection center) to detect the second service node in the first subsystem, which can avoid the service system's dependence on external resources, reduce the deployment difficulty of the service system, and reduce resource waste.

[0170] The present application also provides a service node, which can be each service node in the first subsystem of the embodiments of Figures 4 to 10. As shown in Figure 11, the service node 1100 provided by the present application includes a detection module 1110, a statistics module 1120, a judgment module 1130, and a sending module 1140.

[0171] The sending module 1140 is configured to send a first detection request to an anchor point outside the first subsystem and the second subsystem to determine whether the service node 1100 is alive, thereby determining whether the first subsystem is alive.

[0172] The detection module 1110 is configured to detect whether the second subsystem is reachable.

[0173] The statistics module 1120 is configured to count the number of nodes detected to be reachable from the second subsystem in at least one service node in the first subsystem.

[0174] The judgment module 1130 is used to determine whether the number of nodes reachable from the second subsystem detected in at least one service node in the first subsystem reaches a first threshold. If it is determined that the number of nodes reaches the first threshold, it is determined that the second subsystem is in a surviving state. If it is determined that the number of nodes does not reach the first threshold, it is determined that the second subsystem is in a non-surviving state.

[0175] Alternatively, the statistics module 1120 is configured to count the number of times that at least one service node in the first subsystem detects that the second subsystem is reachable.

[0176] The judgment module 1130 is used to determine whether the number of times at least one service node in the first subsystem detects that the second subsystem is reachable reaches a second threshold. If it is determined that the number of times reaches the second threshold, it is determined that the second subsystem is in a survival state; if it is determined that the number of times does not reach the second threshold, it is determined that the second subsystem is in a non-survival state.

[0177] In a possible embodiment, as shown in FIG11 , the service node 1100 further includes a recording module 1150 and a receiving module 1160. The recording module 1150 is configured to start recording the first time when the sending module 1140 sends the first detection request. The judgment module 1130 is configured to determine that the service node 1100 is alive when the first time recorded by the recording module 1150 does not reach the first time threshold and the receiving module 1160 receives a response returned by the anchor point based on the first detection request, thereby determining that the first subsystem is alive.

[0178] In a possible embodiment, the recording module 1150 is also used to start recording the second time when the sending module 1140 sends a second detection request to the second subsystem, and the judgment module 1130 is used to determine that the second subsystem is reachable when the second time recorded by the recording module 1150 does not reach the second time threshold and the receiving module 1160 receives a response returned by the second subsystem based on the second detection request, and determine that the second subsystem is unreachable when the second time recorded by the recording module 1150 reaches the second time threshold and the receiving module 1160 does not receive a response returned by the second subsystem based on the second detection request.

[0179] In a possible embodiment, after detecting whether the second subsystem is reachable, the detection module 1110 is also used to record the detection results to the first service node, where the first service node is any one of the at least one service nodes in the first subsystem. The statistical module 1120 is used to count the number of nodes in at least one service node that detect that the second subsystem is reachable, or to count the number of times that at least one service node detects that the second subsystem is reachable, based on the detection results on the first service node.

[0180] In a possible embodiment, the detection module 1110 is further used to detect whether the second service node is reachable, wherein the second service node is any service node other than the at least one service node in the first subsystem, and the statistical module 1120 is further used to count the number of nodes in the at least one service node in the first subsystem that are detected to be reachable by the second service node, and the judgment module 1130 is further used to judge whether the number of nodes in the at least one service node that are detected to be reachable by the second service node reaches a third threshold, and when it is determined that the number of nodes reaches the third threshold, it is determined that the second service node is in a live state; and when it is determined that the number of nodes does not reach the third threshold, it is determined that the second service node is in a non-live state; or, the statistical module 1120 is further used to count the number of times at least one service node in the first subsystem detects that the second service node is reachable, and the judgment module 1130 is further used to judge whether the number of times at least one service node detects that the second service node is reachable reaches a fourth threshold, and when it is determined that the number reaches the fourth threshold, it is determined that the second service node is in a live state, and when it is determined that the number does not reach the fourth threshold, it is determined that the second service node is in a non-live state.

[0181] In a possible embodiment, the first subsystem and the second subsystem are backup subsystems for each other.

[0182] In a possible embodiment, the first subsystem and the second subsystem are deployed in the same AZ or different AZs.

[0183] In a specific implementation, the aforementioned modules, such as detection module 1110, statistics module 1120, judgment module 1130, and sending module 1140, can be implemented via software or hardware. For example, the implementation of judgment module 1130 will be described below using judgment module 1130 as an example. Similarly, the implementation of modules such as detection module 1110, statistics module 1120, and sending module 1140 can refer to the implementation of judgment module 1130.

[0184] As an example of a software functional unit, the determination module 1130 may include codes running on a physical host (computing device).

[0185] As an example of a hardware functional unit, the determination module 1130 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0186] It should be noted that, in other embodiments, the judgment module 1130 can be used to execute any step in the above-mentioned liveness detection method, the detection module 1110 can be used to execute any step in the above-mentioned liveness detection method, the statistics module 1120 can be used to execute any step in the above-mentioned liveness detection method, and the other modules shown in Figure 11 can also be used to execute any step in the above-mentioned liveness detection method. The steps that the multiple modules shown in Figure 11 are responsible for implementing can be specified as needed, and all functions of the service node 1100 are realized through these modules.

[0187] It should be understood that the functions of the various modules described above are merely functions that the service node 1100 may have in some embodiments of the present application, and the present application does not limit the functions of the various modules.

[0188] It should also be understood that Figure 11 is an exemplary division method, and the service node 1100 can also include more or fewer modules. The division method of the modules in the service node 1100 can be flexibly adjusted based on the actual business scenario, and this application does not make specific limitations.

[0189] Specifically, the specific implementation of the various operations performed by the service node 1100 can refer to the description in the relevant content of the above-mentioned detection method embodiment, and for the sake of brevity of the description, it will not be repeated here.

[0190] The present application also provides another service node, see Figure 12, which is a structural diagram of another service node 1200 provided by the present application. The service node can be each service node in the first subsystem in the embodiments shown in Figures 4 to 10, and can also be used to deploy the service node 1100 shown in Figure 11.

[0191] As shown in FIG. 12 , a service node 1200 includes a processor 1210 , a memory 1220 , and a communication interface 1230 . The processor 1210 , the memory 1220 , and the communication interface 1230 may be interconnected via a bus 1240 .

[0192] Processor 1210 can read the program code (including instructions) stored in memory 1220 and execute the program code stored in memory 1220, so that service node 1200 performs the steps performed by each service node in at least one service node in the method shown in Figures 4 to 10, or causes service node 1200 to deploy the various modules in service node 1100 shown in Figure 11.

[0193] Processor 1210 can be implemented in various forms, such as a central processing unit (CPU), or a combination of a CPU and a hardware chip. The hardware chip can be an ASIC, a programmable logic device (PLD), or a combination thereof. The PLD can be a CPLD, an FPGA, a GAL, or any combination thereof. Processor 1210 executes various types of digitally stored instructions, such as software or firmware stored in memory 1220, enabling service node 1200 to provide a wide variety of services.

[0194] In a specific implementation, as an embodiment, the processor 1210 includes one or more CPUs.

[0195] In a specific implementation, as an example, service node 1200 also includes multiple processors, each of which can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here refers to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0196] The memory 1220 is used to store program code, and the execution is controlled by the processor 1210. The program code may include one or more software modules, which may be the software modules provided in the embodiment shown in FIG. 11 , such as the detection module 1110, the statistics module 1120, the judgment module 1130, and the sending module 1140.

[0197] The memory 1220 may include a volatile memory, such as a random access memory (RAM); the memory 1220 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); the memory 1220 may also include a combination of the above types.

[0198] The communication interface 1230 may be a wired interface (e.g., an Ethernet interface, a fiber optic interface, or other types of interfaces (e.g., an InfiniBand interface)) or a wireless interface (e.g., a cellular network interface or a wireless local area network interface) for communicating with other service nodes or devices. The communication interface 1230 may utilize a protocol suite based on the Transmission Control Protocol / Internet Protocol (TCP / IP), such as the Remote Function Call (RFC) protocol, the Simple Object Access Protocol (SOAP) protocol, the Simple Network Management Protocol (SNMP) protocol, the Common Object Request Broker Architecture (CORBA) protocol, and distributed protocols.

[0199] Bus 1240 may be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a unified bus (UBus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), or the like. Bus 1240 may be classified as an address bus, a data bus, a control bus, or the like. In addition to a data bus, bus 1240 may also include a power bus, a control bus, and a status signal bus. However, for clarity, various buses are labeled as bus 1240 in the figure. For ease of illustration, FIG12 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.

[0200] The service node 1200 is used to execute the steps performed by each service node in the method shown in Figures 4 to 10. The specific implementation process is detailed in the above method embodiments and will not be repeated here.

[0201] It should be understood that service node 1200 is merely an example provided in the embodiments of the present application, and that service node 1200 may have more or fewer components than those shown in FIG12 , may combine two or more components, or may have different configurations of components. For matters not shown or described in the embodiments of the present application, please refer to the relevant descriptions in the embodiments of FIG1 to FIG11 , and will not be repeated here.

[0202] The present application also provides a node cluster, which can also be called a service node cluster. The node cluster can be the first subsystem in the above embodiment, and the node cluster can be used to implement some or all of the steps performed by the first subsystem in the detection method recorded in the above embodiment.

[0203] The present application also provides a computer-readable storage medium, which can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a high-density digital video disc (DVD), or a semiconductor medium (for example, a solid-state hard disk), etc. The computer-readable storage medium stores instructions, and when the instructions are executed, they can implement some or all of the steps performed by each service node in at least one service node in the first subsystem in the detection method described in the above embodiment.

[0204] The present application also provides a computer program product containing instructions, which may be software or a program product containing instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to execute the detection method described in the above embodiment.

[0205] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0206] In the above embodiments, it can be implemented in whole or in part by software, hardware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium, or a semiconductor medium.

[0207] The above is only a specific embodiment of the present application. Any changes or replacements that can be imagined by those skilled in the art based on the specific embodiment provided in this application should be included in the scope of protection of this application.

Claims

1. A detection method, characterized in that: Applied to a service system, the service system includes a first subsystem and a second subsystem, a plurality of service nodes in the first subsystem are each other's backup nodes, and a plurality of service nodes in the second subsystem are each other's backup nodes, the method includes: The first subsystem sends a first detection request to an anchor point outside the first subsystem and the second subsystem to determine that the first subsystem is in a live state; At least one service node in the first subsystem detects whether the second subsystem is reachable; When the first subsystem detects in the at least one service node that the number of nodes reachable by the second subsystem reaches a first threshold, it determines that the second subsystem is in a live state; when the first subsystem detects in the at least one service node that the number of nodes reachable by the second subsystem does not reach the first threshold, it determines that the second subsystem is in a non-live state; Alternatively, the first subsystem determines that the second subsystem is in a surviving state when the at least one service node detects that the second subsystem is reachable a number of times that reaches a second threshold, and determines that the second subsystem is in a non-surviving state when the at least one service node detects that the second subsystem is reachable a number of times that does not reach the second threshold.

2. The method according to claim 1, characterized in that The first subsystem sends a first detection request to an anchor point outside the first subsystem and the second subsystem to determine that the first subsystem is in a live state, including: When the first subsystem sends a first detection request to an anchor point outside the first subsystem and the second subsystem, the first subsystem starts recording a first time. When the first time does not reach a first time threshold and a response returned by the anchor point based on the first detection request is received, it is determined that the first subsystem is in a survival state.

3. The method according to claim 1 or 2, characterized in that: At least one service node in the first subsystem detects whether the second subsystem is reachable, including: When each service node in at least one service node in the first subsystem sends a second detection request to the second subsystem, it starts to record a second time. When the second time does not reach a second time threshold and a response returned by the second subsystem based on the second detection request is received, it is determined that the second subsystem is reachable. When the second time reaches the second time threshold and a response returned by the second subsystem based on the second detection request is not received, it is determined that the second subsystem is unreachable.

4. The method according to any one of claims 1 to 3, characterized in that: After at least one service node in the first subsystem detects whether the second subsystem is reachable, the method further includes: Each service node in the at least one service node in the first subsystem records a detection result of whether the second subsystem is reachable to a first service node, where the first service node is any one of the at least one service node; The first subsystem counts the number of nodes detected by the at least one service node as being reachable by the second subsystem according to the detection result on the first service node, or counts the number of times the at least one service node detects that the second subsystem is reachable.

5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: At least one service node in the first subsystem detects whether a second service node is reachable, where the second service node is any service node in the first subsystem except the at least one service node; The first subsystem determines that the second service node is in a live state when the at least one service node detects that the number of nodes reachable by the second subsystem reaches a third threshold, and determines that the second service node is in a non-live state when the at least one service node detects that the number of nodes reachable by the second service node does not reach the third threshold; Alternatively, the first subsystem determines that the second service node is in a live state when the at least one service node detects that the second service node is reachable for a number of times reaching a fourth threshold, and determines that the second service node is in a non-live state when the at least one service node detects that the second service node is reachable for a number of times not reaching the fourth threshold.

6. The method according to any one of claims 1 to 5, characterized in that: The first subsystem and the second subsystem are backup subsystems for each other.

7. The method according to any one of claims 1 to 6, characterized in that: The first subsystem and the second subsystem are deployed in different availability zones AZ.

8. A service system, characterized in that: The service system comprises a first subsystem and a second subsystem, a plurality of service nodes in the first subsystem are backup nodes for each other, and a plurality of service nodes in the second subsystem are backup nodes for each other; The first subsystem is used to send a first detection request to an anchor point outside the first subsystem and the second subsystem to determine that the first subsystem is in a live state; The first subsystem is used to detect whether the second subsystem is reachable through at least one service node in the first subsystem; The first subsystem is configured to determine that the second subsystem is in a live state when it is detected in the at least one service node that the number of nodes reachable by the second subsystem reaches a first threshold, and to determine that the second subsystem is in a non-live state when it is detected in the at least one service node that the number of nodes reachable by the second subsystem does not reach the first threshold; Alternatively, the first subsystem is used to determine that the second subsystem is in a surviving state when the at least one service node detects that the second subsystem is reachable a number of times that reaches a second threshold, and to determine that the second subsystem is in a non-surviving state when the at least one service node detects that the second subsystem is reachable a number of times that does not reach the second threshold.

9. The service system according to claim 8, characterized in that: The first subsystem is used to start recording a first time when sending the first detection request to an anchor point outside the first subsystem and the second subsystem, and determine that the first subsystem is in a survival state when the first time does not reach a first time threshold and a response returned by the anchor point based on the first detection request is received.

10. The service system according to claim 8 or 9, characterized in that: Each service node in at least one service node in the first subsystem is used to start recording a second time when sending a second detection request to the second subsystem, and determine that the second subsystem is reachable when the second time does not reach a second time threshold and a response returned by the second subsystem based on the second detection request is received, and determine that the second subsystem is unreachable when the second time reaches the second time threshold and a response returned by the second subsystem based on the second detection request is not received.

11. The service system according to any one of claims 8 to 10, characterized in that: After at least one service node in the first subsystem detects whether the second subsystem is reachable, each service node in the at least one service node in the first subsystem is used to record a detection result of detecting whether the second subsystem is reachable to a first service node, where the first service node is any one of the at least one service node; The first subsystem is used to count the number of nodes detected by the at least one service node as being reachable by the second subsystem based on the detection result on the first service node, or to count the number of times the at least one service node detects that the second subsystem is reachable.

12. The service system according to any one of claims 8 to 11, characterized in that: Each service node of the at least one service node in the first subsystem is used to detect whether a second service node is reachable, where the second service node is any service node in the first subsystem except the at least one service node; The first subsystem is configured to determine that the second service node is in a live state when the at least one service node detects that the number of nodes reachable by the second subsystem reaches a third threshold, and to determine that the second service node is in a non-live state when the at least one service node detects that the number of nodes reachable by the second service node does not reach the third threshold; Alternatively, the first subsystem is configured to determine that the second service node is in a live state if the number of times the at least one service node detects that the second service node is reachable reaches a fourth threshold, and to determine that the second service node is in a non-live state if the number of times the at least one service node detects that the second service node is reachable does not reach the fourth threshold.

13. The service system according to any one of claims 8 to 12, characterized in that: The first subsystem and the second subsystem are backup subsystems for each other.

14. The service system according to any one of claims 8 to 13, characterized in that: The first subsystem and the second subsystem are deployed in different AZs.

15. A service node, characterized in that: The service node is a service node in the service system according to any one of claims 8 to 14.

16. A node cluster, characterized in that: The node cluster is used to implement the method implemented by the first subsystem in any one of claims 1 to 7.

Citation Information

Patent Citations

  • State detection method and apparatus of node in distributed system

    CN109218137A

  • Keep-alive detection method and device, node, storage medium and communication system

    CN110958151A

  • Database detection method and device

    CN114691640A

  • Node activity detection method, device, equipment and medium

    CN115242687A

  • Backup node operation

    US20190238440A1