Server state recovery method and system, storage medium and electronic equipment

By sending the active client set to the inactive server, the problem of metadata server taking a long time to reconnect is solved, achieving rapid recovery and business continuity.

CN120658763APending Publication Date: 2025-09-16JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510806422.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In a multi-MDS cluster, when a metadata server fails and enters the reconnection state, the recovery process takes a long time, affecting business continuity and service quality.

Method used

By sending keep-alive messages to active servers and multiple clients, receiving reply messages and determining the set of active clients, the set is sent to the inactive server to quickly restore it to an active state.

Benefits of technology

This reduces the time spent by the metadata server during the reconnection phase, ensures that inactive servers can quickly handle abnormal clients, and improves business recovery speed and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120658763A_ABST
    Figure CN120658763A_ABST
Patent Text Reader

Abstract

The invention discloses a server state recovery method and system, a storage medium and electronic equipment, relates to the technical field of metadata management, and is applied to active servers in a server cluster. Comprising the following steps: determining an active client set in a plurality of first clients according to timestamps of received reply messages sent by the plurality of first clients based on keep-alive messages; and sending the active client set to the inactive server, so that the inactive server restores the running state from the inactive state to the active state according to the active client set. Namely, the active client set is determined through the active server, and the active client set is sent to the inactive server, so that the inactive server can convert the running state into the active state through the active client set. By means of the method and device, the problem that in the related technology, the time consumption of the metadata server in the reconnection stage (namely the process of recovering from the inactive state to the active state) is long can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of metadata management technology, and in particular to a server status recovery method and system, a storage medium, and an electronic device. Background Art

[0002] Current distributed storage systems utilize metadata servers (MDSs) to manage file system metadata and ensure data consistency and access performance. In a multi-MDS cluster, when an MDS fails and enters the reconnect state, its recovery relies on clients connected to it sending reconnect messages. This process can significantly extend MDS recovery time when the number of clients is large and widely distributed, impacting business continuity and service quality across the entire system.

[0003] Therefore, the problem in the related art that the metadata server takes a long time to reconnect (i.e., the process of recovering from an inactive state to an active state) has not yet been effectively solved. Summary of the Invention

[0004] The present application provides a server state recovery method and system, a storage medium and an electronic device to at least solve the problem in the related art that the metadata server takes a long time in the reconnection phase (i.e., the process of recovering from an inactive state to an active state).

[0005] The present application provides a server state recovery method, which is applied to an active server in a server cluster, comprising: sending a keep-alive message to multiple first clients that are connected to the active server, and receiving a first reply message sent by each first client based on the keep-alive message; determining an active client set from the multiple first clients based on a first timestamp of each received first reply message; and, upon detecting that an inactive server in the server cluster changes from an inactive state to a reconnected state, sending the active client set to the inactive server, so that the inactive server recovers the operating state of the inactive server from an inactive state to an active state based on the active client set.

[0006] The present application also provides a server state recovery method, which is applied to an inactive server in a server cluster, comprising: when the inactive server changes from an inactive state to a reconnected state, receiving an active client set sent by an active server in the server cluster, wherein the active server is used to: send a keep-alive message to multiple first clients that have a connection relationship with the active server, and receive a first reply message sent by each first client based on the keep-alive message; determine the active client set among the multiple first clients according to the first timestamp of the received first reply message, and send the active client set to the inactive server; and restore the operating state of the inactive server from the inactive state to the active state according to the active client set.

[0007] The present application also provides a server state recovery system, comprising: an active server and an inactive server in a server cluster, wherein: the active server is used to send a keep-alive message to multiple first clients that have a connection relationship with the active server, and receive a first reply message sent by each first client based on the keep-alive message; an active client set is determined among the multiple first clients according to the first timestamp of each received first reply message; when it is detected that the inactive server in the server cluster changes from an inactive state to a reconnected state, the active client set is sent to the inactive server; the inactive server is used to receive the active client set sent by the active server in the server cluster, and restore the operating state of the inactive server from an inactive state to an active state according to the active client set.

[0008] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server state recovery methods when executing the computer program.

[0009] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server state recovery methods are implemented.

[0010] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned server state recovery methods when executed by a processor.

[0011] Through this application, keep-alive messages are sent to multiple first clients that have a connection relationship with the active server, and the active client set is determined among the multiple first clients based on the timestamps of the reply messages sent by the multiple first clients based on the keep-alive messages; the active client set is sent to the inactive server, so that the inactive server restores the running state from the inactive state to the active state according to the active client set. In other words, this application determines the active client set through the active server, and sends the active client set to the inactive server, and then the inactive server can change the running state to the active state through the active client set. Through this application, the problem that the metadata server takes a long time in the reconnection stage (that is, the process of recovering from the inactive state to the active state) in the related art can be solved, and the inactive server can quickly process abnormal clients through the active client set during the reconnection stage, reduce the time consumed by the inactive server in the reconnection stage, and restore business as soon as possible. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 This is a hardware structure block diagram of an active server side of a server state recovery method according to an embodiment of the present application;

[0014] Figure 2 This is a flow chart of a method for restoring the state of a server according to an embodiment of the present application (I);

[0015] Figure 3 This is a flow chart (II) of a method for restoring the state of a server according to an embodiment of the present application;

[0016] Figure 4 is an architectural diagram of a keep-alive mechanism according to an optional embodiment of the present application;

[0017] Figure 5 This is an architectural diagram of a server state recovery system according to an embodiment of the present application. DETAILED DESCRIPTION

[0018] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0019] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0020] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0021] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the server state recovery method depends, the specific application environment architecture or specific hardware architecture is described here.

[0022] The method embodiments provided in the embodiments of the present application can be executed in a similar computing device such as a mobile terminal, a computer terminal or an active server. Taking the active server as an example, Figure 1 This is a hardware structure diagram of the active server side of a server state recovery method according to an embodiment of the present application. Figure 1 As shown, the active server may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MPU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. The computer terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above active server. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0023] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining the interactive state in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the active server via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0024] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a communications provider on the active server side. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0025] The embodiment of the present application provides a method for restoring the state of a server. Figure 2 This is a flow chart of a server state recovery method according to an embodiment of the present application (1), which can be applied to Figure 1 Active servers, such as Figure 2 As shown, the process includes the following steps:

[0026] Step S202: Send a keep-alive message to multiple first clients connected to the active server, and receive a first reply message sent by each first client based on the keep-alive message;

[0027] Wherein, a keep-alive message may be sent to each first client via a connection session between the active server and each first client;

[0028] Both the active server and the inactive server in this application can be metadata servers.

[0029] Step S204: determining an active client set from the plurality of first clients according to the first timestamp of each received first reply message;

[0030] Step S206: When it is detected that the inactive server in the server cluster changes from the inactive state to the reconnected state, the active client set is sent to the inactive server, so that the inactive server restores the operating state of the inactive server from the inactive state to the active state according to the active client set.

[0031] The inactive state may be a replay or resolve state.

[0032] Through the state recovery method of the server of the present application, a keep-alive message is sent to multiple first clients that have a connection relationship with the active server, and an active client set is determined among the multiple first clients based on the timestamps of the reply messages sent by the multiple first clients based on the keep-alive message; the active client set is sent to the inactive server, so that the inactive server recovers the running state from the inactive state to the active state according to the active client set. That is to say, the present application determines the active client set through the active server, and sends the active client set to the inactive server, and then the inactive server can change the running state to the active state through the active client set. Through the present application, the problem that the metadata server takes a long time in the reconnection stage (i.e., the process of recovering from the inactive state to the active state) in the related art can be solved, and the inactive server can quickly process abnormal clients through the active client set during the reconnection stage, thereby reducing the time consumed by the inactive server in the reconnection stage and recovering business as soon as possible.

[0033] Optionally, before sending keep-alive messages to multiple first clients that have a connection relationship with the active server in the above step S202, the method also includes: receiving service mapping information corresponding to the server cluster sent by the monitoring node, wherein the monitoring node receives the heartbeat signal sent by each server in the server cluster based on a preset period, and updates the service mapping information based on whether the heartbeat signal sent by each server is received within the preset period, and the service mapping information includes: the operating status of each server; and determining the operating status of each server in the server cluster based on the service mapping information.

[0034] It is understood how the active server uses the service mapping information to determine the operating status of other servers in the server cluster. Specifically:

[0035] Receive service mapping information corresponding to the server cluster from the monitoring node: In a server cluster, the monitoring node is responsible for monitoring the status of the entire cluster. It receives heartbeat signals from each server in the cluster (e.g., MDS servers) to continuously monitor the operational status of each server, including whether it is operating normally and its current state (e.g., active, replay, reconnect, or standby). Heartbeat signals can include information about the health and current processing status of each server in the server cluster.

[0036] For example, suppose there are three MDS servers in a server cluster, labeled MDS1, MDS2, and MDS3. Each MDS server periodically sends heartbeat signals to the monitoring node. These heartbeat signals contain information about the MDS server's operational status. Based on this information, the monitoring node creates and maintains service mapping information.

[0037] The monitoring node receives the heartbeat signal sent by each server in the server cluster based on a preset period: the monitoring node receives the heartbeat signal from each server in the server cluster according to a preset period (for example, every 10 seconds).

[0038] Updating the service mapping information based on whether a heartbeat signal from each server is received within a preset period: The monitoring node updates the service mapping information based on whether a heartbeat signal from each server is received within a preset period. If a heartbeat signal is received, the monitoring node updates the service mapping information, maintaining the server status as normal or current. If no heartbeat signal is received, the monitoring node may mark the corresponding server as potentially faulty and update the service mapping information to reflect this change.

[0039] For example, if the monitoring node does not receive a heartbeat signal from MDS2 within a 10-second period, the monitoring node will update the service mapping information and mark the state of MDS2 from active to replaying or reconnecting, indicating that it may need to recover from the failure.

[0040] The operating status of each server in the server cluster is determined based on the service mapping information: the active MDS receives and parses the service mapping information sent by the monitoring node, and determines the operating status of other MDSs in the cluster based on the service mapping information. For example: MDS1 receives the latest version of the service mapping information from the monitoring node and finds that the status of MDS2 (i.e., the inactive server) has been marked as a reconnected state. This means that MDS2 is undergoing a fault recovery process, and MDS1 (i.e., the active server) can take corresponding actions, such as reallocating the metadata processing tasks that MDS2 is responsible for, or starting a keep-alive mechanism to send keep-alive messages to all clients connected to MDS1 (i.e., multiple first clients) to help MDS2 recover to the active state faster.

[0041] By monitoring the information interaction between nodes and MDS, and updating the service mapping information based on heartbeat signals, it ensures that when a server fails, other servers can take over the task in time, reducing data access delays and improving the overall performance of the system.

[0042] The embodiment of the present application also provides a method for restoring the state of a server. Figure 3 This is a flow chart (II) of a server state recovery method according to an embodiment of the present application, which can be applied to inactive servers, such as Figure 3 As shown, the process includes the following steps:

[0043] Step S302: When the inactive server changes from an inactive state to a reconnected state, receiving an active client set sent by an active server in the server cluster, wherein the active server is configured to: send a keep-alive message to a plurality of first clients connected to the active server, and receive a first reply message sent by each first client based on the keep-alive message; determine the active client set from the plurality of first clients based on a first timestamp of the received first reply message, and send the active client set to the inactive server;

[0044] Step S304: Restoring the operating state of the inactive server from the inactive state to the active state according to the active client set.

[0045] Optionally, the above-mentioned step S304 of restoring the operating state of the inactive server from the inactive state to the active state according to the active client set includes: obtaining a first client set that has a connection session with the inactive server at the current time, wherein the connection session is used to indicate that the clients in the first client set have two-way communication with the inactive server within a first time period before the current time; terminating the connection session between each client in the first client set and the inactive server; sending a reconnection message to each active client in the active client set respectively, and determining whether a second reply message sent by each active client based on the reconnection message is received within a second time period; in the case of determining that a second reply message sent by one or more second clients in the active client set is received within the second time period, determining a second timestamp of receiving each second reply message; determining an establishment order of establishing a connection relationship between the inactive server and each second client according to the second timestamp; and establishing a connection relationship between the inactive server and each second client in sequence according to the establishment order, so as to restore the operating state of the inactive server from the inactive state to the active state.

[0046] It is understood that during the process of recovering the MDS from an inactive state (e.g., replay or reconnect state) to an active state, the connection session can be further optimized and managed to reduce the failure recovery time. Specifically:

[0047] Obtaining the first set of clients that have a connection session with the inactive server at the current time: When an inactive server is in the replay or reconnect state, the inactive server needs to determine the set of all clients that had valid communication (i.e., a connection session) with it before the current time. For example, suppose MDS A enters the reconnect state for some reason. MDS A needs to determine the set of clients that have had valid communication with MDS A. By examining recent communication records, MDS A discovers that Client 1, Client 3, and Client 5 interacted with it in the previous time window (e.g., 30 seconds ago). Clients 1, 3, and 5 constitute the first set of clients.

[0048] Terminating the connection session: After determining the first client set, the inactive server terminates the current connection session with each client in the first client set.

[0049] Reconnect: This means sending a reconnect message to each active client in the active client set. The inactive server sends a reconnect message to all currently active client sets (active client sets), inviting them to reconnect.

[0050] Determining whether a second reply message sent by each active client based on the reconnect message is received within a second time period: Inactive server A sets a second time period (e.g., 45 seconds) during which it waits to receive a second reply message from the set of active clients. The second reply message confirms that the client has received the reconnect invitation and is ready to reestablish the connection.

[0051] Determine a second timestamp of receiving each second reply message: For each client that replies within the second time period, the inactive server A records the timestamp of receiving the reply message.

[0052] The order in which the connection relationship between the inactive server and each second client is established is determined based on the second timestamp: The inactive server A determines the order in which the connection relationships with the responding clients are reestablished based on the timestamps of the received reply messages. This is generally done on a first-come, first-served basis, meaning that the client that received the reply earlier is reconnected first.

[0053] For example, assume that within 45 seconds, MDS server A receives a response from Client 1 first, followed by Client 3, and finally Client 5. Based on the time of the responses received, MDS server A will prioritize establishing connections with Client 1, Client 3, and Client 5 in that order. After completing the reconnection process for all clients and confirming data consistency, MDS server A declares itself active from the reconnected state and becomes a fully functional member of the cluster again.

[0054] Through this application, the efficiency of MDS in the fault recovery process is improved, and the time spent on processing outdated session information and waiting for invalid client connections is reduced.

[0055] Among them, restoring the operating state of the inactive server from an inactive state to an active state according to the active client set also includes: obtaining a first client set that has a connection session with the inactive server at a current time, wherein the connection session is used to indicate that there is two-way communication between the clients in the first client set and the inactive server within a first time period before the current time; determining a third client in the first client set that is inconsistent with the active clients in the active client set; terminating the connection session between the third client and the inactive server, and determining whether there are other clients other than the third client in the first client set; in case it is determined that the other clients exist in the first client set, sending a reconnection message to the other clients, so that the other clients send a third reply message to the inactive server based on the reconnection message; in case the third reply message is received, re-establishing the connection relationship with the other clients to restore the operating state of the inactive server from the inactive state to the active state.

[0056] It's understandable that when an MDS recovers from an inactive state, it can efficiently and accurately manage the connection sessions between the MDS and clients, accelerating the recovery process. Specifically, when the MDS recovers from a failure, it needs to identify which clients maintained active connection sessions with it before the recovery. The existence of these sessions means that there was bidirectional data exchange between these clients and the MDS within a first period before the current time.

[0057] The active client set refers to the set of clients in the entire server cluster that maintain an active connection with at least one MDS server. Tertiary clients are those that are no longer considered active compared to the clients in the active client set. They may have lost communication with the MDS for a period of time due to network outages or other failures.

[0058] After terminating the connection session between the third client and the inactive server, the inactive MDS then checks whether there are any other clients in the first client set besides the one identified as the third client. This is to determine if there are any active clients worth reestablishing a connection with. If other clients are confirmed, the inactive server sends them a reconnection message, inviting them to reconnect with it.

[0059] These other clients will respond to the inactive server's reconnection message with a third confirmation message. After receiving these third confirmation messages, the inactive server will reestablish connections with these clients, gradually restoring its own state from inactive to active.

[0060] Through the above process, MDS can quickly identify and disconnect inactive or abnormal clients during fault recovery, avoiding wasting time on invalid connections, thereby accelerating fault recovery and improving the overall efficiency and reliability of the system.

[0061] After re-establishing the connection relationship with the other client, the method further includes: establishing a first connection session with the other client based on the re-established connection relationship; synchronizing session information in a second connection session to the first connection session to process the session information based on the first connection session, wherein the second connection session is the connection session between the inactive server and the other client before the connection relationship is re-established; and deleting the second connection session when it is determined that the session information has been successfully synchronized to the first connection session.

[0062] It is understandable that during the process of the MDS recovering from the inactive state to the active state, the connection session between the MDS and the client can be managed and updated. Specifically:

[0063] Establishing a first connection session with the other clients based on the re-established connection relationship: When the MDS recovers from the failure, it re-establishes connections with the clients that maintained valid connections with it before the failure.

[0064] Synchronize session information from the second connection session with the first connection session: Before the failure, a connection session may already exist between the MDS and the client. This second connection session may contain the client's requests, data status, and metadata operation history. The inactive server needs to synchronize all session information from the second connection session with the first connection session to ensure that all relevant metadata operations and status information are accurately restored and processed after the failure is recovered.

[0065] Processing the session information based on the first connection session: Once the session information is synchronized to the first connection session, the inactive server can begin processing client requests based on the session information, continuing to perform tasks such as file operations and data reading and writing. Once it is confirmed that all session information in the second connection session has been successfully synchronized to the first connection session and that the inactive server can correctly process client requests based on the session information, the second connection session is no longer needed and can be safely deleted.

[0066] Through the above technical solution, inactive servers can not only recover quickly from failures, but also ensure that they can seamlessly continue to process client metadata requests after recovery, maintaining the high availability and data consistency of the system.

[0067] Optionally, before restoring the operating state of the inactive server from the inactive state to the active state according to the active client set in the above step S304, the method further includes: determining the on / off state of a target switch in the inactive server, wherein the target switch is used to indicate whether a connection relationship is allowed to be established between the inactive server and any client; if it is determined that the target switch is in the on state, prohibiting the establishment of a connection relationship between the inactive server and any active client among the active clients; obtaining a first client set that has a connection session with the inactive server at the current time, wherein the connection session is used to indicate that there is two-way communication between the clients in the first client set and the inactive server within a first time period before the current time; disconnecting the connection session between each client in the first client set and the inactive server to restore the operating state of the inactive server from the inactive state to the active state; if it is determined that the target switch is in the off state, determining that it is allowed to restore the operating state of the inactive server from the inactive state to the active state according to the active client set.

[0068] It is understood that during the process of restoring the MDS from an inactive state to an active state, the connection establishment and disconnection between the MDS and the client can be controlled by managing the on / off status of a specific target switch to optimize the recovery process. Specifically:

[0069] During the recovery process of an MDS from an inactive state, you can check the status of a specific target switch. The target switch status indicates whether the MDS should automatically deny new connections to any clients during the recovery phase. For example, if MDS Y becomes inactive due to an exception, before the recovery process begins, you can check the status of a switch named MDS_deny_all_reconnect (Metadata Server Deny All Reconnections). This target switch determines the connection strategy of MDS Y during the recovery process.

[0070] If the target switch is on, MDS Y will not attempt to establish connections with any active clients during the initial recovery phase. This prevents a large number of connection requests from being processed during the initial recovery phase, shortening the time it takes for the MDS to return to full activity.

[0071] During the recovery process of MDS Y, a first set of clients that had bidirectional communication with MDS Y before the failure may be identified and obtained. The clients in the first set of clients maintained valid connection sessions with MDS Y before the failure.

[0072] After confirming the first set of clients, MDS Y will proactively disconnect existing connection sessions with these clients.

[0073] If the target switch is in the off state, this means that MDS Y will allow the active client set to reestablish a connection relationship. In this case, MDS Y can perform failure recovery and return to the active state based on the client information in the active client set.

[0074] By controlling the state of the target switch, MDS Y can flexibly manage its connection strategy during the fault recovery process, avoiding resource waste and unnecessary delays, ensuring that it can recover to the active state in the most efficient way and continue to provide metadata services to clients.

[0075] In order to better understand the process of the state recovery method of the above-mentioned server, the implementation process of the state recovery method of the above-mentioned server is described below in combination with an optional embodiment, but it is not used to limit the technical solution of the embodiment of this application.

[0076] An optional embodiment of the present application provides a method for quick switching processing based on massive distributed storage metadata services. The optional embodiment of the present application is based on a keep-alive mechanism to reduce the time consumed by the MDS in the reconnection phase. The main idea of ​​the keep-alive mechanism of the distributed storage cluster is to obtain a set of clients in normal status through an active metadata server (Active Metadata Server, referred to as Active MDS for short), and send the set of clients in normal status to a faulty metadata server. The faulty MDS can quickly process abnormal clients through a set of normal clients in the reconnection phase, which can reduce the time consumed by the MDS in the reconnection phase and restore business as soon as possible.

[0077] Figure 4 is an architectural diagram of a keep-alive mechanism according to an optional embodiment of the present application, such as Figure 4 As shown in the figure, the main process of the keep-alive mechanism includes:

[0078] 1) Each MDS in the metadata cluster (i.e., server cluster) can perceive the status changes of other MDSs through the metadata service map (MDS Map, i.e., service mapping information) sent by the monitor node (MON). When an MDS perceives that an MDS in the cluster is in the replay or resolve state (i.e., inactive state), the first unit (unit of the MDS, an MDS has multiple units (e.g., Figure 4 Unit 4, 5, 6, 7), the first unit is Figure 4 Unit 4 in) to all (connection (connection) session (session)) clients (ie multiple first clients, i.e. Figure 4 Clients 1, 2, and 3) send a keep-alive message. After receiving the message, the client responds and records the timestamp (i.e., the first timestamp) in the MDS.

[0079] 2) When the first unit in the active MDS (i.e. active server) senses that an MDS in the cluster is in the reconnect state (reconnect state) through the MDS Map, the active MDS (i.e. Figure 4 The first unit in MDS1) obtains the keep alive client (i.e. active client) and sends it to the MDS in the reconnecting state (i.e. Figure 4 All units of MDS2 in Figure 4 unit0, 1, 2, 3).

[0080] A reconnect message may be sent to the faulty MDS through a function library (Lib) (wherein the reconnect message (ie, the reconnect message) includes an active client set).

[0081] Lib senses that the MDS is in a reconnection state in a faulty state by monitoring the MDS Map sent by the node and that there is a connection session between Lib and the units in the MDS, and then sends a reconnection message to all units in the faulty MDS.

[0082] 3) All units in the reconnected MDS (i.e., inactive clients) receive the reconnect message (i.e., the set of active clients), and terminate (kill) the clients that have abnormal connection sessions with the reconnected MDS, and no longer wait for clients that are not keep alive and have no sessions to send reconnect messages.

[0083] There are two ways to handle a faulty MDS during the reconnection phase:

[0084] (1) Add a metadata server deny all reconnection switch (mds_deny_all_reconnect switch, i.e., the target switch). When the mds_deny_all_reconnect switch is turned on, all units in the failed MDS (i.e., the inactive server) deny connection sessions with all clients, reducing the processing time of the rejoin phase and thus reducing the overall failover time.

[0085] The mds_deny_all_reconnect switch is currently disabled by default.

[0086] (2) All units in the faulty MDS process the reconnection message upon receiving it. If the processing is completed normally, they enter the rejoin state. If all units in the faulty MDS exceed 45 seconds (a preset time, which can be set based on the actual application) in the reconnection phase and do not receive a message from the client (any message is acceptable) for 22 seconds, the unprocessed connection session will be disconnected. The maximum time in the reconnection phase in this way is 67 seconds.

[0087] 4) All units of the faulty MDS enter the rejoining state after processing the reconnection messages of all normal clients.

[0088] MDS can process reconnection messages only during the reconnection phase to prevent multiple rejoin processes.

[0089] The current MDS processes client reconnection messages quickly. The duration of the reconnection phase depends on the client sending reconnection messages to the MDS as quickly as possible. The rejoin phase only begins when all units in the failed MDS have completed processing every client connection session.

[0090] In summary, the MDS in the reconnection phase handles the following connection sessions as follows:

[0091] 1) The connection session does not enter the reconnection state and is directly disconnected;

[0092] 2) MDS does not have a connection session with a certain client and can be ignored;

[0093] 3) For a client with a normal connection session, if it sends a reconnect message within 45 seconds, it will be processed; if it times out, it will be disconnected;

[0094] 4) Non-keep-alive clients (i.e., clients other than the active client set) are ignored.

[0095] Through the optional embodiments of the present application, if all connection sessions between the MDS and the client are normal, there is no optimization; if the client is not keep-alive, the keep-alive mechanism can save 45 seconds of failure time in the best case; if the connection session of the MDS-associated client does not send a reconnection message, the failure time can be saved in the best case (45 seconds - 3 seconds).

[0096] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0097] This embodiment also provides a server state recovery system, which is used to implement the above embodiments and preferred implementations. Details that have already been described will not be repeated. Although the systems described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and contemplated.

[0098] Figure 5 This is an architectural diagram of a server state recovery system according to an embodiment of the present application. Figure 5 As shown, the system includes: an active server and an inactive server in a server cluster, wherein:

[0099] The active server 52 is configured to send a keep-alive message to a plurality of first clients connected to the active server, and receive a first reply message sent by each first client based on the keep-alive message; determine an active client set from the plurality of first clients based on a first timestamp of each received first reply message; and upon detecting that an inactive server in the server cluster changes from an inactive state to a reconnected state, send the active client set to the inactive server;

[0100] The inactive server 54 is configured to receive an active client set sent by an active server in the server cluster, and restore the operating state of the inactive server from an inactive state to an active state according to the active client set.

[0101] Through the state recovery system of the server of the present application, keep-alive messages are sent to multiple first clients that have a connection relationship with the active server, and the active client set is determined among the multiple first clients based on the timestamps of the reply messages sent by the multiple first clients based on the keep-alive messages; the active client set is sent to the inactive server, so that the inactive server recovers the running state from the inactive state to the active state according to the active client set. That is to say, the present application determines the active client set through the active server, and sends the active client set to the inactive server, and then the inactive server can change the running state to the active state through the active client set. Through the present application, the problem that the metadata server takes a long time in the reconnection stage (i.e., the process of recovering from the inactive state to the active state) in the related art can be solved, and the inactive server can quickly process abnormal clients through the active client set during the reconnection stage, thereby reducing the time consumed by the inactive server in the reconnection stage and recovering the business as soon as possible.

[0102] In an exemplary embodiment, the active server 52 is further used to receive service mapping information corresponding to the server cluster sent by a monitoring node, wherein the monitoring node receives a heartbeat signal sent by each server in the server cluster based on a preset period, and updates the service mapping information based on whether the heartbeat signal sent by each server is received within the preset period, and the service mapping information includes: the operating status of each server; and determining the operating status of each server in the server cluster based on the service mapping information.

[0103] In an exemplary embodiment, the inactive server 54 is further used to obtain a first client set that has a connection session with the inactive server at a current time, wherein the connection session is used to indicate that there is two-way communication between the clients in the first client set and the inactive server within a first time period before the current time; terminate the connection session between each client in the first client set and the inactive server; reconnect message each active client in the active client set, and determine whether a second reply message sent by each active client based on the reconnection message is received within a second time period; when it is determined that a second reply message sent by one or more second clients in the active client set is received within the second time period, determine a second timestamp of receiving each second reply message; determine an establishment order of establishing a connection relationship between the inactive server and each second client based on the second timestamp; and establish a connection relationship between the inactive server and each second client in sequence according to the establishment order to restore the operating state of the inactive server from the inactive state to the active state.

[0104] In an exemplary embodiment, the inactive server 54 is further used to obtain a first client set that has a connection session with the inactive server at a current time, wherein the connection session is used to indicate that there is two-way communication between the clients in the first client set and the inactive server within a first time period before the current time; determine a third client in the first client set that is inconsistent with the active clients in the active client set; terminate the connection session between the third client and the inactive server, and determine whether there are other clients in the first client set except the third client; if it is determined that the other clients exist in the first client set, send a reconnection message to the other clients, so that the other clients send a third reply message to the inactive server based on the reconnection message; if the third reply message is received, re-establish the connection relationship with the other clients to restore the operating state of the inactive server from the inactive state to the active state.

[0105] In an exemplary embodiment, the inactive server 54 is further used to establish a first connection session with the other client based on the re-established connection relationship; synchronize session information in the second connection session to the first connection session to process the session information based on the first connection session, wherein the second connection session is the connection session between the inactive server and the other client before the connection relationship is re-established; and delete the second connection session when it is determined that the session information has been successfully synchronized to the first connection session.

[0106] In an exemplary embodiment, the inactive server 54 is further used to determine the on / off state of a target switch in the inactive server, wherein the target switch is used to indicate whether a connection relationship is allowed to be established between the inactive server and any client; when it is determined that the target switch is in the on state, establishing a connection relationship between the inactive server and any active client among the active clients is prohibited; obtaining a first set of clients that have a connection session with the inactive server at the current time, wherein the connection session is used to indicate that there is two-way communication between the clients in the first client set and the inactive server within a first time period before the current time; disconnecting the connection session between each client in the first client set and the inactive server to restore the operating state of the inactive server from the inactive state to the active state; and when it is determined that the target switch is in the off state, determining that it is allowed to restore the operating state of the inactive server from the inactive state to the active state according to the active client set.

[0107] For the description of the features in the embodiment corresponding to the server state recovery system, please refer to the relevant description of the embodiment corresponding to the server state recovery method, which will not be repeated here.

[0108] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps of any of the above-mentioned server state recovery method embodiments.

[0109] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned server state recovery method embodiments when running.

[0110] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0111] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned server state recovery method embodiments are implemented.

[0112] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, implementing the steps of any of the above-mentioned server state recovery method embodiments.

[0113] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0114] The above is a detailed introduction to the state recovery of a server provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core ideas of this application. It should be pointed out that, for those skilled in the art, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

Claims

1. A server status recovery method, characterized in that: Applies to active servers in a server cluster, including: Sending a keep-alive message to a plurality of first clients connected to the active server, and receiving a first reply message sent by each first client based on the keep-alive message; determining an active client set from the plurality of first clients according to the first timestamp of each received first reply message; When it is detected that an inactive server in the server cluster changes from an inactive state to a reconnected state, the active client set is sent to the inactive server, so that the inactive server restores the operating state of the inactive server from the inactive state to the active state according to the active client set.

2. The server status recovery method according to claim 1, characterized in that: Before sending keep-alive messages to a plurality of first clients connected to the active server, the method further includes: Receive service mapping information corresponding to the server cluster sent by a monitoring node, wherein the monitoring node receives a heartbeat signal sent by each server in the server cluster based on a preset period, and updates the service mapping information based on whether the heartbeat signal sent by each server is received within the preset period, and the service mapping information includes: the operating status of each server; An operating status of each server in the server cluster is determined based on the service mapping information.

3. A server status recovery method, characterized in that: Applicable to inactive servers in a server cluster, including: When the inactive server changes from an inactive state to a reconnected state, receiving an active client set sent by an active server in the server cluster, wherein the active server is configured to: send a keep-alive message to a plurality of first clients connected to the active server, and receive a first reply message sent by each first client based on the keep-alive message; determine the active client set from the plurality of first clients according to a first timestamp of the received first reply message, and send the active client set to the inactive server; The operating state of the inactive server is restored from an inactive state to an active state according to the active client set.

4. The server status recovery method according to claim 3, characterized in that: Restoring the operating state of the inactive server from an inactive state to an active state according to the active client set includes: Acquire a first set of clients that have connection sessions with the inactive server at a current time, wherein the connection sessions are used to indicate that clients in the first set of clients have had two-way communication with the inactive server within a first time period before the current time; terminating a connection session between each client in the first set of clients and the inactive server; sending a reconnection message to each active client in the set of active clients respectively, and determining whether a second reply message sent by each active client based on the reconnection message is received within a second time period; In a case where it is determined that a second reply message sent by one or more second clients in the set of active clients is received within a second time period, determining a second timestamp of receiving each second reply message; determining, according to the second timestamp, an order of establishing a connection relationship between the inactive server and each second client; The connection relationship between the inactive server and each of the second clients is established in sequence according to the establishment order, so as to restore the running state of the inactive server from the inactive state to the active state.

5. The server status recovery method according to claim 3, characterized in that: Restoring the operating state of the inactive server from an inactive state to an active state according to the active client set further includes: Acquire a first set of clients that have connection sessions with the inactive server at a current time, wherein the connection sessions are used to indicate that clients in the first set of clients have had two-way communication with the inactive server within a first time period before the current time; determining, in the first client set, a third client that is inconsistent with an active client in the active client set; terminating a connection session between the third client and the inactive server, and determining whether there are other clients in the first client set except the third client; When it is determined that the other client exists in the first client set, sending a reconnection message to the other client, so that the other client sends a third reply message to the inactive server based on the reconnection message; When the third reply message is received, the connection relationship with the other client is re-established to restore the running state of the inactive server from the inactive state to the active state.

6. The server status recovery method according to claim 5, characterized in that: After re-establishing the connection relationship with the other client, the method further includes: Establishing a first connection session with the other client based on the re-established connection relationship; Synchronizing session information in a second connection session to the first connection session, so as to process the session information based on the first connection session, wherein the second connection session is a connection session between the inactive server and the other client before the re-establishment of the connection relationship; If it is determined that the session information has been successfully synchronized to the first connection session, the second connection session is deleted.

7. The server status recovery method according to claim 3, characterized in that: Before restoring the operating state of the inactive server from the inactive state to the active state according to the active client set, the method further includes: determining an on / off state of a target switch in the inactive server, wherein the target switch is used to indicate whether a connection relationship between the inactive server and any client is allowed; When it is determined that the target switch is in the on state, prohibiting establishment of a connection relationship between the inactive server and any active client among the active clients; Acquire a first set of clients that have connection sessions with the inactive server at a current time, wherein the connection sessions are used to indicate that clients in the first set of clients have had two-way communication with the inactive server within a first time period before the current time; disconnecting a connection session between each client in the first client set and the inactive server, so as to restore the running state of the inactive server from the inactive state to the active state; In a case where it is determined that the target switch is in the off state, it is determined to allow the running state of the inactive server to be restored from the inactive state to the active state according to the active client set.

8. A server status recovery system, characterized in that: include: Active servers and inactive servers in a server cluster, where: The active server is configured to send a keep-alive message to a plurality of first clients connected to the active server, and receive a first reply message sent by each first client based on the keep-alive message; determine an active client set from the plurality of first clients based on a first timestamp of each received first reply message; and send the active client set to the inactive server in the server cluster when detecting that the inactive server changes from an inactive state to a reconnected state; The inactive server is configured to receive an active client set sent by an active server in the server cluster, and restore the operating state of the inactive server from an inactive state to an active state according to the active client set.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the server state recovery method according to any one of claims 1 to 2 or 3 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the server state recovery method according to any one of claims 1 to 2 or 3 to 7 are implemented.