Implementation Method and System for High Availability of Stateless Computing Instances in a Distributed Database

By monitoring and analyzing distributed database computing instances, restarting or scheduling offline instances, the high availability complexity problem caused by computing and storage binding is solved, and the high availability of stateless computing instances and the separation of computing and storage are achieved.

CN113886490BActive Publication Date: 2025-06-20BEIJING EASTERN JIN TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111073070.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-14
Publication Date
2025-06-20
Estimated Expiration
2041-09-14

AI Technical Summary

Technical Problem

Computing and storage in distributed databases are bound, and the computing instances are stateful, resulting in complex high availability implementations and insufficient elastic capabilities of computing and storage.

Method used

By monitoring all computing instances in the database, the offline computing instances are determined, and the offline computing instances are restarted or scheduled according to the analyzed state, to achieve high availability of stateless computing instances.

Benefits of technology

It realizes the high availability of the computing layer in the distributed database, avoids the confusion of unexpected termination and scheduling restart of computing instances, satisfies the separation of computing and storage, and improves the application capabilities of distributed databases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113886490B_ABST
    Figure CN113886490B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for realizing high availability of stateless computing instances in a distributed database, which is characterized in that it includes: monitoring the states of all computing instances of the database to determine the offline computing instances; analyzing the states of the computing nodes where the offline computing instances are located, and according to the analyzed states, restarting the offline computing instances or scheduling the offline computing instances to other computing nodes. The present invention can meet the high availability of the computing layer while not affecting the storage layer of the database, and can realize the separation of computing and storage, and can be widely applied in the field of distributed databases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed databases, and particularly to a method and system for realizing high availability of stateless computing instances in a distributed database. Background Art

[0002] Currently, distributed databases usually adopt the architecture as Figure 1 shown, which splits data into several shards, and each shard is processed by a different node with no data sharing between nodes. This can solve the problem that the computing and storage capabilities of distributed databases linearly increase with the number of nodes.

[0003] However, this structure has a major defect, that is, computing and storage are bound, the elastic capabilities of database computing and storage are insufficient, the computing instances are "stateful", and the realization of their high availability is also relatively complex. Summary of the Invention

[0004] In view of the above problems, the object of the present invention is to provide a method and system for realizing high availability of stateless computing instances in a distributed database, so as to achieve high availability of stateless computing instances in a distributed database where computing and storage have been separated.

[0005] To achieve the above object, the present invention adopts the following technical solutions: On the one hand, it provides a method for realizing high availability of stateless computing instances in a distributed database, including:

[0006] Monitor the states of all computing instances of the database to determine the offline computing instances;

[0007] Analyze the state of the computing node where the offline computing instance is located, and according to the analyzed state, restart the offline computing instance or schedule the offline computing instance to other computing nodes.

[0008] Further, the monitoring the states of all computing instances of the database to determine the offline computing instances includes:

[0009] Monitor the states of all computing instances of the database and analyze the execution states of all computing instances;

[0010] Determine the offline computing instances according to the execution states of all computing instances.

[0011] Further, the computing instances include four execution states, where:

[0012] The first execution state: the computing instance is online, and the current computing node of the computing instance is the same as the target computing node;

[0013] The second execution state: the computing instance is online, and the current computing node of the computing instance is different from the target computing node;

[0014] The third execution state: The computing instance is offline, and the current computing node of the computing instance is different from the target computing node;

[0015] The fourth execution state: The computing instance is offline, and the current computing node of the computing instance is the same as the target computing node.

[0016] Furthermore, the state monitoring of the computing instance includes monitoring whether the computing instance is alive and can respond to the probe message in a timely manner.

[0017] Furthermore, the state monitoring can be performed in a probing manner or by actively reporting heartbeat packets.

[0018] Furthermore, analyzing the state of the computing node where the offline computing instance is located, and according to the analyzed state, restarting the offline computing instance or scheduling the offline computing instance to other computing nodes includes:

[0019] A) Monitoring the state of the computing nodes of the database and analyzing the state of the computing node where the offline computing instance is located;

[0020] B) If the state of the computing node is normal, go to step C); otherwise, go to step D);

[0021] C) Restarting the offline computing instance. If the restart count of the computing instance exceeds a preset threshold, go to step D);

[0022] D) Scheduling the offline computing instance to other computing nodes.

[0023] Furthermore, the specific process of step D) is as follows:

[0024] Select a computing node with a normal state in the cluster and set this computing node as the target computing node to which the offline computing instance needs to be scheduled;

[0025] Set the current running node of the offline computing instance as the target computing node to be scheduled;

[0026] Find the port used by the offline computing instance on the target computing node and start the offline computing instance at this port on the target computing node.

[0027] On the other hand, a system for implementing high availability of stateless computing instances in a distributed database is provided, including:

[0028] An offline computing instance determination module, configured to monitor the states of all computing instances of the database and determine the offline computing instances;

[0029] A computing node status analysis module is used to analyze the status of the computing nodes where the offline computing instances are located, and based on the analyzed status, restart the offline computing instances or schedule the offline computing instances to other computing nodes.

[0030] On the other hand, a processing device is provided, including computer program instructions, where when the computer program instructions are executed by the processing device, they are used to implement the steps corresponding to the method for realizing high availability of stateless computing instances in the above-mentioned distributed database.

[0031] On the other hand, a computer-readable storage medium is provided, and computer program instructions are stored on the computer-readable storage medium, where when the computer program instructions are executed by a processor, they are used to implement the steps corresponding to the method for realizing high availability of stateless computing instances in the above-mentioned distributed database.

[0032] Due to the above technical solutions adopted by the present invention, it has the following advantages:

[0033] 1. By monitoring the execution status of computing instances and the status of the computing nodes where the offline computing instances are located, and operating on the offline computing instances according to the monitored status, the present invention can achieve high availability of the computing layer in the distributed database.

[0034] 2. The present invention divides computing instances into four execution statuses, which can avoid confusing the restart of computing instances due to accidental termination with the restart due to scheduling.

[0035] 3. While meeting the high availability of the computing layer, the present invention does not affect the storage layer of the database, can achieve the separation of computing and storage, and can be widely applied in the field of distributed databases. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a schematic diagram of the architecture of a distributed database in the prior art;

[0037] Figure 2 is a flowchart of the method provided by an embodiment of the present invention;

[0038] Figure 3 is a schematic diagram of the scheduling of computing instances provided by an embodiment of the present invention;

[0039] Figure 4 is a schematic diagram of the four execution statuses of computing instances provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.

[0041] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and therefore specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless explicitly indicated as the order of performance. It should also be understood that additional or alternative steps may be used.

[0042] Although the terms first, second, third, etc. may be used herein to describe multiple elements, components, regions, layers, and / or sections, these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms may be used only to distinguish one element, component, region, layer, or section from another. Unless the context clearly indicates otherwise, terms such as "first" and "second" and other numerical terms when used herein do not imply an order or sequence. Thus, a first element, component, region, layer, or section discussed below may be referred to as a second element, component, region, layer, or section without departing from the teachings of the example embodiments.

[0043] Distributed databases usually adopt a method of separating computing from storage to solve problems. The method and system for achieving high availability of stateless computing instances in the distributed database provided by the embodiments of the present invention enable stateless computing instances to achieve high availability in a distributed database where computing and storage have been separated. The present invention monitors the execution status of computing instances and the status of the computing nodes where the offline computing instances are located, and operates on the offline computing instances according to the monitored status. When the computing node is normal, it indicates that the failure of the computing instance is caused by the computing instance itself. At this time, the present invention solves it by restarting the computing instance. To prevent misjudgment of "problems with the instance itself", the present invention increases the limit on the number of restarts. If the number of restarts exceeds the preset threshold, the present invention still solves it by scheduling the offline computing instance to other computing nodes. When it is clearly found that the computing node fails, the computing instance is immediately scheduled to other computing nodes to achieve high availability of the computing layer in the distributed database.

[0044] Embodiment 1

[0045] As Figure 2 shown, this embodiment provides a method for achieving high availability of stateless computing instances in a distributed database, including:

[0046] 1) Monitor the status of all computing instances in the database to determine the offline computing instances. Specifically:

[0047] Monitor the status of all computing instances in the database, analyze the execution status of all computing instances, and determine the offline computing instances.

[0048] Specifically, monitoring the status of computing instances includes monitoring whether the computing instances are alive and whether they can respond to detection messages in a timely manner. More specifically, status monitoring can be carried out in a detection manner or by actively reporting heartbeat packets.

[0049] 2) Analyze the status of the computing nodes where the offline computing instances are located, and according to the analyzed status, restart the offline computing instances or schedule the offline computing instances to other computing nodes. Specifically:

[0050] 2.1) Monitor the status of the computing nodes in the database and analyze the status of the computing nodes where the offline computing instances are located.

[0051] Specifically, monitoring the status of computing nodes includes monitoring whether the computing nodes are online, whether the reading and writing of the database data directory on the computing nodes are normal, etc. More specifically, the status of the computing nodes can be actively set to faulty, for example: when the computing nodes need to be replaced.

[0052] 2.2) If the status of the computing node is normal, that is, the computing node is online and the database data directory on the computing node can be read and written normally, then go to step 2.3); otherwise, go to step 2.4).

[0053] 2.3) Restart the offline computing instance. If the number of restarts of the computing instance exceeds the preset threshold, then go to step 2.4).

[0054] If the status of the computing node is normal, it indicates that the failure is caused by the computing instance itself. Therefore, the offline computing instance can be restarted. To prevent misjudgment of "problems with the computing instance itself", the present invention increases the limit on the number of restarts of the computing instance. If the number of restarts of the computing instance exceeds the preset threshold, the following method of scheduling the computing instance to other computing nodes is still used to solve the problem.

[0055] 2.4) As Figure 3 shown, schedule the offline computing instance to other computing nodes:

[0056] 2.4.1) Select a computing node with normal status in the cluster and set this computing node as the target computing node to which the offline computing instance needs to be scheduled.

[0057] 2.4.2) Set the current running node of the offline computing instance as the target computing node to be scheduled.

[0058] 2.4.3) Find the port used by the offline computing instance on the target computing node and start the offline computing instance at this port on the target computing node.

[0059] In a preferred embodiment, to avoid confusion between the restart of the computing instance due to an unexpected termination and the restart due to scheduling, the database needs to maintain three attributes for each computing instance: the online status, the current computing node, and the scheduling target node. Different combinations of the values of these three attributes can divide the computing instances into four execution states, as Figure 4 shown:

[0060] ① The first execution state: The computing instance is online, and the current computing node of the computing instance is the same as the target computing node;

[0061] ② The second execution state: The computing instance is online, and the current computing node of the computing instance is different from the target computing node;

[0062] ③ The third execution state: The computing instance is offline, and the current computing node of the computing instance is different from the target computing node;

[0063] ④ The fourth execution state: The computing instance is offline, and the current computing node of the computing instance is the same as the target computing node.

[0064] Among the above four execution states, the first execution state is the normal operation state; when a computing instance needs to be scheduled, it is in the second execution state; when a computing instance has stopped and is ready for scheduling, it is in the third execution state; when a computing instance has completed scheduling and is ready to be restarted, it is in the fourth execution state.

[0065] Embodiment 2

[0066] This embodiment provides an implementation system for high availability of stateless computing instances in a distributed database, including:

[0067] An offline computing instance determination module, configured to monitor the status of all computing instances in the database and determine the offline computing instances.

[0068] A computing node status analysis module, configured to analyze the status of the computing node where the offline computing instance is located, and according to the analyzed status, restart the offline computing instance or schedule the offline computing instance to other computing nodes.

[0069] In a preferred embodiment, the offline computing instance determination module includes:

[0070] A first status monitoring unit, configured to monitor the status of all computing instances in the database and analyze the execution status of all computing instances.

[0071] An offline computing instance determination unit, configured to determine the offline computing instances according to the execution status of all computing instances.

[0072] In a preferred embodiment, the computing node status analysis module includes:

[0073] A second status monitoring unit, configured to monitor the status of the computing nodes in the database, analyze the status of the computing node where the offline computing instance is located, if the status of the computing node is normal, enter the computing instance restart unit; otherwise, enter the computing instance scheduling unit.

[0074] A computing instance restart unit, configured to restart the offline computing instance, if the number of restarts of the computing instance exceeds a preset threshold, enter the computing instance scheduling unit.

[0075] A computing instance scheduling unit, configured to schedule the offline computing instance to other computing nodes.

[0076] Embodiment 3

[0077] This embodiment provides a processing device corresponding to the method for implementing high availability of stateless computing instances in the distributed database provided in Embodiment 1. The processing device can be a processing device for a client, such as a mobile phone, a laptop computer, a tablet computer, a desktop computer, etc., to execute the method of Embodiment 1.

[0078] The processing device includes a processor, a memory, a communication interface, and a bus. The processor, the memory, and the communication interface are connected through the bus to complete communication with each other. A computer program that can run on the processing device is stored in the memory. When the processing device runs the computer program, it executes the implementation method of high availability of stateless computing instances in the distributed database provided in Embodiment 1 of the present invention.

[0079] In some implementations, the memory may be a high-speed random access memory (RAM: Random Access Memory), and may also include non-volatile memory, such as at least one disk memory.

[0080] In other implementations, the processor may be various types of general-purpose processors such as a central processing unit (CPU), a digital signal processor (DSP), etc., which are not limited herein.

[0081] Embodiment 4

[0082] This embodiment provides a computer program product corresponding to the implementation method of high availability of stateless computing instances in the distributed database provided in Embodiment 1 of the present invention. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing the implementation method of high availability of stateless computing instances in the distributed database described in Embodiment 1 of the present invention are uploaded.

[0083] The computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above.

[0084] The above embodiments are only used to illustrate the present invention. The structures, connection methods, manufacturing processes, etc. of each component can all be changed. Any equivalent transformation and improvement made on the basis of the technical solution of the present invention should not be excluded from the protection scope of the present invention.

Claims

1. A method for implementing high availability of stateless computing instances in a distributed database, characterized in that, including: monitor the status of all computing instances in the database to determine the offline computing instances, where the status monitoring of the computing instances includes monitoring whether the computing instances are alive and whether they can respond to probe messages in a timely manner; analyze the status of the computing nodes where the offline computing instances are located, and based on the analyzed status, restart the offline computing instances or schedule the offline computing instances to other computing nodes; the monitoring the status of all computing instances in the database to determine the offline computing instances includes: monitor the status of all computing instances in the database and analyze the execution status of all computing instances; the computing instances include four execution statuses, where: the first execution status: the computing instance is online, and the current computing node of the computing instance is the same as the target computing node; the second execution status: the computing instance is online, and the current computing node of the computing instance is different from the target computing node; the third execution status: the computing instance is offline, and the current computing node of the computing instance is different from the target computing node; the fourth execution status: the computing instance is offline, and the current computing node of the computing instance is the same as the target computing node; determine the offline computing instances according to the execution status of all computing instances; the analyzing the status of the computing nodes where the offline computing instances are located, and based on the analyzed status, restarting the offline computing instances or scheduling the offline computing instances to other computing nodes includes: A) Monitor the status of the computing nodes in the database and analyze the status of the computing nodes where the offline computing instances are located; B) If the status of the computing node is normal, go to step C); otherwise, go to step D); C) Restart the offline computing instance. If the restart count of the computing instance exceeds a preset threshold, go to step D); D) Schedule the offline computing instance to other computing nodes. The specific process is: Select a computing node with normal status in the cluster and set this computing node as the target computing node to which the offline computing instance needs to be scheduled; Set the current running node of the offline computing instance as the target computing node to be scheduled; Find the port used by the offline computing instance on the target computing node and start the offline computing instance at this port on the target computing node.

2. The method for implementing high availability of stateless computing instances in a distributed database according to claim 1, characterized in that, The status monitoring can be performed in a probing manner or by actively reporting heartbeat packets.

3. A system for implementing high availability of stateless computing instances in a distributed database, characterized in that, including: an offline computing instance determination module, configured to monitor the status of all computing instances in the database to determine the offline computing instances, where the status monitoring of the computing instances includes monitoring whether the computing instances are alive and whether they can respond to probe messages in a timely manner; a computing node status analysis module, configured to analyze the status of the computing nodes where the offline computing instances are located, and based on the analyzed status, restart the offline computing instances or schedule the offline computing instances to other computing nodes; the monitoring the status of all computing instances in the database to determine the offline computing instances includes: monitor the status of all computing instances in the database and analyze the execution status of all computing instances; the computing instances include four execution statuses, where: the first execution status: the computing instance is online, and the current computing node of the computing instance is the same as the target computing node; The second execution state: The computing instance is online, and the current computing node of the computing instance is different from the target computing node; The third execution state: The computing instance is offline, and the current computing node of the computing instance is different from the target computing node; The fourth execution state: The computing instance is offline, and the current computing node of the computing instance is the same as the target computing node; Determine the offline computing instances according to the execution states of all computing instances; Analyze the state of the computing node where the offline computing instance is located, and according to the analyzed state, restart the offline computing instance or schedule the offline computing instance to other computing nodes, including: A) Monitor the state of the computing nodes in the database and analyze the state of the computing node where the offline computing instance is located; B) If the state of the computing node is normal, go to step C); otherwise, go to step D); C) Restart the offline computing instance. If the number of restart times of the computing instance exceeds the preset threshold, go to step D); D) Schedule the offline computing instance to other computing nodes. The specific process is as follows: Select a computing node with a normal state in the cluster and set this computing node as the target computing node to which the offline computing instance needs to be scheduled; Set the current running node of the offline computing instance as the target computing node to be scheduled; Find the port used by the offline computing instance on the target computing node and start the offline computing instance at the target computing node with this port.

4. A processing device, characterized in that, Including computer program instructions, wherein when the computer program instructions are executed by a processing device, they are used to implement the steps corresponding to the method for realizing high availability of stateless computing instances in the distributed database described in any one of claims 1-2.

5. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, wherein when the computer program instructions are executed by a processor, they are used to implement the steps corresponding to the method for realizing high availability of stateless computing instances in the distributed database described in any one of claims 1-2.

Citation Information

Patent Citations

  • A database failure switching method and a device based on a SAS dual control device

    CN109471759A

  • Instance control method, node, terminal and distributed storage system

    CN111385352A