RabbitMQ cluster automatic recovery method and device and electronic equipment
By determining the first startup node in the RabbitMQ cluster and re-executing the cluster creation operation, the sequential constraint problem at cluster startup is solved, and the robustness and availability of the cluster is improved.
Patent Information
- Application Number
- CN202510112422.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
AI Technical Summary
The RabbitMQ cluster has a startup sequence constraint at startup, resulting in insufficient robustness. If the last closed node cannot be started, the entire cluster will not be restarted.
By detecting the communication between all nodes in the target RabbitMQ cluster and other nodes, the first startup node is determined, and the first startup node is instructed to re-execute the cluster creation operation, unblock the startup sequence constraints, and ensure that the cluster can automatically recover in the event of a failure.
Improves the robustness of the RabbitMQ cluster, ensures that the cluster can be started and run normally in the event of failure, and avoids the problem of cluster unavailability caused by a single node failure.
Smart Images

Figure CN120011144A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of RabbitMQ high availability, and more specifically, to a RabbitMQ cluster automatic recovery method, device and electronic device. Background Art
[0002] RabbitMQ is an open source message queue system that implements the Advanced Message Queuing Protocol (AMQP). It is a message broker that allows applications to communicate through a common messaging protocol without having to worry about network details and locations. Therefore, RabbitMQ is a commonly used message middleware in distributed systems. RabbitMQ supports cluster deployment, which can achieve high availability and load balancing of RabbitMQ.
[0003] However, when all nodes in the RabbitMQ cluster are shut down, there is a startup order constraint when the RabbitMQ cluster is started next time. The last RabbitMQ node shut down needs to be started first, otherwise other RabbitMQ nodes will be blocked when starting and cannot start normally. Therefore, once the last shut down node cannot be started, the RabbitMQ cluster will not be able to restart, and the robustness is insufficient. Summary of the invention
[0004] In view of the defects of the prior art, the purpose of the present application is to provide a RabbitMQ cluster automatic recovery method, device and electronic device, aiming to solve the problem of insufficient robustness of the prior art due to the startup sequence constraints when starting the RabbitMQ cluster.
[0005] To achieve the above objectives, in a first aspect, the present application provides a RabbitMQ cluster automatic recovery method, comprising: Determine the first startup node based on the communication status of all nodes in the target RabbitMQ cluster with other nodes; Instruct all nodes in the target RabbitMQ cluster to start the RabbitMQ service respectively, wherein the first starting node carries the new cluster parameters when starting the RabbitMQ service, and performs the operations of creating a user and creating a vhost.
[0006] In the process of realizing RabbitMQ high availability in RabbitMQ cluster, the present application first determines the first startup node each time the cluster is restarted, and instructs the first startup node to re-execute the cluster creation operation to release the startup order constraint when the RabbitMQ cluster is restarted, avoids the problem that the entire cluster cannot be restarted when a node fails and cannot be started, and improves the robustness of RabbitMQ high availability.
[0007] According to a RabbitMQ cluster automatic recovery method provided by the present application, the first startup node is determined based on the communication status of all nodes in the target RabbitMQ cluster with other nodes, including: For the first node in the target RabbitMQ cluster, the RabbitMQ service port numbers of other nodes are intermittently detected up to a preset number of times. If the RabbitMQ service port number of the second node is communicable within the preset number of times, the second node is determined as the first startup node, otherwise the first node is determined as the first startup node.
[0008] According to a RabbitMQ cluster automatic recovery method provided by the present application, when other nodes in the target RabbitMQ cluster except the first startup node start the RabbitMQ service, a cluster joining operation is performed.
[0009] According to a RabbitMQ cluster automatic recovery method provided by the present application, the method further includes: Detect the RabbitMQ service status of all nodes in the target RabbitMQ cluster; Based on the service status, an error correction operation is performed.
[0010] This application detects problem nodes in real time during cluster operation and removes them. When the nodes are restarted normally, they can automatically join the cluster again, ensuring cluster availability and fault tolerance.
[0011] According to a RabbitMQ cluster automatic recovery method provided by the present application, the detecting the service status of all nodes in the target RabbitMQ cluster includes: Instruct a first startup node in the target RabbitMQ cluster to detect and record RabbitMQ service status of other nodes, and instruct the other nodes to detect and record RabbitMQ service status of the first startup node.
[0012] According to a RabbitMQ cluster automatic recovery method provided by the present application, the error correction operation is performed based on the service status, including: If the RabbitMQ service status of the first startup node is abnormal, a new first startup node is selected through competition, and the original first startup node is removed from the cluster; If the RabbitMQ service status of the other nodes is abnormal, the nodes with the abnormal RabbitMQ service status are directly removed from the cluster.
[0013] In a second aspect, the present application provides a RabbitMQ cluster automatic recovery device, comprising: A determination module, used to determine a first startup node based on the communication status of all nodes in the target RabbitMQ cluster with other nodes respectively; The recovery module is used to instruct all nodes in the target RabbitMQ cluster to start the RabbitMQ service respectively, wherein the first starting node carries the new cluster parameters when starting the RabbitMQ service, and performs the operations of creating a user and creating a vhost.
[0014] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is used to execute the RabbitMQ cluster automatic recovery method described in the first aspect or any possible implementation of the first aspect.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the RabbitMQ cluster automatic recovery method described in the first aspect or any possible implementation of the first aspect.
[0016] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the RabbitMQ cluster automatic recovery method described in the first aspect or any possible implementation of the first aspect.
[0017] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0018] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the prior art: (1) In the process of achieving RabbitMQ high availability in the RabbitMQ cluster, each time the cluster is restarted, the first startup node is first determined and the first startup node is instructed to re-execute the cluster creation operation to release the startup order constraints when the RabbitMQ cluster is restarted. This avoids the problem that the entire cluster cannot be restarted when a node fails and cannot be started, thereby improving the robustness of RabbitMQ high availability.
[0019] (2) Real-time detection of problematic nodes during cluster operation and their removal. When the nodes are restarted, they can automatically join the cluster, thus ensuring cluster availability and fault tolerance. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 It is a flowchart of the RabbitMQ cluster automatic recovery method provided in an embodiment of the present application; Figure 2 This is a schematic diagram of the RabbitMQ node startup process provided in an embodiment of the present application; Figure 3 It is a schematic diagram of the RabbitMQ node detection and error correction process provided by an embodiment of the present application; Figure 4 It is a structural diagram of a RabbitMQ cluster automatic recovery device provided in an embodiment of the present application; Figure 5 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0023] The term "and / or" in this article is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The symbol " / " in this article indicates that the associated objects are in an or relationship, for example, A / B means A or B.
[0024] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0025] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.
[0026] Next, combine Figure 1-Figure 3 The RabbitMQ cluster automatic recovery method provided in the embodiment of the present application is introduced.
[0027] Figure 1 : is a flow chart of the RabbitMQ cluster automatic recovery method provided in the embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps: Step 100, determining a first startup node based on the communication status of all nodes in the target RabbitMQ cluster with other nodes; Optionally, when the RabbitMQ cluster is restarted, a startup module can be used to execute the startup process of the RabbitMQ node. First, it is necessary to determine the first startup node in the RabbitMQ cluster. Specifically, the first startup node is determined by detecting whether all nodes in the RabbitMQ cluster can communicate with other nodes respectively.
[0028] Step 110, instructing all nodes in the target RabbitMQ cluster to start the RabbitMQ service respectively, wherein the first starting node carries the new cluster parameters when starting the RabbitMQ service, and performs the operations of creating a user and creating a vhost.
[0029] After determining the first startup node, start the RabbitMQ service according to the node status of each node. Specifically, when the startup module instructs the first startup node to start the RabbitMQ service, it also needs to instruct it to carry the new cluster parameters and perform the operations of creating users and creating vhosts, so that other nodes can join the cluster created by the first startup node, so as to release the startup order constraints when the RabbitMQ cluster is restarted, and avoid the problem that the entire cluster cannot be restarted when a node fails and cannot be started.
[0030] The present application provides an automatic recovery method for a RabbitMQ cluster. In the process of realizing RabbitMQ high availability in the RabbitMQ cluster, each time the cluster is restarted, the first startup node is first determined, and the first startup node is instructed to re-execute the cluster creation operation to release the startup sequence constraints when the RabbitMQ cluster is restarted, thereby avoiding the problem that the entire cluster cannot be restarted when a node fails and cannot be started, thereby improving the robustness of RabbitMQ high availability.
[0031] In some embodiments, step 100 specifically includes: Step 1001, for the first node in the target RabbitMQ cluster, intermittently detect the RabbitMQ service port numbers of other nodes up to a preset number of times. If the RabbitMQ service port number of the second node can communicate within the preset number of times, the second node is determined as the first startup node, otherwise the first node is determined as the first startup node.
[0032] For each node, you can set its detection interval and detection times, such as setting 10s / time and detection 15 times, each time detecting whether the RabbitMQ service port number of other nodes can communicate. If the RabbitMQ service port number of other nodes is detected to be communicable within the detection times, it will be set as the first startup node, and the first startup node IP will be set as the node IP.
[0033] Optionally, you can use the telnet tool to periodically detect the RabbitMQ service ports of other nodes in the cluster.
[0034] Figure 2 : is a schematic diagram of the RabbitMQ node startup process provided in the embodiment of the present application, such as Figure 2 As shown, in one embodiment of the present application, the RabbitMQ node startup process includes the following steps: 1a, initialization parameters, set the initialization detection times to 0, set the first startup node IP to empty, and set the maximum detection times to 15; 2a. Use the telnet tool to detect the RabbitMQ service port numbers of other nodes in the cluster at regular intervals. If there is a node whose RabbitMQ service port number is accessible, set the first startup node IP to that node IP. Otherwise, increase the detection count by one. 3a, if the first startup node IP is empty and the detection times are less than the maximum detection times, jump to step 2a, otherwise, jump to step 4a; 4a. Set the node status. If the first startup node IP is empty, the current node is marked as the first startup node, otherwise it is marked as another node.
[0035] 5a. Start the RabbitMQ service according to the node status. If the current node is the first startup node, add the new cluster parameters to the startup parameters, and perform operations such as creating a RabbitMQ user and creating a vhost.
[0036] In some embodiments, in step 100, other nodes in the target RabbitMQ cluster except the first startup node perform a cluster joining operation when starting the RabbitMQ service.
[0037] like Figure 2 As shown, after the first startup node is determined, other nodes start the RabbitMQ service normally and perform the cluster joining operation.
[0038] In some embodiments, the method further comprises: Step 120, detecting the RabbitMQ service status of all nodes in the target RabbitMQ cluster; Step 130, performing error correction operations based on the service status.
[0039] Optionally, a dynamic detection module may be responsible for executing the detection and error correction process. Specifically, the dynamic detection module first detects the RabbitMQ service status of all nodes in the RabbitMQ cluster, and then performs error correction operations based on the service status.
[0040] In some embodiments, step 120 specifically includes: Step 1201, instructing a first startup node in a target RabbitMQ cluster to detect and record RabbitMQ service states of other nodes, and instructing other nodes to detect and record RabbitMQ service states of the first startup node.
[0041] The dynamic detection module needs to obtain the node status of the current node. If the current node is the first startup node, the dynamic detection module instructs it to detect the RabbitMQ service status of other nodes. If the current node is other nodes, the dynamic detection module instructs it to detect the RabbitMQ service status of the first startup node.
[0042] Figure 3 : is a schematic diagram of the RabbitMQ node detection and error correction process provided by the embodiment of the present application, such as Figure 3 As shown, in one embodiment of the present application, the RabbitMQ node detection and error correction process is as follows: 1b, get the node status of the current node; 2b. If the current node is the first startup node, the RabbitMQ service of other nodes is periodically detected and the RabbitMQ service status of other nodes is recorded; if the current node is other nodes, the RabbitMQ service status of the first startup node is periodically detected and the RabbitMQ service status of the first startup node is recorded.
[0043] In some embodiments, step 130 specifically includes: Step 1301: if the RabbitMQ service status of the first startup node is abnormal, a new first startup node is selected through competition, and the original first startup node is removed from the cluster; Step 1302: If the RabbitMQ service status of other nodes is abnormal, the nodes with abnormal RabbitMQ service status are directly removed from the cluster.
[0044] The dynamic detection module needs to obtain the node status of the current node. If the current node is the first startup node, other nodes detected to have problems need to be removed from the cluster. If the current node is another node and the RabbitMQ service of the first startup node is detected to have problems, it needs to compete with other nodes for the identity of the first startup node. After the competition is successful, it switches itself to the first startup node and removes the original first startup node from the cluster.
[0045] like Figure 3 As shown, the RabbitMQ node detection and error correction process also includes the following steps: 3b, get the node status of the current node; 4b. If the current node is the first startup node, other nodes with problems will be removed from the cluster. If the current node is another node and the RabbitMQ service of the first startup node is detected to have problems, it needs to compete with other nodes for the identity of the first startup node. After the competition is successful, it will switch itself to the first startup node and remove the original first startup node from the cluster.
[0046] The dynamic detection module can monitor the nodes in the RabbitMQ cluster in real time, and immediately remove problematic nodes when they are found, thus avoiding the problem of cluster services being affected by a single node failure.
[0047] Figure 4 is a structural diagram of a RabbitMQ cluster automatic recovery device provided in an embodiment of the present application, such as Figure 4 As shown, the device includes a determination module 410 and a recovery module 420, wherein: A determination module 410 is used to determine a first startup node based on the communication status of all nodes in the target RabbitMQ cluster with other nodes respectively; The recovery module 420 is used to instruct all nodes in the target RabbitMQ cluster to start the RabbitMQ service respectively, wherein the first starting node carries the new cluster parameters when starting the RabbitMQ service, and performs the operations of creating a user and creating a vhost.
[0048] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method, which will not be repeated here.
[0049] Based on the method in the above embodiment, Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, an embodiment of the present application provides an electronic device, which may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the RabbitMQ cluster automatic recovery method in the above embodiment.
[0050] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the RabbitMQ cluster automatic recovery method described in each embodiment of the present application.
[0051] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the RabbitMQ cluster automatic recovery method in the above embodiment.
[0052] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the RabbitMQ cluster automatic recovery method in the above embodiment.
[0053] It is understandable that the processor in the embodiment of the present application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0054] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0055] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
[0056] It should be understood that the various numerical numbers involved in the embodiments of the present application are only used for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0057] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A RabbitMQ cluster automatic recovery method, characterized in that: include: Determine the first startup node based on the communication status of all nodes in the target RabbitMQ cluster with other nodes; Instruct all nodes in the target RabbitMQ cluster to start the RabbitMQ service respectively, wherein the first starting node carries the new cluster parameters when starting the RabbitMQ service, and performs the operations of creating a user and creating a vhost.
2. The RabbitMQ cluster automatic recovery method according to claim 1, characterized in that: The determining of the first startup node based on the communicability of all nodes in the target RabbitMQ cluster with other nodes respectively includes: For the first node in the target RabbitMQ cluster, the RabbitMQ service port numbers of other nodes are intermittently detected up to a preset number of times. If the RabbitMQ service port number of the second node is communicable within the preset number of times, the second node is determined as the first startup node, otherwise the first node is determined as the first startup node.
3. The RabbitMQ cluster automatic recovery method according to claim 1, characterized in that: When other nodes in the target RabbitMQ cluster except the first startup node start the RabbitMQ service, a cluster joining operation is performed.
4. The RabbitMQ cluster automatic recovery method according to claim 1, characterized in that: The method further comprises: Detect the RabbitMQ service status of all nodes in the target RabbitMQ cluster; Based on the service status, an error correction operation is performed.
5. The RabbitMQ cluster automatic recovery method according to claim 4, characterized in that: The detecting the service status of all nodes in the target RabbitMQ cluster includes: Instruct a first startup node in the target RabbitMQ cluster to detect and record RabbitMQ service status of other nodes, and instruct the other nodes to detect and record RabbitMQ service status of the first startup node.
6. The RabbitMQ cluster automatic recovery method according to claim 5, characterized in that: The performing an error correction operation based on the service status includes: If the RabbitMQ service status of the first startup node is abnormal, a new first startup node is selected through competition, and the original first startup node is removed from the cluster; If the RabbitMQ service status of the other nodes is abnormal, the nodes with the abnormal RabbitMQ service status are directly removed from the cluster.
7. A RabbitMQ cluster automatic recovery device, characterized in that: include: A determination module, used to determine a first startup node based on the communication status of all nodes in the target RabbitMQ cluster with other nodes respectively; The recovery module is used to instruct all nodes in the target RabbitMQ cluster to start the RabbitMQ service respectively, wherein the first starting node carries the new cluster parameters when starting the RabbitMQ service, and performs the operations of creating a user and creating a vhost.
8. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the RabbitMQ cluster automatic recovery method as described in any one of claims 1-6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor executes the RabbitMQ cluster automatic recovery method according to any one of claims 1 to 6.
10. A computer program product, characterized in that When the computer program product runs on a processor, the processor executes the RabbitMQ cluster automatic recovery method according to any one of claims 1 to 6.