Engine instance management method and apparatus, and electronic device
By adjusting the distribution of Engines during failures and recovery in a distributed cluster, the problem of uneven Engine distribution on business nodes is solved, achieving uniform resource utilization and balanced performance pressure.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XINHUASAN INFORMATION TECH CO LTD
- Filing Date
- 2024-06-20
- Publication Date
- 2026-04-21
AI Technical Summary
In a distributed cluster, it is difficult to ensure that the Engines are distributed evenly across business nodes, resulting in uneven resource utilization.
In a distributed cluster, when a business node fails, the auxiliary Engine of the failed node is enabled on the business nodes in normal condition, and the distribution of the Engine is adjusted when the failure is recovered, so that the Engine is evenly distributed on the business nodes in normal condition.
It achieves a balanced distribution of Engines across business nodes in a distributed cluster, ensuring uniform utilization of resources and avoiding uneven distribution of business traffic and performance pressure.
Smart Images

Figure CN118827686B_ABST
Abstract
Description
Technical Field
[0001] This application relates to network communication technology, and in particular to methods, apparatus and electronic devices for managing engine instances (or simply engines). Background Technology
[0002] In a distributed cluster, each business node uses a configured Engine to support various data acceleration services such as data caching and cloud storage data deduplication. For example, ... Figure 1 As shown, the three business nodes in the distributed cluster, namely node1 to node3, each carry their respective data acceleration services through a configured Engine.
[0003] To ensure resource balance among business nodes in a distributed cluster, it is often required that the engines distributed across these nodes be evenly distributed. Once the engines are evenly distributed, each engine will utilize the resources of its respective business node evenly when handling data acceleration services, thus guaranteeing resource balance among the business nodes in the distributed cluster. However, in practical applications, it is often difficult to ensure that the engines are evenly distributed across the business nodes. Summary of the Invention
[0004] This application provides a method, apparatus, and electronic device for managing engine instances to achieve balanced distribution of Engines across business nodes in a distributed cluster.
[0005] This application provides a method for managing engine instances, which is applied in a distributed cluster where the number of subordinate engines N on each business node in the distributed cluster is equal; the method includes:
[0006] If the first business node in the distributed cluster fails, when the number N of the first business node’s auxiliary engines is greater than 1, the auxiliary engines that originally belonged to the first business node will be enabled on at least two business nodes that are currently in normal condition, so that the auxiliary engines that originally belonged to the first business node are evenly distributed on the business nodes that are in normal condition in the distributed cluster.
[0007] If it is discovered that the first service node that failed also carries auxiliary engines that originally belonged to other failed service nodes, then the auxiliary engines that originally belonged to other service nodes and were carried by the first service node will be enabled on at least one service node that is currently in normal condition, so as to balance the distribution of engines on the service nodes in normal condition in the distributed cluster.
[0008] This application provides a method for managing engine instances, applied in a distributed cluster where each business node in the distributed cluster has an equal number of subordinate engines; the method includes:
[0009] If the first service node in the distributed cluster recovers from the failure, the auxiliary Engine of the first service node is enabled on the first service node, and the auxiliary Engines that were originally established on other service nodes in the distributed cluster and belonged to the first service node are disabled.
[0010] For each currently faulty node, perform the following adjustment steps: disable at least one auxiliary engine of the faulty node carried by at least one service node in normal condition other than the first service node, and enable the disabled auxiliary engine in the first service node, so that the auxiliary engines of each faulty node are evenly distributed among the currently normal service nodes.
[0011] This application also provides an electronic device. The electronic device includes: a processor and a machine-readable storage medium;
[0012] The machine-readable storage medium stores machine-executable instructions that can be executed by the processor;
[0013] The processor is used to execute machine-executable instructions to implement the steps of the methods disclosed above.
[0014] As can be seen from the above technical solutions, in this application, when the first business node in the distributed cluster fails, the auxiliary Engines originally belonging to the first business node are first evenly distributed to the business nodes in the distributed cluster that are in normal condition. Then, the auxiliary Engines originally belonging to other business nodes carried by the first business node are evenly distributed to the business nodes in the distributed cluster that are in normal condition, so as to finally achieve the balanced distribution of Engines on the business nodes in the distributed cluster that are in normal condition.
[0015] Furthermore, in this embodiment, if the first service node in the distributed cluster recovers from a fault, the following adjustment steps are performed on each faulty node: at least one auxiliary Engine of the faulty node carried by at least one service node in normal state other than the first service node is disabled, and the disabled auxiliary Engine is enabled on the first service node, so that the auxiliary Engines of each faulty node are evenly distributed among the service nodes in normal state. This ultimately achieves a balanced distribution of Engines carried by each service node in normal state when a service node in the distributed cluster recovers from a fault. Attached Figure Description
[0016] Figure 1 A schematic diagram illustrating how business nodes in a distributed cluster host the Engine;
[0017] Figure 2 A flowchart of a first method provided for an embodiment of this application;
[0018] Figure 3 A flowchart of the second method provided in the embodiments of this application;
[0019] Figure 4 This is a structural diagram of the device provided in the embodiments of this application;
[0020] Figure 5 Another device structure diagram provided for embodiments of this application;
[0021] Figure 6 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0022] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, and to make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, further detailed descriptions are provided below in conjunction with the accompanying drawings.
[0023] See Figure 2 , Figure 2 This is a flowchart of a first method provided in an embodiment of this application. This process is applied to a management controller. The management controller can be deployed on a business node in a distributed cluster; this embodiment does not specifically limit this.
[0024] In this embodiment, each service node in the distributed cluster carries at least one auxiliary engine. Here, the number of auxiliary engines carried by any service node is configured according to business requirements. To ensure resource balance among the service nodes in the distributed cluster, initially, the number of auxiliary engines carried by each service node in the distributed cluster is equal. For example, initially, the number of auxiliary engines carried by each service node in the distributed cluster is N, where N is greater than or equal to 1.
[0025] In actual business operations, business nodes in a distributed cluster often fail (in this case, the failed business node can be simply referred to as the failed node). Once a business node fails, the following will be executed: Figure 2 The process shown is as follows:
[0026] like Figure 2 As shown, the process may include the following steps:
[0027] Step 201: If the first business node in the distributed cluster fails, when the number N of the first business node's auxiliary engines is greater than 1, the auxiliary engines originally belonging to the first business node are enabled on at least two business nodes that are currently in normal condition, so that the auxiliary engines originally belonging to the first business node are evenly distributed on the business nodes in normal condition in the distributed cluster.
[0028] The term "first business node" here refers to any business node in the distributed cluster. It is a name used for ease of description and is not intended to be limiting.
[0029] It should be noted that the auxiliary engines originally belonging to the first business node are evenly distributed across the business nodes in the distributed cluster that are in normal condition. This includes, but is not limited to, the number of auxiliary engines originally belonging to the first business node carried by each business node in normal condition being exactly equal. In this embodiment, as long as the difference in the number of auxiliary engines originally belonging to the first business node carried by each business node in normal condition is within a set range (e.g., the difference is less than or equal to 1), it will avoid the business traffic and performance pressure of each business node in normal condition being uneven due to a large difference in the number of auxiliary engines originally belonging to the first business node carried by each business node in normal condition.
[0030] For example, if business node 3 fails, and business node 3 has 8 auxiliary engines, and there are currently 6 business nodes in normal condition (i.e., business nodes 1 to 2, and business nodes 4 to 7), if only business nodes 1 and 4 in normal condition are assigned to carry the 8 auxiliary engines of business node 3, for example, business node 1 carries 5 auxiliary engines of business node 3 and business node 4 carries 3 auxiliary engines of business node 3, then the business traffic and performance pressure originally carried by business node 3 will be borne by business node 1 for 5 / 8 and by business node 4 for the remaining 3 / 8. This will result in a large difference in the number of auxiliary engines originally belonging to business node 3 carried by each business node in normal condition, leading to uneven business traffic and performance pressure among the business nodes in normal condition. This embodiment avoids this defect through step 201. For example, according to step 201, the eight auxiliary engines of business node 3 can be carried on business nodes 1 to 2 and business nodes 4 to 7 respectively. Specifically, business nodes 1 to 2 each carry two auxiliary engines of business node 3, and business nodes 4 to 7 each carry one auxiliary engine of business node 3. This avoids the uneven distribution of business traffic and performance pressure on business nodes in normal operation caused by a large difference in the number of auxiliary engines originally belonging to business node 3 carried by each business node in normal operation. It is equivalent to evenly distributing the auxiliary engines originally belonging to business node 3 across the business nodes in normal operation in the distributed cluster.
[0031] As for how to determine which service nodes in normal operation should enable the auxiliary engines originally belonging to the first service node when the number N of the auxiliary engines of the first service node is greater than 1, an example will be given below, and will not be elaborated here.
[0032] Step 202: If the first service node also carries auxiliary engines that originally belonged to other service nodes that have failed, then enable the auxiliary engines that originally belonged to other service nodes carried by the first service node on at least one service node that is currently in normal condition, so as to balance the distribution of engines on the service nodes in normal condition in the distributed cluster.
[0033] Here, the Engine distributed on any business node in a normal state in the distributed cluster includes the auxiliary Engine of that business node, as well as the auxiliary Engines of other business nodes that have failed and are carried by that business node.
[0034] It should be noted that the Engine balancing distributed across business nodes in a normal state in a distributed cluster includes, but is not limited to, ensuring that the number of Engines carried by each business node in a normal state is completely equal. In this embodiment, as long as the difference in the number of Engines carried by each business node in a normal state is within a set range (e.g., the difference is less than or equal to 1), it will prevent the large difference in the number of Engines carried by each business node in a normal state from causing uneven business traffic and performance pressure on each business node in a normal state.
[0035] Step 202 ensures that the distribution of Engines across the normally functioning business nodes in the distributed cluster is balanced. How to determine which normally functioning business nodes should activate the auxiliary Engines originally belonging to other business nodes will be described with examples below, and will not be elaborated upon here.
[0036] This concludes the process. Figure 2 The process is shown below.
[0037] As can be seen, through the above steps 201 and 202, when the first business node in the distributed cluster fails, the auxiliary engines originally belonging to the first business node are first evenly distributed to the business nodes in the distributed cluster that are in normal condition. Then, the auxiliary engines originally belonging to other business nodes carried by the first business node are evenly distributed to the business nodes in the distributed cluster that are in normal condition, so as to ultimately achieve the balanced distribution of engines on the business nodes in the distributed cluster that are in normal condition.
[0038] The following describes steps 201 and 202 in detail:
[0039] As an example, step 201 above, enabling the auxiliary Engine originally belonging to the first service node on at least two service nodes currently in normal operation, may include:
[0040] Step a1: If it is found that the N auxiliary engines of the first service node meet the requirement of being evenly distributed to each service node currently in normal status, then enable T1 auxiliary engines originally belonging to the first service node on each service node currently in normal status; otherwise, proceed to step a2.
[0041] Here, the requirement that the N subordinate engines of the first service node are completely and evenly distributed among the service nodes currently in normal status means that the number of subordinate engines N of the first service node is an integer multiple of the total number of service nodes currently in normal status x1.
[0042] In this embodiment, T1 is the quotient of N / X1.
[0043] Step a2: On each service node that is currently in normal condition, enable T1 auxiliary Engines that originally belonged to the first service node, and then execute step a3.
[0044] It should be noted that in steps a1 and a2, in accordance with the principle that the same auxiliary Engine is prohibited from being enabled on at least two service nodes that are currently in normal status, T1 auxiliary Engines that originally belonged to the first service node can be enabled on each service node that is currently in normal status, so as to avoid the same auxiliary Engine of the first service node being enabled repeatedly on at least two service nodes.
[0045] Step a3: For each of the remaining T2 auxiliary engines that originally belonged to the first service node, find a first target node that meets the requirements, and enable the auxiliary engine on the first target node.
[0046] Where T2 is the difference between N and T1.
[0047] Here, the first target node that meets the requirements is the business node found among all business nodes currently in a normal state that satisfies the following criteria: it carries the smallest number of auxiliary engines originally belonging to the first business node, and it currently carries the smallest total number of engines. Following these requirements, the auxiliary engines originally belonging to the first business node will ultimately be evenly distributed across the business nodes in a normal state within the distributed cluster. See the example description in the following embodiment for details.
[0048] Similarly, in step 202 above, enabling the auxiliary Engine originally belonging to other service nodes, which is carried by the first service node, on at least one service node currently in normal operation includes:
[0049] Step b involves finding a suitable second target node for each auxiliary Engine of any other failed service node carried by the first service node, and enabling that auxiliary Engine on the second target node. The second target node should have the minimum number of auxiliary Engines of the failed service node and the minimum total number of all Engines currently carried. This ultimately achieves balanced Engine distribution across service nodes in a normal state within the distributed cluster. See the example description in the following embodiments for details.
[0050] It should be noted that, to facilitate the management of the auxiliary engines of each business node in the distributed cluster, this embodiment introduces the engine instance distribution. The engine instance distribution records the distribution of the auxiliary engines of each business node across all business nodes in the distributed cluster (specifically, the number of auxiliary engines of each business node distributed across all business nodes in the distributed cluster), and the situation where each business node in the distributed cluster hosts the auxiliary engines of all business nodes in the distributed cluster (specifically, the number of auxiliary engines of all business nodes in the distributed cluster hosted by each business node).
[0051] As an example, the distribution of the aforementioned engine instances can be represented by an M*M two-dimensional matrix. M represents the total number of business nodes in the distributed cluster. Each row in this two-dimensional matrix is configured to correspond to one of the business nodes in the distributed cluster. Different rows correspond to different business nodes in the distributed cluster, and the business nodes corresponding to all rows constitute all the business nodes in the distributed cluster. For example, the first row corresponds to node1, the second row corresponds to node2, and so on. Each column in this two-dimensional matrix is also configured to correspond to one of the business nodes in the distributed cluster. Different columns correspond to different business nodes in the distributed cluster, and the business nodes corresponding to all columns constitute all the business nodes in the distributed cluster. For example, the first column corresponds to node1, the second column corresponds to node2, and so on.
[0052] In this embodiment, each row of the two-dimensional matrix represents the distribution of the auxiliary engines on the corresponding business node across all business nodes in the distributed cluster (e.g., the number of auxiliary engines of the corresponding business node carried by all business nodes in the distributed cluster). Each column of the two-dimensional matrix represents the distribution of auxiliary engines of each business node in the distributed cluster carried by the corresponding business node (e.g., the number of auxiliary engines of each business node in the distributed cluster carried by the corresponding business node).
[0053] Example Description: Assume a distributed cluster has 7 business nodes (denoted as node1 to node7), and each business node initially hosts 8 auxiliary engines. The distribution of engine instances can then be represented by a 7x7 two-dimensional matrix. Each row of this matrix is configured with a corresponding business node; for example, the first row corresponds to node1, the second row to node2, and so on. Each column is also configured with a corresponding business node; for example, the first column corresponds to node1, the second column to node2, and so on. Each row of the matrix represents the distribution of auxiliary engines on the corresponding business node across the 7 business nodes in the distributed cluster, and each column represents the distribution of auxiliary engines on the 7 business nodes in the distributed cluster hosted by the corresponding business node. Table 1 below shows the matrix illustrating the initial engine instance distribution of the distributed cluster.
[0054] 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0055] Table 1
[0056] As shown in Table 1, for a 7*7 two-dimensional matrix, an element located by a certain row and a certain column (denoted as table[row][col]) represents the distribution of the subordinate engines of the business node corresponding to that row on the business node corresponding to that col.
[0057] For example, table[1][2] represents the distribution of the subordinate engines of the business node (e.g., node1) in the first row on the business node (e.g., node2) in the second column. If table[1][2] equals 0 in the 7*7 two-dimensional matrix shown in Table 1, it means that the distribution of the subordinate engines of the business node (e.g., node1) in the first row on the business node (e.g., node2) in the second column is as follows: the number of subordinate engines of the business node (e.g., node1) in the first row distributed on the business node (e.g., node2) in the second column is 0.
[0058] For example, table[1][1] represents the distribution of the subordinate Engines of the business node (e.g., node1) in the first row on the business node (e.g., node1) in the first column. Since table[1][1] equals 8 in the 7*7 two-dimensional matrix shown in Table 1, it means that the distribution of the subordinate Engines of the business node (e.g., node1) in the first row on the business node (e.g., node1) in the first column is as follows: the number of subordinate Engines of the business node (e.g., node1) in the first row distributed on the business node (e.g., node1) in the first column is 8.
[0059] Based on the above description, in step 201, if the first service node in the distributed cluster fails, this embodiment will first look up the distribution of the auxiliary engines on the first service node based on the engine instance distribution, such as the two-dimensional matrix described above. Taking Table 1 as an example, if the first service node is the service node corresponding to the third row (denoted as node3), then based on Table 1, the distribution of the auxiliary engines on node3 is as follows: all auxiliary engines on node3 are distributed on node3, and the number of auxiliary engines on node3 is 8. This achieves the goal of first looking up the distribution of the auxiliary engines on the first service node based on the engine instance distribution.
[0060] Similarly, in step 202, this embodiment will also find the situation of the auxiliary engines of other business nodes carried by the first business node based on the above engine instance distribution situation, such as the two-dimensional matrix mentioned above. Taking Table 1 as an example, if the first business node is the business node corresponding to the third row (denoted as node3), then based on Table 1, it is found that node3 does not currently carry auxiliary engines of other business nodes, so the operation of enabling the auxiliary engines originally belonging to other business nodes carried by the first business node on at least one business node currently in a normal state in step 202 will not be executed; however, if in other cases, node3 does not currently carry auxiliary engines of other business nodes, then the operation of enabling the auxiliary engines originally belonging to other business nodes carried by the first business node on at least one business node currently in a normal state in step 202 will be executed.
[0061] The following describes steps a1 to a3 using the above engine instance distribution as an example, such as the two-dimensional matrix:
[0062] As an example, as described in steps a1 to a3 above, this example can first perform a division operation (i.e., N / X1) between the number N of the first service node's subordinate engines and the number X1 of the service nodes currently in normal status.
[0063] If the result has no remainder, it is assumed that the 8 auxiliary engines of the first service node meet the requirement of being evenly distributed among the currently normal service nodes. Therefore, T1 auxiliary engines originally belonging to the first service node are directly enabled on each currently normal service node. T1 is the quotient of N / X1. Here, the same auxiliary engine originally belonging to the first service node is prohibited from being enabled on at least two currently normal service nodes to avoid duplicate creation of the same auxiliary engine on different service nodes. Furthermore, the sum of the T1 auxiliary engines originally belonging to the first service node enabled on each currently normal service node is equal to the total number of auxiliary engines originally belonging to the first service node.
[0064] If the result has a remainder (denoted as T2), it is considered that the N auxiliary engines of the first service node do not meet the requirement of being evenly distributed among the currently normal service nodes. Therefore, T1 auxiliary engines originally belonging to the first service node are first enabled on the currently normal service nodes. Then, for each of the remaining T2 auxiliary engines originally belonging to the first service node, based on the engine instance distribution shown in the two-dimensional matrix above, recording the distribution of the first service node's auxiliary engines across all normally normal service nodes, a first target node that meets the requirements is found. This first target node has the minimum number of auxiliary engines originally belonging to the first service node, and the minimum number of all engines currently supported by this first target node. Then, this auxiliary engine is enabled on this first target node.
[0065] Of course, if N is less than X1, then the remainder T2 will be N. At this time, operations similar to those performed on each of the T2 auxiliary engines that originally belonged to the first business node can be executed.
[0066] Finally, as described above, when the number N of the auxiliary engines of the first business node is greater than 1 in steps a1 to a3, the auxiliary engines originally belonging to the first business node can be enabled on at least two business nodes that are currently in normal state, so that the auxiliary engines originally belonging to the first business node are evenly distributed on the business nodes in normal state in the distributed cluster.
[0067] Correspondingly, in this embodiment, enabling T1 auxiliary engines originally belonging to the first service node on each service node currently in normal operation may further include updating the recorded engine instance distribution. Specifically, the number of auxiliary engines of the first service node carried by each service node currently in normal operation may be adjusted from an initial value, such as 0, to T1 in the recorded engine instance distribution, and the number of auxiliary engines of the first service node carried by the first service node may be adjusted from N to N-X1*T1 in the recorded engine instance distribution.
[0068] Correspondingly, enabling the auxiliary Engine on the first target node further includes updating the recorded engine instance distribution. Specifically, in the recorded engine instance distribution, the number of auxiliary Engines of the first business node carried by the first target node is adjusted from the current value, for example, T1, to T1+1, and the number of auxiliary Engines of the first business node carried by the first business node is adjusted from the current value to the difference between the current value and 1. This ultimately achieves timely updating of the engine instance distribution, such as the two-dimensional matrix described above.
[0069] The following describes step b using the above engine instance distribution, such as the two-dimensional matrix, as an example:
[0070] Regarding step b above, if the first business node also carries auxiliary engines that originally belonged to any other business node that has failed (for ease of description, this other business node can be referred to as the second business node), then in the specific implementation, for each auxiliary engine originally belonging to the second business node carried by the first business node, based on the distribution of engine instances (such as the distribution of the auxiliary engines of the second business node across all business nodes in the distributed cluster recorded in the two-dimensional matrix above), a second target node that meets the requirements will be found. Then, the auxiliary engine will be enabled on the second target node. Here, the second target node is defined as having the smallest number of auxiliary engines originally belonging to the second business node carried, and the smallest total number of all engines currently carried by the second target node. Ultimately, this achieves balanced distribution of engines across business nodes in a normal state within the distributed cluster (here, the engines distributed across any business node include the auxiliary engines of that business node and the auxiliary engines of other business nodes).
[0071] Correspondingly, enabling the auxiliary engine on the second target node further includes updating the recorded engine instance distribution. Specifically, in the recorded engine instance distribution, the number of auxiliary engines of the second business node carried by the second target node is adjusted from the current value to the sum of the current value and 1, and the number of auxiliary engines of the second business node carried by the first business node is adjusted from the current value to the difference between the current value and 1.
[0072] The following examples illustrate this:
[0073] Assume there are 7 business nodes in the distributed cluster (denoted as node1 to node7), and each business node in the distributed cluster initially carries 8 auxiliary engines. The distribution of engine instances is shown in Table 1.
[0074] If node3 fails, the number of business nodes currently in normal operation in the distributed cluster will be 6 (7-1=6). As shown in the third row of Table 1 (corresponding to node3), it records that node3 carries 8 auxiliary engines, which is greater than 1. Therefore, we first calculate 8 / 6 = 1...2. Finding a remainder, we first enable one auxiliary engine originally belonging to node3 on each of the currently normal business nodes, following the principle that the same auxiliary engine of node3 is prohibited from being enabled on at least two business nodes (the purpose is to avoid the same auxiliary engine originally belonging to node3 being repeatedly created on different business nodes). Table 1 can then be updated to Table 2 as follows:
[0075] 8 0 0 0 0 0 0 0 8 0 0 0 0 0 1 1 1 2 1 1 1 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0076] Table 2
[0077] Next, for one of the remaining two auxiliary engines hosted by node3, the distribution of node3's auxiliary engines across all business nodes in the distributed cluster is first determined based on the data recorded in the third row of Table 2 (corresponding to node3). A first target node that meets the requirements is then identified. According to the above definition of the first target node, node1, node2, and nodes4 through 7 are all qualified first target nodes. Any one of these can be chosen; for example, if node1 is selected, the auxiliary engine will be enabled on node1 (at this point, node1 has the two auxiliary engines that originally belonged to node3). Table 2 can then be updated to Table 3.
[0078] 8 0 0 0 0 0 0 0 8 0 0 0 0 0 2 1 1 1 1 1 1 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0079] Table 3
[0080] As of now, the remaining auxiliary engine hosted by node3 has not yet been assigned. Therefore, for this remaining auxiliary engine hosted by node3, based on the distribution of node3's auxiliary engines across all business nodes in the distributed cluster as recorded in the third row of Table 2 (corresponding to node3), a first target node that meets the requirements is found. According to the above definition of the first target node, node2, node4 through node7 are all qualified first target nodes. Any one of them can be chosen; for example, if node2 is chosen, the auxiliary engine will be enabled on node2 (at this point, node2 has the two auxiliary engines originally belonging to node3). Table 3 can then be updated to Table 4.
[0081] 8 0 0 0 0 0 0 0 8 0 0 0 0 0 2 2 0 1 1 1 1 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0082] Table 4
[0083] Ultimately, the auxiliary engines originally belonging to node3 were evenly distributed across the normally functioning business nodes in the distributed cluster. Based on Table 4, it was found that node3 currently does not host auxiliary engines originally belonging to other business nodes, so step 202 above can be skipped.
[0084] If node1 subsequently fails, and node3, which failed previously, has not yet recovered, the number of business nodes currently in a normal state in the distributed cluster will be 5. As shown in the first row of Table 4 (corresponding to node1), the number N of node1's dependent engines is 8, which is greater than 1. Therefore, we first calculate 8 / 5 = 1...3. Finding a remainder, we first enable one dependent engine originally belonging to node1 on each currently normal business node. The dependent engines originally belonging to node1 enabled on each normal business node are different to avoid the same dependent engine originally belonging to node1 being created repeatedly on different business nodes. Table 4 can then be updated to Table 5 as follows:
[0085] 3 1 0 1 1 1 1 0 8 0 0 0 0 0 2 2 0 1 1 1 1 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0086] Table 5
[0087] Next, for one of the remaining three auxiliary engines hosted by node1, the distribution of node1's auxiliary engines across all business nodes in the distributed cluster is first determined based on the data recorded in the first row of Table 5 (corresponding to node1). A suitable first target node is then identified. According to the above definition of the first target node, nodes 4 through 7 are all suitable first target nodes. Any one of them can be chosen; for example, if node4 is selected, the auxiliary engine will be enabled on node4 (at this point, node4 has the two auxiliary engines that originally belonged to node1).
[0088] Table 5 can be updated to Table 6:
[0089] 2 1 0 2 1 1 1 0 8 0 0 0 0 0 2 2 0 1 1 1 1 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0090] Table 6
[0091] As of now, the remaining two auxiliary engines hosted by node1 have not yet been allocated. Therefore, for one of these two auxiliary engines, based on the distribution of node1's auxiliary engines across all business nodes in the distributed cluster as recorded in the first row of Table 6 (corresponding to node1), a suitable first target node is identified. According to the above definition of the first target node, nodes 5 through 7 are all suitable first target nodes. Any one can be chosen; for example, if node5 is selected, the auxiliary engine will be enabled on node5 (at this point, node5 has the two auxiliary engines originally belonging to node1). Table 6 can then be updated to Table 7.
[0092] 1 1 0 2 2 1 1 0 8 0 0 0 0 0 2 2 0 1 1 1 1 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0093] Table 7
[0094] As of now, the remaining auxiliary engine hosted by node1 has not yet been assigned. Based on the distribution of node1's auxiliary engines across all business nodes in the distributed cluster, as recorded in the first row of Table 7 (corresponding to node1), a suitable first target node is identified. According to the above criteria for the first target node, nodes 6 through 7 are all suitable first target nodes. Any one of them can be chosen; for example, selecting node6 will enable the auxiliary engine on node6 (at this point, node6 has the two auxiliary engines originally belonging to node1). Table 7 can then be updated to Table 8.
[0095] 0 1 0 2 2 2 1 0 8 0 0 0 0 0 2 2 0 1 1 1 1 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0096] Table 8
[0097] Ultimately, the auxiliary Engines originally belonging to node1 were evenly distributed across the business nodes in the distributed cluster that were in normal condition.
[0098] Based on the first column of Table 8 (corresponding to node1), it is found that node1 currently hosts two auxiliary engines that originally belonged to other business nodes, namely node3. Therefore, based on the description in step 202 above, for one of the auxiliary engines originally belonging to node3 hosted by node1, based on the distribution of node3's auxiliary engines across all business nodes in the distributed cluster recorded in Table 8 (specifically, the third row corresponding to node3 in Table 8), a suitable second target node is found. According to the constraint of the second target node, node7 is now a suitable second target node. Therefore, the auxiliary engine is enabled on node7 (at this time, node7 has the two auxiliary engines originally belonging to node3). Table 8 can then be updated to Table 9:
[0099] 0 1 0 2 2 2 1 0 8 0 0 0 0 0 1 2 0 1 1 1 2 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0100] Table 9
[0101] Currently, node1 still hosts one auxiliary engine that originally belonged to another business node, node3. Based on the distribution of node3's auxiliary engines across all business nodes in the distributed cluster as recorded in Table 9 (specifically, the third row corresponding to node3 in Table 9), a suitable second target node is found. According to the criteria for the second target node, nodes4 through 6 are all suitable second target nodes. Any one of them can be chosen; for example, if node4 is selected, the auxiliary engine will be enabled on node4 (at this point, node4 has two auxiliary engines that originally belonged to node3). Table 9 can then be updated to Table 10.
[0102]
[0103]
[0104] Table 10
[0105] Ultimately, the auxiliary Engines originally belonging to node3, which were carried by node1, were evenly distributed across the business nodes in the distributed cluster that were in normal condition.
[0106] As can be seen from the above description, in this embodiment, the M*M two-dimensional matrix used to represent the distribution of engine instances can have the following characteristics:
[0107] 1) The sum of all elements in each row is equal. Each element in a row represents the number of auxiliary engines originally belonging to the business node corresponding to that row, which is carried by the business node in the column containing that element. The sum of all elements in each row represents the total number of auxiliary engines for the business node corresponding to that row.
[0108] 2) The difference between all elements in each row is less than or equal to 1 (i.e., does not exceed 1). By ensuring that the difference between all elements in each row is less than or equal to 1, it is possible to achieve uniform distribution of the auxiliary Engine of each business node across all business nodes in normal operation in the distributed cluster.
[0109] 3) The sum of any two columns (the sum of any column refers to the sum of all elements in that column) differs by no more than 1. The sum of any column represents the total number of Engines (including the auxiliary Engines of that business node and the auxiliary Engines originally belonging to other business nodes) carried by the business node corresponding to that column. By ensuring that the sums of any two columns differ by no more than 1, it is possible to guarantee a balanced distribution of Engines (including the auxiliary Engines of that business node and the auxiliary Engines originally belonging to other business nodes) on business nodes in a normal state within the distributed cluster.
[0110] The above describes the failure of a business node in a distributed cluster. The following describes the recovery of a business node from a failure in a distributed cluster:
[0111] See Figure 3 , Figure 3 This is a second flowchart provided for an embodiment of this application. This process applies to the management controller described above. Figure 3 As shown, the process may include the following steps:
[0112] Step 301: If the first service node in the distributed cluster recovers from the failure, then enable the auxiliary Engine of the first service node on the first service node, and disable the auxiliary Engines that were originally established on other service nodes in the distributed cluster that belong to the first service node.
[0113] Step 302: Perform the following adjustment steps for each currently faulty node: disable at least one auxiliary Engine of the faulty node carried by at least one service node in normal state other than the first service node, and enable the disabled auxiliary Engine in the first service node, so that the auxiliary Engines of each faulty node are evenly distributed among the service nodes in normal state.
[0114] As an example, if the number of auxiliary engines carried by each service node in the distributed cluster that was in a normal state before the first service node recovered was the same, then the above adjustment steps can be as follows:
[0115] Step 302a: For each faulty node, select a corresponding replacement node from the service nodes that are in normal condition, excluding the first service node. Different faulty nodes correspond to different replacement nodes. Disable the E1 auxiliary engines originally belonging to the faulty node carried by the replacement node, and enable the E1 auxiliary engines on the first service node to adjust the distribution of auxiliary engines of each faulty node evenly among the service nodes that are currently in normal condition.
[0116] For example, in a distributed cluster, nodes 1 through 3 fail (referred to as failed nodes), while nodes 4 through 7 are in a normal state. Each of the normal nodes 4 through 7 is carrying two auxiliary engines that originally belonged to node 1, two auxiliary engines that originally belonged to node 2, and two auxiliary engines that originally belonged to node 3, respectively. If node 1 recovers from the failure, then the number of auxiliary engines carried by each normal business node in the distributed cluster before node 1's recovery is the same as the number of auxiliary engines carried by the failed nodes. For ease of description, we will refer to this situation as the first case, where each normal business node in the distributed cluster carries the same number of auxiliary engines carried by the failed nodes before the first business node recovers.
[0117] Optionally, in the specific implementation of step 302a above, each faulty node is traversed, and the traversed faulty node is taken as the current faulty node. Additionally, each service node in normal condition (excluding the first service node) is traversed, and the traversed service node is taken as the current replacement node for the current faulty node. The E1 auxiliary engines originally belonging to the current faulty node and carried by the current replacement node are disabled, and the E1 auxiliary engines are enabled in the first service node. If there are still faulty nodes that have not been traversed, the previously untraversed faulty nodes are traversed, and the traversed faulty nodes are taken as the current faulty node. Additionally, each service node in normal condition (excluding the first service node) that has not been traversed is traversed, and the traversed service node is taken as the current replacement node for the current faulty node. The process then returns to the step of disabling the E1 auxiliary engines originally belonging to the current faulty node and enabling the E1 auxiliary engines in the first service node. To achieve the selection of different replacement nodes for different faulty nodes, the E1 auxiliary engines originally belonging to the corresponding faulty node carried by each replacement node are disabled, and the E1 auxiliary engines are enabled on the first service node.
[0118] In this embodiment, E1 is determined based on N and the number of service nodes (including the first service node) currently in normal state X2. For example, E1 is the quotient of N / X2.
[0119] Through step 302a, in the first case, the E1 auxiliary engines originally belonging to the corresponding faulty node and carried on the replacement node corresponding to each faulty node are disabled, and the E1 auxiliary engines are enabled on the first service node.
[0120] As another embodiment, if the number of auxiliary engines of each service node in the distributed cluster that were in normal condition before the first service node recovered was different, then the above adjustment steps can be as follows:
[0121] Step 302b: If PR is greater than or equal to NQ, where PR is the remainder of N and the number of business nodes in normal state before the first business node recovers, multiplied by 2, and NQ is the quotient of N and the number of business nodes in normal state after the first business node recovers, multiplied by 3, then proceed to step 302c. If PR is less than NQ, then proceed to step 302d.
[0122] In this embodiment, if PR is greater than or equal to NQ, it is equivalent to the number of remaining auxiliary engines after the auxiliary engines of any faulty node are evenly distributed among the service nodes in normal condition before the first service node recovers. This number is sufficient to cover the number of auxiliary engines originally belonging to the faulty node that need to be migrated back to the first service node after the first service node recovers. For ease of description, this situation can be referred to as the second case.
[0123] If PR is less than NQ, it means that the number of remaining auxiliary engines after the auxiliary engines of any faulty node are evenly distributed among the service nodes in normal condition before the first service node recovers is insufficient to cover the number of auxiliary engines originally belonging to the faulty node that need to be migrated back to the first service node after the first service node recovers. For ease of description, this situation can be referred to as the third case.
[0124] Step 302c: For each faulty node, determine the business node with the most currently running Engines and the most affiliated Engines of the faulty node from among the currently normal business nodes. Disable the E2 affiliated Engines originally belonging to the faulty node on the business node, and enable the disabled affiliated Engines on the first business node, so that the affiliated Engines of the faulty node are evenly distributed among the normal business nodes in the distributed cluster; E2 is greater than or equal to 1.
[0125] Step 302c enables the following in the second scenario: for each faulty node, one of the business nodes currently in normal operation (excluding the first business node) will be selected to migrate the E2 auxiliary engines of the faulty node back to the first business node, so that the auxiliary engines of each faulty node are evenly distributed across the business nodes (including the first business node) currently in normal operation in the distributed cluster.
[0126] As an example, E2 is NQ.
[0127] Step 302d: From the service nodes currently in normal condition, determine the service node with the largest number of currently carried Engines and the largest number of auxiliary Engines of a faulty node. Disable the E3 auxiliary Engines originally belonging to the faulty node on the service node, and enable the disabled E3 auxiliary Engines on the first service node. For each faulty node, select a corresponding replacement node from the service nodes in normal condition (excluding the first service node). Different faulty nodes correspond to different replacement nodes. Disable the E4 auxiliary Engines originally belonging to the faulty node carried by the replacement node, and enable the E4 auxiliary Engines on the first service node to adjust the distribution of auxiliary Engines of each faulty node evenly among the service nodes currently in normal condition.
[0128] Optionally, in this embodiment, for each currently faulty node, the service node with the largest number of currently running engines and the largest number of associated engines of the faulty node can be determined from among the currently normal service nodes. The E3 associated engines originally belonging to the faulty node on this service node can then be disabled. In this embodiment, E3 is the difference between NQ and PR. By first enabling the E3 associated engines of the faulty node on the first service node, the remaining associated engines after evenly distributing each faulty node to all service nodes except the first service node can cover the number of associated engines originally belonging to the faulty node that need to be migrated back to the first service node after the first service node recovers.
[0129] Next, for each currently failed node, a corresponding replacement node is selected from all the normally functioning service nodes (excluding the first service node). Different failed nodes correspond to different replacement nodes. The E4 auxiliary engines originally belonging to the failed node and hosted on the replacement node are disabled. These E4 auxiliary engines are then enabled on the first service node to adjust the distribution of auxiliary engines from each failed node evenly across the normally functioning service nodes. Here, E4 stands for PR (Publication Release). Ultimately, this ensures that the auxiliary engines originally belonging to the failed node are evenly distributed across the normally functioning service nodes in the distributed cluster.
[0130] This concludes the process. Figure 3 The process is shown below.
[0131] pass Figure 3 The process shown enables the migration of the auxiliary Engine in the first to third scenarios described above.
[0132] It should be noted that, in this embodiment, Figure 3The illustrated process only adjusts the distribution of auxiliary engines from each failed node to the currently functioning service nodes from a local perspective; it does not address the overall distributed cluster. Therefore, to ensure a balanced distribution of engines across all service nodes in the distributed cluster, the following step C is required:
[0133] Step C: If it is found that the number of all Engines currently carried by the first service node does not meet the balance requirement, then at least one auxiliary Engine of at least one faulty node carried by at least one service node in the normal state in the distributed cluster is disabled, and the disabled auxiliary Engine is enabled on the first service node so that the Engines distributed on all service nodes in the normal state in the distributed cluster are balanced.
[0134] As an example, the above-mentioned load balancing requirement can be determined based on the total number of auxiliary engines (N*M) of each service node in the distributed cluster and the total number of service nodes currently in normal condition (X3, including the first service node). N and M are as described above. If the division of N*M and X3 yields a quotient S1, and the number of all engines currently carried by the first service node is less than S1, it indicates that the first service node carries relatively few engines and can continue to share the load of auxiliary engines from the failed node. In this case, it can be considered that the number of all engines currently carried by the first service node does not meet the load balancing requirement. Conversely, if the number of engines carried by the first service node is equal to or greater than S1, it can be considered that the number of all engines currently carried by the first service node meets the load balancing requirement.
[0135] As an example, in step C, disabling at least one auxiliary engine of at least one faulty node carried by at least one service node in a normal state in the distributed cluster, and enabling the disabled auxiliary engine on the first service node may include:
[0136] The process iterates through the faulty nodes, designating each faulty node as the current faulty node. Based on the engine instance distribution, such as the distribution of the faulty node's associated engines across the normally functioning business nodes as recorded in the matrix above, a suitable third target node is found. This third target node carries the largest number of associated engines originally belonging to the current faulty node, and also carries the largest total number of engines. Then, one associated engine of the faulty node is disabled on this third target node, and the disabled associated engine is enabled on the first business node. If the total number of engines carried by the first business node is still less than S1, the process continues iterating through unvisited faulty nodes, returning to the step of designating a visited faulty node as the current faulty node, until the total number of engines carried by the first business node is S1. Following this method, the balanced distribution of engines across the normally functioning business nodes in the distributed cluster is ultimately achieved.
[0137] Step C ultimately achieved balanced distribution of Engines across all business nodes in a normal state throughout the distributed cluster.
[0138] The following example uses an M*M two-dimensional matrix to represent the engine instance distribution, where M represents the total number of business nodes in the distributed cluster. If there are 7 business nodes in the distributed cluster (denoted as node1 to node7), and each business node initially carries 8 auxiliary engines, and if node1, node2, and node3 fail while nodes4 to 7 are functioning normally, then the engine instance distribution is as shown in Table 11.
[0139]
[0140]
[0141] Table 11
[0142] If node1 recovers from the failure, then as described in step 301, the auxiliary Engine for node1 is enabled, and the auxiliary Engines originally belonging to node1 that have been established on other business nodes in the distributed cluster are disabled. The distribution of engine instances at this time is shown in Table 12.
[0143] 8 0 0 0 0 0 0 0 0 0 2 2 2 2 0 0 0 2 2 2 2 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0144] Table 12
[0145] Table 12 shows that before node1 recovered from the failure, the number of auxiliary engines carried by each business node in the distributed cluster that was in normal condition (nodes 4 to 7) was the same. For example, nodes 4 to 7 each carried 2 auxiliary engines that originally belonged to node1, 2 auxiliary engines that originally belonged to node2, and 2 auxiliary engines that originally belonged to node3. This is the first situation mentioned above.
[0146] In this first scenario, as described in step 302, for the faulty node Node2, a business node can be found from the normal business nodes (node4 to node7) excluding node1. This found business node can be any one of node4 to node7, for example, node4. The E1 auxiliary engines originally belonging to node2 and hosted on node4 are disabled (i.e., the E1 auxiliary engines originally belonging to node2 are deleted from node4), and these E1 auxiliary engines are enabled on node1. E1, as described above, has a quotient of 8 / 5 (the quotient is 1). It should be noted that node4 also needs to be marked as a replacement node for node1. The engine instance distribution at this time is shown in Table 13.
[0147]
[0148]
[0149] Table 13
[0150] Next, for the faulty node 3, find one of the normal service nodes (nodes 4 to 7) excluding node 1. Since node 4 has been marked as a replacement node for node 1, select one of the unmarked nodes 5 to 7, for example, node 5. Disable the E1 auxiliary engines originally belonging to node 3 hosted on node 5 (i.e., delete the E1 auxiliary engines originally belonging to node 2 on node 5), and enable the E1 auxiliary engines on node 1. E1, as described above, is a quotient of 8 / 5 (the quotient is 1). It should be noted that node 5 also needs to be marked as a replacement node for node 1. The engine instance distribution at this time is shown in Table 14:
[0151] 8 0 0 0 0 0 0 1 0 0 1 2 2 2 1 0 0 2 1 2 2 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0152] Table 14
[0153] It should be noted that after performing the operations described in step 302 on all faulty nodes, the aforementioned markers on each of the above business nodes must also be deleted.
[0154] Next, as described in step 305, since the total number of auxiliary engines of each business node in the distributed cluster is 7*8, and the total number of business nodes currently in normal state is 5 (X3), the quotient of 7*8 / 5 is 11. This requires that the total number of all engines carried by node1 (including auxiliary engines of node1 and auxiliary engines of faulty nodes) is at least greater than or equal to 11. However, according to Table 14, the total number of all engines carried by node1 (including auxiliary engines of node1 and auxiliary engines of faulty nodes) is 10, which is less than 11. Therefore, the faulty nodes are traversed, and the traversed faulty node, such as node2, is taken as the current faulty node. Based on the distribution of engine instances, such as the distribution of auxiliary engines of the current faulty node, such as node2, in each business node in normal state as recorded in Table 14 above, a third target node that meets the requirements is first found. As mentioned above regarding the limitation of the third target node, the third target node here can be node6 or node7. Taking node6 as an example, disable one of the dependent engines of the failed node, such as node2, on node6 (i.e., delete the E1 dependent engines that originally belonged to node2 on node6), and enable the disabled dependent engine on node1. The distribution of engine instances at this time is shown in Table 15:
[0155] 8 0 0 0 0 0 0 2 0 0 1 2 1 2 1 0 0 2 1 2 2 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0156] Table 15
[0157] Table 15 shows that the total number of Engines hosted by node1 (including the subordinate Engines of node1 and the subordinate Engines of the failed node) is 11, which meets the load balancing requirement, so the current process ends. At this point, the distributed cluster is load-balanced overall.
[0158] Subsequently, if node2 recovers from the failure, as described in step 301, the auxiliary engine for node2 is enabled, and the auxiliary engines originally belonging to node2 that have been established on other service nodes in the distributed cluster are disabled. The distribution of engine instances at this time is shown in Table 16.
[0159] 8 0 0 0 0 0 0 0 8 0 0 0 0 0 1 0 0 2 1 2 2 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0160] Table 16
[0161] Table 16 reveals that before node2 recovered from the fault, the number of auxiliary engines carried by the faulty nodes on each of the normally functioning service nodes in the distributed cluster (nodes 1, 4, to 7) was not the same. NQ was calculated to be 1 (obtained using the formula: 8 / 6 = 1...2), and PR was calculated to be 3 (obtained using the formula: 8 / 5 = 1...3). Since NQ is less than PR, this corresponds to the second scenario described above.
[0162] In this second scenario, as described in step 403, for the faulty node 3, the service node with the largest number of currently running Engines and the largest number of its associated Engines can be determined from among the currently normal service nodes. Table 16 shows that nodes 4, 6, and 7 currently run the largest number of Engines and the largest number of associated Engines of the faulty node 3. Taking node 4 as an example, the E2 associated Engines originally belonging to the faulty node 3 on node 4 are disabled (i.e., the E1 associated Engines originally belonging to node 3 are deleted from node 4), and the disabled associated Engines are enabled on node 2 (i.e., the deleted E1 associated Engines originally belonging to node 2 are enabled from node 2). As described above, E2 (E2 equals NQ), so here E2 is 1. The engine instance distribution at this time is shown in Table 17.
[0163] 8 0 0 0 0 0 0 0 8 0 0 0 0 0 1 1 0 1 1 2 2 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0164] Table 17
[0165] Next, as described in step 305, since the total number of auxiliary engines of each business node in the distributed cluster is 7*8, and the total number of business nodes currently in normal state is 6 (X3), the quotient of 7*8 / 6 is 9. This requires that the total number of all engines carried by node2 (including auxiliary engines of node2 and auxiliary engines of the failed node) be at least greater than or equal to 9. According to Table 17, the total number of all engines carried by node2 (including auxiliary engines of node2 and auxiliary engines of the failed node) is 9, which meets the load balancing requirement, so the current process ends. At this point, the distributed cluster is balanced overall.
[0166] Subsequently, if node3 recovers from the failure, as described in step 301, the auxiliary Engine for node3 is enabled, and the auxiliary Engines originally belonging to node3 that have been established on other service nodes in the distributed cluster are disabled. The distribution of engine instances at this time is shown in Table 18.
[0167] 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 8
[0168] Table 18
[0169] Since there are no more faulty nodes in the entire distributed cluster as of the current time, the current process can be ended. As shown in Table 18, the distribution of engine instances at this time has returned to the initial distribution of engine instances.
[0170] It should be noted that when the above business node, such as the above-mentioned first business node, recovers from a fault to normal, how to determine the affiliated Engine originally belonging to the first business node among all the Engines enabled in other business nodes, or how to determine the affiliated Engine originally belonging to the faulty node. As an embodiment, it can identify the affiliated Engines of different business nodes to distinguish the affiliated Engines of different business nodes.
[0171] As an embodiment, to ensure the identification of the affiliated Engines of different business nodes and distinguish the affiliated Engines of different business nodes, each business node and all Engines (composed of the affiliated Engines of each business node) can be sequentially encoded starting from an initial value C, such as 0. Based on this, for any Engine (denoted as Engine X), the number of the business node to which it belongs is: floor(X / N)+C, where C <= X < M*N. M and N are as described above, and floor represents taking the integer. Similarly, for a node Y, the numbers of its affiliated Engines are [N*Y+C, N*(Y+1)+C), where +C <= Y < N. In this way, the affiliated Engines of different business nodes can be identified in the above manner to distinguish the affiliated Engines of different business nodes.
[0172] The method provided by the embodiments of the present application has been described above. Next, the device provided by the embodiments of the present application will be described:
[0173] 参见 Figure 4 , Figure 4 is the structure diagram of the device provided by the embodiments of the present application. This device is applied in a distributed cluster, and the number N of the affiliated Engines of each business node in the distributed cluster is equal; this device includes:
[0174] A local adjustment unit, configured to, if a first business node in the distributed cluster fails, when the number N of the affiliated Engines of the first business node is greater than 1, enable the affiliated Engines originally belonging to the first business node on at least two business nodes that are currently in a normal state, so that the affiliated Engines originally belonging to the first business node are evenly distributed on the business nodes in the distributed cluster that are in a normal state;
[0175] The global adjustment unit is used to enable the auxiliary engines originally belonging to other business nodes that were originally carried by the first business node on at least one business node that is currently in normal condition if the first business node is found to have the fault. This is to ensure that the engines distributed on the business nodes in normal condition in the distributed cluster are balanced.
[0176] Optionally, enabling the auxiliary Engine originally belonging to the first service node on at least two service nodes currently in normal operation includes:
[0177] If it is found that the N auxiliary engines of the first service node meet the requirement of being completely and evenly distributed among the service nodes currently in normal condition, then T1 auxiliary engines originally belonging to the first service node are activated on each of the service nodes currently in normal condition; otherwise,
[0178] On each service node currently in normal condition, enable T1 auxiliary engines that originally belonged to the first service node. For each of the remaining T2 auxiliary engines that originally belonged to the first service node, find a first target node that meets the requirements and enable the auxiliary engine on the first target node. The first target node has the minimum number of auxiliary engines that originally belonged to the first service node and the minimum number of all engines currently carried by the first target node.
[0179] Among them, the same subordinate Engine that originally belonged to the first business node is prohibited from being enabled on at least two business nodes that are currently in normal status. T1 is the quotient of N / X1, where X1 is the number of business nodes that are currently in normal status; T2 is the difference between N and T1.
[0180] Optionally, enabling T1 auxiliary engines originally belonging to the first service node on each service node currently in normal state further includes: adjusting the number of auxiliary engines of the first service node carried by each service node currently in normal state from the initial value to T1 in the recorded engine instance distribution, and adjusting the number of auxiliary engines of the first service node carried by the first service node from N to N-X1*T1 in the recorded engine instance distribution;
[0181] Enabling the auxiliary Engine on the first target node further includes: adjusting the number of auxiliary Engines of the first business node carried by the first target node from the current value to T1+1 in the recorded engine instance distribution, and adjusting the number of auxiliary Engines of the first business node carried by the first business node from the current value to the difference between the current value and 1 in the recorded engine instance distribution.
[0182] The first target node is determined based on the number of auxiliary engines of the first business node currently carried by each business node in a normal state, as recorded in the engine instance distribution information, and the total number of engines carried by each business node in a normal state, as recorded in the engine instance distribution information.
[0183] Optionally, enabling the auxiliary Engine originally belonging to other service nodes, which is carried by the first service node, on at least one service node currently in normal operation includes:
[0184] For each auxiliary Engine of the failed second service node carried by the first service node, find a second target node that meets the requirements, and enable the auxiliary Engine on the second target node; the number of auxiliary Engines originally belonging to the second service node carried by the second target node is minimized, and the total number of all Engines currently carried by the second target node is minimized; the second service node is any other failed service node carried by the first service node.
[0185] Optionally, enabling the auxiliary engine on the second target node further includes:
[0186] In the recorded engine instance distribution, the number of auxiliary engines of the second business node carried by the second target node is adjusted from the current value to the sum of the current value and 1, and the number of auxiliary engines of the second business node carried by the first business node is adjusted from the current value to the difference between the current value and 1.
[0187] The second target node is determined based on the number of auxiliary engines of the second business node carried by each business node in the distributed cluster that is in a normal state, as recorded by the engine instance distribution, and the total number of engines carried by each business node in a normal state, as recorded by the engine instance distribution.
[0188] Optionally, the distribution of engine instances is represented by an M*M two-dimensional matrix; M represents the total number of business nodes in the distributed cluster;
[0189] Each row in the two-dimensional matrix corresponds to one of the business nodes in the distributed cluster. Different rows correspond to different business nodes in the distributed cluster. The business nodes corresponding to all rows make up all the business nodes in the distributed cluster. Each column in the matrix corresponds to one of the business nodes in the distributed cluster. Different columns correspond to different business nodes in the distributed cluster. The business nodes corresponding to all columns make up all the business nodes in the distributed cluster.
[0190] Each row in the two-dimensional matrix represents the number of auxiliary engines carried by all business nodes in the distributed cluster corresponding to that row; each column in the two-dimensional matrix represents the number of auxiliary engines carried by the business nodes in the distributed cluster corresponding to that column.
[0191] This concludes the process. Figure 4 Structural description of the device shown.
[0192] See Figure 5 , Figure 5 Another device structure diagram provided in this application embodiment. This device is applied in a distributed cluster, wherein each service node in the distributed cluster has an equal number of auxiliary engines; the device includes:
[0193] The first migration unit is used to enable the auxiliary Engine of the first service node on the first service node and disable the auxiliary Engines that were originally established on other service nodes in the distributed cluster if the first service node in the distributed cluster recovers from the failure.
[0194] The second migration unit is used to perform the following adjustment steps for each currently faulty node: disable at least one auxiliary engine of the faulty node carried by at least one service node in normal state other than the first service node, and enable the disabled auxiliary engine in the first service node, so that the auxiliary engines of each faulty node are evenly distributed among the service nodes in normal state.
[0195] Optionally, the adjustment step includes:
[0196] If the number of auxiliary engines of each business node in the distributed cluster that is in normal state before the first business node recovers to normal is the same, then for each faulty node, a corresponding replacement node is selected from the business nodes in normal state other than the first business node. Different faulty nodes correspond to different replacement nodes. The E1 auxiliary engines that originally belonged to the faulty node and were carried by the replacement node are disabled, and the E1 auxiliary engines are enabled in the first business node.
[0197] If PR is greater than or equal to NQ, where PR is the remainder of N multiplied by 2 (the number of service nodes in normal state before the first service node recovers), and NQ is the quotient of N multiplied by 3 (the number of service nodes in normal state after the first service node recovers), then for each faulty node, the service node with the most currently carried Engines and the most attached Engines of the faulty node is determined from among the service nodes currently in normal state. The E2 attached Engines originally belonging to the faulty node on this service node are disabled, and the disabled attached Engines are enabled on the first service node; E2 is greater than or equal to 1.
[0198] If PR is less than NQ, determine the service node with the most currently running Engines and the most affiliated Engines of the faulty node from among the service nodes currently in normal condition. Disable the E3 affiliated Engines of one of the faulty nodes carried by the service node and enable the disabled E3 affiliated Engines on the first service node. For each faulty node, select a corresponding replacement node from among the service nodes in normal condition other than the first service node. Disable the E4 affiliated Engines originally belonging to the faulty node carried by the replacement node and enable the E4 affiliated Engines on the first service node.
[0199] Optionally, after performing the adjustment steps for each currently faulty node, the second migration unit further disables at least one auxiliary engine of at least one faulty node carried by at least one service node in the distributed cluster when it finds that the number of all Engines currently carried by the first service node does not meet the balance requirements, and enables the disabled auxiliary engine on the first service node so that the Engines distributed on all service nodes in the distributed cluster in the normal state are balanced.
[0200] Optionally, the load balancing requirement is determined based on the total number of N*M of the auxiliary engines of each service node in the distributed cluster and the total number of service nodes currently in normal state X3.
[0201] Wherein, if the number of all Engines currently carried by the first service node is less than S1, it means that the number of all Engines currently carried by the first service node does not meet the balance requirement; S1 is the quotient obtained by dividing N*M and X3.
[0202] Optionally, disabling at least one auxiliary engine of at least one faulty node carried by at least one service node in the distributed cluster that is in a normal state, and enabling the disabled auxiliary engine on the first service node includes:
[0203] Traverse the faulty nodes, take the traversed faulty node as the current faulty node, find a third target node that meets the requirements, disable one of the auxiliary engines of the faulty node on the third target node, and enable the disabled auxiliary engine on the first business node; the third target node carries the largest number of auxiliary engines originally belonging to the current faulty node, and the third target node currently carries the largest number of all engines.
[0204] If the number of all Engines currently carried by the first service node does not meet the balance requirement, continue to traverse the untraversed faulty nodes, return to the step of taking the traversed faulty node as the current faulty node, until the number of all Engines currently carried by the first service node meets the balance requirement.
[0205] Correspondingly, embodiments of this application also provide Figure 4 or Figure 5 The hardware structure of the device shown. See also Figure 6 , Figure 6 This is a structural diagram of an electronic device provided in an embodiment of this application. Figure 6 As shown, the hardware structure may include: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above example of this application.
[0206] Based on the same application concept as the above method, this application embodiment also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the method disclosed in the above examples of this application.
[0207] For example, the aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For instance, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.
[0208] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. An engine instance (Engine) management method, characterized by, This method is applied to a distributed cluster, wherein the number N of auxiliary engines of each service node in the distributed cluster is equal; the method includes: If the first business node in the distributed cluster fails, when the number N of the first business node’s auxiliary engines is greater than 1, the auxiliary engines originally belonging to the first business node will be enabled on at least two business nodes that are currently in normal condition, so that the auxiliary engines originally belonging to the first business node are evenly distributed on the business nodes in normal condition in the distributed cluster. If it is subsequently discovered that the first service node that failed also carries auxiliary engines that originally belonged to other failed service nodes, then the auxiliary engines that originally belonged to other service nodes and were carried by the first service node will be enabled on at least one service node that is currently in normal condition. The auxiliary engines that originally belonged to other service nodes and were carried by the first service node will be evenly distributed across the service nodes in normal condition in the distributed cluster, so that the engines distributed across the service nodes in normal condition in the distributed cluster are balanced.
2. The method of claim 1, wherein, The activation of the auxiliary Engine originally belonging to the first service node on at least two service nodes that are currently in normal operation includes: If it is found that the N auxiliary engines of the first service node meet the requirement of being completely and evenly distributed among the service nodes currently in normal condition, then T1 auxiliary engines originally belonging to the first service node are activated on each of the service nodes currently in normal condition; otherwise, On each service node currently in normal condition, enable T1 auxiliary engines that originally belonged to the first service node. For each of the remaining T2 auxiliary engines that originally belonged to the first service node, find a first target node that meets the requirements and enable the auxiliary engine on the first target node. The first target node has the minimum number of auxiliary engines that originally belonged to the first service node and the minimum number of all engines currently carried by the first target node. Where T1 is the quotient of N / X1, X1 is the number of business nodes currently in normal state; T2 is the difference between N and T1.
3. The method of claim 2, wherein, The step of enabling T1 auxiliary engines originally belonging to the first business node on each business node currently in normal state further includes: adjusting the number of auxiliary engines of the first business node carried by each business node currently in normal state from the initial value to T1 in the recorded engine instance distribution, and adjusting the number of auxiliary engines of the first business node carried by the first business node from N to N-X1*T1 in the recorded engine instance distribution; Enabling the auxiliary Engine on the first target node further includes: adjusting the number of auxiliary Engines of the first business node carried by the first target node from the current value to T1+1 in the recorded engine instance distribution, and adjusting the number of auxiliary Engines of the first business node carried by the first business node from the current value to the difference between the current value and 1 in the recorded engine instance distribution. The first target node is determined based on the number of auxiliary engines of the first business node currently carried by each business node in a normal state, as recorded in the engine instance distribution information, and the total number of engines carried by each business node in a normal state, as recorded in the engine instance distribution information.
4. The method of claim 1, wherein, Enabling the auxiliary Engine originally belonging to other service nodes, which is carried by the first service node, on at least one service node currently in normal operation includes: For each auxiliary Engine of a failed second service node carried by a first service node, where the second service node is any other failed service node carried by the first service node, a second target node that meets the requirements is found, and the auxiliary Engine is enabled on the second target node; the number of auxiliary Engines originally belonging to the second service node carried on the second target node is minimized, and the total number of all Engines currently carried by the second target node is minimized.
5. The method of claim 4, wherein, Enabling the auxiliary Engine on the second target node further includes: In the recorded engine instance distribution, the number of auxiliary engines of the second business node carried by the second target node is adjusted from the current value to the sum of the current value and 1, and the number of auxiliary engines of the second business node carried by the first business node is adjusted from the current value to the difference between the current value and 1. The second target node is determined based on the number of auxiliary engines of the second business node carried by each business node in the distributed cluster that is in normal state, as recorded by the engine instance distribution, and the total number of engines carried by each business node in normal state, as recorded by the engine instance distribution.
6. The method according to claim 3 or 5, characterized in that, The distribution of engine instances is represented by an M * M two-dimensional matrix; M represents the total number of business nodes in the distributed cluster; Each row in the two-dimensional matrix corresponds to one of the business nodes in the distributed cluster. Different rows correspond to different business nodes in the distributed cluster. The business nodes corresponding to all rows make up all the business nodes in the distributed cluster. Each column in the matrix corresponds to one of the business nodes in the distributed cluster. Different columns correspond to different business nodes in the distributed cluster. The business nodes corresponding to all columns make up all the business nodes in the distributed cluster. Each row in the two-dimensional matrix represents the number of auxiliary Engines carried by all business nodes in the distributed cluster on the business node corresponding to that row. Each column in the two-dimensional matrix represents the number of auxiliary engines of each business node in the distributed cluster that the corresponding business node carries.
7. A method for managing engine instances, characterized in that, This method is applied to a distributed cluster, in which each business node has an equal number of auxiliary engines; The method includes: If the first service node in the distributed cluster recovers from the failure, the auxiliary Engine of the first service node is enabled on the first service node, and the auxiliary Engines that were originally established on other service nodes in the distributed cluster and belonged to the first service node are disabled. For each currently faulty node, perform the following adjustment steps: disable at least one auxiliary Engine of the faulty node carried by at least one service node in normal condition other than the first service node, and enable the disabled auxiliary Engine in the first service node, so that the auxiliary Engines of each faulty node are evenly distributed among the service nodes in normal condition. If the number of all Engines currently carried by the first service node does not meet the balance requirement, then at least one auxiliary Engine of at least one faulty node carried by at least one service node in the normal state in the distributed cluster is disabled, and the disabled auxiliary Engine is enabled on the first service node, so that the Engines distributed on all service nodes in the normal state in the distributed cluster are balanced.
8. The method of claim 7, wherein, The adjustment steps include: If the number of auxiliary engines of each business node in the distributed cluster that were in normal condition before the first business node recovered was the same, then for each faulty node, a corresponding replacement node is selected from the business nodes in normal condition other than the first business node. Different faulty nodes correspond to different replacement nodes. The E1 auxiliary engines that originally belonged to the faulty node and were carried by the replacement node are disabled, and the E1 auxiliary engines are enabled on the first business node.
9. The method of claim 7, wherein, The adjustment steps include: If PR is greater than or equal to NQ, where PR is the remainder of N multiplied by 2 (the number of service nodes in normal state before the first service node recovers), and NQ is the quotient of N multiplied by 3 (the number of service nodes in normal state after the first service node recovers), then for each faulty node, the service node with the most currently carried Engines and the most attached Engines of the faulty node is determined from among the service nodes currently in normal state. The E2 attached Engines originally belonging to the faulty node on this service node are disabled, and the disabled attached Engines are enabled on the first service node; E2 is greater than or equal to 1. If PR is less than NQ, determine the service node with the most currently running Engines and the most affiliated Engines of the faulty node from among the service nodes currently in normal condition. Disable the E3 affiliated Engines of one of the faulty nodes carried by the service node and enable the disabled E3 affiliated Engines on the first service node. For each faulty node, select a corresponding replacement node from among the service nodes in normal condition other than the first service node. Disable the E4 affiliated Engines originally belonging to the faulty node carried by the replacement node and enable the E4 affiliated Engines on the first service node.
10. The method of claim 7, wherein, The balance requirement is determined based on the total number of N*M of the auxiliary engines of each business node in the distributed cluster and the total number of business nodes currently in normal state X3. Wherein, if the number of all Engines currently carried by the first service node is less than S1, it means that the number of all Engines currently carried by the first service node does not meet the balance requirement; S1 is the quotient obtained by dividing N*M and X3.
11. The method of claim 7, wherein, The method of disabling at least one auxiliary engine of at least one faulty node carried by at least one service node in the control distributed cluster that is in a normal state, and enabling the disabled auxiliary engine on the first service node includes: Traverse the faulty nodes, take the traversed faulty node as the current faulty node, find a third target node that meets the requirements, disable one of the auxiliary engines of the faulty node on the third target node, and enable the disabled auxiliary engine on the first business node; the third target node carries the largest number of auxiliary engines originally belonging to the current faulty node, and the third target node currently carries the largest number of all engines. If the number of all Engines currently carried by the first service node does not meet the balance requirement, continue to traverse the untraversed faulty nodes, return to the step of taking the traversed faulty node as the current faulty node, until the number of all Engines currently carried by the first service node meets the balance requirement.
12. An engine instance (Engine) management apparatus characterized by comprising: This device is used in a distributed cluster, where each service node in the distributed cluster has the same number of auxiliary engines N; the device includes: The local adjustment unit is used to enable the auxiliary engines that originally belonged to the first business node on at least two business nodes that are currently in normal condition if the first business node fails and the number N of the auxiliary engines of the first business node is greater than 1, so that the auxiliary engines that originally belonged to the first business node are evenly distributed on the business nodes that are in normal condition in the distributed cluster. The global adjustment unit is used to enable the auxiliary engines originally belonging to other business nodes that were originally carried by the first business node on at least one business node that is currently in normal condition if it is found that the first business node that failed also carries auxiliary engines originally belonging to other business nodes. This will distribute the auxiliary engines originally belonging to other business nodes carried by the first business node evenly across the business nodes in normal condition in the distributed cluster, so as to make the distribution of engines on the business nodes in normal condition in the distributed cluster balanced.
13. The apparatus of claim 12, wherein, The activation of the auxiliary Engine originally belonging to the first service node on at least two service nodes that are currently in normal operation includes: If it is found that the N auxiliary engines of the first service node meet the requirement of being completely and evenly distributed among the service nodes currently in normal condition, then T1 auxiliary engines originally belonging to the first service node are activated on each of the service nodes currently in normal condition; otherwise, On each service node currently in normal condition, enable T1 auxiliary engines that originally belonged to the first service node. For each of the remaining T2 auxiliary engines that originally belonged to the first service node, find a first target node that meets the requirements and enable the auxiliary engine on the first target node. The first target node has the minimum number of auxiliary engines that originally belonged to the first service node and the minimum number of all engines currently carried by the first target node. Among them, the same subordinate Engine that originally belonged to the first business node is prohibited from being enabled on at least two business nodes that are currently in normal status. T1 is the quotient of N / X1, where X1 is the number of business nodes that are currently in normal status; T2 is the difference between N and T1.
14. The apparatus of claim 13, wherein, The step of enabling T1 auxiliary engines originally belonging to the first business node on each business node currently in normal state further includes: adjusting the number of auxiliary engines of the first business node carried by each business node currently in normal state from the initial value to T1 in the recorded engine instance distribution, and adjusting the number of auxiliary engines of the first business node carried by the first business node from N to N-X1*T1 in the recorded engine instance distribution; Enabling the auxiliary Engine on the first target node further includes: adjusting the number of auxiliary Engines of the first business node carried by the first target node from the current value to T1+1 in the recorded engine instance distribution, and adjusting the number of auxiliary Engines of the first business node carried by the first business node from the current value to the difference between the current value and 1 in the recorded engine instance distribution. The first target node is determined based on the number of auxiliary engines of the first business node currently carried by each business node in a normal state, as recorded in the engine instance distribution information, and the total number of engines carried by each business node in a normal state, as recorded in the engine instance distribution information.
15. The apparatus of claim 12, wherein, Enabling the auxiliary Engine originally belonging to other service nodes, which is carried by the first service node, on at least one service node currently in normal operation includes: For each auxiliary Engine of the failed second service node carried by the first service node, find a second target node that meets the requirements, and enable the auxiliary Engine on the second target node; the number of auxiliary Engines originally belonging to the second service node carried by the second target node is minimized, and the total number of all Engines currently carried by the second target node is minimized; the second service node is any other failed service node carried by the first service node.
16. The apparatus of claim 15, wherein, Enabling the auxiliary Engine on the second target node further includes: In the recorded engine instance distribution, the number of auxiliary engines of the second business node carried by the second target node is adjusted from the current value to the sum of the current value and 1, and the number of auxiliary engines of the second business node carried by the first business node is adjusted from the current value to the difference between the current value and 1. The second target node is determined based on the number of auxiliary engines of the second business node carried by each business node in the distributed cluster that is in a normal state, as recorded by the engine instance distribution, and the total number of engines carried by each business node in a normal state, as recorded by the engine instance distribution.
17. The apparatus of claim 14 or 16, wherein, The distribution of engine instances is represented by an M * M two-dimensional matrix; M represents the total number of business nodes in the distributed cluster; Each row in the two-dimensional matrix corresponds to one of the business nodes in the distributed cluster. Different rows correspond to different business nodes in the distributed cluster. The business nodes corresponding to all rows make up all the business nodes in the distributed cluster. Each column in the matrix corresponds to one of the business nodes in the distributed cluster. Different columns correspond to different business nodes in the distributed cluster. The business nodes corresponding to all columns make up all the business nodes in the distributed cluster. Each row in the two-dimensional matrix represents the number of auxiliary Engines carried by all business nodes in the distributed cluster on the business node corresponding to that row. Each column in the two-dimensional matrix represents the number of auxiliary engines of each business node in the distributed cluster that the corresponding business node carries.
18. An engine instance (Engine) management apparatus characterized by comprising: This device is used in a distributed cluster, where each business node in the distributed cluster has an equal number of auxiliary engines; The device includes: The first migration unit is used to enable the auxiliary Engine of the first service node on the first service node and disable the auxiliary Engines that were originally established on other service nodes in the distributed cluster if the first service node in the distributed cluster recovers from the failure. The second migration unit is used to perform the following adjustment steps for each currently faulty node: disable at least one auxiliary Engine of the faulty node carried by at least one service node in normal state other than the first service node, and enable the disabled auxiliary Engine in the first service node, so that the auxiliary Engines of each faulty node are evenly distributed among the service nodes in normal state. After performing the adjustment steps for each faulty node, the second migration unit further disables at least one auxiliary engine of at least one faulty node carried by at least one service node in the distributed cluster when it finds that the number of all engines currently carried by the first service node does not meet the balance requirements, and enables the disabled auxiliary engine on the first service node so that the engines distributed on all service nodes in the distributed cluster in the normal state are balanced.
19. The apparatus of claim 18, wherein, The adjustment steps include: If the number of auxiliary engines of each business node in the distributed cluster that is in normal state before the first business node recovers to normal is the same, then for each faulty node, a corresponding replacement node is selected from the business nodes in normal state other than the first business node. Different faulty nodes correspond to different replacement nodes. The E1 auxiliary engines that originally belonged to the faulty node and were carried by the replacement node are disabled, and the E1 auxiliary engines are enabled in the first business node. If PR is greater than or equal to NQ, where PR is the remainder of N multiplied by 2 (the number of service nodes in normal state before the first service node recovers), and NQ is the quotient of N multiplied by 3 (the number of service nodes in normal state after the first service node recovers), then for each faulty node, the service node with the most currently carried Engines and the most attached Engines of the faulty node is determined from among the service nodes currently in normal state. The E2 attached Engines originally belonging to the faulty node on this service node are disabled, and the disabled attached Engines are enabled on the first service node; E2 is greater than or equal to 1. If PR is less than NQ, determine the service node with the most currently running Engines and the most affiliated Engines of the faulty node from among the service nodes currently in normal condition. Disable the E3 affiliated Engines of one of the faulty nodes carried by the service node and enable the disabled E3 affiliated Engines on the first service node. For each faulty node, select a corresponding replacement node from among the service nodes in normal condition other than the first service node. Disable the E4 affiliated Engines originally belonging to the faulty node carried by the replacement node and enable the E4 affiliated Engines on the first service node.
20. The apparatus of claim 18, wherein, The balance requirement is determined based on the total number of N*M of the auxiliary engines of each business node in the distributed cluster and the total number of business nodes currently in normal state X3. Wherein, if the number of all Engines currently carried by the first service node is less than S1, it means that the number of all Engines currently carried by the first service node does not meet the balance requirement; S1 is the quotient obtained by dividing N*M and X3.
21. The apparatus of claim 18, wherein, The method of disabling at least one auxiliary engine of at least one faulty node carried by at least one service node in a normal state in a distributed cluster, and enabling the disabled auxiliary engine on the first service node, includes: Traverse the faulty nodes, take the traversed faulty node as the current faulty node, find a third target node that meets the requirements, disable one of the auxiliary engines of the faulty node on the third target node, and enable the disabled auxiliary engine on the first business node; the third target node carries the largest number of auxiliary engines originally belonging to the current faulty node, and the third target node currently carries the largest number of all engines. If the number of all Engines currently carried by the first service node does not meet the balance requirement, continue to traverse the untraversed faulty nodes, return to the step of taking the traversed faulty node as the current faulty node, until the number of all Engines currently carried by the first service node meets the balance requirement.
22. An electronic device, comprising: The electronic device includes: a processor and a machine-readable storage medium; The machine-readable storage medium stores machine-executable instructions that can be executed by the processor; The processor is configured to execute machine-executable instructions to implement the method steps of any one of claims 1-11.
Citation Information
Patent Citations
Fault recovery method and device for distributed file system supporting additional writing
CN116010149A