A computer cluster management method, computer and computer cluster

By enabling computers within a computer cluster to autonomously determine their master and slave units and broadcast cluster data, the complexity and high cost of computer cluster management are resolved, achieving the effects of simplified management and reduced costs.

CN114218035BActive Publication Date: 2026-03-20UBTECH ROBOTICS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Computer cluster management is complex and costly, and existing technologies require the installation of a central controller, which increases complexity and operating costs.

Method used

By enabling computers in a computer cluster to autonomously determine their master and slave units and broadcast cluster data, the computer cluster can achieve autonomous management, avoiding dependence on a central controller.

Benefits of technology

It simplifies the management process of computer clusters, reduces management costs, ensures data consistency across all computers, and supports efficient computer cluster operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114218035B_ABST
    Figure CN114218035B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of computers, and provides a computer cluster management method, a computer and a computer cluster. The method comprises the following steps: after a first computer receives an online message of a second computer, if the first computer is a host in a first computer cluster, the first computer determines a host in a second computer cluster; if the host of the second computer cluster is determined to be the second computer, the first computer broadcasts a first message, so that the second computer broadcasts first cluster data; after the first computer receives the first cluster data, the first computer updates the cluster data stored by the first computer to the first cluster data. According to the application, a general controller is not needed to manage the computer cluster, and the management of the computers in the computer cluster can be completed by using the computers in the computer cluster, so that the management of the computer cluster is simpler, and the cost is lower.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computer, and particularly relates to a computer cluster management method, a computer and a computer cluster. BACKGROUND

[0002] A computer cluster is composed of multiple computers. Different computers in the computer cluster can run to process different data at the same time, thereby improving the speed of data processing.

[0003] Since multiple computers exist in the computer cluster, in order to ensure that each computer runs in order, the computers in the computer cluster need to be uniformly managed. At present, a total controller is often used to manage each computer in the computer cluster, for example, the total controller is used to determine the master and slave computers in the computer cluster. Since the total controller is needed to manage the computer cluster, when the computer cluster is installed, not only each computer but also a total controller need to be installed, thereby increasing the complexity and use cost of the computer cluster. SUMMARY

[0004] The embodiments of the present application provide a computer cluster management method, a computer and a computer cluster, and can solve the problems of complex management and high use cost of the computer cluster.

[0005] In a first aspect, the embodiments of the present application provide a computer cluster management method, comprising:

[0006] After the first computer receives the online message of the second computer, if the first computer is the master in the first computer cluster, the first computer determines the master in the second computer cluster, wherein the second computer cluster is a cluster formed after the second computer joins the first computer cluster;

[0007] If it is determined that the master of the second computer cluster is the second computer, the first computer broadcasts a first message, wherein the first message is used to instruct the second computer to broadcast first cluster data, and the first cluster data comprises information of the master in the second computer cluster and information of the slave in the second computer cluster;

[0008] After the first computer receives the first cluster data, the first computer updates the cluster data stored in the first computer to the first cluster data.

[0009] In a second aspect, the embodiments of the present application provide a computer, comprising:

[0010] The host determining module is configured to, after the first computer receives the online message of the second computer, if the first computer is a host in a first computer cluster, the first computer determines a host in a second computer cluster, wherein the second computer cluster is a cluster formed after the second computer joins the first computer cluster.

[0011] The message broadcasting module is configured to, if the host of the second computer cluster is determined to be the second computer, the first computer broadcasts a first message, wherein the first message is used to instruct the second computer to broadcast first cluster data, and the first cluster data includes information of the host in the second computer cluster and information of slaves in the second computer cluster.

[0012] The first data updating module is configured to, after the first computer receives the first cluster data, update the cluster data stored in the first computer to the first cluster data.

[0013] In a third aspect, an embodiment of the present application provides a computer, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the computer cluster management method in any of the first aspect when executing the computer program.

[0014] In a fourth aspect, an embodiment of the present application provides a computer cluster, including the computer in the third aspect.

[0015] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executable on a processor to implement the computer cluster management method in any of the first aspect.

[0016] In a sixth aspect, an embodiment of the present application provides a computer program product, and when the computer program product is run on a terminal device, the terminal device executes the computer cluster management method in any of the first aspect.

[0017] Compared with the prior art, the first aspect embodiment of the present application has the beneficial effect that: after the first computer receives the online message of the second computer, if the first computer is the master in the first computer cluster, the first computer determines the master in the second computer cluster; if the master of the second computer cluster is determined to be the second computer, the first computer broadcasts the first message, so that the second computer broadcasts the first cluster data; after the first computer receives the first cluster data, the first computer updates the cluster data stored by the first computer to the first cluster data; the present application does not need a total controller to manage the computer cluster, and the management of the computers in the computer cluster can be completed by using the computers in the computer cluster, so that the management of the computer cluster is simpler and the cost is lower.

[0018] It can be understood that the beneficial effects of the second aspect to the fifth aspect can be referred to the related description in the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0020] Figure 1 is a running flowchart of a computer in a computer cluster provided by an embodiment of the present application;

[0021] Figure 2 is a flowchart of a management method of a computer cluster provided by an embodiment of the present application;

[0022] Figure 3 is a flowchart of the first computer determining the slave running state provided by an embodiment of the present application;

[0023] Figure 4 is a structure diagram of a computer provided by an embodiment of the present application;

[0024] Figure 5 is a structure diagram of a computer provided by another embodiment of the present application. DETAILED DESCRIPTION

[0025] In the following description, specific details such as specific system structures, techniques, etc. are presented in order to thoroughly understand the embodiments of the present application. However, it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details that hinder the description of the present application.

[0026] It should be understood that the term "include" as used in the specification and the appended claims indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0027] It should also be understood that the term "and / or" as used in the specification and the appended claims, means any one of the associated listed items, as well as all possible combinations of the items, and includes the combinations thereof.

[0028] In addition, in the description of the specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0029] In the present specification, the phrase "one embodiment" or "some embodiments" etc. means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Therefore, the phrases "in one embodiment", "in some embodiments", "in other some embodiments", "in further some embodiments" etc. appearing in different places in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variants mean "including but not limited to", unless otherwise specifically emphasized.

[0030] Each computer is a separate, self-contained electronic computer machine. Each computer can process an input task and output the processing result independently. However, with the development of big data, the amount of data to be processed is increasing, and it takes a long time to use only one computer to complete the processing task. In order to adapt to the increasing amount of data, computer clusters emerge as the times require. A computer cluster can include multiple computers, and multiple computers can be parallel, that is, multiple computers can process multiple data at the same time to achieve the purpose of quickly processing multiple tasks. After multiple computers complete the output processing result, the processing results output by multiple computers are summarized to obtain the required data. Since the computer cluster includes multiple computers, the multiple computers need to be managed. The present application proposes a computer cluster management method, which does not need to additionally increase a manager, and can complete the management of each computer in the computer cluster using the computers in the computer cluster.

[0031] A schematic diagram of a computer cluster provided by the present application is shown in FIG. 1. As shown in FIG. 1, the computer cluster includes multiple computers 101, 102, 103, 104 and 105. The computer cluster can include more than five computers, and the number of computers in the computer cluster is not limited. Figure 1The running process of the computers in the computer cluster is described as follows:

[0032] The computer cluster can include one or more computers. In this application, the computers that are online or powered on at the current time can be recorded as a first computer cluster. As the current time changes, the computers can have faults, and therefore, the computers in the first computer cluster can also change. In this application, the computer fault can include computer offline, computer power off, and computer running error, etc.

[0033] If there is a new online or powered-on computer at the current time, the first computer cluster can be updated, the new online computer is recorded as a second computer, the second computer is added to the first computer cluster, and the updated first computer cluster, that is, the first computer cluster with the second computer added, is recorded as a second computer cluster.

[0034] For example, if the current time is 3:10, there are computer A, computer B, computer C, and computer D at the current time. Computer A, computer B, and computer C are computers that are powered on at the current time, and computer A, computer B, and computer C form the first computer cluster. If computer E is online at the current time, computer E is recorded as the second computer, and the first computer cluster and the second computer form the second computer cluster.

[0035] If the current time becomes 4:00, computer C and computer D are online at 4:00, the first computer cluster includes computer C and computer D, if computer F is online and computer A is online at 4:00, the second computer cluster includes computer A, computer C, computer D, and computer F.

[0036] In this embodiment, the computers in the computer cluster can be computers in a local area network, and each computer communicates through Transmission Control Protocol (TCP).

[0037] After each computer comes online, it needs to initialize its stored cluster data. This cluster data can exist as a list. Initialization involves deleting other data from the list, leaving only the computer's own data. This data can include its own identity and performance information. Performance information can include CPU performance, available storage space, and computer architecture. After initializing the cluster data, the computer can create a thread to broadcast its online message, allowing other online computers to receive it. Specifically, when broadcasting its online message, the computer can use a watchdog timer. If the watchdog timer expires (indicating a computer malfunction or goes offline), the cluster data initialization is performed again, and the online message is broadcast again. Additionally, the computer can receive external information via TCP / UDP ports.

[0038] In this embodiment, each computer in the second computer cluster can receive external information. For ease of description, this application uses only one computer in the first computer cluster as an example, hereinafter referred to as the first computer. Other computers in the first computer cluster can also implement the functions that the first computer described in this application can perform.

[0039] like Figure 1 As shown, in one possible implementation, if the first computer receives a network message at the current time, the first computer determines whether the network message is host information or a new computer coming online.

[0040] If the network information is a message broadcast by the second computer indicating that the second computer has come online, the first computer first determines whether a host already exists in its own cluster based on its stored cluster data.

[0041] In this embodiment, if there is a host in the first computer cluster, it is determined whether the first computer itself is a host in the first computer cluster. If the first computer itself is not a host in the first computer cluster, the online message of the second computer is ignored, that is, the online message of the second computer is not processed, and the first computer continues to receive other external information.

[0042] In this embodiment, if the first computer is the host of the first computer cluster, the cluster data stored in the first computer at the current time can be updated first. That is, the information of the second computer is added to the cluster data stored at the current time, and the second computer is added to the first cluster data as a slave, so as to obtain the second cluster data, which includes the information of the current host and the information of the slave.

[0043] After determining that the first computer is the master of the first computer cluster, the first computer determines the master of the second computer cluster according to the information of the second computer and the information of the first computer. Specifically, the master of the computer cluster can be determined according to the performance information of the computer or a preset priority. For example, the master of the second computer cluster can be determined according to the performance information of the first computer and the performance information of the second computer. Alternatively, the master of the second computer cluster can be determined according to a preset method, for example, the computer with high CPU performance is taken as the master. Alternatively, the performance information of the computer can be input into a deep learning model to obtain the master of the second computer cluster.

[0044] Specifically, after determining that the first computer is the master of the first computer cluster, if the first computer determines that the master of the second computer cluster is the second computer, the master of the first computer cluster needs to be updated in the cluster data stored by each first computer in the first computer cluster. The first computer broadcasts first information, and the first information includes master update information and an online message of the first computer. The master update information is used to prompt other computers in the first computer cluster to update the master information, and the other computers in the first computer cluster are computers in the first computer cluster except the first computer. If the first computer determines that the master of the second computer cluster is the second computer, the first computer exits the master mode, the first computer initializes the stored cluster data, and the first computer broadcasts the online message of the first computer.

[0045] After receiving the master update information broadcast by the first computer, the slave in the first computer cluster initializes the cluster data stored by the slave, and the slave in the first computer cluster broadcasts the online message of the slave. At this time, there is no master in each computer in the second computer cluster, each computer in the second computer cluster obtains the online message of other computers, and each computer in the second computer cluster determines the master of the second computer cluster. After determination, the master of the second computer cluster determined by each computer in the second computer cluster is the second computer, and the second computer also determines the master of the second computer cluster according to the received online message. After determining that the second computer is the master of the second computer cluster, the second computer enters the master mode, the second computer updates the stored cluster data to obtain first cluster data, and broadcasts the first cluster data. The first cluster data includes the information of the master in the second computer cluster and the information of the slave in the second computer cluster. The slave in the second computer cluster can be referred to as the fourth computer, and the fourth computer is a computer in the second computer cluster except the second computer. After receiving the first cluster data, the fourth computer updates the cluster data stored by the fourth computer to the first cluster data, and enters the slave mode.

[0046] In addition, the second computer can further create a cluster watchdog thread to send a second request at a preset interval after entering the host mode, the second request being used to determine whether the fourth computer is in an offline state, the second request including handshake information. The fourth computer returns second data to the second computer if the fourth computer is in an online state after receiving the second request. The second computer determines that the fourth computer is in an offline state if the second computer does not receive the second data returned by the fourth computer within a third preset time period after sending the second request. The second computer updates the stored cluster data based on the information of the third computer in the offline state to obtain sixth cluster data, the sixth cluster data not including the information of the fourth computer in the offline state. The second computer broadcasts the sixth cluster data, so that the fourth computer in the online state updates the cluster data stored by itself after obtaining the sixth cluster data.

[0047] In the embodiment, after determining that the first computer is the host of the first computer cluster, if the first computer determines that the second computer cluster host is the first computer, the first computer updates the cluster data of the first computer to second cluster data based on the information of the second computer, and broadcasts the second cluster data. After obtaining the second cluster data, the fourth computer updates the stored cluster data to the second cluster data, so as to ensure that the cluster data in each computer in the second computer cluster is the same, and facilitate the host in the computer cluster to manage the slave. The second cluster data includes the information of the host in the second computer cluster and the information of the slave in the second computer cluster, and the slave in the second computer cluster is the computer in the second computer cluster except the host.

[0048] In the embodiment, if the first computer is the host of the second computer cluster, the first computer sends a first request to the fourth computer. If the first computer does not receive the first data returned by the fourth computer within a first preset time period after sending the first request, the first computer determines that the fourth computer fails. The first computer updates the second cluster data based on the information of the fourth computer that fails to obtain third cluster data, and broadcasts the third cluster data. The third computer in the second computer cluster updates the cluster data stored by itself to the third cluster data after receiving the third cluster data. The third computer is the fourth computer in the second computer cluster except the computer that fails.

[0049] In the embodiment, after receiving the online message of the second computer, the first computer determines that there is no host in the first computer cluster, for example, the host in the first computer cluster fails at the current time, or there are only two computers in the second computer cluster. When there is no host in the first computer cluster, each online computer in the first computer cluster broadcasts its online message so that other online computers can receive the online message. The first computer determines the host in the second computer cluster according to its own information and the information of the second computer. If the first computer determines that it is the host of the second computer cluster, the first computer updates the cluster data stored by itself and broadcasts the updated cluster data. If the first computer determines that the second computer is the host of the second computer cluster, the first computer waits to receive the cluster data sent by the second computer. After determining that it is the host of the second computer cluster according to the information of each computer received, the second computer updates the cluster data stored by itself and broadcasts the updated cluster data. The updated cluster data of the second computer includes the information of the host of the second computer cluster and the information of the slave of the second computer cluster. If the host of the second computer cluster is a computer other than the first computer and the second computer, the host can broadcast the cluster data.

[0050] In a possible implementation, if the network information received by the first computer is the host information sent by the host of the first computer cluster or the host of the second computer cluster. If the host information is the host update information, the first computer initializes the stored cluster data and broadcasts its online message. If the host information is the cluster data, the first computer updates the cluster data stored by itself to the received cluster data.

[0051] The computer cluster management method provided by the embodiment of the application is described in detail below.

[0052] Figure 2 A schematic flowchart of the computer cluster management method provided by the application is shown, and the method is described in detail below with reference to the flowchart. Figure 2 The method is described in detail as follows.

[0053] S101, after receiving the online message of the second computer, if the first computer is the host in the first computer cluster, the first computer determines the host in the second computer cluster.

[0054] In the embodiment, the second computer cluster is formed after the second computer is added to the first computer cluster. The computers that are online at the current time form the first computer cluster.

[0055] In the embodiment, after the second computer is online, the cluster data stored by the second computer is initialized, and the information of the second computer is stored in the cluster data. After the second computer is online, the second computer can also broadcast the online message of the second computer, so that the computers in the first computer cluster can receive the online message of the second computer.

[0056] In the embodiment, the first computer is any computer in the first computer cluster. For the convenience of description, the first computer in the first computer cluster is taken as an example for description, and the computer is recorded as the first computer.

[0057] In the embodiment, after the first computer receives the online message broadcast by the second computer, the first computer needs to determine whether the host exists in the first computer cluster according to the cluster data stored by the first computer. If the first computer determines that the host exists in the first computer cluster, the first computer also needs to determine whether the first computer is the host of the first computer cluster. If the first computer determines that the first computer is the host of the first computer cluster, the first computer can continue to determine the host of the second computer cluster according to the information of the first computer and the information of the second computer. Specifically, the host of the second computer cluster can be determined according to the performance information of each computer.

[0058] S102, if it is determined that the host of the second computer cluster is the second computer, the first computer broadcasts the first message.

[0059] In the embodiment, the first message is used to instruct the second computer to broadcast the first cluster data, and the first cluster data includes the information of the host in the second computer cluster and the information of the slave in the second computer cluster.

[0060] Specifically, the implementation process of step S102 can include:

[0061] S1021, the first computer broadcasts host update information, wherein the host update information is used to instruct other computers in the first computer cluster to broadcast the online message of the other computers, the other computers in the first computer cluster are the computers in the first computer cluster except the first computer, and the first message includes the host update information.

[0062] Specifically, if the first computer determines that the host of the second computer cluster is the second computer, since the first computer is set as the host in the cluster data currently stored by each computer in the first computer cluster, the first computer needs to inform the other computers in the first computer cluster that the host needs to be updated, and the computers in the first computer cluster except the first computer are the slaves of the first computer cluster, so that the slaves in the first computer cluster broadcast the online message of the slaves after receiving the host update information.

[0063] In the embodiment, the on-line message broadcasted by each computer can include information of the computer, such as identification information of the computer, identification information of the computer, performance information of the computer, etc.

[0064] In S1022, the first computer broadcasts an on-line message of the first computer.

[0065] In the embodiment, the on-line message is used to indicate that the second computer determines the host of the second computer cluster based on the received on-line message, and when the second computer determines that the host of the second computer cluster is the second computer, generates the first cluster data based on the received on-line message, broadcasts the first cluster data, and the first message includes the on-line message of the first computer.

[0066] Specifically, the second computer can receive the on-line message broadcasted by each computer in the first computer cluster. The second computer determines the host of the second computer cluster according to each received on-line message. After determining that the second computer is the host of the second computer cluster, the second computer generates the first cluster data based on the information of each computer in the first computer cluster and the information of the second computer, and broadcasts the first cluster data, so that each computer in the first computer cluster updates the cluster data stored by itself to the first cluster data.

[0067] In S103, after receiving the first cluster data, the first computer updates the cluster data stored by the first computer to the first cluster data.

[0068] In the embodiment, the first computer can save the first cluster data in the database of the first computer, and when it is necessary to find the cluster data, it is only necessary to find from the database of the first computer.

[0069] In the embodiment, after the first computer receives the on-line message of the second computer, if the first computer is the host of the first computer cluster, the first computer determines the host of the second computer cluster; if the host of the second computer cluster is determined to be the second computer, the first computer broadcasts the first message; after the first computer receives the first cluster data, the first computer updates the cluster data stored by the first computer to the first cluster data; the application does not need a general controller to manage the computer cluster, and the management of the computer cluster can be completed by using the computers in the computer cluster, so that the management of the computer cluster is simpler and the cost is lower. In addition, setting the cluster data in each computer in the computer cluster to the same cluster data can ensure that the cluster data of each computer is the same, and can facilitate the operation and management of the computers in the computer cluster.

[0070] In one possible implementation, after the first computer determines the hosts in the second computer cluster, the method may further include:

[0071] If the host in the second computer cluster is determined to be the first computer, the first computer updates its cluster data to the second cluster data based on the information of the second computer, and broadcasts the second cluster data.

[0072] In this embodiment, the second cluster data includes information about the host in the second computer cluster and information about the slave in the second computer cluster. The slave in the second computer cluster is a computer in the second computer cluster other than the host. The second cluster data is used to instruct the slave in the second computer cluster to update the cluster data stored in the slave in the second computer cluster to the second cluster data.

[0073] In this embodiment, if the first computer is determined to be the host of the second computer cluster, the first computer needs to store the information of the second computer in its own cluster data as a slave, obtain the second cluster data, and notify each slave in the second computer cluster by broadcasting.

[0074] like Figure 3 As shown, in one possible implementation, after the first computer determines the hosts in the second computer cluster, the method may further include:

[0075] S201, the first computer sends a first request to the slave in the second computer cluster.

[0076] In this embodiment, the first request may include handshake information or preset characters, etc. The first computer sends the first request to the slave machines in the second computer cluster to determine whether each slave machine in the second computer cluster is operating normally and whether it is online.

[0077] S202, if the first computer does not receive the first data returned by the slave in the second computer cluster within a first preset time period after the first computer sends the first request, the first computer determines that there is a faulty slave in the second computer cluster.

[0078] In this embodiment, a slave device malfunction may include being offline, powered off, or malfunctioning.

[0079] S203, the first computer updates the second cluster data to third cluster data based on the information of the slave machine that failed in the second computer cluster.

[0080] In the embodiment, the third cluster data is used to instruct the third computer to update the cluster data stored by the third computer to the third cluster data, and the third computer is a slave computer of the second computer cluster except the slave computer which has failed.

[0081] In the embodiment, if there is a slave computer which has failed in the second computer cluster, the first computer needs to mark or remove the information of the slave computer which has failed in the stored cluster data, and therefore the first computer needs to update the stored cluster data to obtain the third cluster data.

[0082] In the embodiment, the first computer determines the slave computer which has failed and updates the cluster data in time, and therefore the computer cluster can be better managed.

[0083] In a possible implementation, the method further includes:

[0084] If the first computer does not receive the online message of the second computer and the first computer receives the fourth cluster data sent by the master computer in the first computer cluster, the first computer stores the cluster data as the fourth cluster data.

[0085] In the embodiment, if there is no computer which has newly logged on at the current time, the first computer is not the master computer of the first computer cluster. After the master computer in the first computer cluster updates the cluster data of the master computer to the fourth cluster data, the first computer can receive the fourth cluster data sent by the master computer in the first computer cluster and update the cluster data stored by the first computer to the fourth cluster data. The master computer in the first computer cluster updates the cluster data of the master computer in two aspects. One aspect is that the master computer updates the cluster data of the master computer after determining that the master computer is the master computer of the first computer cluster. The other aspect is that when there is a slave computer which has failed in the first computer cluster, the master computer of the first computer cluster needs to update the cluster data of the master computer.

[0086] The fourth cluster data includes the information of the master computer of the first computer cluster and the information of the slave computer of the first computer cluster.

[0087] In a possible implementation, after the first computer determines the master computer in the second computer cluster, the method further includes:

[0088] If there is a master computer in the first computer cluster and the first computer is not the master computer of the first computer cluster, the first computer does not process the online message of the second computer.

[0089] If the first computer receives the second information, the first computer broadcasts the online message of the first computer, wherein the second information is broadcasted by the master in the first computer cluster when it is determined that the master in the first computer cluster is not the master in the second computer cluster.

[0090] In the embodiment, if the first computer is not the master in the first computer cluster and the master in the first computer cluster is not the master in the second computer cluster, the first computer can receive the second information sent by the master in the first computer cluster. After receiving the second information, the first computer can broadcast its own online message so that other computers receive the online message of the first computer.

[0091] For example, if the first computer is computer C, the master in the first computer cluster is computer D, and after computer E in the second computer cluster is online, computer D determines that the master in the second computer cluster is not computer D, computer D needs to broadcast the second information. After computer C receives the second information, computer C broadcasts its own online message.

[0092] In a possible implementation, after it is determined that the master in the second computer cluster is the second computer, the method further comprises:

[0093] S301, if the first computer does not receive the second request sent by the second computer within a second preset time period, the first computer determines that the second computer fails.

[0094] In the embodiment, if the master in the second computer cluster is the second computer and the second computer fails, the first computer will not receive the second request sent by the second computer. The second request is sent by the second computer to each slave in the second computer cluster, and is used to determine whether each slave in the second computer cluster fails. The second request can include handshake information.

[0095] If the first computer does not receive the second request sent by the second computer within a preset time interval or a second preset time period, it can be determined that the second computer may have failed, and each computer in the second computer cluster needs to re-determine the master in the second computer cluster. The second preset time period can be set as needed, for example, the second preset time period can be set to 10 seconds, 5 seconds or 3 seconds, etc.

[0096] S302, the first computer determines the master in the second computer cluster after the second computer fails.

[0097] Specifically, the first computer can re-determine the master in the second computer cluster according to the information of each computer in the cluster data stored by the first computer.

[0098] S303, if the first computer is the host of the second computer cluster after the second computer fails, the first computer updates the first cluster data into fifth cluster data based on the information of the second computer, and broadcasts the fifth cluster data.

[0099] In this embodiment, the second computer cluster after the second computer fails can be recorded as a third computer cluster. The fifth cluster data includes the information of the host and the information of the slaves of the third computer cluster. The slaves of the third computer cluster include the computers in the second computer cluster except the host of the fifth computer cluster and the computer that fails.

[0100] In this embodiment, if the first computer determines that the host of the second computer cluster after the second computer fails is not itself, the first computer needs to continue to receive the cluster data sent by the host of the second computer cluster after the second computer fails, and update the cluster data stored by the first computer into the cluster data sent by the host of the second computer cluster.

[0101] For example, if the host of the second computer cluster is computer D, computer D will broadcast the second request regularly if it is normal. If the first computer F cannot receive the second request, it is determined that computer D fails. The first computer F re-determines the host of the second computer cluster according to the online messages broadcast by each computer in the second computer cluster. If the first computer F determines that the host of the second computer cluster is computer G, computer G will broadcast the fifth cluster data. After the first computer F receives the fifth cluster data, it updates the cluster data stored by itself into the fifth cluster data.

[0102] In a possible implementation, the method can further include:

[0103] After the first computer receives the online message of the second computer, if there is no host in the first computer cluster, the first computer determines the host of the second computer cluster. If the first computer determines that it is the host of the second computer cluster, the first computer updates the cluster data stored by itself into the seventh cluster data, and broadcasts the seventh cluster data.

[0104] In this embodiment, the slaves in the second computer cluster update the cluster data stored by themselves into the seventh cluster data after receiving the seventh cluster data.

[0105] After the first computer receives the online message of the second computer, if there is no host in the first computer cluster, the first computer determines the host of the second computer cluster. If the first computer determines that it is not the host of the second computer cluster, the first computer will continue to receive external information. After the first computer receives the eighth cluster data sent by the host of the second computer cluster, the first computer updates the cluster data stored by the first computer to the eighth cluster data.

[0106] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0107] The computer cluster management method corresponding to the above embodiment, Figure 4 The structural block diagram of the computer provided by the embodiments of the present application is shown, and only the parts related to the embodiments of the present application are shown for the convenience of description.

[0108] Referring to Figure 4 The computer 400 can include a host determination module 410, a message broadcast module 420, and a first data update module 430.

[0109] The host determination module 410 is configured to, after the first computer receives the online message of the second computer, if the first computer is the host in the first computer cluster, the first computer determines the host in the second computer cluster, wherein the second computer cluster is a cluster formed after the second computer joins the first computer cluster.

[0110] The message broadcast module 420 is configured to, if it is determined that the host of the second computer cluster is the second computer, the first computer broadcasts a first message, wherein the first message is used to instruct the second computer to broadcast first cluster data, and the first cluster data includes information of the host in the second computer cluster and information of the slave in the second computer cluster.

[0111] The first data update module 430 is configured to, after the first computer receives the first cluster data, update the cluster data stored by the first computer to the first cluster data.

[0112] In a possible implementation, the message broadcast module 420 can be specifically configured to:

[0113] The first computer broadcasts host update information, wherein the host update information is used to instruct other computers in the first computer cluster to broadcast online messages of the other computers, the other computers being computers in the first computer cluster other than the first computer, and the first message comprises the host update information;

[0114] The first computer broadcasts an online message of the first computer, wherein the online message is used to instruct the second computer to determine a host of the second computer cluster based on the received online message, and when the second computer determines that the host of the second computer cluster is the second computer, generate the first cluster data based on the received online message, broadcast the first cluster data, and the first message comprises the online message of the first computer.

[0115] In a possible implementation, the host determination module 410 is further connected to:

[0116] The second data update module is configured to, if it is determined that the host of the second computer cluster is the first computer, update, by the first computer, the cluster data of the first computer based on the information of the second computer to second cluster data, and broadcast the second cluster data, wherein the second cluster data comprises information of the host of the second computer cluster and information of slaves in the second computer cluster, the slaves being computers in the second computer cluster other than the host, and the second cluster data is used to instruct the slaves in the second computer cluster to update the cluster data stored by the slaves in the second computer cluster to the second cluster data.

[0117] In a possible implementation, the second data update module can be specifically configured to:

[0118] The first computer sends a first request to the slaves in the second computer cluster;

[0119] If the first computer does not receive first data returned by the slaves in the second computer cluster within a first preset time period after the first computer sends the first request, the first computer determines that a slave in the second computer cluster has failed, wherein the failed slave is a computer in the second computer cluster that does not return the first data;

[0120] The first computer updates the second cluster data into third cluster data based on information of the slave computer that fails in the second computer cluster, and broadcasts the third cluster data, wherein the third cluster data is used to instruct a third computer to update cluster data stored in the third computer into the third cluster data, and the third computer is a slave computer of the second computer cluster except the slave computer that fails.

[0121] In a possible implementation, the computer 400 further includes:

[0122] The third data updating module is configured to, if the first computer does not receive the online message of the second computer and the first computer receives fourth cluster data sent by a master computer in the first computer cluster, update the cluster data stored in the first computer into the fourth cluster data.

[0123] In a possible implementation, the master computer determining module 410 can be specifically configured to:

[0124] If there is a master computer in the first computer cluster and the first computer is not the master computer of the first computer cluster, the first computer does not process the online message of the second computer.

[0125] In a possible implementation, the message broadcasting module 420 can be specifically configured to:

[0126] If the first computer does not receive the second request sent by the second computer within a second preset time period, the first computer determines that the second computer fails.

[0127] The first computer determines the master computer of the second computer cluster after the second computer fails.

[0128] If the first computer is the master computer of the second computer cluster after the second computer fails, the first computer updates the first cluster data into fifth cluster data based on information of the second computer, and broadcasts the fifth cluster data.

[0129] It should be noted that the information interaction, execution process and the like between the above apparatuses / units are based on the same concept as the method embodiments of the present application, and the specific functions and the technical effects brought by the same can be referred to the method embodiments part, and will not be described here.

[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the above-described functions. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or software. In addition, the specific name of each functional unit and module is only for easy distinction, and does not limit the protection scope of the present application. The specific working process of the unit and module in the system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0131] The embodiment of the present application also provides a computer, referring to Figure 5 The computer 500 can include at least one processor 510, a memory 520, and a computer program stored in the memory 520 and executable on the at least one processor 510, wherein the processor 510 implements the steps in any of the above method embodiments when executing the computer program, for example Figure 2 Steps S101 to S103 in the embodiment shown. Alternatively, the processor 510 implements the functions of each module / unit in the above apparatus embodiments when executing the computer program, for example Figure 4 The functions of the modules 410 to 430 shown.

[0132] For example, the computer program can be divided into one or more modules / units, one or more modules / units are stored in the memory 520 and executed by the processor 510 to complete the present application. The one or more modules / units can be a series of computer program segments capable of completing a specific function, which are used to describe the execution process of the computer program in the computer 500.

[0133] Those skilled in the art can understand that Figure 5 It is only an example of a computer and does not constitute a limitation on the computer, which can include more or fewer components than shown, or combine certain components, or different components, such as input / output devices, network access devices, buses, etc.

[0134] The processor 510 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0135] The memory 520 can be an internal storage unit of the computer, and can also be an external storage device of the computer, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, or the like. The memory 520 is used to store the computer program and other programs and data required by the computer. The memory 520 can also be used to temporarily store data that has been output or will be output.

[0136] The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, the bus in the drawings of the present application does not limit to only one bus or one type of bus.

[0137] The computer provided by the embodiments of the present application can be applied to terminal devices such as tablet computers, notebook computers, netbooks, personal digital assistants (PDAs), and the like. The embodiments of the present application do not make any limitation on the specific type of the terminal device.

[0138] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0139] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware or in a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solutions. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0140] In the embodiments provided in the present application, it should be understood that the disclosed terminal device, apparatus and method can be implemented by other manners. For example, the terminal device embodiments described above are only schematic, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0141] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0142] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0143] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by computer programs instructing related hardware, and the computer programs can be stored in a computer readable storage medium, and the computer programs can realize the steps of each method embodiment when executed by one or more processors.

[0144] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be implemented by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. The computer program, when executed by one or more processors, can implement the steps of each method embodiment.

[0145] Similarly, as a computer program product, when the computer program product runs on the terminal device, it enables the terminal device to implement the steps in each of the above method embodiments.

[0146] The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0147] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for managing a computer cluster, characterized in that, include: After the first computer receives the online message from the second computer, if the first computer is a host in the first computer cluster, the first computer determines the host in the second computer cluster, wherein the second computer cluster is the cluster formed after the second computer is added to the first computer cluster; If the host of the second computer cluster is determined to be the second computer, the first computer broadcasts a first message, wherein the first message is used to instruct the second computer to broadcast first cluster data, and the first cluster data includes information about the host in the second computer cluster and information about the slave in the second computer cluster; After receiving the first cluster data, the first computer updates the cluster data stored in the first computer with the first cluster data; The first computer broadcasts a first message, including: The first computer broadcasts host update information, wherein the host update information is used to instruct other computers in the first computer cluster to broadcast the online message of the other computers, wherein the other computers in the first computer cluster are computers other than the first computer in the first computer cluster, and the first message includes the host update information; The first computer broadcasts a message indicating that it has come online. The message instructs the second computer to determine the host of the second computer cluster based on the received message. When the second computer determines that the host of the second computer cluster is the second computer, it generates the first cluster data based on the received message and broadcasts the first cluster data. The first message includes the message indicating that it has come online.

2. The computer cluster management method as described in claim 1, characterized in that, After the first computer determines the hosts in the second computer cluster, the process includes: If the host in the second computer cluster is determined to be the first computer, the first computer updates its cluster data to the second cluster data based on the information of the second computer, and broadcasts the second cluster data. The second cluster data includes information about the host in the second computer cluster and information about the slave in the second computer cluster. The second cluster data is used to instruct the slave in the second computer cluster to update the cluster data stored in the slave in the second computer cluster to the second cluster data.

3. The computer cluster management method as described in claim 2, characterized in that, If the host in the second computer cluster is determined to be the first computer, the method further includes: The first computer sends a first request to the slave device in the second computer cluster; If the first computer does not receive the first data returned by the slave in the second computer cluster within a first preset time period after the first computer sends the first request, the first computer determines that the slave in the second computer cluster has failed, wherein the slave that has failed is the computer in the second computer cluster that has not returned the first data; The first computer updates the second cluster data to third cluster data based on the information of the slave machine that has failed in the second computer cluster, and broadcasts the third cluster data. The third cluster data is used to instruct the third computer to update the cluster data stored in the third computer to the third cluster data. The third computer is a slave machine in the second computer cluster other than the slave machine that has failed.

4. The computer cluster management method as described in claim 1, characterized in that, The method further includes: If the first computer does not receive the online message from the second computer, and the first computer receives the fourth cluster data sent by the host in the first computer cluster, the first computer updates the cluster data stored in the first computer to the fourth cluster data.

5. The computer cluster management method as described in claim 1, characterized in that, After the first computer receives the online message from the second computer, the method further includes: If there is a host in the first computer cluster, and the first computer is not a host in the first computer cluster, the first computer will not process the online message of the second computer.

6. The computer cluster management method as described in claim 1, characterized in that, If the host of the second computer cluster is determined to be the second computer, the method further includes: If the first computer does not receive the second request from the second computer within the second preset time period, the first computer determines that the second computer has malfunctioned. The first computer determines the host of the second computer cluster after the second computer fails; If the first computer is the host of the second computer cluster after the second computer fails, the first computer updates the first cluster data to the fifth cluster data based on the information of the second computer and broadcasts the fifth cluster data.

7. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the computer cluster management method as described in any one of claims 1 to 6.

8. A computer cluster comprising the computer as claimed in claim 7.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the computer cluster management method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Network system, management computer, cluster management method, and computer program

    CN101031886A

  • Method of managing nodes in computer cluster

    US20080040628A1