Cluster server scheduling method and device, equipment and medium

By configuring multi-service instances and autonomous detection and voting mechanisms in the node server of the cluster server, the problem of inaccurate scheduling in the node server failure of the distributed system is solved, and fast and accurate server scheduling and improved disaster recovery capabilities are achieved.

CN120162200AInactive Publication Date: 2025-06-17SHENZHEN FEIQUAN CLOUD DATA SERVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510650676.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, there is a problem that a distributed system based on a cluster server cannot perform server scheduling quickly and accurately when a node server fails.

Method used

By configuring two service instances in the node server and connecting to the cloud server and the operator backbone network, an autonomous detection and voting mechanism is realized. The specific steps include: generating health status information based on the operating status of the service instance, determining whether it is the main server, receiving or sending delayed update time, voting and feedback the results to the cloud server, and the cloud server determines the target server identification and configures the main node information.

Benefits of technology

It realizes rapid and accurate scheduling of cluster servers in the case of a main server failure, and improves the disaster recovery capabilities and independent decision-making capabilities of distributed systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162200A_ABST
    Figure CN120162200A_ABST
Patent Text Reader

Abstract

The invention discloses a cluster server scheduling method, device and equipment and a medium, and the method comprises the steps: each node server automatically detects a running state and updates state recording information, and if the node server is not a main server, the node server receives delay update time from the main server and updates a detection time point; if the server is not the main server and the delay update time is not received after the receiving duration, the stored state record information is sent and voting is carried out according to the received state record information; and the cloud server determines a target server identifier and sends the target server identifier to each node server to enable each node server to perform main node information configuration, and the main server generates delay update time in a healthy state and sends the delay update time to other node servers. According to the cluster server scheduling method, the health state information can be automatically detected, and autonomous decision making and voting are carried out to determine and obtain the target server identifier under the condition that the main server fails, so that the distributed cluster server is rapidly and accurately scheduled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cluster servers, and particularly to a cluster server scheduling method, device, equipment and medium. Background Art

[0002] Compared with traditional cluster servers, node servers can be respectively arranged in computer rooms at multiple locations and combined through network connections to form a distributed system based on cluster servers. The distributed consistency protocol is an algorithm or protocol used to ensure data consistency among multiple node servers in a distributed system, thereby improving the risk resistance ability of cluster servers through the distributed system architecture.

[0003] To ensure data consistency among node servers, there are various existing mature solutions for disaster recovery of data centers or computer rooms, such as using third-party services or pre-deploying monitoring points at different locations to test the connectivity to two data centers; establishing a robust heartbeat detection mechanism between two data centers; and providing multiple physical path selection for data transmission through the use of BGP (Border Gateway Protocol) or multi-path routing technology. For the application services in computer rooms that have suffered serious failures, it is impossible to automatically and accurately determine whether the failure occurs locally or remotely due to network reasons. Current various solutions can monitor the failed computer rooms and issue alarms; then the failed computer rooms usually require manual decision-making and intervention to perform configuration adjustments to adapt the cluster servers to the current operating state through server scheduling. However, there are problems such as low reliability of autonomous detection and increased costs brought by multiple monitoring points in the current technical methods, resulting in the inability to quickly and accurately perform server scheduling for the distributed system based on cluster servers when node servers fail. Therefore, there is a problem in the prior art methods that server scheduling for cluster servers cannot be quickly and accurately performed. Summary of the Invention

[0004] Embodiments of the present invention provide a cluster server scheduling method, device, equipment and medium, aiming to solve the problem that server scheduling for cluster servers cannot be quickly and accurately performed in the prior art methods.

[0005] In a first aspect, embodiments of the present invention provide a cluster server scheduling method, where the method is applied to a node server of a cluster server, the cluster server includes node servers configured in each computer room and a cloud server configured outside the cloud, and two service instances are configured in each node server; the cloud server and each node server are both communicatively connected to the operator backbone network to implement data information transmission, and the method includes: If the detection time point corresponding to the set delay time is reached, corresponding health status information is generated according to the running status of the service instances configured internally, and the pre-stored status record information is updated accordingly; Judge whether it is the main server according to the configured main node information; If it is not the main server and the delay update time from the main server is received within the preset receiving duration, update the currently set delay time according to the delay update time and judge whether the detection time point corresponding to the set delay time is reached; If it is not the main server and the delay update time from the main server is not received within the preset receiving duration, send the stored status record information to other node servers; Vote according to the preset voting strategy and the status record information of other received node servers, and feedback the voting result to the cloud server; Obtain the target server identifier determined by the cloud server based on the voting result, and configure the main node information according to the target server identifier; If it is the main server and the health status information is healthy, generate a delay update time corresponding to the status record information according to the preset time parameter configuration rule and send it to other node servers and the cloud server; If it is the main server and the health status information is unhealthy, perform a demotion adjustment on the configured main node information.

[0006] In a second aspect, an embodiment of the present invention further provides a cluster server scheduling device, wherein the device is configured in a node server of a cluster server. The cluster server includes node servers configured in each computer room and a cloud server configured externally in the cloud. Two service instances are configured in each of the node servers; the cloud server and each of the node servers are both communicatively connected to the operator backbone network to implement data information transmission. The device is used to execute the cluster server scheduling method described in the first aspect above. The device includes: A status record information update unit, configured to, if the detection time point corresponding to the set delay time is reached, generate corresponding health status information according to the running status of the service instances configured internally, and update the pre-stored status record information accordingly; A judgment unit, configured to judge whether it is the main server according to the configured main node information; A time point judgment unit, configured to, if it is not the main server and the delay update time from the main server is received within the preset receiving duration, update the currently set delay time according to the delay update time and judge whether the detection time point corresponding to the set delay time is reached; A status record information sending unit, configured to, if it is not the master server and no delayed update time from the master server is received within a preset receiving duration, send the stored status record information to each other node server; A voting result feedback unit, configured to vote according to a preset voting policy and the received status record information of other node servers and feedback the voting result to the cloud server; A master node information configuration unit, configured to obtain the target server identifier determined by the cloud server based on the voting result and perform master node information configuration according to the target server identifier; A delayed update time sending unit, configured to, if it is the master server and the health status information is healthy, generate a delayed update time corresponding to the status record information according to a preset time parameter configuration rule and send it to each other node server and the cloud server; A configuration adjustment unit, configured to, if it is the master server and the health status information is unhealthy, perform a demotion adjustment on the configured master node information.

[0007] In a third aspect, an embodiment of the present invention further provides a computer device, where the device includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store a computer program; The processor is configured to, when executing the program stored on the memory, implement the steps of the cluster server scheduling method described in the first aspect above.

[0008] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. Among them, when the computer program is executed by a processor, the steps of the cluster server scheduling method described in the first aspect above are implemented.

[0009] An embodiment of the present invention provides a method, apparatus, device, and medium for scheduling cluster servers. The method includes: each node server independently detects its operating status and correspondingly updates the status record information, and determines whether it is the master server. If it is not the master server, it receives the delayed update time from the master server and updates the detection time point. If it is not the master server and does not receive the delayed update time after exceeding the receiving duration, it sends the stored status record information and votes according to the received status record information, and feeds back the voting result to the cloud server. The cloud server determines the target server identifier and sends it to each node server so that each node server can configure the master node information. The master server generates the delayed update time in a healthy state and sends it to other node servers. The above-mentioned method for scheduling cluster servers can independently detect the health status information, make autonomous decisions in the case of master server failure, and vote to determine the target server identifier, so as to quickly and accurately schedule the distributed cluster servers. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0011] Figure 1 It is a flowchart of the method for scheduling cluster servers provided by the embodiment of the present invention; Figure 2 It is a schematic diagram of the application scenario of the method for scheduling cluster servers provided by the embodiment of the present invention; Figure 3 It is a schematic block diagram of the device for scheduling cluster servers provided by the embodiment of the present invention; Figure 4 It is a schematic block diagram of the computer device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0013] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0014] It should also be understood that the terminology used in the specification of the present invention is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0015] It should be further understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0016] Embodiments of the present invention application provide a method for scheduling a cluster server. The method is applied to a node server 10 of a cluster server. The node server 10 executes a stored software program to implement the above-mentioned method for scheduling a cluster server; as Figure 2As shown in the figure, the cluster server 20 includes node servers 10 configured in each computer room and cloud servers 30 configured in an external cloud. Two service instances 11 are configured in each of the node servers 10; both the cloud server 30 and each of the node servers 10 are communicatively connected to an operator backbone network 40 to achieve data information transmission. The cloud server 30 can be a public cloud ECS (Elastic Compute Service), and a Zookeeper instance is configured in the public cloud ECS. To improve the disaster tolerance of the above distributed cluster server, the two service instances 11 of each node server 10 can be respectively connected to a network path in the operator backbone network 40. The two network paths can be respectively connected to two operator networks. The two service instances 11 of each node server 10 perform data interaction internally to ensure data consistency between the two service instances 11; the cloud server is simultaneously connected to the two network paths in the operator backbone network 40; then when a network failure occurs in any one of the service instances 11 in a certain node server 10, it will not cause the node server 10 to disconnect from the operator backbone network 40; since two active service instances 11 are configured in the node server 10 in the computer room, the computer room where the node server 10 is configured is also called a dual-active computer room. The service instance 11 can be a Zookeeper instance. A main Key management service and a local memory are also configured in each node server 10. The main Key management service is communicatively connected to the local memory and the two service instances 11 respectively to manage each service instance 11 through the main Key management service, such as configuring and managing the main node information of each service instance 11 and performing data synchronization management on the two service instances 11. The main Key management service is also communicatively connected to a database service configured in the node server. The database service can be a virtual service module built based on a Mysql database or a Redis database. The database service can provide services such as reading and writing data information in the database.

[0017] As Figure 1 shown, the method includes steps S110 to S180.

[0018] S110. If the detection time point corresponding to the set delay time is reached, generate corresponding health status information according to the running status of the service instances configured internally and update the pre-stored status record information correspondingly.

[0019] When the node server reaches the detection time point, it can obtain the running status of the internally configured service instances, generate health status information based on the running status of the service instances, and update the pre-stored status record information. Among them, the detection time point is determined corresponding to the set delay time. Taking the time point of the set delay time as the starting time, the time point after the set delay time is the detection time point.

[0020] In a specific embodiment, step S110 includes sub-steps: obtaining the running status of each of the two configured service instances to obtain running status information; obtaining the response delay and error count information corresponding to the current running period of the service instance; the current running period is the time period between the current detection time point and the previous detection time point; generating corresponding health status information according to the running status information, the response delay and the error count information; and updating the current status record information according to the running status information, the response delay and the error count information.

[0021] Specifically, the running status of each of the two configured service instances can be obtained. Each service instance can run independently and manage data synchronization through the main Key management service to ensure data consistency inside the node server. The running status of the service instance can be normal or abnormal, so the obtained running status information includes the running status of each of the two service instances.

[0022] Furthermore, the response delay and error count information corresponding to the current running period of the service instance can be obtained. The current running period is also the time period between the current detection time point and the previous detection time point. The response delay is the time taken for the service instance to process and respond to external application services. The longer the response delay, the longer the time taken for the application service to be processed. The error count information is the number of errors that occur in the internal program of the service instance during operation.

[0023] Generate corresponding health status information according to the running status information, the response delay and the error count information. Specifically, determine whether the running statuses of the two service instances in the running status information are both normal; if at least one of the service instances has an abnormal running status, the generated health status information is unhealthy. If the running statuses of the two service instances are both normal, further determine whether the response delays of each service instance are not greater than the preset delay threshold, and whether the error count information is not greater than the preset count threshold. If the response delays of each service instance are not greater than the delay threshold, and the number of errors in the error count information is not greater than the count threshold, the generated health status information is healthy; otherwise, the generated health status information is unhealthy.

[0024] Further, update the currently stored status record information to obtain the updated status record information. The status record information includes the health status information, response delay, and error count information corresponding to each detection time point. Add the health status information, response delay, and error count information corresponding to the current detection time point as a new set of information to the status record information for storage, thereby realizing the update of the status record information.

[0025] In a specific embodiment, the obtaining of the response delay and error count information corresponding to the service instance during the current running period includes: obtaining the average response duration of the application service processed by the service instance during the current running period as the corresponding response delay; obtaining the running log corresponding to the service instance during the current running period; matching the error codes in the running log according to a preset code configuration table to statistically obtain the occurrence times information corresponding to each error code; and obtaining the cumulative quantity of each error type as the corresponding error count information according to the code classification information in the code configuration table and the occurrence times information.

[0026] Specifically, the average value of the response duration of the application service processed by the service instance during the current running period can be obtained. The response duration of the application service is also the time interval between the time point when the service instance receives the application service and the time point when the service instance processes it and feeds back the processing result. Obtaining the average value of the response duration of the application service processed during the current running period can be used as the response delay of the service instance.

[0027] Further, the running log corresponding to the service instance during the current running period can be obtained, such as intercepting the log content within the current running period from the log file of the service instance as the running log. Match the error codes in the running log according to the code configuration table, thereby statistically obtaining the hit times of each error code in the running log to obtain the occurrence times information corresponding to each error code.

[0028] The code configuration table includes multiple error types, and each error type corresponds to at least one error code. The error codes included in the error type are the code classification information. For example, the error type can be exception, warning, ordinary error, serious error, etc. The occurrence times of each error type can be accumulated according to the code classification information and the occurrence times information obtained in the above steps, thereby obtaining the cumulative times of each error type. The cumulative times of each error type can be used as the corresponding error count information.

[0029] S120. Determine whether it is the main server according to the configured main node information.

[0030] The node server determines the master node information configured internally, and can determine whether the master node information matches the server identifier (ID identification code) of the node server itself. If the master node information matches its own server identifier, it is determined as the master server; if the master node information does not match its own server identifier, it is determined not to be the master server.

[0031] S130. If it is not the master server and a delayed update time from the master server is received within the preset reception duration, the currently set delay time is updated according to the delayed update time, and it is determined whether the detection time point corresponding to the set delay time is reached.

[0032] If it is not the master server, it is further determined whether a delayed update time from the master server is received within the preset reception duration; there is exactly one master server in the distributed server cluster, and other node servers that are not the master server can be slave servers. The slave servers perform data synchronization by receiving data synchronization messages from the master server, and the slave servers can perform data write operations; the slave servers can also receive external access requests and execute data read operations to obtain corresponding data information for feedback. The reception duration is the time information pre-configured in the node server, and the corresponding reception time point is determined based on the detection time point and the reception duration. For example, if the reception duration is 2 seconds, it is determined whether a delayed update time from the master server is received within 2 seconds after the detection time point.

[0033] If a delayed update time is received, the currently set delay time is updated according to the received delayed update time, that is, the delayed update time is set as the new delay time. After updating the delay time, it can be determined whether the detection time point corresponding to the set delay time is reached, and step S110 is returned for execution.

[0034] S140. If it is not the master server and a delayed update time from the master server is not received within the preset reception duration, the stored status record information is sent to other node servers.

[0035] If the node server is not the master server and a delayed update time from the master server is not received within the preset duration, the currently stored status record information is sent to other node servers. If the network service of the node server is not interrupted, even if the health status information of the node server is unhealthy, the node server can still send the status record information to other node servers through the network. If the network service of the node server is interrupted, the node server is in an offline state. At this time, it is necessary to wait for the network connection to be restored before reconnecting to the cluster server and performing data information synchronization, and then it can be networked and run.

[0036] S150: Vote according to the preset voting strategy and the received status record information of other node servers, and feedback the voting result to the cloud server.

[0037] The node server can receive the status record information from other node servers, then vote according to the voting strategy within the node server, and each node server can feedback the obtained voting result to the cloud server; the cloud server can receive the voting results obtained by multiple node servers voting respectively.

[0038] In a specific embodiment, step S150 includes sub-steps: obtain the node status coefficients of each node server according to the voting strategy and the status record information; sort each node server according to the node status coefficients to obtain the corresponding sorting result; obtain the server identifier of a node server at the front in the sorting result as the corresponding voting result and feedback it to the cloud server.

[0039] Obtain the node status coefficients of each node server according to the voting strategy and the status record information. The node status coefficient can quantitatively represent the status of the node server. The larger the node status coefficient, the worse the status of the node server; the smaller the node status coefficient, the better the status of the node server. If the node server receives n groups of status record information, and combines it with its own status record information, n + 1 groups of status record information can be obtained, that is, n + 1 node status coefficients can be correspondingly obtained.

[0040] The node server further sorts the node servers according to the obtained node status coefficients to obtain the sorting result; specifically, it can be sorted in ascending order according to the node status coefficients, obtain the server identifier of a node server at the front in the sorting result as the corresponding voting result, and feedback the voting result to the cloud server.

[0041] In a specific embodiment, the obtaining the node status coefficients of each node server according to the voting strategy and the status record information includes: calculating the error quantity information in the status record information according to the error coefficient calculation formula in the voting strategy to obtain the error coefficient values of each node server; calculating the response delay in the status record information according to the delay coefficient calculation formula in the voting strategy to obtain the delay coefficient values of each node server; performing a combined calculation on the error coefficient values and the delay coefficient values according to the combined calculation formula in the voting strategy to obtain the node status coefficients of each node server.

[0042] The specific process of obtaining the node status coefficient includes: calculating the error quantity information in the status record information according to the error coefficient calculation formula in the voting strategy. The error coefficient calculation formula can be used to calculate the error coefficient value corresponding to the error quantity information. Then, a set of error coefficient values can be calculated corresponding to the error quantity information in each group of status record information. The error coefficient calculation formula can be expressed by formula (1): (1); Where, X is the calculated error coefficient value, m is the total number of error types, z1 is the total number of errors of the first service instance in the status record information, z2 is the total number of errors of the second service instance in the status record information, r1 is the coefficient value corresponding to the first error type, r2 is the coefficient value corresponding to the second error type, r m is the coefficient value corresponding to the m-th error type; c 1,1 is the number of errors of the first service instance corresponding to the first error type, c 2,1 is the number of errors of the second service instance corresponding to the first error type, c 1,m is the number of errors of the first service instance corresponding to the m-th error type, c 2,m is the number of errors of the second service instance corresponding to the m-th error type. In the specific application process, r i can be set to e i-m , e is the base of the natural logarithm, i is an integer belonging to [1, m]; when i = 1, the value of r1 can be calculated as e 1-m .

[0043] Further, calculate the response delay in the status record information according to the delay coefficient calculation formula in the voting strategy. The delay coefficient calculation formula can be used to calculate the delay coefficient value corresponding to the response delay. Then, a set of delay coefficient values can be calculated corresponding to the response delay in each group of status record information. The delay coefficient calculation formula can be expressed by formula (2): (2); Where, T is the calculated delay coefficient value, t1 is the response delay of the first service instance in the status record information, t2 is the response delay of the second service instance in the status record information, is the average value of the response delays of the two service instances.

[0044] Further, the error coefficient value and the delay coefficient value can be combined and calculated according to the combination calculation formula in the voting strategy; specifically, the combination calculation formula can be expressed by formula (3): (3); Among them, J is the node state coefficient obtained by performing combined calculation.

[0045] S160. Obtain the target server identifier determined by the cloud server based on the voting result, and configure the primary node information according to the target server identifier.

[0046] The cloud server obtains the voting results of each node server, and determines the target server identifier based on the voting result. The cloud server can obtain the server identifiers in each voting result and record the voting times of each server identifier, and obtain the server identifier with the highest voting times as the target server identifier and feedback it to each node server.

[0047] Then the node server can receive the target server identifier fed back by the cloud server, and configure the primary node information according to the target server identifier, that is, obtain the configuration information such as the node address and transmission protocol corresponding to the target server identifier to configure the primary node information. Here, when configuring the primary node information in the node server, it is necessary to configure the primary node information of each of the two service instances in the node server separately.

[0048] In a specific embodiment, after step S160, the following steps are further included: setting a delay time according to a pre-stored default time; determining whether the detection time point corresponding to the set delay time is reached.

[0049] After completing the configuration of the primary node information, the delay time can be set according to the default time pre-stored in the node server, that is, setting the new delay time as the pre-stored default time. Further determine whether the detection time point corresponding to the set delay time is reached, and return to execute step S110.

[0050] S170. If it is the primary server and the health status information is healthy, generate a delay update time corresponding to the status record information according to a preset time parameter configuration rule and send it to each other node server and the cloud server.

[0051] If the node server is the primary server and the health status information is healthy, it indicates that the node server is suitable to continue running as the primary server. A delay update time corresponding to the status record information can be generated according to the pre-set time parameter configuration rule, and the generated delay update time is sent to other nodes and the cloud server. Other nodes and the cloud server can update the delay time correspondingly based on the received delay update time.

[0052] In a specific embodiment, step S170 includes sub-steps: obtaining the average delay time corresponding to each time period from the status record information according to the time period information in the parameter configuration rule; calculating the average delay time according to the time setting formula in the parameter configuration rule to obtain the corresponding delay update time.

[0053] Specifically, the time period information contains multiple time periods. The delay time corresponding to each time period in the status record information of the node server itself can be obtained according to the time period information in the parameter configuration rule, and the average value of the delay times of each time period can be calculated to obtain the average delay time of each time period. For example, four time periods can be set in the time period information. For example, the first time period is [-5min, 0], the second time period is [-15min, -5min), the third time period is [-30min, -15min), and the second time period is [-60min, -30min). 0 is the current time, and -5min is the time point five minutes before the current time; the average delay times of the four time periods are y1, y2, y3, and y4 respectively. Further calculate the average delay time of each time period according to the time setting formula in the parameter configuration rule. The time setting formula is shown in formula (4): (4); Where S is the calculated delay update time, y1, y2, y3, and y4 are the average delay times corresponding to the first time period, the second time period, the third time period, and the fourth time period respectively, and y0 is the pre-stored default time. For example, y0 = 30 seconds can be set.

[0054] Through the above method, the delay update time corresponding to the average delay time of each time period can be calculated. In addition to sending the calculated delay update time outwards, the node server also sets its own delay time according to the delay update time.

[0055] S180. If it is the master server and the health status information is unhealthy, perform a demotion adjustment on the configured master node information.

[0056] If the node server is the master server and the health status information is unhealthy, a demotion adjustment can be performed on the master node information configured by itself, that is, the information configured in the master node information is deleted and a master-slave switch is performed. The master-slave switch is to switch the main label of the server identifier of the service instance in the node server to the slave label.

[0057] In the cluster server scheduling method disclosed in the foregoing embodiments, the method includes: each node server detects its own operating status and correspondingly updates the status record information, and determines whether it is the master server. If it is not the master server, it receives the delayed update time from the master server and updates the detection time point. If it is not the master server and does not receive the delayed update time within the received duration, it sends the stored status record information and votes according to the received status record information, and feeds back the voting result to the cloud server. The cloud server determines the target server identifier and sends it to each node server so that each node server can perform the master node information configuration. The master server generates the delayed update time in a healthy state and sends it to other node servers. The foregoing cluster server scheduling method can detect the health status information by itself, make autonomous decisions in the case of master server failure, and vote to determine the target server identifier, so as to quickly and accurately schedule the distributed cluster servers.

[0058] An embodiment of the present invention further provides a cluster server scheduling device, which can be configured in the node server of the cluster server. The cluster server scheduling device is used to execute any one of the foregoing embodiments of the cluster server scheduling method. Specifically, please refer to Figure 3 , Figure 3 which is a schematic block diagram of the cluster server scheduling device provided by the embodiment of the present invention.

[0059] As Figure 3 shown, the cluster server scheduling device 100 includes a status record information update unit 110, a judgment unit 120, a time point judgment unit 130, a status record information sending unit 140, a voting result feedback unit 150, a master node information configuration unit 160, a delayed update time sending unit 170, and a configuration adjustment unit 180.

[0060] The status record information update unit 110 is used to, if the detection time point corresponding to the set delay time is reached, generate corresponding health status information according to the operating status of the service instance configured internally and correspondingly update the pre-stored status record information.

[0061] The judgment unit 120 is used to judge whether it is the master server according to the configured master node information.

[0062] The time point judgment unit 130 is used to, if it is not the master server and receives the delayed update time from the master server within the preset received duration, update the currently set delay time according to the delayed update time and judge whether the detection time point corresponding to the set delay time is reached.

[0063] A status record information sending unit 140 is configured to, if it is not the master server and has not received the delayed update time from the master server within a preset receiving duration, send the stored status record information to other node servers.

[0064] A voting result feedback unit 150 is configured to vote according to a preset voting policy and the received status record information of other node servers and feedback the voting result to the cloud server.

[0065] A master node information configuration unit 160 is configured to obtain the target server identifier determined by the cloud server based on the voting result and configure the master node information according to the target server identifier.

[0066] A delayed update time sending unit 170 is configured to, if it is the master server and the health status information is healthy, generate a delayed update time corresponding to the status record information according to a preset time parameter configuration rule and send it to other node servers and the cloud server.

[0067] A configuration adjustment unit 180 is configured to, if it is the master server and the health status information is unhealthy, perform a demotion adjustment on the configured master node information.

[0068] In the cluster server scheduling device provided by the embodiment of the present invention, the above cluster server scheduling method is applied. Each node server independently detects its running status and correspondingly updates the status record information, and determines whether it is the master server. If it is not the master server, it receives the delayed update time from the master server and updates the detection time point. If it is not the master server and has not received the delayed update time after exceeding the receiving duration, it sends the stored status record information and votes according to the received status record information, and feeds back the voting result to the cloud server. The cloud server determines the target server identifier and sends it to each node server so that each node server configures the master node information. The master server generates a delayed update time in a healthy state and sends it to other node servers. The above cluster server scheduling method can independently detect the health status information, make an autonomous decision in the case of a master server failure, and vote to determine the target server identifier, so as to quickly and accurately perform server scheduling for distributed cluster servers.

[0069] The above cluster server scheduling device can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 4 shown.

[0070] Please refer to Figure 4 , Figure 4It is a schematic block diagram of a computer device provided by an embodiment of the present invention. The computer device may be a node server for executing a cluster server scheduling method to schedule and manage a distributed cluster server.

[0071] Referring to Figure 4 , the computer device 500 includes a processor 502, a memory, and a communication interface 505 connected through a communication bus 501. Among them, the memory may include a storage medium 503 and an internal memory 504.

[0072] The storage medium 503 can store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 can be made to execute the cluster server scheduling method. Among them, the storage medium 503 can be a volatile storage medium or a non-volatile storage medium.

[0073] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0074] The internal memory 504 provides an environment for the operation of the computer program 5032 in the storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be made to execute the cluster server scheduling method.

[0075] The communication interface 505 is used for network communication, such as providing the transmission of data information, etc. Those skilled in the art can understand that Figure 4 the structure shown in

[0076] is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device 500 to which the solution of the present invention is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0077] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the corresponding functions in the above-mentioned cluster server scheduling method. Figure 4 Those skilled in the art can understand that Figure 4 the embodiment of the computer device shown in

[0078] It should be understood that in the embodiments of the present invention, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0079] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps included in the above-mentioned cluster server scheduling method are implemented.

[0080] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0081] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, or units with the same function may be aggregated into a unit. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings, direct couplings, or communication connections to each other may be indirect couplings or communication connections through some interfaces, devices, or units, and may also be electrical, mechanical, or other forms of connection.

[0082] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.

[0083] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0084] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned computer-readable storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes.

[0085] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A cluster server scheduling method, characterized in that: The method is applied to a node server of a cluster server, the cluster server includes a node server configured in each computer room and a cloud server configured in an external cloud, each of the node servers is configured with two service instances; the cloud server and each of the node servers establish a communication connection with an operator backbone network to realize the transmission of data information, the method includes: If the detection time point corresponding to the set delay time is reached, the corresponding health status information is generated according to the running status of the internally configured service instance and the pre-stored status record information is updated accordingly; Determine whether it is the primary server based on the configured primary node information; If it is not the main server and the delay update time is received from the main server within the preset receiving time, the currently set delay time is updated according to the delay update time and it is determined whether the detection time point corresponding to the set delay time has been reached; If it is not the main server and the delayed update time is not received from the main server within the preset receiving time, the stored status record information is sent to other node servers; Voting is performed according to the preset voting strategy and the status record information received from other node servers, and the voting results are fed back to the cloud server; Obtaining a target server identifier determined by the cloud server based on the voting result, and configuring master node information according to the target server identifier; If it is the main server and the health status information is healthy, a delayed update time corresponding to the status record information is generated according to a preset time parameter configuration rule and sent to each other node server and the cloud server; If it is a primary server and the health status information is unhealthy, the configured primary node information is downgraded.

2. The cluster server scheduling method according to claim 1, characterized in that: The method generates corresponding health status information according to the running status of the internally configured service instance and updates the pre-stored status record information accordingly, including: Get the running status of the two configured service instances and obtain the running status information; Obtaining the response delay and error number information corresponding to the service instance in the current running time period; the current running time period is the time period between the current detection time point and the previous detection time point; Generate corresponding health status information according to the operation status information, the response delay and the error quantity information; The current status record information is updated according to the operation status information, the response delay and the error quantity information.

3. The cluster server scheduling method according to claim 2, characterized in that: The obtaining of the response delay and error quantity information corresponding to the service instance in the current running period includes: Obtaining an average response time of the application service processed by the service instance in the current running period as the corresponding response delay; Obtain the operation log corresponding to the service instance during the current operation period; Match the error codes in the operation log according to the preset code configuration table to obtain the number of times each error code corresponds to; According to the code classification information and the number information in the code configuration table, the cumulative number of each error type is obtained as the corresponding error number information.

4. The cluster server scheduling method according to any one of claims 1 to 3, characterized in that: The voting according to the preset voting strategy and the received status record information of other node servers and feeding back the voting result to the cloud server includes: Obtaining the node status coefficient of each node server according to the voting strategy and the status record information; Sort the node servers according to the node status coefficient to obtain a corresponding sorting result; The server identifier of a node server that is at the front in the sorting result is obtained as the corresponding voting result and fed back to the cloud server.

5. The cluster server scheduling method according to claim 4, characterized in that: The node status coefficient of each node server is obtained according to the voting strategy and the status record information, including: Calculate the error quantity information in the state record information according to the error coefficient calculation formula in the voting strategy to obtain the error coefficient value of each node server; Calculate the response delay in the state record information according to the delay coefficient calculation formula in the voting strategy to obtain the delay coefficient value of each node server; The error coefficient value and the delay coefficient value are combined and calculated according to the combined calculation formula in the voting strategy to obtain the node state coefficient of each node server.

6. The cluster server scheduling method according to any one of claims 1 to 3, characterized in that: The generating the delayed update time corresponding to the status record information according to the preset time parameter configuration rule includes: Obtaining the average delay time corresponding to each time period from the state record information according to the time period information in the parameter configuration rule; The average delay time is calculated according to the time setting formula in the parameter configuration rule to obtain the corresponding delayed update time.

7. The cluster server scheduling method according to any one of claims 1 to 3, characterized in that: After configuring the master node information according to the target server identifier, the method further includes: Set the delay time according to the pre-stored default time; It is determined whether the detection time point corresponding to the set delay time has been reached.

8. A cluster server scheduling device, characterized in that: The device is configured in a node server of a cluster server, the cluster server includes a node server configured in each computer room and a cloud server configured in an external cloud, each of the node servers is configured with two service instances; the cloud server and each of the node servers establish a communication connection with an operator backbone network to realize the transmission of data information, the device is used to execute the cluster server scheduling method according to any one of claims 1 to 7, and the device includes: A status record information updating unit, configured to generate corresponding health status information according to the running status of the internally configured service instance and update the pre-stored status record information accordingly if a detection time point corresponding to the set delay time is reached; A judging unit, used to judge whether it is a master server according to the configured master node information; A time point judgment unit, for updating the currently set delay time according to the delay update time and judging whether a detection time point corresponding to the set delay time has been reached if the delay update time is received from the main server within a preset receiving time period and if the delay update time is not received from the main server; A status record information sending unit, used for sending the stored status record information to other node servers if it is not a main server and has not received the delayed update time from the main server within a preset receiving time; A voting result feedback unit, used to vote according to a preset voting strategy and the received status record information of other node servers and to feed back the voting result to the cloud server; A master node information configuration unit, configured to obtain a target server identifier determined by the cloud server based on the voting result, and to configure the master node information according to the target server identifier; A delayed update time sending unit, for generating a delayed update time corresponding to the status record information according to a preset time parameter configuration rule and sending the delayed update time to each other node server and the cloud server if the server is a main server and the health status information is healthy; The configuration adjustment unit is used to downgrade the configured master node information if it is a master server and the health status information is unhealthy.

9. A computer device, characterized in that: The device includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory, used to store computer programs; The processor is used to implement the steps of the cluster server scheduling method described in any one of claims 1 to 7 when executing the program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the cluster server scheduling method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method for selecting service main node of power grid monitoring system based on real-time state perception

    CN113489149A

  • Master server arbitration election method and server system

    CN116437321A

  • Method and system for adjusting state of service cluster

    CN116723088A

  • Method for automatically deploying kubernetes worker node, device, terminal apparatus, and readable storage medium

    WO2019184164A1