Server node status control method, control device, electronic equipment and medium

By obtaining and sorting the time series status parameters of server nodes, judging anomalies based on thresholds and controlling node status, the problem of identifying and isolating abnormal nodes in the server cluster is solved, ensuring the continuity and real-time performance of the server cluster's computing power.

CN116708138BActive Publication Date: 2025-09-05INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310659597.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-06
Publication Date
2025-09-05
Estimated Expiration
2043-06-06

AI Technical Summary

Technical Problem

In a server cluster, how to accurately identify abnormal server nodes and isolate them to ensure a healthy and orderly isolation process, while ensuring that the server cluster provides sufficient computing power support during isolation to avoid the application being unable to provide services or reducing computing efficiency due to too many abnormal nodes.

Method used

By obtaining the timing status parameters within the preset time period, time sorting is performed to determine the proportion of node flip states, and the isolation status of the server node is determined based on the difference volatility threshold and the cluster isolation threshold. When necessary, instructions are sent to control the node to suspend or resume the data loading function.

Benefits of technology

It achieves the continuity and real-time computing power of the server cluster, avoids service interruption or reduced computing efficiency due to too many abnormal nodes, and ensures the stable operation of the server cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116708138B_ABST
    Figure CN116708138B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, control device, electronic device, and medium for controlling the state of a server node, which can be applied to the fields of server control and financial technology. The method comprises: obtaining state data generated within a preset time period; performing time sorting on multiple state parameters corresponding to the server node extracted from the state data to obtain sorted data; determining, based on each parameter group in the sorted data, a first state flip result of the server node corresponding to the multiple parameter groups; if the proportion of the number of node flipped states meets a differential volatility threshold, preliminarily determining the state of the server node as an isolated state; if the number of server nodes in the isolated state meets a cluster isolation threshold, updating the state of a preset number of server nodes determined to be in the isolated state to a target isolation state, wherein the preset number is determined based on the cluster isolation threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of server control and finance, and in particular to a method for controlling the status of a server node, a control device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the continuous development of Internet technology, server clusters, as an important part of Internet technology, have an increasingly large number of server nodes. Especially after the rapid development of distributed technology, various businesses in business systems are often provided in a cluster mode.

[0003] When a server cluster provides business support, abnormalities in one or several server nodes within the cluster will become a normal operation and maintenance situation. However, how to accurately identify which server nodes in the server cluster have abnormalities can provide support for the rapid isolation of abnormal nodes. At the same time, when isolating server nodes, how to ensure that the server cluster provides sufficient computing power support during the healthy and orderly isolation process is an important guarantee for the application to continue to provide services. Summary of the Invention

[0004] In view of the above problems, the present disclosure provides a method for controlling the status of a server node, a control device, an electronic device, a computer-readable storage medium, and a computer program product.

[0005] According to a first aspect of the present disclosure, a method for controlling a server node state is provided, comprising:

[0006] Acquire status data generated within a preset time period, wherein the status data includes a plurality of time series status parameters arranged in a time series, and the time series status parameters represent the status of the plurality of server nodes;

[0007] For each of the server nodes, performing time sorting on a plurality of state parameters corresponding to the server node extracted from the state data to obtain sorted data;

[0008] Determine, based on each parameter group in the sorted data, a first state reversal result of the server node corresponding to the plurality of parameter groups, wherein the parameter group includes two adjacent state parameters in the sorted data, and the first state reversal result includes a node reversal state or a node normal state corresponding to each parameter group, wherein the node reversal state indicates that the two adjacent state parameters of the server node are different, and the node normal state indicates that the two adjacent state parameters of the server node are the same;

[0009] If the percentage of the node flipping states meets the difference volatility threshold, the server node is initially determined to be in an isolated state.

[0010] When the number of server nodes in the above-mentioned isolation state meets the cluster isolation threshold, the status of a preset number of server nodes among the above-mentioned server nodes determined to be in the isolation state will be updated to the target isolation state, wherein the above-mentioned preset number is determined based on the above-mentioned cluster isolation threshold, and the target isolation state indicates that the server nodes need to be restricted in data loading function.

[0011] According to an embodiment of the present disclosure, after updating the status of a preset number of server nodes among the server nodes determined to be in the isolated state to the target isolation state, the method further includes:

[0012] Determine, based on each timing group, a cluster rollover sub-result corresponding to each timing group, wherein the timing group includes two timing state parameters that are adjacent in time sequence, and the cluster rollover sub-result includes difference data or normal data, wherein the difference data indicates that a preset number of server nodes in the timing group have had their states changed;

[0013] Generate a cluster flip result based on the multiple cluster flip sub-results;

[0014] When the proportion of the difference data in the above cluster rollover result meets the preset cluster rollover threshold, the state of the server node in the target isolation state is updated to the normal state, and the normal state indicates that the server node can provide data loading function.

[0015] According to an embodiment of the present disclosure, the method for controlling the server node status further includes:

[0016] When the number of the server nodes in the isolation state does not meet the cluster isolation threshold, the state of the server nodes determined to be in the isolation state is updated to the target isolation state.

[0017] According to an embodiment of the present disclosure, the method for controlling the server node status further includes:

[0018] When the state of the server node is determined to be the target isolation state, a first instruction is sent to the server node, where the first instruction is used to control the server node to suspend data loading of the server node.

[0019] According to an embodiment of the present disclosure, the method for controlling the server node status further includes:

[0020] When the target isolation state of the server node is determined to be a normal state, a second instruction is sent to the server node, where the second instruction is used to control the loading of recovery data of the server node.

[0021] According to an embodiment of the present disclosure, determining, based on each parameter group in the sorted data, first state reversal results of the server nodes corresponding to the plurality of parameter groups includes:

[0022] For each of the parameter groups, determining a sub-flip result of the parameter group according to whether the two adjacent state parameters are consistent;

[0023] The first state reversal result is generated according to the plurality of sub-reversal results corresponding to the plurality of parameter groups.

[0024] According to an embodiment of the present disclosure, determining the sub-flip result of the parameter group according to whether the two adjacent state parameters are consistent includes:

[0025] In the case where the two state parameters are inconsistent, the sub-flip result is determined to be the node flip state;

[0026] When the two state parameters are consistent, the sub-flip result is determined to be a normal state of the node.

[0027] According to an embodiment of the present disclosure, the method for controlling the server node status further includes:

[0028] The state data is stored in a state control server so that the state control server can control the state of the server node according to the state data.

[0029] A second aspect of the present disclosure provides a device for controlling a server node state, comprising:

[0030] An acquisition module, configured to acquire status data generated within a preset time period, wherein the status data includes a plurality of time series status parameters arranged in a time series, and the time series status parameters represent the status of the server node;

[0031] a sorting module, configured to perform time sorting on the plurality of time series state parameters corresponding to the server node extracted from the state data for each of the server nodes, to obtain sorted data;

[0032] a first determining module, configured to determine, based on each parameter group in the sorted data, a first state reversal result of the server node corresponding to the plurality of parameter groups, wherein the parameter group includes two adjacent state parameters in the sorted data, and the first state reversal result includes a node reversal state or a node normal state corresponding to each parameter group, wherein the node reversal state indicates that the two adjacent state parameters of the server node are different, and the node normal state indicates that the two adjacent state parameters of the server node are the same;

[0033] A second determination module is configured to preliminarily determine the state of the server node as an isolated state when the proportion of the number of node flip states meets a difference volatility threshold;

[0034] An update module is used to update the status of a preset number of server nodes among the above-mentioned server nodes determined to be in the isolated state to a target isolation state when the number of server nodes in the above-mentioned isolation state meets the cluster isolation threshold, wherein the above-mentioned preset number is determined based on the above-mentioned cluster isolation threshold, and the target isolation state indicates that the server nodes need to be restricted in data loading function.

[0035] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the method.

[0036] A fourth aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above method.

[0037] The fifth aspect of the present disclosure further provides a computer program product, comprising a computer program, which implements the above method when executed by a processor.

[0038] According to an embodiment of the present disclosure, by obtaining the node status within a preset time period in real time, it is determined whether each server node has an abnormality based on the sorted data after time sorting, and by determining the number of node flip states of this server node within the preset time period, it is preliminarily determined whether it is set to an isolated state, and when the number of server nodes in the isolated state meets the cluster isolation threshold, the status of a preset number of server nodes determined to be in an isolated state is updated to a target isolated state. Since the server cluster determines its status through the corresponding timing state parameters during operation and determines whether the number of server nodes in the isolated state exceeds the cluster isolation threshold from an overall perspective, the status of some server nodes determined to be in an isolated state is updated to a target isolated state, and other server nodes determined to be in an isolated state provide services normally, thereby avoiding the situation where the application cannot provide services or the computing efficiency of the server cluster is reduced due to the large number of server nodes in the target isolation state in the server cluster, and ensuring the continuity and real-time nature of the computing power provided by the server cluster to the application. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0040] Figure 1 A diagram schematically illustrates an application scenario of a method for controlling a server node state according to an embodiment of the present disclosure;

[0041] Figure 2 The flowchart of the method for controlling the server node status according to the embodiment of the present disclosure is schematically shown;

[0042] Figure 3 A schematic diagram schematically illustrates status data according to an embodiment of the present disclosure;

[0043] Figure 4 The flowchart schematically shows a method for controlling the server node status according to another embodiment of the present disclosure;

[0044] Figure 5 A block diagram schematically illustrates a structure of a device for controlling a server node state according to an embodiment of the present disclosure; and

[0045] Figure 6 A block diagram of an electronic device suitable for implementing a method for controlling a server node state according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0046] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0047] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0048] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0049] When expressions such as "at least one of A, B and C, etc." are used, they should generally be interpreted in accordance with the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0050] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of the data involved (including but not limited to user personal information) comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0051] The embodiments of the present disclosure provide a method for controlling the state of a server node, a control device, an electronic device, a computer-readable storage medium, and a computer program product. The method includes: obtaining state data generated within a preset time period, wherein the state data includes a plurality of time-series state parameters arranged in a time sequence, and the time-series state parameters characterize the states of a plurality of server nodes; for each server node, performing time sorting on a plurality of state parameters corresponding to the server node extracted from the state data to obtain sorted data; determining a first state flip result of the server node corresponding to the plurality of parameter groups based on each parameter group in the sorted data, wherein the parameter group includes two adjacent state parameters in the sorted data, and the first state flip result includes a node flip state or a node normal state corresponding to each parameter group, and the node flip state characterizes that the two adjacent state parameters of the server node are different; when the proportion of the number of node flip states meets a difference volatility threshold, the state of the server node is preliminarily determined to be an isolated state; when the number of server nodes in the isolated state meets a cluster isolation threshold, the state of a preset number of server nodes determined to be in the isolated state is updated to a target isolation state, wherein the preset number is determined based on the cluster isolation threshold.

[0052] Figure 1 The application scenario diagram of the server node status control method according to an embodiment of the present disclosure is schematically shown.

[0053] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a server cluster that provides computing power support for a banking application. A network 104 is used as a medium to provide a communication link between a first terminal device 101, a second terminal device 102, a third terminal device 103, and a server cluster 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0054] The user may use at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server cluster 105 via the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0055] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0056] The server cluster 105 may include server nodes that provide various services, such as server nodes that support websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (for example only). The server cluster 105 also includes a state control server 106, which is connected to other servers and can execute the server node state control method of the present disclosure in response to a server node state control request and control the server node state based on the processing result.

[0057] It should be noted that the method for controlling the state of a server node provided in the embodiment of the present disclosure can generally be executed by the state control server 106. Accordingly, the device for controlling the state of a server node provided in the embodiment of the present disclosure can generally be set in the state control server 106. The method for controlling the state of a server node provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server cluster 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server cluster 105. Accordingly, the device for controlling the state of a server node provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server cluster 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server cluster 105.

[0058] It should be understood that Figure 1 The number of terminal devices, networks, server clusters, and server nodes in the embodiment is merely illustrative. Any number of terminal devices, networks, server clusters, and server nodes may be provided as required.

[0059] The following will be based on Figure 1 The scene described by Figures 2 to 4The method for controlling the server node status of the disclosed embodiment is described in detail.

[0060] Figure 2 The flowchart of the method for controlling the status of a server node according to an embodiment of the present disclosure is schematically shown. Figure 3 The figure schematically shows a schematic diagram of status data according to an embodiment of the present disclosure.

[0061] like Figure 2 As shown, the method for controlling the server node status of this embodiment includes operations S210 to S250.

[0062] In operation S210, status data generated within a preset time period is acquired, wherein the status data includes a plurality of time series status parameters arranged in a time series, and the time series status parameters represent the status of the plurality of server nodes;

[0063] In operation S220, for each server node, a plurality of state parameters corresponding to the server node extracted from the state data are time-sorted to obtain sorted data;

[0064] In operation S230, a first state flip result of a server node corresponding to the plurality of parameter groups is determined based on each parameter group in the sorted data, wherein the parameter group includes two adjacent state parameters in the sorted data, and the first state flip result includes a node flip state or a node normal state corresponding to each parameter group, wherein the node flip state indicates that the two adjacent state parameters of the server node are different, and the node normal state indicates that the two adjacent state parameters of the server node are the same.

[0065] In operation S240 , if the percentage of the number of node flip states meets the difference volatility threshold, the state of the server node is preliminarily determined to be an isolated state;

[0066] In operation S250, when the number of server nodes in the isolated state meets the cluster isolation threshold, the status of a preset number of server nodes among the server nodes determined to be in the isolated state is updated to a target isolation state, wherein the preset number is determined based on the cluster isolation threshold, and the target isolation state indicates that the server nodes need to be restricted in data loading function.

[0067] According to an embodiment of the present disclosure, the preset time period can be specifically set according to actual needs, for example, it can be five minutes, and the status of each server node is obtained according to the preset frequency, wherein the preset frequency can be 10 seconds. Therefore, the time series status parameters of the present disclosure with 10 seconds as a node include the status of each server node at the current moment, thereby obtaining status data about multiple server nodes, such as Figure 3As shown, a timing state parameter can be expressed as "010,1111111110", where the first three digits represent the 10th second, and the last ten digits are the state parameters of the ten server nodes respectively. It should be noted that the last ten digits are the same as the number of server nodes, and the number of server nodes is related to the actual situation on site. Therefore, the last ten digits in the above expression can be modified according to the actual situation.

[0068] According to an embodiment of the present disclosure, for each server node, such as the first node among the ten server nodes mentioned above, the status parameter "1" at the tenth second, the status parameter at the 20th second... the status parameter at the 450th second (i.e., the fifth minute) are extracted, and sorted by time to obtain sorted data.

[0069] According to an embodiment of the present disclosure, every two adjacent timing state parameters in the sorted data are determined as a parameter group, and the first state flip results of the server nodes corresponding to the multiple parameter groups are determined based on the multiple parameter groups. In the first state flip results, the first flip identifier can be used to represent the node flip state, and the second flip identifier can be used to represent the node normal state.

[0070] According to an embodiment of the present disclosure, if the proportion of the first flip identifiers in the first state flip result meets a differential volatility threshold, the server node is preliminarily determined to be in an isolated state; otherwise, the node is determined to be in a normal state. The differential volatility threshold can be set based on actual conditions, for example, 50%.

[0071] According to an embodiment of the present disclosure, after determining the status of each server node, if the number of server nodes determined to be in an isolated state exceeds the total number of server nodes in the server cluster and satisfies the cluster isolation threshold, the status of a preset number of server nodes determined to be in an isolated state is updated to a target isolation state. The cluster isolation threshold can be set based on actual circumstances, for example, 50% of the total number of server nodes, and the preset number can be equal to or less than the cluster isolation threshold.

[0072] According to an embodiment of the present disclosure, the target isolation state or the normal state representation can perform corresponding operations on the server node, for example, the data loading function can be suspended for the server node in the target isolation state.

[0073] According to an embodiment of the present disclosure, by obtaining the node status within a preset time period in real time, it is determined whether each server node has an abnormality based on the sorted data after time sorting, and by determining the number of node flip states of this server node within the preset time period, it is preliminarily determined whether it is set to an isolated state, and when the number of server nodes in the isolated state meets the cluster isolation threshold, the status of a preset number of server nodes determined to be in an isolated state is updated to a target isolated state. Since the server cluster determines its status through the corresponding timing state parameters during operation and determines whether the number of server nodes in the isolated state exceeds the cluster isolation threshold from an overall perspective, the status of some server nodes determined to be in an isolated state is updated to a target isolated state, and other server nodes determined to be in an isolated state provide services normally, thereby avoiding the situation where the application cannot provide services or the computing efficiency of the server cluster is reduced due to the large number of server nodes in the target isolation state in the server cluster, and ensuring the continuity and real-time nature of the computing power provided by the server cluster to the application.

[0074] Figure 4 The flowchart schematically shows a method for controlling the status of a server node according to another embodiment of the present disclosure.

[0075] like Figure 4 As shown, after the states of a preset number of server nodes among the server nodes determined to be in the isolated state are updated to the target isolated state, operations S401 to S403 are further included.

[0076] In operation S401, a cluster flip sub-result corresponding to each timing group is determined according to each timing group, wherein the timing group includes two timing state parameters that are adjacent in time sequence, and the cluster flip sub-result includes difference data or normal data, where the difference data indicates that the states of a preset number of server nodes in the timing group have changed;

[0077] In operation S402 , a cluster flipping result is generated according to the plurality of cluster flipping sub-results.

[0078] In operation S403 , when the ratio of the amount of difference data in the cluster rollover result meets the preset cluster rollover threshold, the state of the target isolated server node is updated to a normal state, where the normal state indicates that the server node can provide a data loading function.

[0079] According to an embodiment of the present disclosure, the preset cluster flip threshold may be 70%.

[0080] In an exemplary embodiment, assuming that there are 10 servers in the server cluster and the cluster isolation threshold is assumed to be 5, when the number of server nodes preliminarily determined to be in an isolated state is 6, the status of any 5 of the 6 server nodes is updated to the target isolation state, and the other server node remains in a normal state so that it can load data to provide services.

[0081] According to an embodiment of the present disclosure, after determining the target isolation state, each two adjacent timing state parameters can be determined as a timing group. For each timing group, the two timing state parameters are XORed to obtain the overall difference between multiple server nodes in the two time periods, for example Figure 3 The first timing state parameter is "000, 11111111111", and the second timing state parameter is "010, 11111111110". Performing the XOR operation yields 1111111111 xor 1111111110 = 00000000001. This indicates that the number of difference nodes is 1, the total number of nodes is 10, and the difference rate is 10%. A threshold of 60% is set. Since 10% < 60%, the cluster flip sub-result corresponding to this timing group is normal data and can be identified with "1". If abnormal, it is identified with "0".

[0082] According to an embodiment of the present disclosure, a cluster flip result is generated based on multiple cluster flip sub-results. For example, the cluster flip result is 00010000100001001100010000011, thereby determining that the proportion of the number of difference data is 21 / 29=72.4%>70%. At this time, the status of all server nodes updated to the target isolation status can be updated to the normal status, so that the server node can restore the data loading function and continue to provide services.

[0083] According to an embodiment of the present disclosure, a cluster flip sub-result is determined based on the timing state parameters of each two adjacent nodes, and a cluster flip result composed of the cluster flip sub-results is judged to determine whether the proportion of the number of difference data in the cluster flip result meets a preset cluster flip threshold, thereby determining whether the status of all server nodes updated to the target isolation state is updated to the normal state, so that the server nodes in the target isolation state can resume service, further avoiding the occurrence of the phenomenon that the application cannot provide services normally due to the isolation of more server nodes.

[0084] According to an embodiment of the present disclosure, the method for controlling the server node status further includes:

[0085] When the number of server nodes in the isolated state does not meet the cluster isolation threshold, the state of the server nodes determined to be in the isolated state is updated to the target isolation state.

[0086] According to an embodiment of the present disclosure, if 4 server nodes out of 10 server nodes are initially determined to be in an isolated state, which is significantly less than the cluster isolation threshold of 5, all server nodes initially determined to be in an isolated state can be determined to be in a target isolation state.

[0087] According to an embodiment of the present disclosure, the method for controlling the server node status further includes:

[0088] When the state of the server node is determined to be the target isolation state, a first instruction is sent to the server node, where the first instruction is used to control the server node to suspend data loading of the server node.

[0089] According to an embodiment of the present disclosure, when it is determined that the server node belongs to the target isolation state, a first instruction may be sent to the server node so that the server node can suspend the data loading function in response to the first instruction to implement isolation measures.

[0090] It should be noted that the isolation measures may be implemented by the corresponding server node, or the isolation of the server node may be implemented by other devices.

[0091] According to an embodiment of the present disclosure, the method for controlling the server node status further includes:

[0092] When the target isolation state of the server node is determined to be a normal state, a second instruction is sent to the server node, where the second instruction is used to control the loading of recovery data of the server node.

[0093] According to an embodiment of the present disclosure, if it is later determined that the target isolation state of the server node needs to be restored to a normal state, a second instruction can be sent to the server node, and the server node or other device performs a data loading function in response to the second instruction.

[0094] According to an embodiment of the present disclosure, determining, based on each parameter group in the sorted data, first state flip results of server nodes corresponding to the plurality of parameter groups includes:

[0095] For each parameter group, the sub-flip result of the parameter group is determined based on whether the two adjacent state parameters are consistent;

[0096] A first state flip result is generated according to a plurality of sub-flip results corresponding to a plurality of parameter groups.

[0097] According to an embodiment of the present disclosure, if five minutes is used as the preset time period and the frequency is 10 seconds, 30 sets of timing state parameters can be obtained, and the state parameters of each server node within the preset time period are extracted, and a sorted data with a length of 30 digits can be obtained in chronological order.

[0098] According to an embodiment of the present disclosure, in the sorted data, two adjacent parameters are regarded as a parameter group. There are 29 parameter groups at this time. The XOR operation is used to determine whether the state in each parameter group has changed, so as to obtain the sub-flip result of a server node.

[0099] According to an embodiment of the present disclosure, determining a sub-flip result of a parameter group according to whether two adjacent state parameters are consistent includes:

[0100] In the case that the two state parameters are inconsistent, the sub-flip result is determined to be the node flip state;

[0101] When the two state parameters are consistent, the sub-flip result is determined to be the normal state of the node.

[0102] In an exemplary embodiment, if the two state parameters within a parameter are "0" and "1" respectively, the number "0" is used to represent the sub-flip result of the state change; if the two state parameters are both "0" or "1", the number "1" can be used to represent the sub-flip result of the state not changing. Based on multiple "0"s or "1", the first state flip result can be obtained, so the first state flip result can be expressed as "11011111111111111111111111111111".

[0103] According to an embodiment of the present disclosure, in the first state flip result, 28 are normal and 1 is abnormal, and the number of node flip states accounts for 1 / 29, or 3.4%, which is significantly smaller than the difference volatility threshold of 50%. At this time, it can be determined that the state of the server node is normal.

[0104] According to an embodiment of the present disclosure, the method for controlling the server node status further includes:

[0105] The state data is stored in the state control server so that the state control server can control the state of the server node according to the state data.

[0106] According to an embodiment of the present disclosure, the above-mentioned method for controlling the status of a server node may be executed by a status control server, thus including the above-mentioned first instruction, second instruction, and measures for isolating or releasing the isolation of the server node.

[0107] According to an embodiment of the present disclosure, the method for controlling the state of a server node further includes overwriting the state data generated in the next preset time period with the state data in the previous preset time period by data overwriting. The state control server can save the state rollover results determined in each preset time period.

[0108] According to an embodiment of the present disclosure, the method for controlling the server node status further includes:

[0109] Provides a visual display of the status of server nodes, the percentage of node flip states, the number of server nodes in isolated states, etc.

[0110] It should be noted that when the preset time period of the present disclosure is exemplified as 5 minutes, the preset time period is 5 minutes as a time window for acquiring state data. For example, when the control method of the present disclosure is performed within the next preset time period, the state data obtained is the same as the state data obtained by the control method. Figure 3 The difference between the status data in the example is that the time sequence status parameter at time "000" is removed, while the time sequence status parameter at time "460" is added after time "450". Thus, the status of multiple server nodes is controlled using the method disclosed in this disclosure every ten seconds. The "ten seconds" interval can also be set as needed, for example, to 5 seconds, 20 seconds, etc.

[0111] Based on the above-mentioned server node status control method, the present disclosure also provides a server node status control device. Figure 5 The device is described in detail.

[0112] Figure 5 The structural block diagram of the device for controlling the status of a server node according to an embodiment of the present disclosure is schematically shown.

[0113] like Figure 5 As shown, the control device 500 of the server node status of this embodiment is communicatively connected to multiple server clusters 105 through the network 104, wherein the control device 500 includes an acquisition module 510, a sorting module 520, a first determination module 530, a second determination module 540, and an update module 550.

[0114] An acquisition module 510 is configured to acquire status data generated within a preset time period, wherein the status data includes a plurality of time series status parameters arranged in a time sequence, and the time series status parameters represent the status of the plurality of server nodes;

[0115] A sorting module 520 is configured to perform time sorting on a plurality of state parameters corresponding to each server node extracted from the state data for each server node to obtain sorted data;

[0116] A first determining module 530 is configured to determine, based on each parameter group in the sorted data, a first state flip result of a server node corresponding to the plurality of parameter groups, wherein a parameter group includes two adjacent state parameters in the sorted data, and the first state flip result includes a node flip state or a node normal state corresponding to each parameter group, wherein the node flip state indicates that the two adjacent state parameters of the server node are different, and the node normal state indicates that the two adjacent state parameters of the server node are the same;

[0117] The second determining module 540 is configured to preliminarily determine the state of the server node as an isolated state if the proportion of the number of nodes in the flipped state meets the difference volatility threshold;

[0118] The update module 550 is used to update the status of a preset number of server nodes in the isolated state to a target isolation state when the number of server nodes in the isolated state meets the cluster isolation threshold, wherein the preset number is determined based on the cluster isolation threshold, and the target isolation state indicates that the server nodes need to be restricted in data loading function.

[0119] According to an embodiment of the present disclosure, by obtaining the node status within a preset time period in real time, it is determined whether each server node has an abnormality based on the sorted data after time sorting, and by determining the number of node flip states of this server node within the preset time period, it is preliminarily determined whether it is set to an isolated state, and when the number of server nodes in the isolated state meets the cluster isolation threshold, the status of a preset number of server nodes determined to be in an isolated state is updated to a target isolated state. Since the server cluster determines its status through the corresponding timing state parameters during operation and determines whether the number of server nodes in the isolated state exceeds the cluster isolation threshold from an overall perspective, the status of some server nodes determined to be in an isolated state is updated to a target isolated state, and other server nodes determined to be in an isolated state provide services normally, thereby avoiding the situation where the application cannot provide services or the computing efficiency of the server cluster is reduced due to the large number of server nodes in the target isolation state in the server cluster, and ensuring the continuity and real-time nature of the computing power provided by the server cluster to the application.

[0120] According to an embodiment of the present disclosure, the control device 500 further includes a second determining module, a generating module, and a second updating module.

[0121] A second determining module is configured to determine, based on each timing group, a cluster flip sub-result corresponding to each timing group, wherein a timing group includes two timing state parameters that are adjacent in time sequence, and the cluster flip sub-result includes difference data or normal data, wherein the difference data indicates that a preset number of server nodes in the timing group have had their states changed;

[0122] A generation module, used for generating a cluster flipping result according to a plurality of cluster flipping sub-results;

[0123] The second updating module is used to update the state of the target isolated server node to a normal state when the proportion of the amount of difference data in the cluster rollover result meets the preset cluster rollover threshold. The normal state indicates that the server node can provide data loading function.

[0124] According to an embodiment of the present disclosure, the control device 500 further includes a third determining module.

[0125] The third determining module is configured to update the status of the server nodes determined to be in the isolated state to a target isolation state when the number of the server nodes in the isolated state does not meet the cluster isolation threshold.

[0126] According to an embodiment of the present disclosure, the control device 500 further includes a first sending module.

[0127] The first sending module is used to send a first instruction to the server node when the state of the server node is determined to be a target isolation state, where the first instruction is used to control the server node to suspend data loading of the server node.

[0128] According to an embodiment of the present disclosure, the control device 500 further includes a second sending module.

[0129] The second sending module is used to send a second instruction to the server node when the target isolation state of the server node is determined to be a normal state, and the second instruction is used to control the loading of recovery data of the server node.

[0130] According to an embodiment of the present disclosure, the first determining module 530 includes a judging unit and a generating unit.

[0131] a judgment unit, configured to determine, for each parameter group, a sub-flip result of the parameter group based on whether two adjacent state parameters are consistent;

[0132] The generating unit is configured to generate a first state flip result according to a plurality of sub-flip results corresponding to a plurality of parameter groups.

[0133] According to an embodiment of the present disclosure, the judgment unit includes a first determination subunit and a second determination subunit.

[0134] A first determining subunit is configured to determine, when two state parameters are inconsistent, that the sub-flip result is a node flip state;

[0135] The second determining subunit is configured to determine that the sub-flip result is a normal node state when the two state parameters are consistent.

[0136] According to an embodiment of the present disclosure, the control device 500 further includes a storage module.

[0137] The storage module is used to store the state data in the state control server so that the state control server can control the state of the server node according to the state data.

[0138] According to an embodiment of the present disclosure, any multiple modules among the acquisition module 510, the sorting module 520, the first determination module 530, the second determination module 540, and the update module 550 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the acquisition module 510, the sorting module 520, the first determination module 530, the second determination module 540, and the update module 550 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in an appropriate combination of any of them. Alternatively, at least one of the acquisition module 510 , the sorting module 520 , the first determination module 530 , the second determination module 540 , and the update module 550 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.

[0139] Figure 6 A block diagram of an electronic device suitable for implementing a method for controlling a server node state according to an embodiment of the present disclosure is schematically shown.

[0140] like Figure 6 As shown, the electronic device 600 according to an embodiment of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include an onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for executing different actions of the method flow according to an embodiment of the present disclosure.

[0141] Various programs and data required for the operation of the electronic device 600 are stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.

[0142] According to an embodiment of the present disclosure, electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to bus 604. Electronic device 600 may further include one or more of the following components connected to I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or a modem. Communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in drive 610 as needed, so that computer programs read from the removable media can be installed into storage section 608 as needed.

[0143] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.

[0144] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, such as but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 602 and / or RAM 603 described above and / or one or more memories other than ROM 602 and RAM 603.

[0145] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiments of the present disclosure.

[0146] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the processor 601 executes the computer program. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0147] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0148] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from a removable medium 611. When the computer program is executed by the processor 601, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0149] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0151] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways, even if such combinations and / or couplings are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or couplings are intended to fall within the scope of this disclosure.

[0152] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A method for controlling a server node state, comprising: Acquire status data generated within a preset time period, wherein the status data includes a plurality of time series status parameters arranged in a time series, and the time series status parameters represent the status of the plurality of server nodes; For each of the server nodes, performing time sorting on a plurality of state parameters corresponding to the server node extracted from the state data to obtain sorted data; Determining, based on each parameter group in the sorted data, a first state flip result of the server node corresponding to the plurality of parameter groups, wherein the parameter group includes two adjacent state parameters in the sorted data, and the first state flip result includes a node flip state or a node normal state corresponding to each parameter group, the node flip state indicating that two adjacent state parameters of the server node are different, and the node normal state indicating that two adjacent state parameters of the server node are the same; When the proportion of the number of node flip states meets the difference volatility threshold, the state of the server node is preliminarily determined to be an isolated state; When the number of server nodes in the isolation state meets the cluster isolation threshold, the status of a preset number of server nodes among the server nodes determined to be in the isolation state is updated to a target isolation state, wherein the preset number is determined based on the cluster isolation threshold, and the target isolation state indicates that the server nodes need to be restricted in data loading function.

2. The method according to claim 1, wherein After updating the states of a preset number of server nodes among the server nodes determined to be in the isolated state to the target isolation state, the method further includes: Determine, according to each timing group, a cluster flip sub-result corresponding to each timing group, wherein the timing group includes two timing state parameters that are adjacent in time sequence, and the cluster flip sub-result includes difference data or normal data, wherein the difference data indicates that a preset number of server nodes in the timing group have had their states changed; generating a cluster flipping result according to the plurality of cluster flipping sub-results; When the proportion of the amount of difference data in the cluster rollover result meets the preset cluster rollover threshold, the state of the server node in the target isolation state is updated to a normal state, and the normal state indicates that the server node can provide a data loading function.

3. The method according to claim 1, further comprising: When the number of server nodes in the isolated state does not meet the cluster isolation threshold, the state of the server nodes determined to be in the isolated state is updated to the target isolation state.

4. The method according to claim 1, further comprising: When the state of the server node is determined to be the target isolation state, a first instruction is sent to the server node, where the first instruction is used to control the server node to suspend data loading of the server node.

5. The method according to claim 2, further comprising: When the target isolation state of the server node is determined to be a normal state, a second instruction is sent to the server node, where the second instruction is used to control the loading of recovery data of the server node.

6. The method according to claim 1, wherein Determining, according to each parameter group in the sorted data, first state flip results of the server nodes corresponding to the plurality of parameter groups, comprising: For each parameter group, determining a sub-flip result of the parameter group according to whether the two adjacent state parameters are consistent; The first state reversal result is generated according to a plurality of the sub-reversal results corresponding to a plurality of the parameter groups.

7. The method according to claim 6, wherein: Determining a sub-flip result of the parameter group according to whether the two adjacent state parameters are consistent includes: In the case where the two state parameters are inconsistent, determining the sub-flip result as a node flip state; When the two state parameters are consistent, the sub-flip result is determined to be a normal node state.

8. The method according to claim 1, further comprising: The state data is stored in a state control server, so that the state control server can control the state of the server node according to the state data.

9. A device for controlling the status of a server node, comprising: An acquisition module, configured to acquire status data generated within a preset time period, wherein the status data includes a plurality of time series status parameters arranged in a time series, and the time series status parameters represent the status of the server node; a sorting module, configured to perform time sorting on a plurality of the time series state parameters corresponding to the server node extracted from the state data for each of the server nodes, to obtain sorted data; a first determining module, configured to determine, based on each parameter group in the sorted data, a first state flip result of the server node corresponding to the plurality of parameter groups, wherein the parameter group includes two adjacent state parameters in the sorted data, and the first state flip result includes a node flip state or a node normal state corresponding to each parameter group, wherein the node flip state indicates that two adjacent state parameters of the server node are different, and the node normal state indicates that two adjacent state parameters of the server node are the same; A second determining module is configured to preliminarily determine the state of the server node as an isolated state if the proportion of the number of node flip states meets a difference volatility threshold; An update module is used to update the status of a preset number of server nodes among the server nodes determined to be in an isolated state to a target isolation state when the number of server nodes in the isolated state meets the cluster isolation threshold, wherein the preset number is determined based on the cluster isolation threshold, and the target isolation state indicates that the server nodes need to be restricted in data loading function.

10. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to execute the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Traffic isolation method, device and system based on distributed service architecture

    CN113542027A

  • Network health monitoring method and device based on fault domain detection and medium

    CN115250225A