A cluster management method, system, apparatus, and medium
By dynamically adjusting node groups through a cluster management platform, based on traffic characteristics and risk values, the high availability and security issues of the cluster management system in the face of failures and attacks are solved, achieving more efficient traffic processing and risk response.
Patent Information
- Application Number
- CN202510508567.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Existing cluster management systems struggle to identify and prevent negative impacts such as hardware and software failures and network attacks in a timely manner, resulting in insufficient high availability and security in cluster management.
By using the cluster management platform to determine the number of target cluster nodes and temporary node groups based on the traffic characteristics and risk values of temporary traffic data, the node groups are dynamically adjusted using node activity, load information, and configuration information to achieve traffic allocation and transfer.
It improves the high availability and robustness of the cluster, ensures the efficiency and flexibility of traffic processing, can cope with potential risks, and improves the security and stability of the system.
Smart Images

Figure CN120416249B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the field of data processing, and in particular to a cluster management method, system, device and medium. BACKGROUND
[0002] With the development of cloud computing and cloud platform technologies, the flexibility, high performance and cost advantage of cluster management are gradually reflected. The cluster system includes a cluster scheduler, a cluster application manager and a plurality of cluster nodes. After receiving an application deployment request, the cluster scheduler can deploy the application to different cluster nodes. However, the cluster nodes may fail due to negative effects such as hardware and software failures, network attacks, etc. How to timely identify and prevent such negative effects, and then timely allocate and transfer traffic, is a problem that must be solved to ensure the effectiveness of cluster management.
[0003] Therefore, it is necessary to provide a cluster management method, system, device and medium to improve the high availability, security and robustness of the cluster. SUMMARY
[0004] One of the embodiments of the present specification provides a cluster management method, the method is executed by a cluster management platform, the cluster management platform controls a plurality of cluster nodes, and the method comprises: determining a traffic risk value based on temporary traffic features of temporary traffic data, the temporary traffic features comprising traffic size and traffic frequency of the temporary traffic data; determining a target cluster node quantity based on the temporary traffic features and the traffic risk value; determining a temporary node group based on the traffic risk value, the target cluster node quantity and node features of the plurality of cluster nodes, the node features comprising node activity, node load information and node configuration information.
[0005] One of the embodiments of the present specification provides a cluster management system, the system comprises a cluster management platform and a plurality of cluster nodes, the cluster management platform is configured to control a plurality of cluster nodes; the cluster management platform comprises a first determination module, a second determination module and a third determination module, wherein the first determination module is configured to determine a traffic risk value based on temporary traffic features of temporary traffic data, the temporary traffic features comprising traffic size and traffic frequency of the temporary traffic data; the second determination module is configured to determine a target cluster node quantity based on the temporary traffic features and the traffic risk value; the third determination module is configured to determine a temporary node group based on the traffic risk value, the target cluster node quantity and node features of the plurality of cluster nodes, the node features comprising at least node activity, node load information and node configuration information.
[0006] One of the embodiments of the present specification provides a cluster management apparatus, comprising at least one storage medium and at least one processor; the at least one storage medium is used to store computer instructions; the at least one processor is used to execute the computer instructions to realize the cluster management method described above.
[0007] One of the embodiments of the present specification provides a computer readable storage medium, the storage medium stores computer instructions, when the computer instructions are executed by a processor, the cluster management method described above is realized. BRIEF DESCRIPTION OF DRAWINGS
[0008] The present specification will be further illustrated in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same numbers represent the same structures, wherein:
[0009] Figure 1 is a schematic diagram of an application scenario of a cluster management system according to some embodiments of the present specification;
[0010] Figure 2 is a schematic diagram of a module of a cluster management system according to some embodiments of the present specification;
[0011] Figure 3 is an exemplary flowchart of a cluster management method according to some embodiments of the present specification;
[0012] Figure 4 is an exemplary flowchart of a method for determining the node activity of an effective cluster node according to some embodiments of the present specification;
[0013] Figure 5 is an exemplary flowchart of a method for determining a traffic risk value according to some embodiments of the present specification. DETAILED DESCRIPTION
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present specification, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some examples or embodiments of the present specification, and for those skilled in the art, the present specification can also be applied to other similar scenarios without creative labor. Unless it is obvious from the language environment or otherwise stated, the same reference numbers in the drawings represent the same structures or operations.
[0015] It should be understood that the "system", "apparatus", "unit" and / or "module" used herein is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the words can be replaced by other expressions.
[0016] Unless the context clearly indicates an exception, words such as "a," "an," "a kind," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0017] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0018] Figure 1 This is a schematic diagram illustrating an application scenario of a cluster management system based on some embodiments of this specification.
[0019] like Figure 1 As shown, the application scenario 100 of the cluster management system may include a cluster management platform 110, cluster nodes 120, storage devices 130, network 140, and terminal 150.
[0020] A cluster is a distributed system composed of multiple computer devices connected through a network. Clusters include, but are not limited to, load balancing clusters, high-availability clusters, high-performance computing clusters, and storage clusters. Each computer device in a cluster can be called a cluster node, server / server node, or computer node. Cluster nodes can include management nodes and worker nodes. Management nodes are cluster nodes used to manage and schedule worker nodes, while worker nodes are cluster nodes used to process tasks.
[0021] Cluster management platform 110 refers to a platform that can be used to manage and / or control a cluster (such as cluster 120). In some embodiments, cluster management platform 110 consists of one or more management nodes, which can be in the form of a single server or a group of servers.
[0022] The cluster management platform 110 can receive temporary traffic data that needs to be processed. Temporary traffic data includes access requests (such as network requests). In some embodiments, the cluster management platform 110 can distribute access requests to cluster nodes 120. In some embodiments, the cluster management platform 110 can obtain data and / or information from cluster nodes 120, storage devices 130, and terminals 150 via network 140. For example, the cluster management platform 110 can obtain node activity, node load information, and / or node configuration information of one or more cluster nodes in cluster nodes 120 via network 140. As another example, the cluster management platform 110 can send alarm information to terminal 150 via network 140 based on traffic risk values.
[0023] In some embodiments, the cluster management platform 110 can be used to control and / or manage one or more cluster nodes in the management cluster nodes 120. For example, the cluster management platform 110 can determine a temporary node group based on traffic risk values, the number of target cluster nodes, and the node characteristics of multiple cluster nodes. More related content can be found elsewhere in this specification (e.g., Figure 3 ).
[0024] Cluster node 120 can include multiple nodes, such as Figure 1 As shown, cluster nodes 120 include cluster nodes 120-1, 120-2, 120-3, ..., 120-n. In some embodiments, cluster nodes 120 can interact with one or more components of application scenario 100 (e.g., cluster management platform 110, storage device 130, and terminal 150) via network 140. For example, each cluster node in cluster nodes 120 can report its connection status information with other cluster nodes to cluster management platform 110.
[0025] Storage device 130 may store data and / or instructions. In some embodiments, storage device 130 may store data obtained from cluster management platform 110, cluster node 120, and / or terminal 150. For example, storage device 130 may store data obtained from cluster management platform 110 (such as traffic risk values). In some embodiments, storage device 130 may store data and / or instructions for performing the exemplary methods described herein. For example, storage device 130 may store instructions from cluster management platform 110 to perform the methods shown in the flowcharts. In some embodiments, storage device 130 may include a mass storage device, a removable storage device, volatile read-write memory, read-only memory (ROM), etc., or any combination thereof. In some embodiments, storage device 130 may be implemented on a cloud platform. In some embodiments, storage device 130 may be part of cluster management platform 110 and cluster node 120.
[0026] Network 140 may include any suitable network that facilitates information and / or data exchange within application scenario 100 of the cluster management system. In some embodiments, one or more components of application scenario 100 (e.g., cluster management platform 110, cluster nodes 120, storage devices 130, and terminals 150) may transmit information and / or data with one or more other components of application scenario 100 via network 140. For example, cluster management platform 110 may send temporary traffic data to one or more cluster nodes in cluster node 120 via network 140.
[0027] In some embodiments, network 140 can be any one or more of wired or wireless networks. In some embodiments, the network can be various topologies such as point-to-point, shared, centralized, or a combination of multiple topologies.
[0028] Terminal 150 may include mobile device 150-1, tablet computer 150-2, laptop computer 150-3, etc., or any combination thereof. In some embodiments, terminal 150 can interact with other components in application scenario 100 via network 140. For example, terminal 150 can receive information and / or instructions input by the user and send the received information and / or instructions to cluster management platform 110 via network 140. In some embodiments, terminal 150 can send and / or present risk warning information to the user based on the traffic risk value of temporary traffic data. The risk warning information includes, but is not limited to, voice, text, images, video, etc. More information about temporary traffic data and traffic risk values can be found elsewhere in this specification (e.g., Figure 3 ).
[0029] The above description is for illustrative purposes only, and actual application scenarios may vary.
[0030] It should be noted that application scenario 100 is provided for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can make various modifications or variations based on the description in this specification. However, these modifications and variations will not depart from the scope of this specification.
[0031] Figure 2 This is a schematic diagram of the modules of a cluster management system according to some embodiments of this specification.
[0032] like Figure 2 As shown, the cluster management system 200 may include a first determining module 210, a second determining module 220, and a third determining module 230. In some embodiments, the first determining module 210, the second determining module 220, and the third determining module 230 may be deployed in a cluster management platform (such as cluster management platform 110).
[0033] The first determining module 210 is configured to determine a traffic risk value based on temporary traffic features of the temporary traffic data, the temporary traffic features including traffic size and traffic frequency of the temporary traffic data.
[0034] In some embodiments, the temporary traffic features further include a first risk value and a second risk value, and the traffic risk value is determined based on the first risk value and the second risk value.
[0035] In some embodiments, the first determining module 210 is further configured to: determine a first sampling parameter based on the first risk value; sample the temporary traffic data based on the first sampling parameter to determine sampled traffic data; determine state slice data based on the test node group and the sampled traffic data; determine a sampling attack degree corresponding to the sampled traffic data based on the state slice data; and determine the second risk value based on the first sampling parameter and the sampling attack degree.
[0036] In some embodiments, the first risk value is determined based on a risk prediction model, and the risk prediction model is a machine learning model.
[0037] In some embodiments, the first determining module 210 is further configured to: determine the second risk value based on the first sampling parameter, a second sampling parameter, and the sampling attack degree.
[0038] In some embodiments, the first determining module 210 is further configured to: determine a first weight corresponding to the first risk value and a second weight corresponding to the second risk value, wherein the first weight and the second weight are related to the first sampling parameter and the second sampling parameter, respectively; and determine the traffic risk value based on the first risk value, the first weight, the second risk value, and the second weight.
[0039] In some embodiments, the first determining module 210 is further configured to: determine the second sampling parameter based on the test node group and the sampled traffic data; and sample the node state data based on the second sampling parameter to determine the state slice data.
[0040] The second determining module 220 is configured to determine a target cluster node quantity based on the temporary traffic features and the traffic risk value.
[0041] The third determining module 230 is configured to determine a temporary node group based on the traffic risk value, the target cluster node quantity, and node features of the plurality of cluster nodes.
[0042] In some embodiments, the node features at least include node activity, node load information, and node configuration information.
[0043] In some embodiments, the third determining module 230 is further configured to: monitor, for each cluster node, request queue information of network requests of the cluster node, the request queue information comprising a number of network requests; and determine, based on the request queue information, the number of cluster nodes in the temporary node group.
[0044] In some embodiments, the determining module 230 is further configured to: control at least two cluster nodes among the plurality of cluster nodes to send detection data packets to each other; determine, based on response data between the at least two cluster nodes, the plurality of effective cluster nodes and an effective node graph corresponding to the plurality of effective cluster nodes; determine, based on node load information of the effective cluster nodes, node importance of the effective cluster nodes; and determine, based on the effective node graph and the node importance, node activity of the effective cluster nodes.
[0045] In some embodiments, the determining module 230 is further configured to: control the at least two cluster nodes to send the detection data packets to each other based on a first sending frequency.
[0046] In some embodiments, the determining module 230 is further configured to: control the at least two effective cluster nodes to send the detection data packets to each other based on a second sending frequency.
[0047] In some embodiments, the second sending frequency is related to the importance of the at least two effective cluster nodes.
[0048] In some embodiments, the determining module 230 is further configured to: determine, based on node features of the plurality of cluster nodes, a plurality of node groups from the plurality of cluster nodes, any two cluster nodes in the plurality of node groups being non-overlapping; determine an anti-risk coefficient corresponding to each of the plurality of node groups; and determine the temporary node group based on the traffic risk value, the target cluster node number, and the anti-risk coefficient corresponding to each of the plurality of node groups.
[0049] In some embodiments, the determining module 230 is further configured to: determine, based on node features of the cluster nodes, a risk processing capability value of each cluster node; and determine the plurality of node groups based on the risk processing capability value of each of the plurality of cluster nodes.
[0050] It should be noted that the above description of the cluster management system and its modules is for the convenience of description only, and cannot limit the scope of the present specification to the embodiments described. It can be understood that, for those skilled in the art, after understanding the principles of the system, various modules can be combined or connected to form a subsystem without departing from the principles. For example, the first determination module 210, the second determination module 220 and the third determination module 230 can be different modules, or one module can implement the functions of two or more modules described above. For another example, each module can share a storage module, and each module can have its own storage module. Such variations are within the scope of the present specification.
[0051] Figure 3 is an exemplary flowchart of a cluster management method according to some embodiments of the present specification.
[0052] In some embodiments, the flow 300 can be performed by a cluster management platform. As shown in Figure 3 the flow 300 includes the following steps.
[0053] Step 310, determining a traffic risk value based on a temporary traffic feature of temporary traffic data.
[0054] Temporary traffic data refers to traffic data that needs to be analyzed and / or processed by the cluster management platform, including but not limited to various types of traffic data such as text, audio / video, image or instruction.
[0055] The temporary traffic data can include external traffic data and internal traffic data of the cluster.
[0056] The external traffic data includes communication data between the outside and the cluster. For example, it can be data of network interaction (such as network request) between the outside (such as network user, client or third-party service) and the cluster (such as cluster node in the cluster).
[0057] The internal traffic data includes communication data between cluster nodes. For example, it can be data transmitted between multiple cluster nodes (such as service access between cluster nodes).
[0058] In some embodiments, the cluster management platform is configured to receive and preprocess the temporary traffic data to realize the monitoring of the external traffic data and the management and coordination of the internal traffic data. Exemplarily, the preprocessing includes but is not limited to data verification, filtering, data feature extraction, etc.
[0059] The temporary traffic feature is used to reflect various properties of the temporary traffic data. The temporary traffic feature includes traffic size and traffic frequency of the temporary traffic data.
[0060] The traffic size of the temporary traffic data reflects the byte amount of the temporary traffic data in a preset time period (e.g., 10 minutes, 1 hour). The traffic frequency reflects the frequency of transmission of the temporary traffic data in the preset time period.
[0061] In some embodiments, the cluster management platform can monitor and analyze the temporary traffic data to determine the temporary traffic features corresponding to the temporary traffic. In some embodiments, the cluster management platform can determine the traffic size and the traffic frequency by performing traffic statistics on the temporary traffic data. For example, the number of bytes of the temporary traffic data and the number of network requests in the preset time period are counted respectively as the traffic size and the traffic frequency.
[0062] The temporary traffic features can be various features or indicators set according to actual needs (e.g., load balancing needs, concurrency needs, attack defense needs, etc.). In some embodiments, the temporary traffic features further include traffic basic features. For example, data sources (e.g., IP addresses of senders), destinations (e.g., IP addresses of receivers), access ports, protocol types (e.g., HTTP, TCP, etc.). The temporary traffic features further include traffic timing features. For example, periodicity of the traffic data, time point distribution of the traffic, etc.
[0063] In some embodiments, the cluster management platform can perform packet analysis by packet capturing and the like to determine the basic features and the timing features corresponding to each temporary traffic data.
[0064] In some embodiments, the temporary traffic features further include a first risk value. The first risk value is used to reflect the probability of abnormal situations of the cluster nodes (e.g., cluster nodes) after the temporary traffic data is actually processed by the cluster.
[0065] In some embodiments, the cluster management platform can monitor the resource usage of the cluster nodes to determine the first risk value. For example, the first risk value can be determined according to the current CPU occupancy, memory usage and the like resource usage information of the cluster nodes. For example, when the CPU occupancy rate of one or more cluster nodes is greater than a preset CPU usage threshold, it indicates that the first risk value is high.
[0066] In some embodiments, the cluster management platform can determine the first risk value based on a risk prediction model.
[0067] The risk prediction model refers to a model used to predict the first risk value. In some embodiments, the risk prediction model is a trained machine learning model. For example, a neural network model (Neural Network, NN), etc.
[0068] In some embodiments, the input of the risk prediction model 710 includes the temporary traffic features 701, and the output includes the first risk value 702.
[0069] In some embodiments, the initial model can be iteratively trained by multiple training samples to obtain the trained risk prediction model 710. Each training sample can include sample temporal traffic features corresponding to sample temporal traffic data in a historical time period (e.g., the past month, half a year, etc.). The training label can be determined according to the sample cluster anomaly information of the sample temporal traffic data after actual processing by the cluster, which can be manually labeled or labeled in other ways. It should be noted that the training label can be the value after normalization of various sample cluster anomaly information, for example, it can be a numerical value in the interval [0, 1].
[0070] The sample cluster anomaly information can be determined by the abnormal situation of one or more sample cluster nodes.
[0071] The sample cluster anomaly information can be various indicators preset according to actual needs, for example, resource usage anomaly information, security risk information, etc. of the sample cluster node. Exemplarily, the resource usage anomaly information includes CPU usage rate anomaly (e.g., usage rate greater than a preset CPU usage rate threshold), memory anomaly (e.g., occupancy rate greater than a preset memory occupancy rate threshold, memory leak), port anomaly (e.g., port occupancy rate greater than a threshold), etc.
[0072] The security risk information includes abnormal login attempts (e.g., multiple failed login attempts), permission changes (e.g., unauthorized permission changes), etc.
[0073] It should be noted that the cluster management platform can record the cluster anomaly information in the form of logs, texts, etc. for analysis and processing.
[0074] During training, the value of the loss function can be determined based on the difference between the output of the initial model and the training label. The parameters of the initial model can be iteratively updated based on the value of the loss function until the training termination condition is met (e.g., the loss function converges, a certain number of iterations are performed, etc.). The updated initial model can be used as the trained risk prediction model.
[0075] In some embodiments of the present specification, through the multiple risk prediction models, the correlation between the features of the temporal traffic data and the cluster anomaly information can be learned.
[0076] The traffic risk value is used to represent the degree of potential risk that the temporal traffic data brings to one or more components (e.g., the cluster management platform 110, the cluster node 120) in the cluster management system. The potential risk includes but is not limited to high load, existence of malicious attacks, etc. The greater the traffic risk value, the greater the potential risk. The cluster management platform can evaluate the traffic risk value corresponding to the temporal traffic data according to the temporal traffic features.
[0077] In some embodiments, the cluster management platform can detect the temporary traffic features of the temporary traffic data by deploying a network threat detection engine (such as Suricata, etc.), to determine the traffic risk value. For example, the cluster management platform can determine the traffic risk value corresponding to the temporary traffic data according to the intrusion information or records detected by the network threat detection engine, through a preset threat risk relationship table, and generate an alarm.
[0078] In some embodiments, the cluster management platform is deployed with a vector database (such as Milvus), which includes a plurality of reference records constructed based on historical traffic data. The reference records include reference traffic features (for example, reference first risk values) corresponding to the reference historical traffic data and corresponding reference traffic risk values. The cluster management platform can perform vector matching processing in the vector database based on the temporary traffic features, and take the reference traffic risk value corresponding to the reference record with the maximum similarity between the temporary traffic features and the reference traffic features as the traffic risk value corresponding to the temporary traffic data. Wherein, the maximum similarity can be the minimum vector distance.
[0079] In some embodiments, the temporary traffic features further include a second risk value, and the cluster management platform can further determine the traffic risk value based on the first risk value and the second risk value. For more information about the second risk value and the determination of the traffic risk value based on the first risk value and the second risk value, please refer to Figure 5 and the description thereof.
[0080] Step 320, determining the target cluster node quantity based on the temporary traffic features and the traffic risk value.
[0081] The target cluster node quantity refers to the number of cluster nodes required to process the temporary traffic. In some embodiments, the cluster management platform can pre-set a node quantity reference table. The node quantity reference table includes temporary traffic features, traffic risk values, and corresponding reference cluster node quantities. Wherein, the node quantity reference table can be obtained according to historical data. The cluster management platform can match the corresponding reference cluster node quantity in the node quantity reference table according to the temporary traffic features and the traffic risk value, as the target cluster node quantity.
[0082] Step 330, determining the temporary node group based on the traffic risk value, the target cluster node quantity, and the node features of the plurality of cluster nodes.
[0083] In some embodiments, the node features include node activity, node load information, and node configuration information.
[0084] The node activity is used to reflect the frequency of data processing (such as network request response) of the cluster node. For example, the more network requests a cluster node responds to, the more frequently it processes data, and the higher its corresponding node activity.
[0085] The node load information is used to reflect the resource usage pressure of the cluster nodes. The node load information includes, but is not limited to, CPU load (such as CPU usage), memory load (such as memory usage), disk I / O load (such as disk read-write time proportion), network load (such as bandwidth occupancy), and the like.
[0086] The node configuration information refers to the hardware configuration information of the cluster nodes. For example, the number of CPUs, the size of memory, the capacity of disks, the network bandwidth, and the like.
[0087] The temporary node group refers to a node group composed of cluster nodes for processing temporary traffic data. For example, the cluster management platform can select a target number of cluster nodes from a plurality of cluster nodes to form a temporary node group. The selection methods include, but are not limited to, random selection, priority selection of cluster nodes with smaller processing load according to the processing load of the cluster nodes, and the like.
[0088] In some embodiments, the cluster management platform can determine a plurality of node groups from a plurality of cluster nodes based on the node characteristics of the plurality of cluster nodes, and determine an anti-risk coefficient corresponding to each of the node groups, and then determine a temporary node group based on the traffic risk value, the target number of cluster nodes, and the anti-risk coefficient corresponding to each of the node groups.
[0089] In some embodiments, the cluster management platform can select a preset number of cluster nodes with node relevance from a plurality of cluster nodes to form a plurality of node groups according to the node characteristics of the plurality of cluster nodes. The node relevance can be determined according to actual needs. For example, a plurality of cluster nodes with the same geographic area, similar business or task types (such as storage processing type, business processing type) can be taken as a node group.
[0090] In some embodiments, any two cluster nodes in the plurality of node groups do not overlap, that is, any cluster node exists in only one node group.
[0091] In some embodiments, the cluster management platform can determine the risk processing capability value of the cluster nodes based on the node characteristics of the cluster nodes, and determine a plurality of node groups based on the risk processing capability values of the plurality of cluster nodes.
[0092] In some embodiments, the cluster management platform can preprocess the node characteristics of each cluster node. The preprocessing can include, but is not limited to, dimensionless processing by various algorithms such as normalization, reverse normalization, intervalization, Min-Max normalization, and the like, to obtain dimensionless data, so that the data structure, unit, number system or format of the node characteristic value corresponding to the node characteristics are kept uniform, thereby enabling subsequent uniform operations (such as arithmetic operations, etc.).
[0093] In some embodiments, the cluster management platform can perform weighted summation based on the pre-processed node features of each cluster node to obtain a risk processing capability value. The weighting weight can be pre-set based on experience. The risk processing capability value of the cluster node can be used to reflect the robustness or reliability of the cluster node for data processing.
[0094] In some embodiments, the cluster management platform can perform gradient division according to the risk processing capability value of each cluster node to obtain a gradient distribution of the risk processing capability values of the plurality of cluster nodes, for example, the gradient division can be according to the size of the risk processing capability value in descending order. In some embodiments, the cluster management platform can select a pre-set number of cluster nodes (such as 10) from the plurality of cluster nodes after gradient division to generate a plurality of node groups. It should be noted that the pre-set number of groups can be the same as or different from the target number of cluster nodes.
[0095] The risk resistance coefficient refers to the stability of the cluster nodes in the node group after processing the temporary traffic data. The higher the risk resistance coefficient, the higher the stability of the cluster nodes after processing the temporary traffic data.
[0096] In some embodiments, the risk resistance coefficient corresponding to the node group can be determined according to the average of the risk processing capability values of the plurality of cluster nodes in the node group.
[0097] In some embodiments, the cluster management system can select a node group from the plurality of node groups as a target node group according to the traffic risk value. The greater the traffic risk value, the greater the risk resistance coefficient of the selected target node group.
[0098] In some embodiments, the cluster management system can also adjust the target node group according to the target number of cluster nodes to match the number of cluster nodes of the target node group with the target number of cluster nodes, thereby obtaining a temporary node group with the target number of cluster nodes. For example, in response to the number of cluster nodes of the target node group being less than the target number of cluster nodes, one or more cluster nodes with larger risk processing capability values can be selected from node groups with similar risk resistance coefficients to join the target node group. In response to the number of cluster nodes of the target node group being greater than the target number of cluster nodes, one or more cluster nodes with smaller risk processing capability values can be removed from the target node group.
[0099] In some embodiments, for each cluster node in the temporary node group, the cluster management platform can monitor the request queue information of the network request of the cluster node, and determine the number of cluster nodes in the temporary node group based on the request queue information.
[0100] The request queue information can be used to reflect the situation of network requests that need to be processed in the cluster nodes. The network requests that need to be processed include network requests that are being processed and network requests that are waiting to be processed.
[0101] In some embodiments, the request queue information includes a number of network requests. The larger the number of network requests indicates that the concurrent capability of the cluster nodes is insufficient or the load is too large.
[0102] In some embodiments, the cluster management platform can perform expansion processing or contraction processing on the temporary node group according to the request queue information corresponding to each cluster node. In response to the number of network requests being greater than a preset expansion threshold, the cluster management platform can append one or more cluster nodes to the temporary node group to realize expansion of the temporary node group. In response to the number of network requests that need to be processed being less than a preset contraction threshold, the cluster management platform can remove one or more cluster nodes (such as idle cluster nodes) in the temporary node group to realize contraction of the temporary node group. The preset expansion threshold and the preset contraction threshold can be preset thresholds.
[0103] In some embodiments of the present specification, by monitoring the request queue information of the cluster nodes in the temporary group, the temporary node group can be automatically expanded or contracted, thereby realizing load balancing of multiple cluster nodes.
[0104] In some embodiments of the present specification, considering the temporary traffic characteristics of the temporary traffic data, the software and hardware resources of the cluster nodes in the cluster can be coordinated and managed, and the efficiency or flexibility of processing the temporary traffic data can be improved. At the same time, considering the traffic risk value of the temporary traffic data, the cluster can cope with potential risks, thereby improving the high availability and robustness of the cluster.
[0105] Figure 4 is an exemplary flowchart of a method for determining the node activity of an effective cluster node according to some embodiments of the present specification.
[0106] In some embodiments, the flow 400 can be performed by the cluster management platform. As shown in Figure 4 , the flow 400 includes the following steps.
[0107] Step 410, control at least two cluster nodes among the plurality of cluster nodes to send detection data packets to each other.
[0108] The detection data packet refers to data used to determine whether communication between any two cluster nodes can be performed. For example, a preset lightweight data (such as a heartbeat signal).
[0109] In some embodiments, the cluster management platform can control at least two cluster nodes to send detection data packets to each other based on a first sending frequency.
[0110] The first sending frequency is used to represent the frequency of sending the detection data packet between any two cluster nodes. The first sending frequency can be a value pre-set based on experience, for example, the first sending frequency is 20 seconds per interval.
[0111] In some embodiments, the first sending frequency can be determined according to the node characteristics of the cluster nodes. For example, the greater the node activity, the greater the first sending frequency.
[0112] In step 420, based on the response data between at least two cluster nodes, a plurality of effective cluster nodes and a plurality of effective node graphs corresponding to the plurality of effective cluster nodes are determined.
[0113] The response data refers to the data received after the cluster node sends the detection data. When any two cluster nodes can send detection data packets to each other and receive response data, it means that any one of the two cluster nodes is an effective cluster node, which is also called an active node.
[0114] The effective node graph refers to a knowledge graph constructed according to the communication connection information of a plurality of effective cluster nodes in the cluster, which can be used to reflect the active information and change information of the plurality of effective cluster nodes.
[0115] In some embodiments, the effective node graph includes a plurality of graph nodes and a plurality of graph edges, and the graph nodes and the graph edges have attributes. Each graph node corresponds to an effective cluster node, and the graph edge connects two graph nodes that have communication connection, which represents that the two effective cluster nodes can be connected in communication.
[0116] In some embodiments, the graph node attribute includes at least one of the node activity, the node load information, the node configuration information, etc. of the effective cluster node corresponding to the graph node. In some embodiments, the graph node attribute can be determined based on the node characteristics of the effective cluster node corresponding to the graph node.
[0117] In some embodiments, the graph node attribute further includes node importance. For more details, see step 430.
[0118] It should be noted that for a certain effective cluster node, the graph node attribute corresponding to the effective cluster node can be dynamically updated according to the actual situation of the effective cluster node. For example, the cluster management platform can update the node load information in the graph node attribute according to the load information of the effective cluster node.
[0119] In some embodiments, the graph edge attribute includes a connection state, which is used to represent the communication connection state or disconnection state between the two effective cluster nodes corresponding to the graph edge.
[0120] In some embodiments, the cluster management platform can construct and / or update the effective node graph based on the communication connection status of any two cluster nodes.
[0121] At step 430, the node importance of each effective cluster node is determined based on the node load information of each effective cluster node.
[0122] The node importance of an effective cluster node is used to reflect the importance of the effective cluster node for processing the temporary traffic data. In some embodiments, the node importance of an effective cluster node can be determined based on the load information of the effective cluster node. For example, the node importance can be determined based on the ratio of the load of the effective cluster node to the average load of the cluster. The load can be determined based on the traffic size and / or traffic frequency of the temporary traffic data to be processed by the effective cluster node.
[0123] In some embodiments, the cluster management platform can control at least two effective cluster nodes to send detection data packets to each other based on a second sending frequency, the second sending frequency being related to the importance of each of the effective cluster nodes.
[0124] The second sending frequency is used to represent the frequency of sending detection data between any two effective cluster nodes. In some embodiments, the second sending frequency is different from the first sending frequency. The second sending frequency can be set to be greater than the first sending frequency to ensure the timeliness of the communication connection status detection between the effective cluster nodes.
[0125] In some embodiments, the edge attribute further includes the second sending frequency. For any graph edge, the cluster management platform can adjust the second sending frequency corresponding to the graph edge according to the importance of the two effective cluster nodes connected by the graph edge. Thus, the two effective cluster nodes are controlled to send detection data packets to each other based on the adjusted second sending frequency. The greater the average of the importance of the two effective cluster nodes, the higher the corresponding second sending frequency.
[0126] In some embodiments of the present specification, the effective node graph enables the cluster management platform to more timely and efficiently manage the communication connection status of the effective cluster nodes. Meanwhile, setting the second sending frequency according to the importance of different effective cluster nodes enables better tracking and protection of the communication connection status between the more important effective cluster nodes, thereby ensuring the processing of the temporary traffic data.
[0127] At step 440, the node activity of each effective cluster node is determined based on the effective node graph and the node importance of each effective cluster node.
[0128] The node activity of an effective cluster node is used to reflect the relative activity of the effective cluster node in the local network in which the effective cluster node is located.
[0129] In some embodiments, the local network in which the effective cluster node is located can be a network formed by all the effective cluster nodes in the effective node graph that have a communication connection with the effective cluster node. For example, for a certain effective cluster node A, the local network corresponding to the effective cluster node A is a network A formed by all the effective cluster nodes that have a communication connection with the effective cluster node A. For another effective cluster node B, the local network corresponding to the effective cluster node B is a network B formed by all the effective cluster nodes that have a communication connection with the effective cluster node B.
[0130] The node activity degree can be determined based on the following formula (1):
[0131]
[0132] In formula (1), r i represents the activity degree of the i-th effective cluster node, σ represents the standard deviation of the node importance of all the effective cluster nodes in the local network in which the effective cluster node is located, a represents the average value of the node importance of all the effective cluster nodes, e represents the number of edges in which the effective cluster node exists, and E represents the total number of graph edges in the effective node graph.
[0133] wherein, The value of the node activity degree can be used to reflect the concentration degree of the effective cluster node in the temporary traffic data processing of the local network. The greater the value, the more concentrated it is, and the greater the corresponding node activity degree is.
[0134] In some embodiments of the present specification, the node activity degree is introduced to consider the relative concentration degree of the temporary traffic processing of different effective cluster nodes, which can better evaluate the load balancing of different effective cluster nodes, thereby making the scheduling of the cluster management platform more accurate.
[0135] In some embodiments, the local network in which the effective cluster node is located can be a network formed by all the effective cluster nodes in the effective node graph. Considering that all the cluster nodes in the effective node graph perform actual task processing (such as load balancing cooperation, etc.) as a whole in actual application scenarios, the network formed by all the effective cluster nodes in the effective node graph can be used as the local network corresponding to each effective cluster node in the effective node graph, and the node activity degree of each effective cluster node can be calculated based on the above formula (1).
[0136] In some embodiments of the present specification, all the effective cluster nodes in the effective node graph are used as the local network in which any effective cluster node is located, which can reduce the calculation amount of constructing the local network corresponding to each effective cluster node.
[0137] Figure 5is an exemplary flowchart of a method of determining a traffic risk value according to some embodiments of the present specification.
[0138] At step 510, a first sampling parameter is determined based on the first risk value.
[0139] The first sampling parameter refers to a parameter for extracting the temporary traffic data by the cluster management platform. In some embodiments, the first sampling parameter comprises a first sampling frequency.
[0140] The first sampling frequency is used to reflect how frequently the temporary traffic data is extracted or collected by the cluster management platform. For example, the temporary traffic data can be collected based on a preset time period or time interval (e.g., every 10 seconds, every minute). The shorter the preset time interval, the greater the first sampling frequency (i.e., the more frequently).
[0141] In some embodiments, the cluster management platform can adjust the first sampling frequency according to the first risk value. When the first risk value is greater, it indicates that the potential risk of the cluster is greater, and the first sampling frequency can be set greater to enhance the sampling intensity, thereby improving the strength of the detection.
[0142] In some embodiments of the present specification, the first sampling frequency is adjusted considering the first risk value, which improves the collection efficiency while avoiding excessive frequency leading to waste of resources.
[0143] At step 520, the temporary traffic data is sampled based on the first sampling parameter to determine sampled traffic data.
[0144] The sampled traffic data refers to the temporary traffic data collected within a sampling time period based on the first sampling parameter. The sampling time period can be a preset time period (e.g., the past 1 minute, 10 minutes) up to the current time.
[0145] The sampled traffic data can be the temporary traffic data based on a time-sliced sequence within the sampling time period, and the time slice can be a preset time interval or time period (e.g., 5 seconds). For example, it can be the temporary traffic data every 5 seconds within the past 1 minute. Every 5 seconds is a time slice.
[0146] At step 530, state slice data is determined based on the test node group and the sampled traffic data.
[0147] The test node group refers to a node group composed of one or more cluster nodes for risk testing of the sampled traffic data. The cluster nodes in the test node group can be referred to as test cluster nodes.
[0148] The risk testing can be used to simulate, reproduce, track, and analyze the risk of the sampled traffic data to the cluster.
[0149] In some embodiments, the risk test can be used for external risk analysis. For example, it can include, but is not limited to, attack hunting (such as collection of attack types such as DDoS attack, Trojan implant, etc.), attack behavior (such as destruction behavior, data theft, permission) analysis and / or attack chain tracking (such as cross-node attack chain) and the like.
[0150] It should be noted that the risk test can be determined according to actual needs. For example, the risk test can also be used for internal risk analysis. For example, the potential load balancing caused by the data transmission between the simulation cluster nodes and the like.
[0151] In some embodiments, the test cluster node is set to be isolated from the non-test cluster node environment. For example, the cluster management platform can isolate the test node group from the non-test cluster node through network isolation (such as independent subnets), network access permission (limit network communication), set up firewall and the like.
[0152] In some embodiments of the present specification, by setting the test node group, the risk brought by the sampling temporary traffic data can be analyzed, and at the same time, the influence on the data processing (such as business processing) of the non-test cluster node is avoided, and the normal operation of the entire cluster is ensured.
[0153] The state slice data refers to the node state data corresponding to a time slice of the test cluster node in the sampling time period. For example, for a test cluster node, it can include multiple state slice data, and each state slice data corresponds to a time slice.
[0154] The node state data is used to reflect the node running state of the test cluster node after processing the sampling traffic data. The node running state can be determined based on the value of a plurality of preset evaluation indexes.
[0155] For example, the evaluation index includes one or more combinations of a plurality of running states such as heartbeat signal loss rate, heartbeat signal delay rate, access failure rate, and abnormal source connection proportion.
[0156] The heartbeat signal loss rate refers to the proportion of failed heartbeat signals sent or received. The heartbeat signal delay rate refers to the proportion of delayed heartbeat signals. The access failure rate refers to the proportion of timeout or failed access requests. The abnormal source connection proportion refers to the proportion of access connections of abnormal access sources (such as known threat source IP, unexpected business geographic area / position, etc.).
[0157] When the value of the evaluation index is greater than the preset proportion threshold value corresponding to the evaluation index, the evaluation index can be referred to as an abnormal index. The greater the value of the abnormal index, the worse the node running state. The state slice data with the abnormal index can be referred to as abnormal state slice data. The test cluster node with the abnormal state slice data can be referred to as an abnormal test cluster node.
[0158] In some embodiments, the cluster management platform can determine a second sampling parameter based on the test node group and the sampled traffic data; and sample the node state data based on the second sampling parameter to determine the state slice data.
[0159] The second sampling parameter reflects how frequently the cluster management platform extracts or collects the node state data. In some embodiments, the second sampling parameter comprises a second sampling frequency.
[0160] In some embodiments, the cluster management platform can determine the second sampling frequency according to the configuration information of the test node (such as CPU resources, memory resources), the temporary traffic characteristics (such as traffic size) of the sampled traffic data, and the node state data.
[0161] For example, the second sampling frequency can be positively correlated with the CPU resource performance (such as computing power, core number). For example, the better the CPU resource performance, the larger the second sampling frequency can be to cover more state slice data, so that the subsequent analysis result is more accurate; when the CPU resource performance is poor, the second sampling frequency can be set smaller to reduce the computing pressure of the test cluster node.
[0162] For another example, the second sampling frequency can be negatively correlated with the sampled traffic data. For example, the less or sparser the sampled traffic is, the sparser the corresponding node state data can be, and a larger second sampling frequency can be set to be able to collect sufficient state slice data.
[0163] Step 540, determining a sampling attack degree corresponding to the sampled traffic data based on the state slice data.
[0164] The sampling attack degree is used to reflect the degree of the test cluster node in a negative running state after processing the sampled traffic data. The negative running state refers to an unintended node running state. For example, for the evaluation index of the heartbeat signal loss rate, if the heartbeat signal loss rate is greater than 10% of the preset heartbeat signal loss rate, it indicates that the test cluster node is in a negative running state.
[0165] In some embodiments, the cluster management platform can determine the sampling attack degree corresponding to the sampled traffic data based on one or more node running states in the state slice data. For example, the sampling attack degree can be determined according to one or more combinations of the heartbeat signal loss rate, the heartbeat signal delay rate, the access failure rate, and the abnormal source connection proportion of the test cluster node. The attack degree can be represented in various preset forms, such as mild, moderate, and severe.
[0166] In some embodiments, the cluster management platform can determine the sampling attack degree based on the abnormal test cluster node and the abnormal state slice data of the abnormal test cluster node.
[0167] For example, the greater the number of abnormal test cluster nodes, the greater the sampling attack degree; the greater the number of abnormal state slice data, the greater the attack degree; the greater the value of the abnormal index in the abnormal state slice data, the greater the sampling attack degree.
[0168] In some embodiments, the cluster management platform can set an attack index table, which includes the number of abnormal test cluster nodes, the number of abnormal state slice data, and the value of each abnormal index, and the corresponding attack degree. The cluster management platform can match the target attack degree corresponding to the sampling temporary traffic by table lookup, as the sampling attack degree.
[0169] In some embodiments of the present specification, considering the complexity of the sampling temporary traffic data (such as timeliness, concealment, etc.), considering the number of abnormal test cluster nodes, the influence range of the attack degree can be evaluated; considering the number of abnormal state slice data, the frequency of the attack degree can be evaluated; considering the value of the abnormal index, the severity of the attack degree can be evaluated.
[0170] Step 550, determining a second risk value based on the first sampling parameter and the sampling attack degree.
[0171] The second risk value can be used to evaluate the strength of the potential attack risk of the temporary traffic data to the cluster node. The higher the second risk value, the greater the potential attack risk. For example, the attack strength of the actual temporary traffic data can be reflected by the sampling traffic data corresponding to the temporary traffic data.
[0172] In some embodiments, the second risk value is positively correlated with the first sampling parameter and the sampling attack degree. For example, the higher the first sampling frequency and the sampling attack degree, the greater the second risk value. The lower the first sampling parameter and the second sampling parameter, the lower the coverage of the sampling, which may have certain contingency, and the higher the sampling attack degree, the higher the second risk value.
[0173] In some embodiments, the cluster management platform can also determine the second risk value based on the first sampling parameter, the second sampling parameter, and the sampling attack degree. For example, the second risk value is positively correlated with the first sampling frequency, the second sampling frequency, and the sampling attack degree.
[0174] In some embodiments of the present specification, the first sampling parameter and the second sampling parameter are comprehensively considered, which can fully consider the sampling coverage of the temporary traffic data and the sampling coverage of the state slice data, so that the second risk value is more accurate.
[0175] At step 560, a traffic risk value is determined based on the first risk value and the second risk value.
[0176] In some embodiments, the cluster management platform can determine a first weight corresponding to the first risk value and a second weight corresponding to the second risk value; and determine the traffic risk value based on the first risk value, the first weight, the second risk value and the second weight.
[0177] The first weight and the second weight can be preset values, for example, the first weight and the second weight are both 0.5, or other values set according to experience (such as 0.4 and 0.6 respectively).
[0178] In some embodiments, the first weight and the second weight are respectively related to (for example, positively related to) the first sampling parameter and the second sampling parameter. For example, when the first sampling parameter (such as the first sampling frequency) increases and the second sampling parameter remains unchanged, it indicates that the sampling coverage of the temporary traffic data increases, and the first weight can be adjusted to be higher, and the second weight can be adjusted to be lower.
[0179] In some embodiments of the present specification, the relationship between the first sampling parameter and the second sampling parameter is considered to adjust the first weight and the second weight, so as to balance the consumption of resources by the sampling coverage of the temporary traffic data and the sampling coverage of the state slice data while ensuring the accuracy of the second risk value.
[0180] It should be noted that the above description of the process is only for example and illustration, and does not limit the scope of the present specification. Those skilled in the art can make various modifications and changes to the process under the guidance of the present specification. However, these modifications and changes are still within the scope of the present specification.
[0181] The above has described the basic concept, and it is obvious that the above detailed disclosure is only as an example and does not constitute a limitation on the present specification. Although it is not explicitly stated here, those skilled in the art can make various modifications, improvements and corrections to the present specification. Such modifications, improvements and corrections are suggested in the present specification, so such modifications, improvements and corrections still belong to the spirit and scope of the exemplary embodiments of the present specification.
[0182] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.
[0183] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on existing servers or mobile devices.
[0184] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.
[0185] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values are set as precisely as feasible.
[0186] Each patent, patent application, patent publication, and other material cited in this specification is hereby incorporated by reference in its entirety herein for the teachings relevant to the sentence and / or paragraph in which the reference is presented. Document histories, to the extent not inconsistent with the pertinent U.S. patent application file history, are also incorporated by reference herein. To the extent that material incorporated by reference contradicts or contradicts any portion of this specification, including definition, the portion of the material incorporated by reference prevails. Note, however, that in the event of inconsistencies between any such material and the present specification, including definitions, the present specification, including definitions, will control.
[0187] Finally, it should be understood that the embodiments described herein are merely exemplary of the principles of the present description. Other embodiments can be devised without departing from the scope of the present description. Accordingly, the embodiments described herein are not intended to limit the scope of the present description, but rather are intended to be exemplary thereof.
Claims
1. A cluster management method, wherein the method is executed by a cluster management platform, the cluster management platform controlling multiple cluster nodes, the method comprising: Based on the temporary traffic characteristics of temporary traffic data, a traffic risk value is determined, wherein the temporary traffic characteristics include the traffic size and traffic frequency of the temporary traffic data; Based on the temporary traffic characteristics and the traffic risk value, the number of target cluster nodes is determined; Based on the node characteristics of the multiple cluster nodes, the risk handling capability value of the cluster nodes is determined. The node characteristics include node activity, node load information, and node configuration information. Based on the risk handling capability values of each of the multiple cluster nodes, multiple node groups are determined, and no two cluster nodes in the multiple node groups overlap. Determine the risk resistance coefficient corresponding to each of the multiple node groups; Based on the traffic risk value, the number of target cluster nodes, and the risk resistance coefficient, a temporary node group is determined.
2. The method according to claim 1, characterized in that, The method further includes: Control at least two of the multiple cluster nodes to send detection data packets to each other; Based on the response data between at least two of the cluster nodes, a plurality of valid cluster nodes and a valid node graph corresponding to the plurality of valid cluster nodes are determined. Based on the node load information of the effective cluster nodes, the node importance of the effective cluster nodes is determined; Based on the effective node graph and the node importance, the node activity of the effective cluster nodes is determined.
3. The method according to claim 1, characterized in that, The temporary traffic characteristics also include a first risk value and a second risk value. The traffic risk value is determined based on the first risk value and the second risk value. The method for determining the second risk value includes: Based on the first risk value, the first sampling parameters are determined; Based on the first sampling parameter, the temporary traffic data is sampled to determine the sampled traffic data; Based on the test node group and the sampled traffic data, determine the state slice data; Based on the state slice data, the degree of attack on the sampled traffic data is determined; A second risk value is determined based on the first sampling parameters and the degree of attack on the sampling.
4. A cluster management system, characterized in that, The system includes a cluster management platform and multiple cluster nodes. The cluster management platform is configured to control the multiple cluster nodes. The cluster management platform includes a first determining module, a second determining module, and a third determining module, wherein... The first determining module is configured to determine a traffic risk value based on the temporary traffic characteristics of the temporary traffic data, wherein the temporary traffic characteristics include the traffic size and traffic frequency of the temporary traffic data; The second determining module is configured to determine the number of target cluster nodes based on the temporary traffic characteristics and the traffic risk value; The third determining module is configured to determine the risk handling capability value of the cluster nodes based on the node characteristics of the multiple cluster nodes, wherein the node characteristics include node activity, node load information and node configuration information. Based on the risk handling capability values of each of the multiple cluster nodes, multiple node groups are determined, and no two cluster nodes in the multiple node groups overlap. Determine the risk resistance coefficient corresponding to each of the multiple node groups; Based on the traffic risk value, the number of target cluster nodes, and the risk resistance coefficient, a temporary node group is determined.
5. The system according to claim 4, characterized in that, The third determining module is further configured to: Control at least two of the multiple cluster nodes to send detection data packets to each other; Based on the response data between at least two of the cluster nodes, a plurality of valid cluster nodes and a valid node graph corresponding to the plurality of valid cluster nodes are determined. Based on the node load information of the effective cluster nodes, the node importance of the effective cluster nodes is determined; Based on the effective node graph and the node importance, the node activity of the effective cluster nodes is determined.
6. A cluster management device, characterized in that, Includes at least one storage medium and at least one processor; The at least one storage medium is used to store computer instructions; The at least one processor is used to execute the computer instructions to implement the cluster management method as described in any one of claims 1 to 3.
7. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions, which, when executed by a processor, implement the cluster management method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Distributed service cluster runtime parameter adaptive processing method, device and system
CN113032233A
Distributed service cluster load adaptive processing method, device and system
CN113055479A