Cluster management method, system and device and medium
Through the cluster management platform, analyzing temporary traffic data and node characteristics is formed to form a temporary node group, which solves the problem of untimely traffic management in the cluster system when facing negative impacts, and improves the high availability and robustness of the cluster.
Patent Information
- Application Number
- CN202510508567.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-22
AI Technical Summary
When existing cluster management systems face negative impacts such as software and hardware failures and network attacks, it is difficult to identify and prevent them in a timely manner, resulting in untimely allocation and transfer of traffic, affecting the high availability and security of cluster management.
Through the cluster management platform, the traffic risk value and the number of target cluster nodes are determined based on the characteristics of temporary traffic data, and combined with node characteristics, a temporary node group is formed to improve the high availability and robustness of the cluster.
It realizes efficient traffic management of the cluster system, improves the high availability and robustness of the cluster, can respond to potential risks in a timely manner, and ensures the flexibility and security of traffic allocation.
Smart Images

Figure CN120416249A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of data processing, and particularly to a cluster management method, system, device, and medium. Background Art
[0002] With the development of technologies such as cloud computing and cloud platforms, the flexibility, high performance, and cost advantages of cluster management have gradually emerged. A cluster system includes a cluster scheduler, a cluster application manager, and multiple cluster nodes. After receiving an application deployment request, the cluster scheduler can deploy the application to different cluster nodes. However, cluster nodes may experience failure problems due to negative impacts such as software and hardware failures and network attacks. How to timely identify and prevent such negative impacts, and then timely perform traffic allocation and transfer, is a problem that must be solved to ensure the effectiveness of cluster management.
[0003] Therefore, it is necessary to provide a cluster management method, system, device, and medium to improve the high availability, security, and robustness of the cluster. Summary of the Invention
[0004] One embodiment of this specification provides a cluster management method, which is executed by a cluster management platform that controls multiple cluster nodes. The method includes: determining a traffic risk value based on the temporary traffic characteristics of temporary traffic data, where the temporary traffic characteristics include the traffic volume and traffic frequency of the temporary traffic data; determining the number of target cluster nodes based on the temporary traffic characteristics and the traffic risk value; and determining a temporary node group based on the traffic risk value, the number of target cluster nodes, and the node characteristics of the multiple cluster nodes, where the node characteristics include node activity, node load information, and node configuration information.
[0005] One embodiment of this specification provides a cluster management system, which includes a cluster management platform and multiple cluster nodes. The cluster management platform is configured to control multiple cluster nodes. The cluster management platform includes a first determination module, a second determination module, and a third determination module. Among them, the first determination module is configured to determine a traffic risk value based on the temporary traffic characteristics of temporary traffic data, where the temporary traffic characteristics include the traffic volume and traffic frequency of the temporary traffic data; the second determination module is configured to determine the number of target cluster nodes based on the temporary traffic characteristics and the traffic risk value; the third determination module is configured to determine a temporary node group based on the traffic risk value, the number of target cluster nodes, and the node characteristics of the multiple cluster nodes, where the node characteristics at least include node activity, node load information, and node configuration information.
[0006] One embodiment of the present specification provides a cluster management device, including at least one storage medium and at least one processor; the at least one storage medium is used to store computer instructions; the at least one processor is used to execute the computer instructions to implement the above-mentioned cluster management method.
[0007] One embodiment of the present specification provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a processor, the above-mentioned cluster management method is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] This specification will be further described by way of exemplary embodiments, which will be described in detail through the accompanying drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:
[0009] Figure 1 is a schematic diagram of an application scenario of a cluster management system according to some embodiments of this specification;
[0010] Figure 2 is a schematic diagram of modules of a cluster management system according to some embodiments of this specification;
[0011] Figure 3 is an exemplary flowchart of a cluster management method according to some embodiments of this specification;
[0012] Figure 4 is an exemplary flowchart of a method for determining the node activity of valid cluster nodes according to some embodiments of this specification;
[0013] Figure 5 is an exemplary flowchart of a method for determining a traffic risk value according to some embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of this specification. For those of ordinary skill in the art, without creative efforts, this specification can also be applied to other similar scenarios based on these drawings. Unless obvious from the language context or otherwise stated, the same reference numerals in the drawings represent the same structure or operation.
[0015] It should be understood that the "system", "device", "unit" and / or "module" used herein is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the said words can be replaced by other expressions.
[0016] Unless the context clearly indicates an exception, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. A method or device may also include other steps or elements.
[0017] Flowcharts are used in this specification to illustrate the operations performed by the system according to the embodiments of this specification. It should be understood that the previous or subsequent operations do not necessarily need to be executed precisely in sequence. On the contrary, the steps can be processed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or more steps can be removed from these processes.
[0018] Figure 1 is a schematic diagram of the application scenario of the cluster management system shown in some embodiments of this specification.
[0019] As Figure 1 shown, the application scenario 100 of the cluster management system may include a cluster management platform 110, cluster nodes 120, a storage device 130, a network 140, and a terminal 150.
[0020] A cluster refers to a distributed system composed of multiple computer devices connected through a network. Clusters include, but are not limited to, load balancing clusters, high availability clusters, high performance computing clusters, and storage clusters. Each computer device in a cluster can be referred to as a cluster node, server / server node, or computer node. Cluster nodes can include management nodes and working nodes, etc. Among them, a management node refers to a cluster node used to manage and schedule working nodes, and a working node refers to a cluster node used to process tasks.
[0021] The cluster management platform 110 refers to a platform that can be used to manage and / or control a cluster (such as cluster 120). In some embodiments, the cluster management platform 110 is composed of one or more management nodes, which can be in the form of a single server or a server group.
[0022] The cluster management platform 110 can receive temporary traffic data that needs to be processed. The temporary traffic data includes access requests (such as network requests), etc. In some embodiments, the cluster management platform 110 can distribute the access requests to the cluster nodes 120. In some embodiments, the cluster management platform 110 can obtain data and / or information from the cluster nodes 120, the storage device 130, and the terminal 150 through the network 140. For example, the cluster management platform 110 can obtain the node activity, node load information, and / or node configuration information of one or more of the cluster nodes in the cluster nodes 120 through the network 140. For another example, the cluster management platform 110 can send warning information to the terminal 150 based on the traffic risk value through the network 140.
[0023] In some embodiments, the cluster management platform 110 can be used to control and / or manage one or more of the cluster nodes in the managed cluster nodes 120. For example, the cluster management platform 110 can determine a temporary node group based on the traffic risk value, the target number of cluster nodes, and the node characteristics of multiple cluster nodes. More relevant content can be found elsewhere in this specification (such as Figure 3 ).
[0024] There can be multiple cluster nodes 120, such as Figure 1 As shown, the cluster nodes 120 include cluster node 120-1, cluster node 120-2, cluster node 120-3,..., cluster node 120-n. In some embodiments, the cluster nodes 120 can interact with one or more components of the application scenario 100 (such as the cluster management platform 110, the storage device 130, and the terminal 150) through the network 140. For example, each cluster node in the cluster nodes 120 can report its connection status information with other cluster nodes to the cluster management platform 110.
[0025] The storage device 130 can store data and / or instructions. In some embodiments, the storage device 130 can store the data obtained from the cluster management platform 110, the cluster nodes 120, and / or the terminal 150. For example, the storage device 130 can store the data (such as the traffic risk value) obtained from the cluster management platform 110, etc. In some embodiments, the storage device 130 can store the data and / or instructions for executing the exemplary methods described in this specification. For example, the storage device 130 can store the instructions for the cluster management platform 110 to execute the methods shown in each flowchart. In some embodiments, the storage device 130 can include a mass storage device, a removable storage device, a volatile read-write memory, a read-only memory (ROM), etc., or any combination thereof. In some embodiments, the storage device 130 can be implemented on a cloud platform. In some embodiments, the storage device 130 can be a part of the cluster management platform 110 and the cluster nodes 120.
[0026] Network 140 may include any suitable network that facilitates the exchange of information and / or data for the application scenario 100 of the cluster management system. In some embodiments, one or more components of the application scenario 100 (e.g., the cluster management platform 110, the cluster nodes 120, the storage device 130, and the terminal 150) may transmit information and / or data to one or more other components of the application scenario 100 via the network 140. For example, the cluster management platform 110 may send temporary traffic data to one or more of the cluster nodes in the cluster nodes 120 via the network 140.
[0027] In some embodiments, the network 140 may be any one or more of a wired network or a wireless network. In some embodiments, the network may be of various topological structures such as point-to-point, shared, centralized, etc., or a combination of multiple topological structures.
[0028] The terminal 150 may include a mobile device 150-1, a tablet computer 150-2, a laptop computer 150-3, etc., or any combination thereof. In some embodiments, the terminal 150 may interact with other components in the application scenario 100 via the network 140. For example, the terminal 150 may receive information and / or instructions input by the user and send the received information and / or instructions to the cluster management platform 110 via the network 140. In some embodiments, the terminal 150 may send and / or present risk warning information to the user according to the traffic risk value of the temporary traffic data, and the risk warning information includes but is not limited to information such as voice, text, image, video, etc. More content about the temporary traffic data and the traffic risk value can be found elsewhere in this specification (e.g., Figure 3 )
[0029] The above description is for illustrative purposes only, and actual application scenarios may vary in various ways.
[0030] It should be noted that the application scenario 100 is provided only for illustrative purposes and is not intended to limit the scope of this specification. For those of ordinary skill in the art, various modifications or changes can be made according to the description of this specification. However, these changes and modifications will not deviate from the scope of this specification.
[0031] Figure 2 is a schematic diagram of the modules of the cluster management system shown in some embodiments of this specification.
[0032] As Figure 2 shown, the cluster management system 200 may include a first determination module 210, a second determination module 220, and a third determination module 230. In some embodiments, the first determination module 210, the second determination module 220, and the third determination module 230 may be deployed in the cluster management platform (such as the cluster management platform 110).
[0033] The first determination module 210 is configured to determine a traffic risk value based on the temporary traffic characteristics of the temporary traffic data, where the temporary traffic characteristics include the traffic volume and traffic frequency of the temporary traffic data.
[0034] In some embodiments, the temporary traffic characteristics further include a first risk value and a second risk value, and the traffic risk value is determined based on the first risk value and the second risk value.
[0035] In some embodiments, the first determination module 210 is further configured to: determine a first sampling parameter based on the first risk value; sample the temporary traffic data based on the first sampling parameter to determine sampled traffic data; determine status slice data based on the test node group and the sampled traffic data; determine the sampled attack degree corresponding to the sampled traffic data based on the status slice data; determine the second risk value based on the first sampling parameter and the sampled attack degree.
[0036] In some embodiments, the first risk value is determined based on a risk prediction model, and the risk prediction model is a machine learning model.
[0037] In some embodiments, the first determination module 210 is further configured to: determine the second risk value based on the first sampling parameter, the second sampling parameter, and the sampled attack degree.
[0038] In some embodiments, the first determination module 210 is further configured to: determine a first weight corresponding to the first risk value and a second weight corresponding to the second risk value, where the first weight and the second weight are respectively related to the first sampling parameter and the second sampling parameter; determine the traffic risk value based on the first risk value, the first weight, the second risk value, and the second weight.
[0039] In some embodiments, the first determination module 210 is further configured to: determine the second sampling parameter based on the test node group and the sampled traffic data; sample the node status data based on the second sampling parameter to determine status slice data.
[0040] The second determination module 220 is configured to determine the number of target cluster nodes based on the temporary traffic characteristics and the traffic risk value.
[0041] The third determination module 230 is configured to determine a temporary node group based on the traffic risk value, the number of target cluster nodes, and the node characteristics of multiple cluster nodes.
[0042] In some embodiments, the node characteristics at least include node activity, node load information, and node configuration information.
[0043] In some embodiments, the third determination module 230 is further configured to: for each cluster node, monitor the request queue information of the network requests of the cluster node, where the request queue information includes the number of network requests; based on the request queue information, determine the number of cluster nodes in the temporary node group.
[0044] In some embodiments, the determination module 230 is further configured to: control at least two of the multiple cluster nodes to send detection data packets to each other; based on the response data between the at least two cluster nodes, determine multiple valid cluster nodes and a valid node map corresponding to the multiple valid cluster nodes; based on the node load information of the valid cluster nodes, determine the node importance of the valid cluster nodes; based on the valid node map and the node importance, determine the node activity of the valid cluster nodes.
[0045] In some embodiments, the determination module 230 is further configured to: based on the first transmission frequency, control at least two cluster nodes to send the detection data packets to each other.
[0046] In some embodiments, the determination module 230 is further configured to: based on the second transmission frequency, control at least two valid cluster nodes to send the detection data packets to each other.
[0047] In some embodiments, the second transmission frequency is related to the importance of at least two valid cluster nodes.
[0048] In some embodiments, the determination module 230 is further configured to: based on the node characteristics of the multiple cluster nodes, determine multiple node groups from the multiple cluster nodes, where any two cluster nodes in the multiple node groups do not overlap; determine the risk resistance coefficient corresponding to each of the multiple node groups; based on the traffic risk value, the target number of cluster nodes, and the risk resistance coefficient corresponding to each of the multiple node groups, determine the temporary node group.
[0049] In some embodiments, the determination module 230 is further configured to: based on the node characteristics of the cluster nodes, determine the risk handling ability value of the cluster nodes; based on the risk handling ability values of the multiple cluster nodes, determine multiple node groups.
[0050] It should be noted that the above description of the cluster management system and its modules is only for convenience of description and does not limit this specification to the scope of the exemplified embodiments. It can be understood that for those skilled in the art, after understanding the principle of the system, they may, without departing from this principle, make any combination of the various modules, or form a subsystem and connect it with other modules. For example, the first determination module 210, the second determination module 220, and the third determination module 230 may be different modules, or a single module may implement the functions of two or more of the above-mentioned modules. Another example is that the various modules may share a storage module, or each module may have its own storage module. Such variations are all within the scope of protection of this specification.
[0051] Figure 3 is an exemplary flowchart of a cluster management method according to some embodiments of this specification.
[0052] In some embodiments, process 300 may be executed by a cluster management platform. As Figure 3 shown, process 300 includes the following steps.
[0053] Step 310, determining a traffic risk value based on the temporary traffic characteristics of the temporary traffic data.
[0054] Temporary traffic data refers to the traffic data that the cluster management platform needs to analyze and / or process, including but not limited to various types of traffic data such as text, audio and video, images, or instructions.
[0055] The temporary traffic data may include external traffic data and internal traffic data of the cluster.
[0056] The external traffic data includes communication data between the outside and the cluster. For example, it may be data of network interactions (such as network requests) between the outside (such as network users, clients, or third-party services) and the cluster (such as cluster nodes in the cluster).
[0057] The internal traffic data includes communication data between cluster nodes. For example, it may be data transmitted during the interaction between multiple cluster nodes (such as business service access between cluster nodes).
[0058] In some embodiments, the cluster management platform is configured to receive the temporary traffic data and perform preprocessing to achieve monitoring of the external traffic data and management and coordination of the internal traffic data. Exemplarily, the preprocessing includes but is not limited to data verification, filtering, data feature extraction, etc.
[0059] The temporary traffic characteristics are used to reflect various attributes of the temporary traffic data. The temporary traffic characteristics include the traffic volume and traffic frequency of the temporary traffic data.
[0060] The traffic volume of the temporary traffic data reflects the number of bytes of the temporary traffic data within a preset time period (such as 10 minutes, 1 hour). The traffic frequency reflects the frequency of transmission of the temporary traffic data within the preset time period.
[0061] In some embodiments, the cluster management platform can monitor and parse the temporary traffic data to determine the temporary traffic characteristics corresponding to the temporary traffic. In some embodiments, the cluster management platform can perform traffic statistics on the temporary traffic data to determine the traffic volume and traffic frequency. For example, respectively count the number of bytes of the temporary traffic data and the number of network requests within the preset time period as the traffic volume and traffic frequency.
[0062] The temporary traffic characteristics can be various characteristics or metrics set according to actual requirements (such as load balancing requirements, concurrency requirements, attack defense requirements, etc.). In some embodiments, the temporary traffic characteristics also include traffic basic characteristics. For example, the data source (such as the IP address of the sender), the destination (such as the IP address of the receiver), the access port, the protocol type (such as HTTP, TCP, etc.). The temporary traffic characteristics also include traffic timing characteristics. For example, the periodicity of the traffic data, the time point distribution of the traffic, etc.
[0063] In some embodiments, the cluster management platform can perform packet analysis through packet capture and other means to determine the basic characteristics and timing characteristics corresponding to each temporary traffic data.
[0064] In some embodiments, the temporary traffic characteristics also include a first risk value. The first risk value is used to reflect the probability of an abnormal situation occurring in the cluster node after the temporary traffic data passes through the actual processing of the cluster (such as the cluster node).
[0065] In some embodiments, the cluster management platform can monitor the resource usage of the cluster node to determine the first risk value. For example, the first risk value can be determined according to the resource usage information such as the current CPU occupancy rate and memory usage rate of the cluster node. Exemplarily, when the CPU occupancy rate of one or more cluster nodes is greater than the preset CPU usage threshold, it indicates that the first risk value is high.
[0066] In some embodiments, the cluster management platform can determine the first risk value based on a risk prediction model.
[0067] The risk prediction model refers to a model used to predict the first risk value. In some embodiments, the risk prediction model is a trained machine learning model. For example, a neural network model (Neural Network, NN), etc.
[0068] In some embodiments, the input of the risk prediction model 710 includes the temporary traffic characteristics 701, and the output includes the first risk value 702.
[0069] In some embodiments, the initial model can be iteratively trained with multiple training samples to obtain a trained risk prediction model 710. Each training sample can include sample temporary traffic characteristics corresponding to the sample temporary traffic data within a historical time period (such as the past month, half a year, etc.). The training label can be determined based on the sample cluster anomaly information after the actual processing of the sample temporary traffic data by the cluster, and it can be manually labeled or labeled by other means. It should be noted that the training label can be a value obtained by normalizing various types of sample cluster anomaly information. For example, it can be a numerical value within the interval [0, 1].
[0070] The sample cluster anomaly information can be determined based on the anomalies of one or more sample cluster nodes.
[0071] The sample cluster anomaly information can be various types of indicators preset according to actual needs. For example, the resource usage anomaly information of the sample cluster node, security risk information, etc. Exemplarily, the resource usage anomaly information includes CPU usage anomaly (such as the usage rate is greater than the preset CPU usage threshold), memory anomaly (such as the occupancy rate is greater than the preset memory occupancy threshold, memory leak), port anomaly (such as the port occupancy rate is greater than the threshold), etc.
[0072] The security risk information includes abnormal login attempts (such as multiple failed login attempts), permission changes (such as unauthorized permission changes), etc.
[0073] It should be noted that the cluster management platform can record the cluster anomaly information in the form of logs, texts, etc. for analysis and processing.
[0074] During training, based on the difference between the output of the initial model and the training label, the value of the loss function can be determined. The parameters of the initial model can be iteratively updated based on the value of the loss function until the training termination condition is met (such as the loss function converges, a specific number of iterations are performed, etc.). The updated initial model can be used as the trained risk prediction model.
[0075] In some embodiments of this specification, through the multi-risk prediction model, the correlation between the characteristics of the temporary traffic data and the cluster anomaly information can be learned.
[0076] The traffic risk value is used to characterize the degree of potential risk brought by the temporary traffic data to one or more components (such as the cluster management platform 110, cluster node 120) in the cluster management system. The potential risks include but are not limited to overloading, malicious attacks, etc. The larger the traffic risk value, the greater the potential risk. The cluster management platform can evaluate the traffic risk value corresponding to the temporary traffic data according to the temporary traffic characteristics.
[0077] In some embodiments, the cluster management platform can detect the temporary traffic characteristics of the temporary traffic data by deploying a network threat detection engine (such as Suricata, etc.) to determine the traffic risk value. Exemplarily, the cluster management platform can determine the traffic risk value corresponding to the temporary traffic data according to the intrusion information or records detected by the network threat detection engine through a preset threat risk relationship table and issue an alarm.
[0078] In some embodiments, the cluster management platform deploys a vector database (such as Milvus), which includes multiple reference records constructed based on historical traffic data. The reference records include the reference traffic characteristics corresponding to the reference historical traffic data (for example, the reference first risk value) and their corresponding reference traffic risk values. The cluster management platform can perform vector matching processing in the vector database based on the temporary traffic characteristics and use the reference traffic risk value corresponding to the reference record with the maximum similarity between the temporary traffic characteristics and the reference traffic characteristics as the traffic risk value corresponding to the temporary traffic data. Among them, the maximum similarity can be the minimum vector distance.
[0079] In some embodiments, the temporary traffic characteristics further include a second risk value, and the cluster management platform can also determine the traffic risk value based on the first risk value and the second risk value. For more information about the second risk value and the determination of the traffic risk value based on the first risk value and the second risk value, see Figure 5 and its description.
[0080] Step 320: Determine the target cluster node number based on the temporary traffic characteristics and the traffic risk value.
[0081] The target cluster node number refers to the number of cluster nodes required to process the temporary traffic. In some embodiments, the cluster management platform can preset a node number reference table in advance. The node number reference table includes the temporary traffic characteristics, the traffic risk value, and the corresponding reference cluster node number. Among them, the node number reference table can be obtained according to historical data. The cluster management platform can match the corresponding reference cluster node number in the node number reference table based on the temporary traffic characteristics and the traffic risk value as the target cluster node number.
[0082] Step 330: Determine the temporary node group based on the traffic risk value, the target cluster node number, and the node characteristics of multiple cluster nodes.
[0083] In some embodiments, the node characteristics include node activity, node load information, and node configuration information.
[0084] The node activity is used to reflect the frequency of data processing of the cluster node (such as network request response). For example, the more network requests a cluster node responds to and the more frequently it processes data, the higher its corresponding node activity.
[0085] Node load information is used to reflect the resource usage pressure of cluster nodes. The node load information includes, but is not limited to, CPU load (such as CPU usage rate), memory load (presence usage rate), disk I / O load (such as the proportion of disk read / write time), network load (such as bandwidth occupancy rate), etc.
[0086] Node configuration information refers to the hardware configuration information of cluster nodes. For example, the number of CPUs, memory size, disk capacity, network bandwidth, etc.
[0087] A temporary node group refers to a node group composed of cluster nodes used to process temporary traffic data. For example, the cluster management platform can select a number of target cluster nodes from multiple cluster nodes to form a temporary node group. The selection methods include, but are not limited to, random selection, preferentially selecting cluster nodes with a smaller processing load according to the processing load of cluster nodes, etc.
[0088] In some embodiments, the cluster management platform can determine multiple node groups from multiple cluster nodes based on the node characteristics of the multiple cluster nodes, and determine the risk resistance coefficient corresponding to each of the node groups. Furthermore, based on the traffic risk value, the number of target cluster nodes, and the risk resistance coefficient corresponding to each node group, the temporary node group is determined.
[0089] In some embodiments, the cluster management platform can screen a preset number of cluster nodes with node relevance from multiple cluster nodes according to the node characteristics of the multiple cluster nodes to form multiple node groups. Among them, the node relevance can be determined according to actual needs. For example, multiple cluster nodes with the same geographical area, similar business or task types (such as storage processing type, business processing type) can be used as a node group.
[0090] In some embodiments, any two cluster nodes in the multiple node groups do not overlap, that is, any one cluster node exists in and only exists in one node group.
[0091] In some embodiments, the cluster management platform can determine the risk handling ability value of cluster nodes based on the node characteristics of the cluster nodes; and determine multiple node groups based on the risk handling ability values of the multiple cluster nodes.
[0092] In some embodiments, the cluster management platform can preprocess the node characteristics of each cluster node. The preprocessing can include, but is not limited to, dimension normalization processing through various algorithms such as forwardization, reverseization, intervalization, Min - Max normalization, etc., to obtain dimension - normalized data, so that the data structure, unit, number system, or format, etc. of the node characteristic values corresponding to the node characteristics are unified, so as to be able to perform subsequent unified operations (such as arithmetic operations, etc.).
[0093] In some embodiments, the cluster management platform can perform a weighted summation based on the pre-processed node features corresponding to each cluster node to obtain a risk handling capability value. The weights can be preset based on experience. The risk handling capability value of a cluster node can be used to reflect the robustness or reliability of the cluster node in data processing.
[0094] In some embodiments, the cluster management platform can perform gradient partitioning based on the risk handling capability value of each cluster node, thereby obtaining a gradient distribution of the risk handling capability values of multiple cluster nodes. For example, the gradient partitioning can be performed by sorting the risk handling capability values (e.g., in descending order). In some embodiments, the cluster management platform can select a preset number of grouped cluster nodes (e.g., 10) from the multiple cluster nodes after the gradient partitioning to generate multiple node groups. It should be noted that the preset number of grouped cluster nodes can be the same as or different from the target number of cluster nodes.
[0095] The risk resistance coefficient refers to the stability of the cluster nodes in the node group after processing temporary traffic data. A higher risk resistance coefficient indicates a higher stability of the cluster nodes after processing temporary traffic data.
[0096] In some embodiments, the risk resistance coefficient corresponding to the node group may be determined according to an average of the risk handling capability values of the plurality of cluster nodes in the node group.
[0097] In some embodiments, the cluster management system may select a node group from multiple node groups as a target node group based on the traffic risk value, wherein the larger the traffic risk value, the greater the risk resistance coefficient of the selected target node group.
[0098] In some embodiments, the cluster management system may also adjust the target node group based on the target number of cluster nodes so that the number of cluster nodes of the target node group matches the target number of cluster nodes, thereby obtaining a temporary node group with the target number of cluster nodes. Exemplarily, in response to the number of cluster nodes of the target node group being less than the target cluster node, one or more cluster nodes with larger risk handling capability values may be selected from the node group with a similar risk resistance coefficient to the target node group and added to the target node group. In response to the number of cluster nodes of the target node group being greater than the target cluster node, one or more cluster nodes with smaller risk handling capability values in the target node group may be removed from the target node group.
[0099] In some embodiments, for each cluster node in the temporary node group, the cluster management platform may monitor request queue information of network requests of the cluster node; and determine the number of cluster nodes in the temporary node group based on the request queue information.
[0100] The request queue information can be used to reflect the status of network requests that need to be processed in the cluster nodes. The network requests that need to be processed include those that are being processed and those that are waiting to be processed.
[0101] In some embodiments, the request queue information includes the number of network requests. A larger number of network requests indicates that the concurrency capacity of the cluster node is insufficient or the load is too heavy.
[0102] In some embodiments, the cluster management platform can expand or shrink the temporary node group based on the request queue information corresponding to each cluster node. In response to the number of network requests being greater than a preset expansion threshold, the cluster management platform can add one or more cluster nodes to the temporary node group to achieve expansion of the temporary node group. In response to the number of network requests that need to be processed being less than a preset shrinkage threshold, the cluster management platform can remove one or more cluster nodes (such as idle cluster nodes) in the temporary node group to achieve shrinkage of the temporary node group. The preset expansion threshold and the preset shrinkage threshold can be pre-set thresholds.
[0103] In some embodiments of the present specification, by monitoring the request queue information of the cluster nodes in the temporary group, the temporary node group can be automatically expanded or reduced, thereby achieving load balancing of multiple cluster nodes.
[0104] In some embodiments of this specification, considering the temporary traffic characteristics of temporary traffic data can coordinate and manage the software and hardware resources of cluster nodes in the cluster, improving the efficiency or flexibility of processing temporary traffic data. Furthermore, considering the traffic risk value of temporary traffic data enables the cluster to respond to potential risks, thereby improving the high availability and robustness of the cluster.
[0105] Figure 4 This is an exemplary flow chart of a method for determining node activity of valid cluster nodes according to some embodiments of this specification.
[0106] In some embodiments, process 400 may be performed by a cluster management platform. Figure 4 As shown, process 400 includes the following steps.
[0107] Step 410: Control at least two cluster nodes among the plurality of cluster nodes to send detection data packets to each other.
[0108] A detection data packet is data used to determine whether any two cluster nodes can communicate, for example, preset lightweight data (such as a heartbeat signal).
[0109] In some embodiments, the cluster management platform may control at least two cluster nodes to send detection data packets to each other based on the first sending frequency.
[0110] The first transmission frequency is used to characterize the frequency at which any two cluster nodes send detection data packets to each other. It can be a value preset based on experience. For example, the first transmission frequency is to send a detection data packet every 20 seconds.
[0111] In some embodiments, the first transmission frequency can be determined according to the node characteristics of the cluster nodes. For example, when the node activity is greater, the first transmission frequency can be greater.
[0112] Step 420: Based on the response data between at least two cluster nodes, determine multiple valid cluster nodes and the valid node graph corresponding to the multiple valid cluster nodes.
[0113] The response data refers to the data received after the cluster nodes send detection data. When any two cluster nodes can send detection data packets to each other and receive response data, it means that any one of the any two cluster nodes is a valid cluster node, and the valid cluster node is also called an active node.
[0114] The valid node graph is a knowledge graph constructed according to the communication connection information of multiple valid cluster nodes in the cluster, and it can be used to reflect the active information and its change information of the multiple valid cluster nodes.
[0115] In some embodiments, the valid node graph includes multiple graph nodes and multiple graph edges, and the graph nodes and graph edges have attributes. Each graph node corresponds to a valid cluster node, and the graph edge connects two graph nodes with a communication connection, indicating that the two valid cluster nodes can communicate with each other.
[0116] In some embodiments, the graph node attributes include at least one of the node activity, node load information, node configuration information, etc. of the valid cluster node corresponding to the graph node. In some embodiments, the graph node attributes can be determined based on the node characteristics of the valid cluster node corresponding to the graph node.
[0117] In some embodiments, the graph node attributes further include node importance. For more details, see step 430.
[0118] It should be noted that for a certain valid cluster node, the corresponding graph node attributes can be dynamically updated according to the actual situation of the valid cluster node. For example, the cluster management platform can update the node load information in the corresponding graph node attributes according to the load information of the valid cluster node.
[0119] In some embodiments, the graph edge attributes include the connection status, which is used to indicate the communication connection status or disconnection status between the two valid cluster nodes corresponding to the graph edge.
[0120] In some embodiments, the cluster management platform may construct and / or update a valid node map based on the communication connection status of any two cluster nodes.
[0121] Step 430 : Determine the node importance of each valid cluster node based on the node load information of each valid cluster node.
[0122] The node importance reflects the importance of an active cluster node to temporary traffic data processing. In some embodiments, the node importance of an active cluster node can be determined based on the load information of the active cluster node. For example, the node importance can be determined based on the ratio of the load of the active cluster node to the average load of the cluster. The load can be determined based on the traffic volume and / or traffic frequency of the temporary traffic data to be processed by the active cluster node.
[0123] In some embodiments, the cluster management platform may control at least two valid cluster nodes to send detection data packets to each other based on a second sending frequency, where the second sending frequency is related to the importance of each of the valid cluster nodes.
[0124] The second transmission frequency is used to indicate how frequently any two valid cluster nodes transmit detection data to each other. In some embodiments, the second transmission frequency is different from the first transmission frequency. The second transmission frequency can be set to be higher than the first transmission frequency to ensure the timeliness of connectivity detection between valid cluster nodes.
[0125] In some embodiments, edge attributes also include a second transmission frequency. For any graph edge, the cluster management platform can adjust the second transmission frequency corresponding to the graph edge based on the importance of the two valid cluster nodes connected to the graph edge. This controls the two valid cluster nodes to send detection data packets to each other based on the adjusted second transmission frequency. The greater the average importance of the two valid cluster nodes, the higher the corresponding second transmission frequency.
[0126] In some embodiments of the present specification, the effective node map can enable the cluster management platform to manage the communication connection status of the effective cluster nodes in a more real-time and efficient manner. At the same time, the second sending frequency is set considering the importance of different effective cluster nodes, which can better track and ensure the communication connection status between more important effective cluster nodes, thereby ensuring the processing of temporary traffic data.
[0127] Step 440 : Determine the node activity of each valid cluster node based on the valid node map and the node importance of each valid cluster node.
[0128] The node activity of an effective cluster node is used to reflect the relative activity level of the effective cluster node in the local network where the effective cluster node is located.
[0129] In some embodiments, the local network of an active cluster node can be the network consisting of all active cluster nodes in the active node graph that have a communication connection with the active cluster node. For example, for an active cluster node A, its corresponding local network is Network A, consisting of all active cluster nodes that have a communication connection with the active cluster node. For another active cluster node B, its corresponding local network is Network B, consisting of all active cluster nodes that have a communication connection with the active cluster node.
[0130] The node activity can be determined based on the following formula (1):
[0131] In formula (1), r i represents the activity of the i-th valid cluster node, σ represents the standard deviation of the node importance of all valid cluster nodes in the local network where the valid cluster node is located, a represents the average node importance of all valid cluster nodes, e represents the number of edges existing in the valid cluster node, and E represents the total number of graph edges in the valid node graph.
[0132] in, The value of can be used to reflect the concentration of temporary traffic data processing of the effective cluster node in the local network. The larger the value, the more concentrated it is, and the more active the corresponding node is.
[0133] Some embodiments of this specification introduce node activity to consider the relative temporary traffic processing concentration of different valid cluster nodes, which can better evaluate the load balancing of different valid cluster nodes, thereby making the scheduling of the cluster management platform more accurate.
[0134] In some embodiments, the local network where the valid cluster node is located can be a network composed of all valid cluster nodes in the valid node map. Considering that in actual application scenarios, all cluster nodes in the valid node map process actual tasks as a whole (such as load balancing collaboration, etc.), the network composed of all valid cluster nodes in the valid node map can be used as the local network corresponding to each valid cluster node in the valid node map, wherein the node activity of each valid cluster node can be calculated based on the above formula (1).
[0135] In some embodiments of this specification, all valid cluster nodes in the valid node map are used as the local network where any valid cluster node is located, which can reduce the amount of calculation for constructing the local network corresponding to each valid cluster node.
[0136] Figure 5It is an exemplary flowchart of a method for determining a traffic risk value according to some embodiments of this specification.
[0137] Step 510: Based on the first risk value, determine the first sampling parameter.
[0138] The first sampling parameter refers to the parameter for the cluster management platform to extract temporary traffic data. In some embodiments, the first sampling parameter includes the first sampling frequency.
[0139] The first sampling frequency is used to reflect the frequency of the cluster management platform to extract or collect temporary traffic data. For example, temporary traffic data can be collected based on a preset time period or time interval (such as every 10 seconds, every minute). The shorter the preset time interval, the greater the first sampling frequency (i.e., the more frequent).
[0140] In some embodiments, the cluster management platform can adjust the first sampling frequency according to the first risk value. When the first risk value is larger, it indicates that the potential risk of the cluster is greater, and the first sampling frequency can be set larger to enhance the sampling intensity, thereby improving the detection strength.
[0141] In some embodiments of this specification, considering the first risk value to adjust the first sampling frequency can improve the collection efficiency while avoiding waste of resources caused by excessive frequency.
[0142] Step 520: Based on the first sampling parameter, sample the temporary traffic data to determine the sampled traffic data.
[0143] The sampled traffic data refers to the temporary traffic data collected within the sampling time period based on the first sampling parameter. The sampling time period can be a preset time period up to the current moment (such as the past 1 minute, 10 minutes).
[0144] The sampled traffic data can be the temporary traffic data in a sequence based on time slices within the sampling time period. The time slice can be a preset time interval or time period (such as 5 seconds). For example, it can be the temporary traffic data every 5 seconds within the past minute. Every 5 seconds is a time slice.
[0145] Step 530: Based on the test node group and the sampled traffic data, determine the status slice data.
[0146] The test node group refers to a node group composed of one or more cluster nodes used to perform risk tests on the sampled traffic data. The cluster nodes in the test node group can be called test cluster nodes.
[0147] Risk tests can be used to simulate, reproduce, track, and analyze the risks brought by the sampled traffic data to the cluster.
[0148] In some embodiments, risk testing can be used for external risk analysis. For example, it can include, but is not limited to, attack hunting (such as collection of attack types like DDoS attacks, Trojan implantation, etc.), attack behavior (such as sabotage behavior, data theft, permissions) analysis, and / or attack chain tracing (such as cross-node attack chains), etc.
[0149] It should be noted that risk testing can be determined according to actual needs. For example, risk testing can also be used for internal risk analysis. For example, simulating potential load balancing caused by data transmission between cluster nodes.
[0150] In some embodiments, the test cluster nodes are set to be isolated from the non-test cluster node environment. For example, the cluster management platform can isolate the test node group from the non-test cluster nodes through network isolation (such as independent subnets), network access permissions (restricting network communication), setting up firewalls, etc.
[0151] In some embodiments of this specification, by setting up the test node group, it is possible to analyze the risks brought by sampling temporary traffic data while avoiding affecting the data processing (such as business processing) of non-test cluster nodes and ensuring the normal operation of the entire cluster.
[0152] State slice data refers to the node state data corresponding to a certain time slice within the sampling time period of the test cluster nodes. For example, for a test cluster node, it can include multiple state slice data, and each state slice data corresponds to a time slice.
[0153] The node state data is used to reflect the node operating state after the test cluster nodes process the sampled traffic data. The node operating state can be determined based on the values of multiple preset evaluation metrics.
[0154] Exemplarily, the evaluation metrics include one or a combination of multiple operating states such as the heartbeat signal loss rate, heartbeat signal delay rate, access failure rate, and abnormal source connection ratio, etc.
[0155] The heartbeat signal loss rate refers to the proportion of heartbeat signals with sending or receiving failures. The heartbeat signal delay rate refers to the proportion of delayed heartbeat signals. The access failure rate refers to the proportion of access requests with timeouts or failures. The abnormal source connection ratio refers to the proportion of access connections from abnormal access sources (such as known threat source IPs, unexpected business geographical regions / locations, etc.).
[0156] When the value of an evaluation metric is greater than its corresponding preset ratio threshold, this evaluation metric can be called an abnormal metric. The larger the value of the abnormal metric, the worse the node operating state. The state slice data with abnormal metrics can be called abnormal state slice data. The test cluster nodes with abnormal state slice data can be called abnormal test cluster nodes.
[0157] In some embodiments, the cluster management platform may determine a second sampling parameter based on a test node group and sampled traffic data; and sample the node status data based on the second sampling parameter to determine status slice data.
[0158] The second sampling parameter reflects the frequency at which the cluster management platform extracts or collects node status data. In some embodiments, the second sampling parameter includes a second sampling frequency.
[0159] In some embodiments, the cluster management platform may determine the second sampling frequency according to the configuration information of the test nodes (such as CPU resources, memory resources), the temporary traffic characteristics of the sampled traffic data (such as traffic volume), and the node status data.
[0160] For example, the second sampling frequency may be positively correlated with the CPU resource performance (such as computing power, number of cores). Exemplarily, the better the CPU resource performance, the larger the second sampling frequency can be, so as to cover more status slice data, thereby making the subsequent analysis results more accurate; when the CPU resource performance is poor, the second sampling frequency can be set smaller to reduce the computing pressure on the test cluster nodes.
[0161] Again, for example, the second sampling frequency may be negatively correlated with the sampled traffic data. For example, when the sampled traffic is less or sparser, the corresponding node status data may be sparser, and a larger second sampling frequency can be set to be able to collect sufficient status slice data.
[0162] Step 540, determine the sampled attack degree corresponding to the sampled traffic data based on the status slice data.
[0163] The sampled attack degree is used to reflect the degree of the test cluster node being in a negative operating state after processing the sampled traffic data. The negative operating state refers to an unexpected node operating state. Exemplarily, for the evaluation index of the heartbeat signal loss rate, if the heartbeat signal loss rate is greater than the preset heartbeat signal loss rate of 10%, it indicates that the test cluster node is in a negative operating state.
[0164] In some embodiments, the cluster management platform may determine the sampled attack degree corresponding to the sampled traffic data based on one or more node operating states in the status slice data. For example, the sampled attack degree may be determined according to a combination of one or more of the heartbeat signal loss rate, heartbeat signal delay rate, access failure rate, and abnormal source connection ratio of the test cluster nodes. The attack degree may be represented in various preset forms, such as mild, moderate, and severe, etc.
[0165] In some embodiments, the cluster management platform may determine the sampled attack degree based on the abnormal test cluster nodes and the abnormal state slice data of the abnormal test cluster nodes.
[0166] For example, the larger the number of abnormal test cluster nodes, the greater the sampled attack degree; the more the number of abnormal state slice data, the greater the attack degree; the larger the value of the abnormal index in the abnormal state slice data, the greater the sampled attack degree.
[0167] In some embodiments, the cluster management platform may set an attack index table, which includes the number of abnormal test cluster nodes, the number of abnormal state slice data, the values of each abnormal index, and the corresponding attack degree. The cluster management platform may, by looking up the table, match the target attack degree corresponding to the sampled temporary traffic as the sampled attack degree.
[0168] In some embodiments of this specification, considering the complexity of the sampled temporary traffic data (such as timeliness, concealment, etc.), by considering the number of abnormal test cluster nodes, the influence range of the attack degree can be evaluated; by considering the number of abnormal state slice data, the frequency of the attack degree can be evaluated; by considering the value of the abnormal index, the severity of the attack degree can be evaluated.
[0169] Step 550, determine a second risk value based on the first sampling parameter and the sampled attack degree.
[0170] The second risk value can be used to evaluate the intensity of the potential attack risk of the temporary traffic data to the cluster nodes. The higher the second risk value, the greater the potential attack risk. For example, the attack intensity of the actual temporary traffic data can be reflected by the sampled traffic data corresponding to the temporary traffic data.
[0171] In some embodiments, the second risk value is positively correlated with the first sampling parameter and the sampled attack degree. For example, when the first sampling frequency is higher and the sampled attack degree is higher, the second risk value is greater. The lower the means of the first sampling parameter and the second sampling parameter, it indicates that the sampling coverage is not high and there may be a certain contingency. At this time, the higher the sampled attack degree, the higher the second risk value.
[0172] In some embodiments, the cluster management platform may also determine the second risk value based on the first sampling parameter, the second sampling parameter, and the sampled attack degree. For example, the second risk value is positively correlated with the first sampling frequency, the second sampling frequency, and the sampled attack degree.
[0173] In some embodiments of this specification, by comprehensively considering the first sampling parameter and the second sampling parameter, the sampling coverage of the temporary traffic data and the sampling coverage of the state slice data can be fully considered, making the second risk value more accurate.
[0174] Step 560: Determine the traffic risk value based on the first risk value and the second risk value.
[0175] In some embodiments, the cluster management platform may determine a first weight corresponding to the first risk value and a second weight corresponding to the second risk value; and determine the traffic risk value based on the first risk value, the first weight, the second risk value, and the second weight.
[0176] The first weight and the second weight may be preset values. For example, both the first weight and the second weight are 0.5, or other values set according to experience (such as 0.4 and 0.6 respectively).
[0177] In some embodiments, the first weight and the second weight are respectively related to the first sampling parameter and the second sampling parameter (such as positively correlated). For example, when the first sampling parameter (such as the first sampling frequency) increases and the second sampling parameter remains unchanged, it indicates that the sampling coverage rate of the temporary traffic data increases. The first weight can be increased while the second weight is decreased.
[0178] In some embodiments of this specification, considering the relationship between the first sampling parameter and the second sampling parameter to adjust the first weight and the second weight can balance the resource consumption of the sampling coverage rate of the temporary traffic data and the sampling coverage rate of the status slice data while ensuring the accuracy of the second risk value.
[0179] It should be noted that the above description of the process is only for illustration and example, and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the process under the guidance of this specification. However, these modifications and changes are still within the scope of this specification.
[0180] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only an example and does not constitute a limitation to this specification. Although not explicitly stated here, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are proposed in this specification, so such modifications, improvements, and corrections still belong to the spirit and scope of the exemplary embodiments of this specification.
[0181] Meanwhile, this specification uses specific terms to describe the embodiments of this specification. For example, "an embodiment", "one embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that the "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more at different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.
[0182] In addition, unless clearly stated in the claims, the order of the processing elements and sequences, the use of numerical letters, or the use of other names in this specification are not used to limit the order of the processes and methods in this specification. Although some currently considered useful embodiments of the invention are discussed through various examples in the above disclosure, it should be understood that such details only serve the purpose of illustration. The appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that conform to the essence and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only through software solutions, such as installing the described system on existing servers or mobile devices.
[0183] Similarly, it should be noted that, in order to simplify the expression of the disclosure in this specification and thus help the understanding of one or more embodiments of the invention, in the previous description of the embodiments of this specification, sometimes multiple features are merged into one embodiment, drawing, or description thereof. However, this disclosure method does not mean that the features required by the object of this specification are more than those mentioned in the claims. In fact, the features of the embodiments are fewer than all the features of the individual embodiments disclosed above.
[0184] In some embodiments, numbers are used to describe the components and the quantity of attributes. It should be understood that such numbers used for the description of embodiments are, in some examples, modified by the modifiers "about", "approximate", or "substantially". Unless otherwise stated, "about", "approximate", or "substantially" indicate that the said numbers allow a variation of ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, and these approximate values can change according to the characteristics required by individual embodiments. In some embodiments, the numerical parameters should consider the specified significant digits and adopt the method of retaining the general number of digits. Although the numerical ranges and parameters used in some embodiments of this specification to confirm the breadth of their scope are approximate values, in specific embodiments, such numerical settings are made as precise as possible within the feasible range.
[0185] For each patent, patent application, patent application publication, and other materials cited in this specification, such as articles, books, specifications, publications, documents, etc., their entire contents are hereby incorporated by reference into this specification. This excludes application history files that are inconsistent with or conflict with the content of this specification, as well as files that limit the broadest scope of the claims of this specification (currently or subsequently appended to this specification). It should be noted that if there are any inconsistencies or conflicts between the descriptions, definitions, and / or uses of terms in the supplementary materials of this specification and the content described in this specification, the descriptions, definitions, and / or uses of terms in this specification shall prevail.
[0186] Finally, it should be understood that the embodiments described in this specification are only used to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly introduced and described in this specification.
Claims
1. A cluster management method, which is executed by a cluster management platform that controls multiple cluster nodes. The method includes: Determining a traffic risk value based on the temporary traffic characteristics of temporary traffic data, where the temporary traffic characteristics include the traffic volume and traffic frequency of the temporary traffic data; Determining the target number of cluster nodes based on the temporary traffic characteristics and the traffic risk value; Determining a temporary node group based on the traffic risk value, the target number of cluster nodes, and the node characteristics of the multiple cluster nodes, where the node characteristics include node activity, node load information, and node configuration information.
2. The method according to claim 1, wherein The method further includes: Controlling at least two of the multiple cluster nodes to send detection data packets to each other; Determining multiple effective cluster nodes and the corresponding effective node map based on the response data between at least two of the cluster nodes; Determining the node importance of the effective cluster nodes based on the node load information of the effective cluster nodes; Determining the node activity of the effective cluster nodes based on the effective node map and the node importance.
3. The method according to claim 1, characterized in that, The determining of the temporary node group includes: Determining multiple node groups from the multiple cluster nodes based on the node characteristics of the multiple cluster nodes, where any two of the multiple node groups do not overlap; Determining the anti-risk coefficient corresponding to each of the multiple node groups; Determining the temporary node group based on the traffic risk value, the target number of cluster nodes, and the anti-risk coefficient corresponding to each of the multiple node groups.
4. The method according to claim 3, wherein Determining multiple node groups from the multiple cluster nodes based on the node characteristics of the multiple cluster nodes includes: Determining the risk handling ability value of the cluster nodes based on the node characteristics of the cluster nodes; Determining the multiple node groups based on the risk handling ability values of the multiple cluster nodes.
5. The method according to claim 1, wherein The temporary traffic characteristics further include a first risk value and a second risk value, and the traffic risk value is determined based on the first risk value and the second risk value. The determination method of the second risk value includes: Determining a first sampling parameter based on the first risk value; Sampling the temporary traffic data based on the first sampling parameter to determine the sampled traffic data; Determining the state slice data based on the test node group and the sampled traffic data; Determining the sampled attack degree corresponding to the sampled traffic data based on the state slice data; Determining the second risk value based on the first sampling parameter and the sampled attack degree.
6. A cluster management system, characterized in that, The system includes a cluster management platform and multiple cluster nodes. The cluster management platform is configured to control multiple cluster nodes. The cluster management platform includes a first determination module, a second determination module, and a third determination module, where The first determination module is configured to determine a traffic risk value based on the temporary traffic characteristics of temporary traffic data, where the temporary traffic characteristics include the traffic volume and traffic frequency of the temporary traffic data; The second determination module is configured to determine the target number of cluster nodes based on the temporary traffic characteristics and the traffic risk value; The third determination module is configured to determine a temporary node group based on the traffic risk value, the number of target cluster nodes, and the node characteristics of the multiple cluster nodes, where the node characteristics at least include node activity, node load information, and node configuration information.
7. The system according to claim 6, wherein The third determination module is further configured to: Control at least two of the multiple cluster nodes to send detection data packets to each other; Based on the response data between at least two of the cluster nodes, determine multiple valid cluster nodes and the corresponding valid node map of the multiple valid cluster nodes; Based on the node load information of the valid cluster nodes, determine the node importance of the valid cluster nodes; Based on the valid node map and the node importance, determine the node activity of the valid cluster nodes.
8. The system according to claim 6, wherein The third determination module is further configured to: Based on the node characteristics of the multiple cluster nodes, determine multiple node groups from the multiple cluster nodes, and any two of the multiple node groups do not overlap; Determine the risk resistance coefficient corresponding to each of the multiple node groups; Based on the traffic risk value, the number of target cluster nodes, and the risk resistance coefficient corresponding to each of the multiple node groups, determine the temporary node group.
9. A cluster management device, characterized in that, Comprising at least one storage medium and at least one processor; The at least one storage medium is used to store computer instructions; The at least one processor is used to execute the computer instructions to implement the cluster management method according to any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions, and when the computer instructions are executed by the processor, the cluster management method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Data processing method and system
CN105991333A
Load balancing method and system for service cluster
CN111930523A
Distributed service cluster runtime parameter adaptive processing method, device and system
CN113032233A
Distributed service cluster load adaptive processing method, device and system
CN113055479A
Method and device for predicting running state of storage cluster
CN115686381A