A server management method and electronic device

By creating virtual control instances in the server cluster, dynamically scheduling resources and monitoring node operation information, the rigid resource allocation and high expansion costs in traditional server management are solved, achieving efficient and flexible resource management and load optimization.

CN120892293BActive Publication Date: 2026-01-23INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511404878.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-01-23
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Traditional server management solutions rely on physical controllers, resulting in rigid resource allocation, high expansion costs, inability to dynamically adjust according to load, and complex operation and maintenance.

Method used

By creating virtual control instances in a server cluster, resources can be dynamically scheduled and node operation information can be monitored, enabling flexible resource management and load optimization, thus replacing traditional physical controllers.

Benefits of technology

It improves resource utilization, reduces expansion costs and operational complexity, and enables efficient and flexible server cluster management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892293B_ABST
    Figure CN120892293B_ABST
Patent Text Reader

Abstract

The application discloses a kind of server management method and electronic equipment, it is related to server management technical field, including: according to the resource information of server cluster, determine the cluster resource scheduling total amount of server cluster;Based on cluster resource scheduling total amount, create multiple virtual control instances and allocate cluster resource quota;With the virtual control instance of allocated cluster resource quota is scheduled to server node bearing operation and activates, to execute server cluster management task;Monitoring the node running information of multiple server nodes, determine abnormal node and the virtual control instance of abnormal node bearing, adjust the cluster resource quota of abnormal node virtual control instance to reduce the load of abnormal node.The application can manage server cluster by virtual control instance, realize dynamic resource scheduling and high scalability server management method, solve the problem that resource allocation is rigid and expansion cost is high in related art due to server cluster management relying on independent hardware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server management technology, and in particular to a server management method and electronic device. Background Technology

[0002] With the rapid development of cloud computing and edge computing, the scale of server clusters is constantly expanding. The management of traditional server management clusters usually relies on physical controllers, such as BMCs (Baseboard Management Controllers). These typically use fixed hardware configurations and are deployed individually for each server to carry out management tasks, enabling out-of-band management, system monitoring, and firmware control functions.

[0003] However, this server cluster management solution that relies on physical controllers has obvious limitations in practical applications. For example, on the one hand, the hardware resources are fixed and cannot be dynamically adjusted according to the load, resulting in delays in management tasks under high load and a large number of resources being idle under low load. On the other hand, each server needs to be equipped with independent hardware, resulting in high costs and complex operation and maintenance of ultra-large-scale clusters. Summary of the Invention

[0004] This application provides a server management method that enables the management of server clusters through virtual control instances, thereby achieving dynamic resource scheduling and high scalability. This method aims to at least solve the problems of rigid resource allocation and high expansion costs caused by server cluster management relying on independent hardware in related technologies.

[0005] This application provides a server management method applied to a management controller in a server cluster, the server cluster also including multiple server nodes, the method comprising:

[0006] Based on the resource information of the server cluster, determine the total amount of cluster resource scheduling for the server cluster.

[0007] Based on the total cluster resource scheduling amount, create multiple virtual control instances and allocate cluster resource quotas;

[0008] Schedule the virtual control instance with allocated cluster resources to the server node to run, and activate the virtual control instance to perform server cluster management tasks;

[0009] Monitor the node operation information of multiple server nodes and perform load assessment to identify abnormal nodes and the virtual control instances of the abnormal nodes they support;

[0010] Based on the instance operation information of the abnormal node virtual control instance, and in combination with the abnormal node operation information, adjust the cluster resource quota of the abnormal node virtual control instance to reduce the load on the abnormal node.

[0011] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing any of the above-described server management methods when executing the computer program.

[0012] This application achieves improved overall resource utilization by creating virtual control instances in the server cluster instead of traditional physical controllers and performing dynamic resource scheduling based on instance and node operating information. This allows the virtual control instances to flexibly adjust resource allocation according to real-time load conditions, thus avoiding the rigid resource allocation problem caused by the fixed hardware configuration of physical controllers. Furthermore, since the virtual control instances do not require linear expansion of physical hardware with server nodes, the expansion cost and operational complexity of the cluster are reduced. Therefore, compared to existing technologies that rely on independent physical controllers for server cluster management, this application enables efficient and flexible server cluster management, improving the resource utilization and scalability of the server cluster. Attached Figure Description

[0013] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a schematic diagram illustrating the application environment of a server management method provided in an embodiment of this application.

[0015] Figure 2 This is a schematic diagram of a server management method provided in an embodiment of this application;

[0016] Figure 3 This application provides a schematic diagram of server management steps as an embodiment.

[0017] Figure 4 A flowchart of a server management method provided in an embodiment of this application;

[0018] Figure 5 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0020] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0021] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] This application provides a server management method that can be applied to, for example... Figure 1 The application environment is shown. Server cluster 102 comprises multiple server nodes, each providing computing, storage, and network resources. These nodes are interconnected via a high-speed network to achieve resource sharing and task collaboration. One or more management controllers are configured within server cluster 102 to centrally manage resource scheduling, load monitoring, and the scheduling and management of virtual control instances across the entire cluster. The management controller can be deployed on a selected management node among the multiple server nodes, meaning it runs within the control resources of a specific server node, or it can be deployed independently outside the cluster, thus achieving centralized control and management of the server cluster. The management controller is responsible for monitoring the operating status of each server node, assessing load conditions, scheduling virtual control instances, and performing resource adjustments and migration operations for abnormal nodes.

[0023] One or more virtual control instances run on the management controller. Each virtual control instance can undertake server cluster management tasks, including but not limited to load monitoring, task scheduling, resource allocation, and exception handling. The management controller can dynamically schedule virtual control instances to run on different server nodes or adjust the cluster resource quotas of virtual control instances based on the cluster resource status, thereby achieving elastic resource allocation and load optimization. The virtual control instance can be a vBMC instance, i.e., a virtualized baseboard management controller instance, which can be created via KVM, using a customized OpenBMC image, pre-integrated with Redfish / IPMI services, bound to an independent VLAN, and with SR-IOV virtualization enabled.

[0024] Terminal 101 communicates with server cluster 102 via a network. Terminal 101 can be a personal computer, laptop, smartphone, tablet, or other wearable device, used to send management commands to the server cluster or query the cluster management status. After receiving the commands from terminal 101, the management controller, based on the cluster resource information and the running status of each virtual control instance, calls the corresponding virtual control instance to execute the corresponding management tasks, thereby achieving efficient, dynamic, and reliable management of the server cluster.

[0025] In one embodiment, such as Figure 2 As shown, this application provides a server management method, applied to... Figure 1 The management controller in the server cluster 102 shown includes:

[0026] Step 201: Determine the total amount of cluster resource scheduling for the server cluster based on the resource information of the server cluster.

[0027] Step 202: Based on the total cluster resource scheduling amount, create multiple virtual control instances and allocate cluster resource quotas;

[0028] Step 203: Schedule the virtual control instance with allocated cluster resource quota to the server node to run, and activate the virtual control instance to perform server cluster management tasks;

[0029] Step 204: Monitor the node operation information of multiple server nodes and perform load assessment to identify abnormal nodes and the virtual control instances of the abnormal nodes they support.

[0030] Step 205: Based on the instance operation information of the abnormal node virtual control instance, and in combination with the abnormal node operation information, adjust the cluster resource quota of the abnormal node virtual control instance to reduce the load on the abnormal node.

[0031] This embodiment provides a server management method that replaces traditional physical controllers by creating virtual control instances in a server cluster. Dynamic resource scheduling is performed based on instance and node operating information, enabling the virtual control instances to flexibly adjust resource allocation according to real-time load conditions. This improves overall resource utilization and avoids the rigid resource allocation problem caused by the fixed hardware configuration of physical controllers. Furthermore, since the virtual control instances do not require linear expansion of physical hardware with server nodes, the expansion cost and operational complexity of the cluster are reduced. Compared to existing technologies that rely on independent physical controllers for server cluster management, this method achieves efficient and flexible server cluster management, improving resource utilization and scalability.

[0032] In one embodiment, determining the total cluster resource scheduling amount for the server cluster based on the server cluster's resource information includes:

[0033] The server cluster's resource information is analyzed to obtain the server cluster's historical resource utilization rate, and the lowest historical resource utilization rate and the corresponding lowest historical resource utilization time are determined within the preset historical observation period.

[0034] The maximum available resources of the cluster are obtained based on the remaining available resources of the server cluster at the lowest historical resource utilization time.

[0035] Get the remaining available resources of the server cluster at the current moment, and obtain the current available resources of the cluster;

[0036] Using the current available resources of the cluster as a baseline, subtract the current available resources of the cluster from the maximum available resources of the cluster to obtain the elastic resource difference;

[0037] Get the elastic resource release rate within a preset historical observation period, generate the average elastic resource release rate, and set the elastic compensation ratio based on the average elastic resource release rate.

[0038] Multiply the elastic compensation ratio by the elastic resource difference to obtain the elastic compensation resource amount, and add it to the current available resources of the cluster to obtain the total cluster resource scheduling amount.

[0039] Specifically, in this embodiment, by parsing the resource information of the server cluster and combining it with historical observation periods to obtain the minimum resource utilization rate and the maximum available resource quantity, and then introducing the elastic resource release rate to set the compensation ratio, it is possible not only to obtain the available resource quantity at the current moment, but also to fully consider the historical fluctuation characteristics of resources during the scheduling process. As a result, the total amount of cluster resource scheduling is not a simple static estimate, but a dynamic calculation result that combines historical and real-time information. This makes the subsequent resource allocation closer to the actual load level, avoids over-allocation or resource waste, and improves the accuracy and flexibility of the overall scheduling.

[0040] In one embodiment, based on the total cluster resource scheduling amount, multiple virtual control instances are created and cluster resource quotas are allocated, including:

[0041] Based on the operational requirements of the virtual control instance, the minimum resource requirements of the virtual control instance are preset.

[0042] The upper limit of the number of virtual control instances is determined based on the ratio of the total cluster resource scheduling amount to the minimum resource requirement, and the initial number of virtual control instances is generated in combination with the preset security creation ratio.

[0043] Create multiple virtual control instances based on the initial number of virtual control instances;

[0044] From the currently available resources in the cluster, allocate a base resource amount equal to the minimum resource requirement to multiple virtual control instances in sequence, and obtain the current remaining available resources in the cluster;

[0045] The schedulable cluster resource quota is obtained based on the current remaining available resources and elastic compensation resources of the cluster.

[0046] Based on the preset task importance weights, cluster resource quotas are allocated sequentially from the schedulable cluster resource quotas to multiple virtual control instances.

[0047] Specifically, in this embodiment, by pre-setting the minimum resource requirements when creating virtual control instances and determining the initial number by combining the total cluster resource scheduling amount and the safety ratio, this application can reasonably control the total number of instances while ensuring that each instance has the basic operating conditions. Furthermore, in the allocation process, the basic resource amount is allocated first, and then the remaining resources are used for secondary allocation in combination with the task importance weight. This realizes a two-layer strategy of "basic guarantee + weight allocation", which enables resource allocation to cover the survival needs of all instances and reflect the differences in task priority, thereby taking into account both fairness and efficiency and improving the responsiveness of critical tasks.

[0048] In one embodiment, scheduling a virtual control instance with allocated cluster resource quotas to a server node for operation and activating the virtual control instance includes:

[0049] Determine the resource requirements of the virtual control instance based on the cluster resource quota allocated to it.

[0050] Based on resource demand, multiple virtual control instances are sorted in descending order to generate an instance scheduling sequence, and the virtual control instance at the beginning of the instance scheduling sequence is selected as the current instance to be scheduled.

[0051] Obtain the current available resource information of multiple server nodes and determine the load capacity of multiple server nodes;

[0052] If the load capacity of one or more server nodes is greater than or equal to the resource requirement of the currently scheduled instance, the server node with the largest load capacity is selected as the target node, the currently scheduled instance is scheduled to the target node and activated. If the activation is successful, the currently scheduled instance is removed from the instance scheduling sequence and the load capacity of the target node is updated. If the activation fails, the scheduling is canceled and a new target node is selected.

[0053] If no server node has a load capacity greater than or equal to the resource requirement of the currently scheduled instance, the currently scheduled instance is suspended as a standby instance. The scheduling of the virtual control instance preceding the current instance in the instance scheduling sequence is performed until a server node with a load capacity that meets the resource requirement of the standby instance is detected. In this case, the standby instance is woken up and scheduled.

[0054] Specifically, in this embodiment, by arranging multiple virtual control instances in descending order based on resource demand during instance scheduling and matching them one by one with the load capacity of server nodes, this application can prioritize scheduling high-demand instances to the most suitable nodes, avoiding resource fragmentation or situations where large instances cannot run due to low-demand instances occupying slots. At the same time, a suspension and wake-up mechanism is set up so that instances that cannot be allocated temporarily can continue to be scheduled when resource conditions are met, ensuring the continuity and flexibility of the overall system, thereby improving the utilization efficiency of server node resources and the success rate of scheduling.

[0055] In one embodiment, monitoring the node operation information of multiple server nodes and performing load assessment to identify abnormal nodes includes:

[0056] Based on preset operational indicators, multiple operational indicator parameters are obtained by parsing node operational information.

[0057] The operating indicator parameters are compared with the preset operating indicator safety thresholds. If one or more operating indicator parameters are greater than the corresponding operating indicator safety threshold, the server node is determined to be an abnormal node, and the corresponding abnormal operating indicator parameters are recorded.

[0058] Specifically, in this embodiment, by parsing node operation information and comparing it based on preset indicator thresholds, this application can quickly discover potential abnormal nodes and locate the specific operation indicators that cause the abnormality. Compared with existing detection methods that rely on a single monitoring parameter, the cross-validation of multiple operation indicators can more comprehensively reflect the health status of nodes, thereby improving the accuracy and timeliness of fault detection and providing a reliable basis for subsequent instance adjustments or migrations.

[0059] In one embodiment, such as Figure 3 As shown, based on preset task importance weights, cluster resource quotas are allocated sequentially from the schedulable cluster resource quota to multiple virtual control instances, and the method also includes:

[0060] Step 301: Obtain the upper limit of the instance quantity elasticity based on the ratio of the schedulable cluster resource quota to the minimum resource requirement;

[0061] Step 302: Determine the number of instance replicas based on the initial number of virtual control instances and the upper limit of the instance number elasticity;

[0062] Step 303: Set the virtual control instance as the primary instance, and create multiple secondary instances for the primary instance to meet the required number of instance replicas;

[0063] Step 304: Allocate cluster resource quotas from the schedulable cluster resource quotas to the master instance and the corresponding multiple slave instances in sequence.

[0064] Specifically, in this embodiment, the elastic upper limit of the number of instances is calculated based on the ratio of the minimum resource requirement to the schedulable cluster resource quota. On this basis, several slave instances, i.e. replicas, are configured for each master instance. The slave instances serve as backups for the master instance, enabling them to quickly take over management tasks when the master instance fails or experiences abnormal load. This ensures that the master instance can maintain necessary operational guarantees when resources are scarce, while the slave instances are in a pre-configured or on-demand ready state to achieve takeover within seconds, thereby significantly reducing availability loss caused by single point of failure in the management plane. In addition, limiting and dynamically configuring the number of replicas based on the schedulable resource quota helps to obtain the required high availability capabilities with minimal additional consumption under limited resources, thereby improving system reliability and reducing hardware and operation and maintenance costs without relying on hardware redundancy.

[0065] In one embodiment, after allocating cluster resource quotas from the schedulable cluster resource quotas to the primary instance and the corresponding multiple secondary instances in sequence, the method further includes:

[0066] Schedule the primary instance to the first server node and activate it;

[0067] The corresponding multiple instances are scheduled to one or more second server nodes in the server cluster, which are different from the first server node;

[0068] Based on a preset heartbeat cycle, the master instance is controlled to send heartbeat data packets to multiple corresponding slave instances;

[0069] If a slave instance fails to receive a heartbeat data packet within multiple consecutive heartbeat cycles, the master instance is deemed to have failed, and a leader instance is elected from among the existing slave instances based on a preset election rule.

[0070] Activate the leader instance and synchronize the pending server cluster management tasks of the primary instance so that the leader instance can continue to execute them.

[0071] Specifically, in this embodiment, the master instance and its slave instances are deployed on different server nodes, and the readiness monitoring of slave instances and the election of a leader instance are realized based on heartbeat detection and preset election rules. This enables the rapid election and activation of slave instances to take over the task without manual intervention when the master instance fails, ensuring the continuity of management functions. Furthermore, unlike redundancy schemes based on additional physical controllers and other hardware, the virtualized master-slave mechanism provided in this embodiment saves hardware while achieving high availability switching through ready replicas and automatic election mechanisms, reducing the risk of single point of failure and shortening the fault recovery time, thereby improving the robustness and operational efficiency of the cluster management plane.

[0072] In one embodiment, based on the instance runtime information of the abnormal node virtual control instance, and in conjunction with the abnormal node runtime information, the cluster resource quota of the abnormal node virtual control instance is adjusted to reduce the load on the abnormal node, including:

[0073] Parse the instance runtime information to obtain the instance's current resource utilization, instance task queue length, instance historical task response latency, and instance historical resource consumption records;

[0074] Analyze the abnormal node's operational information to obtain the node's current resource utilization and historical load fluctuation records;

[0075] Select the action to be evaluated from the preset instance adjustment actions. The instance adjustment actions include at least expansion, shrinkage, and migration.

[0076] Based on the instance's current resource utilization, instance task queue length, instance's historical resource consumption records, and node's current resource utilization, generate a predicted improvement value for the node's resource utilization for the action to be evaluated.

[0077] Based on the current cluster resource quota of the abnormal node virtual control instance and the remaining resources of the abnormal node, generate the action prediction cost value of the action to be evaluated.

[0078] Based on the instance task queue length, instance historical task response latency, and node historical load fluctuation records, generate a service quality prediction loss value for the action to be evaluated.

[0079] Based on the predicted improvement value of node resource utilization, the predicted cost value of actions, and the predicted loss value of service quality, an action score is generated for the action to be evaluated.

[0080] From multiple actions to be evaluated, select the action with the highest score as the target instance to adjust the action;

[0081] In response to the target instance adjustment action being either expansion or reduction, the cluster resource quota adjustment amount for the abnormal node virtual control instance is determined based on the action score of the target instance adjustment action, and the abnormal node virtual control instance is adjusted accordingly.

[0082] In response to the target instance adjustment action being migration, a target migration node that is distinct from the abnormal node is selected from multiple server nodes, and the virtual control instance of the abnormal node is migrated to the target migration node.

[0083] Specifically, in this embodiment, during the process of adjusting the virtual control instance of abnormal nodes, a predicted improvement value for resource utilization, an action overhead value, and a service quality loss value are generated and quantified into an action score. This enables precise comparison and evaluation among various adjustment actions. Unlike existing manual experience-based adjustments, a prediction and scoring model is introduced, allowing decisions such as scaling up, scaling down, or migration to be made based on data. This ensures that the selected actions can alleviate node load without introducing excessive additional overhead or service quality degradation, thus achieving intelligent and optimized adjustment of abnormal node resources.

[0084] In one embodiment, the action to be evaluated is denoted as a, and a corresponding set of action prediction parameters P is introduced. a It is used to predict the effect of action execution. The action prediction parameter set is not the final execution parameter, but an idealized parameter used for scoring calculation. For example, the expansion action can assume that the resource quota is increased by a fixed percentage, and the migration action can assume that the target node is migrated to a resource-sufficient target node.

[0085] Based on the instance's current resource utilization, instance task queue length, instance historical resource consumption records, and node's current resource utilization, a predicted improvement value for the node's resource utilization for the action to be evaluated is generated, expressed as:

[0086] ;

[0087] Wherein, ∆R(a,P a The ) indicates the predicted improvement in node resource utilization after the virtual control instance of the abnormal node in the abnormal node performs the action to be evaluated, . This represents the resource utilization rate of the virtual control instance of the abnormal node at historical time t. α represents the instance task queue length of the virtual control instance of the abnormal node, and α is the weighting coefficient of the task queue length on resource pressure. k represents the historical resource utilization rate of this abnormal node. R (a,P a The adjustment coefficient is the action adjustment factor, which depends on the action to be evaluated and the set of action prediction parameters. For example, expansion or contraction can use a linear or exponential adjustment function, and migration actions can be determined based on the resource availability ratio of the target node.

[0088] Based on the current cluster resource quota of the abnormal node's virtual control instance and the remaining resources of the abnormal node, the predicted cost value of the action to be evaluated is generated, expressed as:

[0089] ;

[0090] Wherein, C(a,P) a R represents the action prediction cost value of the virtual control instance of the abnormal node executing the action to be evaluated, a. pred (a,P a The ) represents the additional cluster resource quota required after the virtual control instance of the abnormal node executes action 'a' under the action prediction parameters. This quota is determined by the current cluster resource quota and is specified by R. avail This represents the remaining resources of the abnormal node. The larger the action prediction cost value, the higher the pressure the action puts on the node's resources, which is reflected as a potential negative impact in the action scoring.

[0091] Based on the instance task queue length, the instance's historical task response latency, and the node's historical load fluctuation records, a service quality prediction loss value is generated for the action to be evaluated, expressed as follows:

[0092] ;

[0093] Among them, L q (a,P a This indicates the predicted service quality loss value for the virtual control instance of the abnormal node that performed the action to be evaluated. This indicates the predicted task response time of the virtual control instance of the abnormal node under the action to be evaluated. It is predicted by combining the instance task queue length, the instance's historical task response latency, and the node's historical load fluctuation records with action parameters. It is preferably generated using a machine learning model. T target For the target response time, T pred To predict the length of the time window;

[0094] Based on the predicted improvement value of node resource utilization, the predicted cost value of actions, and the predicted loss value of service quality, an action score is generated for the action to be evaluated, represented as follows:

[0095] , where S(a,P a The value indicates the action score of the virtual control instance of the abnormal node that was executed for the action to be evaluated.

[0096] In a further embodiment, after selecting the action with the highest action score from a plurality of actions to be evaluated as the target instance for action adjustment, the method further includes:

[0097] Obtain the set of action prediction parameters for the target instance adjustment action, and use it as the reference action set. The reference action parameter set includes parameters such as resource quota adjustment amount, adjustment execution time, migration target node, and migration execution time.

[0098] When the target instance's adjustment action is to expand or shrink, the cluster resource adjustment amount of the abnormal node's virtual control instance is determined based on the resource quota adjustment amount in the reference action parameter set and the weight coefficient determined by the action score, and the adjustment is executed at the adjustment execution time specified in the reference action parameter set.

[0099] Based on the target node and migration execution time in the reference action parameter set, the virtual control instance of the abnormal node is migrated to the target node, and the migration operation is completed at the specified time.

[0100] Specifically, in this embodiment, by introducing a set of action prediction parameters and a prediction-based scoring mechanism, the impact of various possible actions on node resource utilization, action overhead, and service quality can be assessed in advance when abnormal node virtual control instances experience resource pressure or a declining service quality trend. This allows for the selection of the optimal action for execution. Compared to traditional methods that adjust based solely on current resources or static rules, this embodiment enables dynamic and intelligent resource management and load control. It reduces the risk of node overload, ensures the stability of service quality, improves the overall resource utilization and task processing efficiency of the cluster, and enhances the accuracy and controllability of operations by adjusting the action execution intensity and time based on the action score.

[0101] In one embodiment, in response to a target instance adjustment action being a migration, a target migration node, distinct from the abnormal node, is selected from multiple server nodes, and the virtual control instance of the abnormal node is migrated to the target migration node, including:

[0102] Based on the node operation information of multiple server nodes, the server node with the lowest current load index that is below the preset security reception threshold is identified as a candidate node.

[0103] Obtain historical load fluctuation records of candidate nodes and determine whether the historical load fluctuation frequency of candidate nodes is lower than the preset stability threshold. If so, the candidate node is selected as the target migration node; otherwise, the candidate node is deemed invalid and reselected.

[0104] Suspend the task process of the abnormal node virtual control instance and migrate the abnormal node virtual control instance to the target migration node;

[0105] Reactivate the abnormal node virtual control instance on the target migration node. If activation is successful, the abnormal node virtual control instance will be determined as a normal operating instance. If activation fails, the abnormal node virtual control instance will be rolled back to the abnormal node.

[0106] Specifically, in this embodiment, by first selecting the node with the lowest current load as a candidate during the instance migration process, and then combining historical load fluctuation records for stability judgment, the instance can be avoided from being migrated to a node that is idle but fluctuates frequently, thereby improving the operational stability after migration. At the same time, a pause, activation, and rollback mechanism is set up during the migration process to ensure that even if the migration fails, it can be quickly restored to the original node, reducing the risk of task interruption. This mechanism not only improves the reliability of the migration operation, but also enhances the cluster's adaptive ability to deal with anomalies.

[0107] In one embodiment, such as Figure 4 As shown, in a server management method provided in this application, after scheduling a virtual control instance with allocated cluster resource quota to a server node for operation and activating the virtual control instance to perform server cluster management tasks, the method includes:

[0108] The management controller periodically monitors the node operation information of multiple server nodes and performs load assessment based on preset operation indicators. The node operation information includes current resource utilization, historical load fluctuation records, etc. The instance operation information of the virtual control instance includes the instance's current resource utilization, task queue length, historical task response latency, and resource consumption records.

[0109] The management controller determines the node status based on the load assessment results. If the node load is normal, the cluster resource quota of the current virtual control instance is maintained. If the node load is abnormal, the abnormal node and its virtual control instance are dynamically adjusted.

[0110] For abnormal nodes, the management controller combines the abnormal node's operation information with the instance operation information of the abnormal node's virtual control instance, and selects the action to be evaluated from the preset instance adjustment actions, wherein the instance adjustment actions include expansion, reduction and migration.

[0111] For each action to be evaluated, the predicted improvement value of node resource utilization, the predicted cost value of action, and the predicted loss value of service quality are calculated respectively. Based on these intermediate indicators, an action score is generated. The management controller selects the action with the highest score from multiple actions to be evaluated as the target instance to adjust the action.

[0112] When the target instance's adjustment action is to expand or shrink, the management controller determines the amount of cluster resource quota adjustment for the virtual control instance of the abnormal node based on the action score of the target instance's adjustment action, and adjusts the cluster resource quota of the virtual control instance accordingly.

[0113] When the target instance is adjusted to migrate, the management controller selects the server node with the lowest current load and below the preset safe reception threshold as a candidate node based on the node operation information of multiple server nodes. It further determines whether the historical load fluctuation frequency of the candidate node is lower than the preset stable threshold. If the condition is met, it is determined as the target migration node; otherwise, a candidate node is selected again.

[0114] The management controller suspends the task process of the virtual control instance on the abnormal node and migrates it to the target migration node. On the target migration node, the virtual control instance is reactivated. If the activation is successful, the virtual control instance is determined to be a normal operating instance. If the activation fails, the virtual control instance is rolled back to the original abnormal node.

[0115] Specifically, in this embodiment, through specific monitoring and instance adjustment processes, dynamic monitoring and evaluation of node load can be achieved. When the load is abnormal, the virtual control instance can be scaled up, down, or migrated, thereby effectively reducing node load and improving cluster resource utilization. At the same time, when the load is normal, the existing resource quota is maintained to ensure the stable operation of the server cluster.

[0116] In one embodiment, the server cluster includes multiple server nodes, each server node deploying at least one virtual control instance. A management controller is responsible for the unified scheduling and management of these virtual control instances to achieve node isolation, task scheduling, and access control. The server management method executed by the management controller further includes:

[0117] The system receives access requests from server nodes and forwards them to the corresponding virtual control instance. The virtual control instance performs preliminary authentication, which includes node identity verification, access credential verification, and timestamp confirmation. Specifically, identity verification uses two-way authentication based on X.509 certificates, with a certificate validity period of 90 to 180 days and a timestamp tolerance of ±5 seconds.

[0118] Based on the request type and task priority, access requests are allocated to different isolation environments through virtual control instances. The isolation environments are divided into three categories: high-level, medium-level, and low-level. Each type of isolation environment corresponds to independent virtual resources. High-level environments are allocated 4 to 8 CPU cores and 8 to 16GB of memory, medium-level environments are allocated 2 to 4 CPU cores and 4 to 8GB of memory, and low-level environments are allocated 1 to 2 CPU cores and 2 to 4GB of memory.

[0119] Furthermore, encryption policies are set for different levels of requests through virtual control instances. High-level requests enable AES-256-GCM encryption channels with a key update cycle of 24 hours and generate dynamic random salt values. Medium-level requests enable AES-128-CBC encryption with a key update cycle of 72 hours. Low-level requests are allowed to transmit in plaintext, but all write operations are logged by the virtual control instance and reported to the management controller once per hour.

[0120] It periodically receives access behavior information reported by each virtual control instance. When five consecutive abnormal operations are detected, such as access frequency exceeding the threshold or authentication failure, it executes an isolation policy through the virtual control instance to migrate the corresponding node into the isolation environment. The isolation time is 10 to 30 minutes, which can be dynamically adjusted according to the administrator's policy.

[0121] Specifically, in this embodiment, the management controller implements logical isolation and security control over each server node. The virtual control instance serves as the specific carrier for scheduling and execution by the management controller, ensuring refined management of task scheduling, access control, and security monitoring between nodes. This improves the security, controllability, and real-time access to core business functions of the server cluster, while reducing the risk of single points of failure and attack propagation.

[0122] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0123] like Figure 5 As shown, embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described server management method embodiments.

[0124] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0125] The server management method and electronic device provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A server management method, characterized in that, The method includes: Based on the resource information of the server cluster, determine the total amount of cluster resource scheduling for the server cluster; Based on the total cluster resource scheduling amount, create multiple virtual control instances and allocate cluster resource quotas; The process of scheduling the virtual control instance, which has been allocated the cluster resource quota, to server nodes for operation and activating the virtual control instance to perform server cluster management tasks includes: determining the resource requirement of the virtual control instance based on the allocated cluster resource quota; sorting multiple virtual control instances in descending order based on the resource requirement to generate an instance scheduling sequence, and selecting the virtual control instance at the beginning of the instance scheduling sequence as the current instance to be scheduled; obtaining the current available resource information of multiple server nodes and determining the load capacity of multiple server nodes; and responding when the load capacity of one or more server nodes is greater than or equal to the current instance to be scheduled. If the resource requirement is not met, the server node with the largest load capacity is selected as the target node. The currently scheduled instance is scheduled to the target node and activated. If activation is successful, the currently scheduled instance is removed from the instance scheduling sequence, and the load capacity of the target node is updated. If activation fails, the scheduling is canceled and the target node is reselected. If there is no server node whose load capacity is greater than or equal to the resource requirement of the currently scheduled instance, the currently scheduled instance is suspended as a standby instance. The scheduling of the virtual control instance before the current instance in the instance scheduling sequence is executed until a server node with a load capacity that meets the resource requirement of the standby instance is detected, at which point the standby instance is woken up and scheduled. Monitor the node operation information of multiple server nodes and perform load assessment to identify abnormal nodes and the virtual control instances of the abnormal nodes they support; Based on the instance operation information of the abnormal node virtual control instance, and in conjunction with the abnormal node operation information, the cluster resource quota of the abnormal node virtual control instance is adjusted to reduce the load on the abnormal node.

2. The server management method according to claim 1, characterized in that, The step of determining the total cluster resource scheduling amount for the server cluster based on the server cluster resource information includes: The resource information of the server cluster is analyzed to obtain the historical resource utilization rate of the server cluster, and the lowest historical resource utilization rate and the corresponding lowest historical resource utilization time of the server cluster within a preset historical observation period are determined. The maximum available resources of the cluster are obtained based on the remaining available resources of the server cluster at the lowest historical resource utilization time. Obtain the remaining available resources of the server cluster at the current moment to get the current available resources of the cluster; Using the current available resources of the cluster as a baseline, subtract the current available resources of the cluster from the maximum available resources of the cluster to obtain the elastic resource difference; Obtain the elastic resource release rate within the preset historical observation period, generate the average elastic resource release rate, and set the elastic compensation ratio based on the average elastic resource release rate. Multiplying the elastic compensation ratio by the elastic resource difference yields the elastic compensation resource amount, which is then added to the currently available cluster resources to obtain the total cluster resource scheduling amount.

3. The server management method according to claim 2, characterized in that, The step of creating multiple virtual control instances and allocating cluster resource quotas based on the total cluster resource scheduling amount includes: Based on the operational requirements of the virtual control instance, the minimum resource requirements of the virtual control instance are preset. Based on the ratio of the total cluster resource scheduling amount to the minimum resource requirement, the upper limit of the number of virtual control instances is determined, and the initial number of virtual control instances is generated in combination with the preset security creation ratio. Based on the initial number of virtual control instances, multiple virtual control instances are created; From the currently available resources of the cluster, allocate a base resource amount equal to the minimum resource requirement to each of the virtual control instances in sequence, and obtain the current remaining available resources of the cluster; The schedulable cluster resource quota is obtained based on the current remaining available resources of the cluster and the elastic compensation resource amount; Based on preset task importance weights, the cluster resource quota is sequentially allocated from the schedulable cluster resource quota to multiple virtual control instances.

4. The server management method according to claim 1, characterized in that, The monitoring of node operation information of multiple server nodes and the performance of load assessment to identify abnormal nodes include: Based on preset operating indicators, the node operating information is analyzed to obtain multiple operating indicator parameters; The operation indicator parameters are compared with the preset operation indicator safety thresholds. If one or more operation indicator parameters are greater than the corresponding operation indicator safety threshold, the server node is determined to be an abnormal node, and the corresponding abnormal operation indicator parameters are recorded.

5. The server management method according to claim 3, characterized in that, The method of allocating cluster resources from the schedulable cluster resource quota to multiple virtual control instances based on preset task importance weights further includes: The upper limit of the instance number elasticity is obtained based on the ratio of the schedulable cluster resource quota to the minimum resource requirement; The number of instance replicas is determined based on the initial number of virtual control instances and the elastic upper limit of the number of instances; Set the virtual control instance as the master instance, and create multiple slave instances for the master instance to meet the required number of instance replicas; The cluster resource quota is allocated sequentially from the schedulable cluster resource quota to the master instance and the corresponding multiple slave instances.

6. The server management method according to claim 5, characterized in that, After allocating the cluster resource quota from the schedulable cluster resource quota to the master instance and the corresponding plurality of slave instances in sequence, the method further includes: The primary instance is scheduled to the first server node and activated. The corresponding multiple instances are scheduled to one or more second server nodes in the server cluster that are distinct from the first server node; Based on a preset heartbeat cycle, the master instance is controlled to send heartbeat data packets to multiple corresponding slave instances; If the slave instance fails to receive the heartbeat data packet within multiple consecutive heartbeat cycles, the master instance is determined to be invalid, and a leader instance is elected from the existing slave instances based on a preset election rule. The leader instance is activated, and the pending server cluster management tasks of the master instance are synchronized for the leader instance to continue execution.

7. The server management method according to claim 1, characterized in that, The step of adjusting the cluster resource quota of the abnormal node virtual control instance to reduce the load on the abnormal node, based on the instance operation information of the abnormal node virtual control instance and in combination with the abnormal node operation information, includes: Parse the instance runtime information to obtain the instance's current resource utilization rate, instance task queue length, instance historical task response latency, and instance historical resource consumption records; The abnormal node's operational information is analyzed to obtain the node's current resource utilization and historical load fluctuation records. Select the action to be evaluated from the preset instance adjustment actions, wherein the instance adjustment actions include at least expansion, reduction and migration; Based on the instance's current resource utilization rate, the instance's task queue length, the instance's historical resource consumption records, and the node's current resource utilization rate, generate a predicted improvement value for the node's resource utilization rate for the action to be evaluated. Based on the current cluster resource quota of the abnormal node virtual control instance and the remaining resources of the abnormal node, generate the action prediction cost value of the action to be evaluated. Based on the instance task queue length, the instance historical task response latency, and the node historical load fluctuation record, generate the service quality prediction loss value of the action to be evaluated. Based on the predicted improvement value of node resource utilization, the predicted cost value of action, and the predicted loss value of service quality, an action score is generated for the action to be evaluated. From the multiple actions to be evaluated, select the action with the highest action score as the target instance to adjust the action; In response to the target instance adjustment action being either expansion or reduction, the cluster resource quota adjustment amount of the abnormal node virtual control instance is determined based on the action score of the target instance adjustment action, and the abnormal node virtual control instance is adjusted accordingly. In response to the target instance adjustment action being migration, a target migration node that is distinct from the abnormal node is selected from the plurality of server nodes, and the virtual control instance of the abnormal node is migrated to the target migration node.

8. The server management method according to claim 7, characterized in that, In response to the target instance adjustment action being migration, a target migration node, distinct from the abnormal node, is selected from the plurality of server nodes, and the virtual control instance of the abnormal node is migrated to the target migration node, including: Based on the node operation information of multiple server nodes, the server node with the lowest current load index and below the preset security reception threshold is determined as a candidate node. Obtain the historical load fluctuation records of the candidate node, and determine whether the historical load fluctuation frequency of the candidate node is lower than a preset stability threshold. If so, the candidate node is selected as the target migration node; otherwise, the candidate node is deemed invalid and reselected. Suspend the task process of the abnormal node virtual control instance and migrate the abnormal node virtual control instance to the target migration node; The abnormal node virtual control instance is reactivated on the target migration node. If the activation is successful, the abnormal node virtual control instance is determined to be a normal operating instance. If the activation fails, the abnormal node virtual control instance is rolled back to the abnormal node.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for implementing the server management method as described in any one of claims 1 to 8 when executing the computer program.

Citation Information

Patent Citations

  • Method and system for dynamic scheduling of virtual resources in cloud computing network

    CN102170474A

  • Virtual server Virtual CPU resource monitoring and dynamic allocation method

    CN103729254A