Server management method and electronic equipment
By creating virtual control instances in the server cluster and dynamically scheduling resources, the problems of resource rigidity and high cost caused by traditional physical controllers are solved, achieving efficient and flexible server management and improving resource utilization and scalability.
Patent Information
- Application Number
- CN202511404878.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-09-29
AI Technical Summary
Traditional server management solutions rely on physical controllers, resulting in rigid resource allocation, high expansion costs, inability to dynamically adjust according to load, and complex operation and maintenance.
By creating virtual control instances in the server cluster, resources are dynamically scheduled and node load is monitored, enabling flexible resource management and load optimization. Virtualized baseboard management controller instances (vBMC) are used to replace traditional physical controllers.
It improved resource utilization, reduced expansion costs and operational complexity, and enabled efficient and flexible server cluster management, thereby improving resource utilization and scalability.
Smart Images

Figure CN120892293A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of server management, and particularly relates to a server management method and an electronic device. BACKGROUND
[0002] With the rapid development of cloud computing and edge computing, the scale of server clusters is continuously expanding. The management of traditional server management clusters usually relies on physical controllers such as BMC (Baseboard Management Controller), which are generally separately deployed for each server with fixed hardware configurations, and are used to carry out management tasks, implement out-of-band management, system monitoring and firmware control functions.
[0003] However, this server cluster management scheme relying on physical controllers has obvious limitations in actual application, for example: on the one hand, the hardware resources are fixed and cannot be dynamically adjusted according to the load, resulting in management task delays in high-load situations, while a large amount of resources are idle in low-load situations; on the other hand, each server needs to be equipped with independent hardware, resulting in high costs and complex operation and maintenance of super-large-scale clusters. SUMMARY
[0004] The present application provides a server management method capable of managing a server cluster through a virtual control instance, thereby realizing dynamic resource scheduling and high scalability of server management, to at least solve the problem of rigid resource allocation and high expansion cost caused by the dependence of server cluster management on independent hardware in the related art.
[0005] The present application provides a server management method applied to a management controller in a server cluster, the server cluster further comprising a plurality of server nodes, the method comprising: determining a cluster resource scheduling total amount of the server cluster according to resource information of the server cluster; creating a plurality of virtual control instances and allocating a cluster resource quota based on the cluster resource scheduling total amount; scheduling the virtual control instances with the allocated cluster resource quota to the server nodes for carrying and running, and activating the virtual control instances to perform server cluster management tasks; monitoring node running information of the plurality of server nodes and performing load evaluation to determine an abnormal node and an abnormal node virtual control instance carried by the abnormal node; adjusting the cluster resource quota of the abnormal node virtual control instance according to instance running information of the abnormal node virtual control instance in combination with the abnormal node running information to reduce the load of the abnormal node.
[0006] The present application also provides an electronic device comprising a memory for storing a computer program and a processor for executing the computer program to implement any of the above server management methods.
[0007] Through the present application, since a virtual control instance is created in the server cluster instead of a traditional physical controller, and dynamic resource scheduling is performed based on instance running information and node running information, the virtual control instance can flexibly adjust resource quotas according to real-time load conditions, thereby improving overall resource utilization and avoiding the resource allocation rigidity problem caused by the fixed hardware configuration of the physical controller. At the same time, since the virtual control instance does not need to linearly expand physical hardware with the server nodes, the expansion cost and operation and maintenance complexity of the cluster are reduced. Therefore, compared with the prior art that relies on independent physical controllers for server cluster management, the present application can achieve efficient and flexible server cluster management, improve the resource utilization and scalability of the server cluster. BRIEF DESCRIPTION OF DRAWINGS
[0008] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0009] Figure 1 A server management method application environment schematic diagram is provided for the embodiments of the present application. Figure 2 A server management method schematic diagram is provided for the embodiments of the present application. Figure 3 A server management step schematic diagram is provided for the embodiments of the present application. Figure 4 A server management method flow chart is provided for the embodiments of the present application. Figure 5 An electronic device schematic diagram is provided for the embodiments of the present application. DETAILED DESCRIPTION
[0010] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0011] It should be noted that in the description of the present application, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or further include elements inherent in such processes, methods, articles or devices. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0012] In order to enable those skilled in the art to better understand the present application, the present application will be further described in detail below in combination with the drawings and specific embodiments.
[0013] The server management method provided by the present application can be applied to the application environment as shown in the figure. Figure 1 The server cluster 102 includes a plurality of server nodes, each server node provides computing, storage and network resources, and the server nodes are interconnected through a high-speed network to realize resource sharing and task cooperation. One or more management controllers are arranged in the server cluster 102, which are used to uniformly manage the resource scheduling, load monitoring and scheduling and management of virtual control instances of the entire server cluster. The management controller can be deployed on the management node selected from the plurality of server nodes, that is, the management controller runs in the control resource of a certain server node, or it can be deployed outside the cluster independently of the server node, thereby realizing centralized control and management of the server cluster. The management controller is responsible for monitoring the running state of each server node, evaluating the load condition, scheduling the virtual control instance, and performing resource adjustment and migration operation of the abnormal node.
[0014] One or more virtual control instances are run on the management controller, and each virtual control instance can carry server cluster management tasks, including but not limited to load monitoring, task scheduling, resource allocation and exception handling. The management controller can dynamically schedule the virtual control instance to different server nodes according to the cluster resource state, or adjust the cluster resource quota of the virtual control instance, thereby realizing elastic allocation of resources and load optimization. The virtual control instance can be a vBMC instance, that is, a virtualized baseboard management controller instance, which can be created through KVM, use OpenBMC customized image, pre-integrate Redfish / IPMI service, bind independent VLAN, and enable SR-IOV virtualization.
[0015] The terminal 101 communicates with the server cluster 102 through a network. The terminal 101 can be a personal computer, a notebook computer, a smart phone, a tablet computer or other wearable devices, which is used to initiate a management instruction or query a cluster management state to the server cluster. After receiving the instruction of the terminal 101, the management controller calls the corresponding virtual control instance to perform the corresponding management task according to the cluster resource information and the running state of each virtual control instance, so as to realize efficient, dynamic and reliable management of the server cluster.
[0016] In one embodiment, as shown in Figure 2 The present application provides a server management method, which is applied to a management controller in a server cluster 102 as shown in Figure 1 The present application provides a server management method, which is applied to a management controller in a server cluster 102 as shown in Step 201, determining a cluster resource scheduling total amount of the server cluster according to resource information of the server cluster; Step 202, creating a plurality of virtual control instances and allocating a cluster resource amount based on the cluster resource scheduling total amount; Step 203, scheduling the virtual control instance with the allocated cluster resource amount to a server node for carrying and activating the virtual control instance to perform a server cluster management task; Step 204, monitoring node running information of a plurality of server nodes and performing load evaluation to determine an abnormal node and an abnormal node virtual control instance carried by the abnormal node; Step 205, adjusting the cluster resource amount of the abnormal node virtual control instance according to instance running information of the abnormal node virtual control instance and in combination with the abnormal node running information to reduce the load of the abnormal node.
[0017] The server management method provided in this embodiment replaces the traditional physical controller by creating a virtual control instance in the server cluster, and performs dynamic resource scheduling based on the instance running information and the node running information, so that the virtual control instance can flexibly adjust the resource amount according to the real-time load condition, thereby improving the overall resource utilization and avoiding the resource allocation rigidity problem caused by the fixed hardware configuration of the physical controller. At the same time, since the virtual control instance does not need to linearly expand the physical hardware with the server node, the expansion cost and operation and maintenance complexity of the cluster are reduced, so that compared with the prior art which relies on independent physical controllers to manage the server cluster, efficient and flexible server cluster management can be realized, and the resource utilization and scalability of the server cluster are improved.
[0018] In one embodiment, determining a cluster resource scheduling total amount of the server cluster according to resource information of the server cluster comprises: The resource information of the server cluster is parsed, historical resource utilization of the server cluster is obtained, and the lowest historical resource utilization and the corresponding lowest historical resource utilization time of the server cluster in a preset historical observation period are determined; The maximum available resource amount of the cluster is obtained according to the remaining available resource amount of the server cluster at the lowest historical resource utilization time; The remaining available resource amount of the server cluster at the current time is obtained, and the current available resource amount of the cluster is obtained; The current available resource amount of the cluster is taken as a baseline, the current available resource amount of the cluster is subtracted from the maximum available resource amount of the cluster, and the flexible resource difference is obtained; The flexible resource release rate in the preset historical observation period is obtained, and the average flexible resource release rate is generated, and the flexible compensation ratio is set based on the average flexible resource release rate; The flexible compensation resource amount is obtained by multiplying the flexible compensation ratio and the flexible resource difference, and the cluster resource scheduling total amount is obtained by adding the current available resource amount of the cluster.
[0019] Specifically, in this embodiment, by parsing the resource information of the server cluster, and combining the lowest resource utilization and the maximum available resource amount obtained in the historical observation period, and introducing the flexible resource release rate to set the compensation ratio, not only the available resource amount at the current time can be obtained, but also the historical fluctuation characteristics of the resource can be fully considered in the scheduling process, so that the cluster resource scheduling total amount is not a simple static estimation, but a comprehensive calculation result dynamically combining history and real-time information, so that the subsequent resource allocation is closer to the real load level, avoiding excessive allocation or resource waste, and improving the accuracy and flexibility of the overall scheduling.
[0020] In one embodiment, based on the cluster resource scheduling total amount, a plurality of virtual control instances are created and the cluster resource amount is allocated, including: Based on the running demand of the virtual control instance, the lowest resource demand amount of the virtual control instance is pre-set; According to the ratio of the cluster resource scheduling total amount and the lowest resource demand, the upper limit number of the virtual control instance is determined, and the initial number of the virtual control instance is generated in combination with the preset safety creation ratio; Based on the initial number of the virtual control instance, a plurality of virtual control instances are created; From the current available resource amount of the cluster, the basic resource amount equal to the lowest resource demand amount is allocated to the plurality of virtual control instances in turn, and the current remaining available resource amount of the cluster is obtained; According to the current remaining available resource amount of the cluster and the flexible compensation resource amount, the schedulable cluster resource amount is obtained; Based on the preset task importance weight, the cluster resource amount is allocated to the plurality of virtual control instances from the schedulable cluster resource amount in turn.
[0021] Specifically, in the present embodiment, by pre-setting the minimum resource requirement amount when creating a virtual control instance, and combining the cluster resource scheduling total amount with the safety ratio to determine the initial number, the application can reasonably control the total number of instances under the premise of ensuring that each instance has basic operating conditions; further, in the allocation process, the basic resource amount is allocated first, and then the remaining resources are allocated again in combination with the task importance weight, realizing the double-layer strategy of "basic guarantee + weight allocation", so that the resource allocation can cover the survival needs of all instances and reflect the differences in task priority, thus balancing fairness and efficiency and improving the response capability of critical tasks.
[0022] In one embodiment, the virtual control instance with the allocated cluster resource quota is scheduled to a server node for carrying and activating the virtual control instance, including: According to the cluster resource quota allocated to the virtual control instance, determining the resource requirement amount of the virtual control instance; Based on the resource requirement amount, the multiple virtual control instances are arranged in descending order to generate an instance scheduling sequence, and the virtual control instance at the front of the instance scheduling sequence is selected as the current to-be-scheduled instance; Obtaining the current available resource information of the multiple server nodes and determining the loadable amount of the multiple server nodes; In response to the loadable amount of one or more server nodes being greater than or equal to the resource requirement amount of the current to-be-scheduled instance, the server node with the largest loadable amount is selected as the target node, the current to-be-scheduled instance is scheduled to the target node and activated, if the activation is successful, the current to-be-scheduled instance is removed from the instance scheduling sequence, and the loadable amount of the target node is updated, if the activation fails, the scheduling is cancelled and the target node is reselected; In response to there being no server node with a loadable amount greater than or equal to the resource requirement amount of the current to-be-scheduled instance, the current to-be-scheduled instance is suspended as a standby instance, the scheduling of the next-to-front virtual control instance in the instance scheduling sequence is performed, and the standby instance is awakened and scheduled when it is detected that there is a server node with a loadable amount meeting the resource requirement amount of the standby instance.
[0023] Specifically, in the present embodiment, by arranging the multiple virtual control instances in descending order according to the resource requirement amount when scheduling the instances, and matching them one by one in combination with the loadable amount of the server nodes, the application can preferentially schedule high-demand instances to the most suitable nodes, avoiding the situation that resource fragmentation or low-demand instances occupying space prevent large instances from running; at the same time, the suspension and awakening mechanism is set, so that instances that cannot be allocated temporarily can continue to be scheduled when the resource conditions are met, ensuring the continuity and flexibility of the overall system, thereby improving the utilization efficiency of the server node resources and the success rate of the scheduling.
[0024] In one embodiment, node running information of a plurality of server nodes is monitored and load evaluation is performed to determine abnormal nodes, including: Based on preset running indicators, node running information is analyzed to obtain a plurality of running indicator parameters; The running indicator parameters are compared with preset running indicator safety thresholds, and if one or more running indicator parameters are greater than the corresponding running indicator safety thresholds, the server node is determined to be an abnormal node, and the corresponding abnormal running indicator parameters are recorded.
[0025] Specifically, in this embodiment, by analyzing node running information and comparing based on preset indicator thresholds, the application can quickly find potential abnormal nodes and locate the specific running indicators that cause the abnormality. Compared with the existing detection method which relies on a single monitoring parameter, the use of multiple running indicators for cross-validation can more comprehensively reflect the health status of the node, thereby improving the accuracy and timeliness of fault detection and providing a reliable basis for subsequent instance adjustment or migration.
[0026] In one embodiment, as shown in Figure 3 Based on the preset task importance weight, the cluster resource quota is sequentially allocated to the plurality of virtual control instances from the schedulable cluster resource quota, further including: Step 301, obtaining an instance quantity elasticity upper limit according to the ratio of the schedulable cluster resource quota and the minimum resource requirement quantity; Step 302, determining an instance replica quantity according to the initial quantity of virtual control instances and the instance quantity elasticity upper limit; Step 303, setting the virtual control instance as a master instance, and creating a plurality of slave instances satisfying the instance replica quantity for the master instance; Step 304, sequentially allocating the cluster resource quota to the master instance and the corresponding plurality of slave instances from the schedulable cluster resource quota.
[0027] Specifically, in this embodiment, the instance quantity elasticity upper limit is calculated based on the ratio of the minimum resource requirement quantity and the schedulable cluster resource quota, and a certain number of slave instances, i.e., replicas, are configured for each master instance. The slave instances serve as backups for the master instances, which can quickly take over management tasks when the master instances fail or have abnormal loads. This allows the master instances to retain necessary running guarantees when resources are scarce, while the slave instances are in a pre-set or on-demand ready state to achieve second-level takeover, thereby significantly reducing the loss of availability caused by single-point failures of the management plane. In addition, limiting and dynamically configuring the number of replicas based on the schedulable resource quota helps to obtain the required high availability capability with minimal additional consumption under limited resources, thereby improving system reliability and reducing hardware and operation and maintenance costs without relying on hardware redundancy.
[0028] In one embodiment, after sequentially allocating the cluster resource quota from the schedulable cluster resource quota to the master instance and the corresponding multiple slave instances, further comprising: scheduling the master instance to a first server node and activating; scheduling the corresponding multiple slave instances to one or more second server nodes in the server cluster distinguished from the first server node; controlling the master instance to send heartbeat data packets to the corresponding multiple slave instances based on a preset heartbeat period; in response to the slave instance not receiving the heartbeat data packet within a plurality of consecutive heartbeat periods, determining that the master instance fails, and electing a leader instance from the existing multiple slave instances based on a preset election rule; activating the leader instance and synchronizing the to-be-executed server cluster management tasks of the master instance for the leader instance to continue executing.
[0029] Specifically, in the present embodiment, by deploying the master instance and its slave instances on different server nodes respectively, and based on heartbeat detection and a preset election rule, the readiness monitoring of the slave instances and the election of the leader instance are realized, so that when the master instance fails, the slave instance can be quickly elected and activated to take over the task without human intervention, ensuring the continuity of the management function; and unlike the redundancy scheme based on additional physical controllers and the like hardware, the virtual master-slave mechanism provided in the present embodiment saves hardware while realizing high-availability switching through the ready copy and automatic election mechanism, reducing the risk of single-point failure and shortening the fault recovery time, and improving the robustness and operation efficiency of the cluster management plane.
[0030] In one embodiment, according to the instance running information of the abnormal node virtual control instance, in combination with the abnormal node running information, the cluster resource quota of the abnormal node virtual control instance is adjusted to reduce the load of the abnormal node, comprising: parsing the instance running information to obtain the current resource usage rate of the instance, the instance task queue length, the instance historical task response delay and the instance historical resource consumption record; parsing the abnormal node running information to obtain the current resource utilization rate of the node and the historical load fluctuation record of the node; selecting an to-be-evaluated action from preset instance adjustment actions, the instance adjustment actions at least including expansion, contraction and migration; generating a node resource utilization rate prediction improvement value of the to-be-evaluated action according to the current resource usage rate of the instance, the instance task queue length, the instance historical resource consumption record and the current resource utilization rate of the node; generating an action predicted cost value of the to-be-evaluated action according to the current cluster resource quota of the abnormal node virtual control instance and the remaining resources of the abnormal node; According to the instance task queue length, the instance historical task response delay and the node historical load fluctuation record, a service quality prediction loss value of the to-be-evaluated action is generated; According to the node resource utilization prediction promotion value, the action prediction overhead value and the service quality prediction loss value, an action score of the to-be-evaluated action is generated; From the plurality of to-be-evaluated actions, a to-be-evaluated action corresponding to the maximum action score is selected as a target instance adjustment action; In response to the target instance adjustment action being expansion or contraction, a cluster resource quota adjustment amount of the abnormal node virtual control instance is determined according to the action score of the target instance adjustment action, and the abnormal node virtual control instance is adjusted correspondingly; In response to the target instance adjustment action being migration, a target migration node different from the abnormal node is selected from the plurality of server nodes, and the abnormal node virtual control instance is migrated to the target migration node.
[0031] Specifically, in this embodiment, the resource utilization prediction promotion value, the action overhead value and the service quality loss value are generated respectively in the process of adjusting the abnormal node virtual control instance, and are quantified as the action score, which can accurately compare and evaluate between a plurality of adjustment actions. Unlike the existing manual experience adjustment, the prediction and scoring model is introduced, so that the expansion, contraction or migration decision can be completed based on data driving, thereby ensuring that the selected action can both relieve the node load and not introduce too high additional overhead or service quality decline, and realizing the intelligentization and optimization of the abnormal node resource adjustment.
[0032] In one embodiment, the to-be-evaluated action is denoted as a, and a corresponding action prediction parameter set P a is introduced for predicting the action execution effect, wherein the action prediction parameter set is not the final execution parameter, but an idealized parameter for scoring calculation. For example, the expansion action can be assumed to increase a fixed proportion of resource quota, and the migration action can be assumed to migrate to a target node with sufficient resources. According to the instance current resource usage, the instance task queue length, the instance historical resource consumption record and the node current resource utilization, a node resource utilization prediction promotion value of the to-be-evaluated action is generated, which is denoted as: ; Wherein, ΔR(a, P a ) represents the node resource utilization prediction promotion value of the abnormal node virtual control instance in the abnormal node after the to-be-evaluated action a is executed, R(t) is the resource usage of the abnormal node virtual control instance at the historical time t, Q is the instance task queue length of the abnormal node virtual control instance, and a is the weight coefficient of the task queue length to the resource pressure, node historical resource utilization of the abnormal node, k R (a, P a ) is an action adjustment coefficient, which depends on the action to be evaluated and the action prediction parameter set, for example, the expansion or contraction can adopt a linear or exponential adjustment function, and the migration action can be determined according to the target node resource availability ratio; According to the current cluster resource quota of the abnormal node virtual control instance and the remaining resources of the abnormal node, an action prediction overhead value of the action to be evaluated is generated, which is expressed as: ; Wherein, C(a, P a ) represents the action prediction overhead value of the abnormal node virtual control instance executed by the action to be evaluated a, R pred (a, P a ) represents the additional cluster resource quota required after the abnormal node virtual control instance is executed by the action a under the action prediction parameter, which is determined by the current cluster resource quota, and R avail (a, P a ) represents the remaining resources of the abnormal node, and the larger the action prediction overhead value means the higher the action pressure on the node resources, so as to reflect the potential negative impact in the action score; According to the instance task queue length, the instance historical task response delay and the node historical load fluctuation record, a service quality prediction loss value of the action to be evaluated is generated, which is expressed as: ; Wherein, L q (a, P a ) represents the service quality prediction loss value of the abnormal node virtual control instance executed by the action to be evaluated a, T target (a, P pred ) represents the predicted task response time of the abnormal node virtual control instance executed by the action to be evaluated, which is predicted by combining the action parameters, the instance task queue length, the instance historical task response delay and the node historical load fluctuation record, and is preferably generated by using a machine learning model, T target is the target response time, and T pred is the prediction time window length; According to the node resource utilization prediction promotion value, the action prediction overhead value and the service quality prediction loss value, an action score of the action to be evaluated is generated, which is expressed as: , wherein S(a, P a ) represents the action score of the abnormal node virtual control instance executed by the action to be evaluated.
[0033] In a further embodiment, from the plurality of actions to be evaluated, the action to be evaluated corresponding to the maximum action score is selected as the target instance adjustment action, and then the method further comprises: obtain the action prediction parameter set of the target instance adjustment action as the reference action set, and the reference action parameter set includes parameters such as resource quota adjustment amount, adjustment execution time, migration target node, migration execution time, etc.; When the target instance adjustment action is expansion or contraction, the cluster resource adjustment amount of the abnormal node virtual control instance is determined according to the resource quota adjustment amount in the reference action parameter set combined with the weight coefficient determined by the action score, and the adjustment is performed at the adjustment execution time specified in the reference action parameter set; According to the migration target node and the migration execution time in the reference action parameter set, the abnormal node virtual control instance is migrated to the target node, and the migration operation is completed at the specified time.
[0034] Specifically, in this embodiment, by introducing the action prediction parameter set and the score mechanism based on the prediction, when the abnormal node virtual control instance appears resource pressure or service quality decline trend, the influence of multiple possible actions on node resource utilization, action overhead and service quality can be evaluated in advance, so as to select the optimal action for execution. Compared with the traditional method of adjusting only based on the current resource or static rule, the embodiment can realize dynamic and intelligent resource management and load regulation, which can not only reduce the risk of node overload, but also ensure the stability of service quality, improve the overall resource utilization and task processing efficiency of the cluster, and at the same time, the action execution strength and time can be adjusted according to the action score, which enhances the accuracy and controllability of the operation.
[0035] In one embodiment, in response to the target instance adjustment action being migration, a target migration node different from the abnormal node is selected from the plurality of server nodes, and the abnormal node virtual control instance is migrated to the target migration node, including: According to the node running information of the plurality of server nodes, a server node with the lowest current load index and lower than a preset safe receiving threshold is determined as a candidate node; Obtain the historical load fluctuation record of the candidate node, and determine whether the historical load fluctuation frequency of the candidate node is lower than a preset smooth threshold, if yes, the candidate node is selected as the target migration node, if not, the candidate node is determined as invalid and reselected; Suspend the task process of the abnormal node virtual control instance, and migrate the abnormal node virtual control instance to the target migration node; Reactivate the abnormal node virtual control instance on the target migration node, if the reactivation is successful, the abnormal node virtual control instance is determined as a normal running instance, if the reactivation fails, the abnormal node virtual control instance is rolled back to the abnormal node.
[0036] Specifically, in the present embodiment, by first screening the node with the lowest current load as a candidate during instance migration, and then combining the historical load fluctuation record to make a stability judgment, it can avoid migrating the instance to a node that is idle but has frequent fluctuations, thereby improving the running stability after migration. At the same time, by setting the pause, activation and rollback mechanism during migration, it ensures that even if the migration fails, it can quickly recover to the original node, reducing the risk of task interruption. This mechanism not only improves the reliability of migration operation, but also enhances the adaptive ability of the cluster when responding to exceptions.
[0037] In one embodiment, as shown in Figure 4 The server management method provided by the present application comprises the following steps: The management controller periodically monitors the node running information of the plurality of server nodes, and performs load evaluation based on preset running indicators, wherein the node running information comprises current resource utilization, historical load fluctuation record, etc., and the instance running information of the virtual control instance comprises current resource usage, task queue length, historical task response delay and resource consumption record; The management controller determines the node state according to the load evaluation result, and if the node load is normal, maintains the cluster resource quota of the current virtual control instance, and if the node load is abnormal, further dynamically adjusts the abnormal node and the abnormal node virtual control instance carried by the abnormal node; For the abnormal node, the management controller selects a to-be-evaluated action from the preset instance adjustment actions in combination with the abnormal node running information and the instance running information of the abnormal node virtual control instance, wherein the instance adjustment actions include expansion, contraction and migration; For each to-be-evaluated action, the node resource utilization prediction improvement value, the action prediction overhead value and the service quality prediction loss value are calculated respectively, and the action score is generated according to these intermediate indicators, and the management controller selects the to-be-evaluated action with the highest action score as the target instance adjustment action from the plurality of to-be-evaluated actions; When the target instance adjustment action is expansion or contraction, the management controller determines the cluster resource quota adjustment amount of the abnormal node virtual control instance according to the action score of the target instance adjustment action, and adjusts the cluster resource quota of the virtual control instance accordingly; When the target instance adjustment action is migration, the management controller selects a server node with the lowest current load and lower than a preset safe receiving threshold as a candidate node according to the node running information of the plurality of server nodes, and further judges whether the historical load fluctuation frequency of the candidate node is lower than a preset stable threshold, if the condition is met, the candidate node is determined as the target migration node, otherwise the candidate node is reselected; The management controller suspends the task process of the abnormal node virtual control instance, migrates it to a target migration node, reactivates the virtual control instance on the target migration node, and if the reactivation is successful, determines that the virtual control instance is a normally running instance, and if the reactivation fails, rolls back the virtual control instance to the original abnormal node.
[0038] Specifically, in this embodiment, through specific monitoring and instance adjustment processes, dynamic monitoring and evaluation of node load can be realized, and virtual control instances can be scaled up or down or migrated when the load is abnormal, thereby effectively reducing node load and improving cluster resource utilization, while maintaining existing resource quotas when the load is normal to ensure stable operation of the server cluster.
[0039] In one embodiment, the server cluster includes a plurality of server nodes, at least one virtual control instance is deployed on each server node, and a management controller is responsible for unified scheduling and management of the virtual control instances to realize node isolation, task scheduling, and access control. The server management method executed by the management controller further includes: receiving an access request from a server node, forwarding the access request to a corresponding virtual control instance, and performing preliminary authentication by the virtual control instance, wherein the authentication includes node identity verification, access credential verification, and timestamp confirmation. Specifically, the identity verification adopts two-way authentication based on an X.509 certificate, the certificate validity period is 90 to 180 days, and the timestamp tolerance is ±5 seconds; According to the request type and task priority, the access request is distributed to different isolated environments through the virtual control instance, wherein the isolated environments are divided into three categories: high, medium, and low, each category of isolated environment corresponds to independent virtual resources, the high level is allocated 4 to 8 CPU cores and 8 to 16 GB of memory, the medium level is allocated 2 to 4 CPU cores and 4 to 8 GB of memory, and the low level is allocated 1 to 2 CPU cores and 2 to 4 GB of memory; Further, the virtual control instance sets encryption policies for requests of different levels, wherein the high-level request enables an AES-256-GCM encryption channel, the key update period is 24 hours and a dynamic random salt value is generated, the medium-level request enables AES-128-CBC encryption, the key update period is 72 hours, and the low-level request allows plaintext transmission, but all write operations are recorded by the virtual control instance to access logs, which are summarized every hour and reported to the management controller; Periodically receive access behavior information reported by each virtual control instance, and when detecting five consecutive abnormal operations, such as access frequency exceeding a threshold or identity verification failure, execute an isolation strategy through the virtual control instance to migrate the corresponding node into an isolated environment, and the isolation time is 10 to 30 minutes, which can be dynamically adjusted according to the administrator's strategy.
[0040] Specifically, in the present embodiment, the management controller realizes logical isolation and security control of each server node, the virtual control instance is taken as a specific carrier scheduled and executed by the management controller, fine management of task scheduling, access control and security monitoring between nodes is ensured, security, controllability and real-time access of core business of the server cluster are improved, and the risk of single point failure and attack propagation is reduced.
[0041] Through the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment.
[0042] As shown in the above Figure 5 The embodiments of the present application also provide an electronic device, including a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above server management method embodiments.
[0043] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0044] The above describes in detail a server management method and an electronic device provided by the present application. The principles and embodiments of the present application are described by applying specific examples. The above description of the embodiments is only applicable to help understand the method and its core idea of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A server management method, characterized in that, The method includes: Based on the resource information of the server cluster, determine the total amount of cluster resource scheduling for the server cluster; Based on the total cluster resource scheduling amount, create multiple virtual control instances and allocate cluster resource quotas; The virtual control instance, which has been allocated the cluster resource quota, is scheduled to run on the server node, and the virtual control instance is activated to perform server cluster management tasks. Monitor the node operation information of multiple server nodes and perform load assessment to identify abnormal nodes and the virtual control instances of the abnormal nodes they support; Based on the instance operation information of the abnormal node virtual control instance, and in conjunction with the abnormal node operation information, the cluster resource quota of the abnormal node virtual control instance is adjusted to reduce the load on the abnormal node.
2. The server management method according to claim 1, characterized in that, The step of determining the total cluster resource scheduling amount for the server cluster based on the server cluster resource information includes: The resource information of the server cluster is analyzed to obtain the historical resource utilization rate of the server cluster, and the lowest historical resource utilization rate and the corresponding lowest historical resource utilization time of the server cluster within a preset historical observation period are determined. The maximum available resources of the cluster are obtained based on the remaining available resources of the server cluster at the lowest historical resource utilization time. Obtain the remaining available resources of the server cluster at the current moment to get the current available resources of the cluster; Using the current available resources of the cluster as a baseline, subtract the current available resources of the cluster from the maximum available resources of the cluster to obtain the elastic resource difference; Obtain the elastic resource release rate within the preset historical observation period, generate the average elastic resource release rate, and set the elastic compensation ratio based on the average elastic resource release rate. Multiplying the elastic compensation ratio by the elastic resource difference yields the elastic compensation resource amount, which is then added to the currently available cluster resources to obtain the total cluster resource scheduling amount.
3. The server management method according to claim 2, characterized in that, The step of creating multiple virtual control instances and allocating cluster resource quotas based on the total cluster resource scheduling amount includes: Based on the operational requirements of the virtual control instance, the minimum resource requirements of the virtual control instance are preset. Based on the ratio of the total cluster resource scheduling amount to the minimum resource requirement, the upper limit of the number of virtual control instances is determined, and the initial number of virtual control instances is generated in combination with the preset security creation ratio. Based on the initial number of virtual control instances, multiple virtual control instances are created; From the currently available resources of the cluster, allocate a base resource amount equal to the minimum resource requirement to each of the virtual control instances in sequence, and obtain the current remaining available resources of the cluster; The schedulable cluster resource quota is obtained based on the current remaining available resources of the cluster and the elastic compensation resource amount; Based on preset task importance weights, the cluster resource quota is sequentially allocated from the schedulable cluster resource quota to multiple virtual control instances.
4. The server management method according to claim 1, characterized in that, The step of scheduling the virtual control instance, which has been allocated the cluster resource quota, to the server node for operation and activating the virtual control instance includes: The resource requirements of the virtual control instance are determined based on the cluster resource quota allocated to the virtual control instance. Based on the resource demand, the multiple virtual control instances are sorted in descending order to generate an instance scheduling sequence, and the virtual control instance at the beginning of the instance scheduling sequence is selected as the current instance to be scheduled. Obtain the current available resource information of multiple server nodes, and determine the load capacity of multiple server nodes; In response to the fact that the load capacity of one or more of the server nodes is greater than or equal to the resource requirement of the currently scheduled instance, the server node with the largest load capacity is selected as the target node, the currently scheduled instance is scheduled to the target node and activated, if the activation is successful, the currently scheduled instance is removed from the instance scheduling sequence and the load capacity of the target node is updated, if the activation fails, the scheduling is canceled and the target node is reselected. In response to the absence of a server node whose load capacity is greater than or equal to the resource requirement of the currently scheduled instance, the currently scheduled instance is suspended as a standby instance, and the scheduling of the virtual control instance preceding the instance in the instance scheduling sequence is executed until a server node whose load capacity meets the resource requirement of the standby instance is detected, at which point the standby instance is woken up and scheduled.
5. The server management method according to claim 1, characterized in that, The monitoring of node operation information of multiple server nodes and the performance of load assessment to identify abnormal nodes include: Based on preset operating indicators, the node operating information is analyzed to obtain multiple operating indicator parameters; The operation indicator parameters are compared with the preset operation indicator safety thresholds. If one or more operation indicator parameters are greater than the corresponding operation indicator safety threshold, the server node is determined to be an abnormal node, and the corresponding abnormal operation indicator parameters are recorded.
6. The server management method according to claim 3, characterized in that, The method of allocating cluster resources from the schedulable cluster resource quota to multiple virtual control instances based on preset task importance weights further includes: The upper limit of the instance number elasticity is obtained based on the ratio of the schedulable cluster resource quota to the minimum resource requirement; The number of instance replicas is determined based on the initial number of virtual control instances and the elastic upper limit of the number of instances; Set the virtual control instance as the master instance, and create multiple slave instances for the master instance to meet the required number of instance replicas; The cluster resource quota is allocated sequentially from the schedulable cluster resource quota to the master instance and the corresponding multiple slave instances.
7. The server management method according to claim 6, characterized in that, After allocating the cluster resource quota from the schedulable cluster resource quota to the master instance and the corresponding plurality of slave instances in sequence, the method further includes: The primary instance is scheduled to the first server node and activated. The corresponding multiple instances are scheduled to one or more second server nodes in the server cluster that are distinct from the first server node; Based on a preset heartbeat cycle, the master instance is controlled to send heartbeat data packets to multiple corresponding slave instances; If the slave instance fails to receive the heartbeat data packet within multiple consecutive heartbeat cycles, the master instance is determined to be invalid, and a leader instance is elected from the existing slave instances based on a preset election rule. The leader instance is activated, and the pending server cluster management tasks of the master instance are synchronized for the leader instance to continue execution.
8. The server management method according to claim 1, characterized in that, The step of adjusting the cluster resource quota of the abnormal node virtual control instance to reduce the load on the abnormal node, based on the instance operation information of the abnormal node virtual control instance and in combination with the abnormal node operation information, includes: Parse the instance runtime information to obtain the instance's current resource utilization rate, instance task queue length, instance historical task response latency, and instance historical resource consumption records; The abnormal node's operational information is analyzed to obtain the node's current resource utilization and historical load fluctuation records. Select the action to be evaluated from the preset instance adjustment actions, wherein the instance adjustment actions include at least expansion, reduction and migration; Based on the instance's current resource utilization rate, the instance's task queue length, the instance's historical resource consumption records, and the node's current resource utilization rate, generate a predicted improvement value for the node's resource utilization rate for the action to be evaluated. Based on the current cluster resource quota of the abnormal node virtual control instance and the remaining resources of the abnormal node, generate the action prediction cost value of the action to be evaluated. Based on the instance task queue length, the instance historical task response latency, and the node historical load fluctuation record, generate the service quality prediction loss value of the action to be evaluated. Based on the predicted improvement value of node resource utilization, the predicted cost value of action, and the predicted loss value of service quality, an action score is generated for the action to be evaluated. From the multiple actions to be evaluated, select the action with the highest action score as the target instance to adjust the action; In response to the target instance adjustment action being either expansion or reduction, the cluster resource quota adjustment amount of the abnormal node virtual control instance is determined based on the action score of the target instance adjustment action, and the abnormal node virtual control instance is adjusted accordingly. In response to the target instance adjustment action being migration, a target migration node that is distinct from the abnormal node is selected from the plurality of server nodes, and the virtual control instance of the abnormal node is migrated to the target migration node.
9. The server management method according to claim 8, characterized in that, In response to the target instance adjustment action being migration, a target migration node, distinct from the abnormal node, is selected from the plurality of server nodes, and the virtual control instance of the abnormal node is migrated to the target migration node, including: Based on the node operation information of multiple server nodes, the server node with the lowest current load index and below the preset security reception threshold is determined as a candidate node. Obtain the historical load fluctuation records of the candidate node, and determine whether the historical load fluctuation frequency of the candidate node is lower than a preset stability threshold. If so, the candidate node is selected as the target migration node; otherwise, the candidate node is deemed invalid and reselected. Suspend the task process of the abnormal node virtual control instance and migrate the abnormal node virtual control instance to the target migration node; The abnormal node virtual control instance is reactivated on the target migration node. If the activation is successful, the abnormal node virtual control instance is determined to be a normal operating instance. If the activation fails, the abnormal node virtual control instance is rolled back to the abnormal node.
10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the server management method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Method and system for dynamic scheduling of virtual resources in cloud computing network
CN102170474A
Resource management method and multiple-node cluster device
CN103501242A
Virtual server Virtual CPU resource monitoring and dynamic allocation method
CN103729254A
A shared virtual resource pool share scheduling method and system
CN109597674A
Method and device for simulating cloud physical host by virtual machine based on cloud platform
CN112667363A