Method, device and storage medium for running adjustment of server

By receiving and packaging the device operation data of the server and generating a cluster dynamic adjustment strategy, the problem that the single-node mode cannot cope with complex scenario scheduling is solved, global resource scheduling and adaptation to complex scenarios are realized, the communication rate and security are improved, the bandwidth occupancy rate is reduced, and the cooling efficiency is improved.

CN120492173BActive Publication Date: 2025-10-14INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510954819.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-14
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

The existing technology monitors the operating status of the local server in a single-node mode, which is only suitable for local server resource scheduling and cannot meet the scheduling requirements of complex scenarios.

Method used

By receiving the equipment operation data collected by the baseboard management controller of each server, the data is packaged and processed to generate an operation data set, and the cluster dynamic adjustment strategy is generated using the policy analysis engine, and the server's operation status is adjusted through the policy execution engine.

Benefits of technology

It realizes global server resource scheduling, can cope with the scheduling needs of complex scenarios, improves the communication rate and data transmission security, reduces bandwidth occupancy, and improves cooling efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492173B_ABST
    Figure CN120492173B_ABST
Patent Text Reader

Abstract

The application discloses a server operation adjustment method and device, equipment and a storage medium, and relates to the technical field of servers, and comprises the following steps: receiving equipment operation data collected by a baseboard management controller of each server; performing packing processing on the equipment operation data of each server to obtain corresponding operation data sets; inputting the operation data sets corresponding to the servers into a strategy analysis engine, so that a strategy generation module deployed in the strategy analysis engine generates a cluster dynamic adjustment strategy according to the operation data sets corresponding to the servers; and inputting the cluster dynamic adjustment strategy into a strategy execution engine, so that a strategy execution module deployed in the strategy execution engine adjusts the operation states of the servers according to the cluster dynamic adjustment strategy, so that the server resource scheduling suitable for the whole world can cope with the scheduling requirements of complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of servers, and particularly relates to a server running adjustment method and device, equipment and a storage medium. BACKGROUND

[0002] Server resource scheduling refers to, in a server system or a computer system, reasonably allocating and scheduling various resources according to resource requirements of tasks and resource conditions of the system, so as to realize efficient task execution and resource management. The resource scheduling of the server is mainly to realize the allocation and scheduling of resources by adjusting the running of the server, and therefore, the running adjustment of the server is particularly important in the server resource scheduling.

[0003] In the related art, the current server running adjustment method mainly monitors the running state of the local server through the local baseboard management controller of the server, and adjusts the running state of the local server according to the running state, so as to realize the server resource scheduling, which belongs to a single-node control mode. However, in the related art, the way of monitoring the running state of the local server through the single-node mode is only suitable for local server resource scheduling and cannot meet the scheduling requirements of complex scenarios. SUMMARY

[0004] The present application provides a server running adjustment method, device, equipment and a storage medium to at least solve the problem that the way of monitoring the running state of the local server through the single-node mode in the related art is only suitable for local server resource scheduling and cannot meet the scheduling requirements of complex scenarios.

[0005] The present application provides a server running adjustment method, comprising:

[0006] sending a data collection instruction to a baseboard management controller corresponding to each server, so that the baseboard management controller collects device running data of a plurality of devices in the local server;

[0007] receiving the device running data collected by the baseboard management controller of each server in real time;

[0008] packaging the device running data of each server every preset time period to obtain a corresponding running data set;

[0009] inputting the running data set corresponding to each server into a strategy analysis engine, so that a strategy generation module deployed in the strategy analysis engine generates a cluster dynamic adjustment strategy according to the running data set corresponding to each server;

[0010] inputting the cluster dynamic adjustment strategy into a strategy execution engine, so that a strategy execution module deployed in the strategy execution engine adjusts the running state of each server according to the cluster dynamic adjustment strategy.

[0011] The application further provides a server operation adjustment device, comprising:

[0012] a sending module configured to send data collection instructions to the baseboard management controllers corresponding to the servers, so that the baseboard management controllers collect device operation data of a plurality of devices in the local servers;

[0013] a receiving module configured to receive the device operation data collected by the baseboard management controllers of the servers in real time;

[0014] a packing module configured to pack the device operation data of the servers every preset time period to obtain corresponding operation data sets;

[0015] a first input module configured to input the operation data sets corresponding to the servers into a strategy analysis engine, so that a strategy generation module deployed in the strategy analysis engine generates a cluster dynamic adjustment strategy according to the operation data sets corresponding to the servers;

[0016] a second input module configured to input the cluster dynamic adjustment strategy into a strategy execution engine, so that a strategy execution module deployed in the strategy execution engine adjusts the operation states of the servers according to the cluster dynamic adjustment strategy.

[0017] The application further provides an electronic device, comprising a memory configured to store a computer program and a processor configured to execute the computer program to implement the steps of the server operation adjustment method.

[0018] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the server operation adjustment method.

[0019] The server operation adjustment method, device, equipment and storage medium provided by the application can receive the device operation data collected by the baseboard management controllers of the servers, pack the device operation data of the servers to obtain corresponding operation data sets, input the operation data sets corresponding to the servers into a strategy analysis engine, so that a strategy generation module deployed in the strategy analysis engine generates a cluster dynamic adjustment strategy according to the operation data sets corresponding to the servers, input the cluster dynamic adjustment strategy into a strategy execution engine, so that a strategy execution module deployed in the strategy execution engine adjusts the operation states of the servers according to the cluster dynamic adjustment strategy, determine the operation data sets of the servers, and determine the cluster dynamic adjustment strategy according to the operation data sets, so as to adjust the operation states of the server cluster, so that the server resource scheduling suitable for the whole cluster can cope with the scheduling requirements of complex scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described in the following are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0021] Figure 1 The application scenario of the server running adjustment method provided by the embodiments of the present application is shown in the figure.

[0022] Figure 2 The flow of the server running adjustment method provided by the embodiments of the present application is shown in the figure. Figure 1

[0023] Figure 3 The flow of the server running adjustment method provided by the embodiments of the present application is shown in the figure. Figure 2

[0024] Figure 4 The structure of the server running adjustment device provided by the embodiments of the present application is shown in the figure.

[0025] Figure 5 The hardware structure of the electronic device provided by the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION

[0026] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort fall within the protection scope of the present application.

[0027] It should be noted that in the description of the present application, the terms “include”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, not to describe a specific order or sequence.

[0028] ​​Server resource scheduling refers to the reasonable allocation and scheduling of various resources in a server system or computer system based on the resource requirements of the task and the resource situation of the system, so as to achieve efficient task execution and resource management. The resource scheduling of the server is mainly achieved by adjusting the operation of the server to achieve resource allocation and scheduling. Therefore, the operation adjustment of the server is particularly important in server resource scheduling. In the related art, the current server operation adjustment method mainly monitors the operation status of the local server through the local baseboard management controller of the server, and adjusts the operation status of the local server according to the operation status to achieve server resource scheduling, which belongs to a single-node control mode. However, in the related art, the method of monitoring the operation status of the local server through the single-node mode is only suitable for local server resource scheduling and cannot meet the scheduling needs of complex scenarios.

[0029] In order to solve the above technical problems, the embodiments of the present application propose the following technical concepts: the inventor takes into account the device operation data received from each server, packages and processes the device operation data of each server, obtains the corresponding operation data set, uses the policy generation module in the policy analysis engine to process the operation data set of each server, generates a cluster dynamic adjustment strategy, uses the policy execution module in the policy execution engine to adjust the operation status of each server according to the cluster dynamic adjustment strategy, determines the operation data set of each server, and determines the cluster dynamic adjustment strategy according to each operation data set to adjust the operation status of the server cluster, so that it is suitable for global server resource scheduling and can meet the scheduling requirements of complex scenarios.

[0030] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0031] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the server operation adjustment method depends, the specific application environment architecture or specific hardware architecture is described here. Figure 1 , Figure 1 Schematic diagram of an application scenario of the server operation adjustment method provided in an embodiment of the present application.

[0032] like Figure 1 As shown, the scenario includes: an electronic device 10 and servers 20.

[0033] The electronic device 10 may be a data center terminal or a main server.

[0034] Each server 20 is a cluster composed of multiple servers, wherein any server 20 includes a corresponding baseboard management controller 201 and multiple devices 202 .

[0035] The electronic device 10 sends a data collection instruction to the baseboard management controller 201 corresponding to each server 20, so that the baseboard management controller 201 collects device running data of a plurality of devices 202 in the local server; receives the device running data collected by the baseboard management controller 201 of each server 20 in real time; every preset time period, the device running data of each server 20 is packaged to obtain the corresponding running data set; the running data set corresponding to each server 20 is input to the strategy analysis engine, so that the strategy generation module deployed in the strategy analysis engine generates a cluster dynamic adjustment strategy according to the running data set corresponding to each server 20; the cluster dynamic adjustment strategy is input to the strategy execution engine, so that the strategy execution module deployed in the strategy execution engine adjusts the running state of each server 20 according to the cluster dynamic adjustment strategy. The following will be described in detail by using detailed embodiments.

[0036] Figure 2 Flowchart of the server running adjustment method provided by the embodiments of the application Figure 1 As shown in the method, the embodiments of the application provide a server running adjustment method, which is described in detail as follows. Figure 2

[0037] S201: Send a data collection instruction to the baseboard management controller corresponding to each server, so that the baseboard management controller collects device running data of a plurality of devices in the local server.

[0038] In this embodiment, the data collection instruction can be an energy efficiency query request, or other instructions.

[0039] Exemplarily, the energy efficiency query request is an interface endpoint for obtaining cluster power supply data, running parameters of a central processing unit, running parameters of a graphics processing unit, running parameters of a memory bank, total power consumption of the cluster, average power supply efficiency, and a cooling efficiency detail array.

[0040] In the cooling efficiency detail array, each object includes a rack identifier and a cooling efficiency score.

[0041] The baseboard management controller can be integrated into a non-volatile memory.

[0042] Exemplarily, the device running data of the plurality of devices is the power consumption of the central processing unit, the power consumption of the graphics processing unit, the power consumption of the memory bank, the power consumption of the whole machine, the core temperature of the central processing unit, the inlet and outlet temperature of the rack, the utilization rate of the central processing unit, the utilization rate of the graphics processing unit, and the network bandwidth.

[0043] Exemplarily, the power consumption, temperature and utilization rate of the central processing unit are collected by using a high-frequency collection method, and the other device running data is collected by using a low-frequency collection method. ​

[0044] The high-frequency acquisition mode can be any one of 1 second / time, 2 seconds / time or 3 seconds / time, or other frequencies.

[0045] The low-frequency acquisition mode can be any one of 5 seconds / time, 10 seconds / time or 15 seconds / time, or other frequencies.

[0046] In addition, high-frequency monitoring data is also included, for example, using UDP protocol to transmit data, and transmitting data fields with a change of more than 5%.

[0047] The UDP protocol is a user datagram protocol, which is a connectionless transmission layer protocol.

[0048] S202: Real-time receiving of equipment operation data collected by the baseboard management controller of each server.

[0049] Specifically, step S202 specifically includes:

[0050] S2021: Setting a two-way authentication transmission key control instruction of a preset hardware management interface and each server.

[0051] In this embodiment, the preset hardware management interface can be Redfish API, or other management interfaces.

[0052] The Redfish API refers to an open standard interface developed by the Distributed Management Task Force, which realizes unified remote management of hardware devices such as servers and storage through the baseboard management controller.

[0053] In this embodiment, the two-way authentication is HTTPS two-way authentication.

[0054] In this embodiment, the transmission key control instruction refers to a preset parameter set for precise control of equipment, systems or communication links.

[0055] S2022: Authenticating the transmission key control instruction to obtain the authenticated transmission key control instruction.

[0056] Specifically, the transmission key control instruction is authenticated by dynamic token and hardware-level root of trust to obtain the authenticated transmission key control instruction.

[0057] Illustratively, dynamic token authentication generates a one-time Token with a validity period of 60 seconds to prevent replay attacks.

[0058] Illustratively, the hardware-level root of trust authenticates by signing the transmission data with hardware to ensure data integrity.

[0059] S2023: Real-time receiving, according to the authenticated transmission key control instruction, the equipment running data collected by the baseboard management controller of each server through the preset hardware management interface.

[0060] S203: Packaging the equipment running data of each server every preset time period to obtain the corresponding running data set.

[0061] Specifically, step S203 specifically includes:

[0062] S2031: Compressing and encoding the equipment running data of each server every preset time period by using a preset data compression and encoding tool to obtain the processed equipment running data.

[0063] In this embodiment, the preset time period can be any of 1 hour, 5 hours or 10 hours, or other time.

[0064] Exemplarily, the central processor utilization rate, the image processor utilization rate and the memory bar occupancy rate are compressed by a preset ratio every 1 hour by using a preset data compression and encoding tool to obtain the compressed central processor utilization rate, the image processor utilization rate and the memory bar occupancy rate; and the compressed central processor utilization rate, the image processor utilization rate and the memory bar occupancy rate are encoded and combined to obtain an 8-bit composite field.

[0065] The preset ratio can be 1 / 3, 1 / 2 or other ratio.

[0066] S2032: Packaging the processed equipment running data into the corresponding running data set.

[0067] S204: Inputting the running data set corresponding to each server into the strategy analysis engine, so that the strategy generation module deployed in the strategy analysis engine generates a cluster dynamic adjustment strategy according to the running data set corresponding to each server.

[0068] In this embodiment, the strategy generation module includes a load prediction model and a dynamic strategy generation model; accordingly, step S204 specifically includes:

[0069] S2041: Inputting the running data set corresponding to each server into the strategy analysis engine, so that the strategy analysis engine performs the following steps:

[0070] S2042: Generating a global resource state map according to the running data set corresponding to each server.

[0071] Specifically, step S2042 specifically includes:

[0072] S20421: Perform graph database storage processing on the running data set corresponding to each server to obtain one or more graph information.

[0073] In this embodiment, each graph information can be node topology relationship, rack heat distribution coefficient, real-time global state, and other graph information.

[0074] S20422: Generate a global resource state graph according to each graph information.

[0075] S2043: Input the running data set corresponding to each server into a load prediction model for processing to obtain a cluster predicted load.

[0076] In this embodiment, the load prediction model can be a long short-term memory network model, or other models.

[0077] In addition, historical load data, temperature trend, and business cycle can also be used as auxiliary data and input into the load prediction model to assist in obtaining the cluster predicted load.

[0078] The cluster predicted load can be a load predicted within 5 minutes, 10 minutes, or other time periods.

[0079] S2044: Perform cooling efficiency evaluation according to the cluster predicted load to obtain a cooling efficiency evaluation value.

[0080] Specifically, perform cooling efficiency evaluation according to thermodynamic parameters and the cluster predicted load to obtain a cooling efficiency evaluation value.

[0081] The thermodynamic parameters can be inlet and outlet air temperature, air flow organization, or other parameters.

[0082] S2045: Input the global resource state graph, the cluster predicted load, and the cooling efficiency evaluation value into a dynamic strategy generation model for processing to generate a cluster dynamic adjustment strategy.

[0083] In addition, temperature gradient and electricity price time-sharing signal can also be used as auxiliary data and input into the dynamic strategy generation model to assist in obtaining the cluster dynamic adjustment strategy.

[0084] S205: Input the cluster dynamic adjustment strategy into a strategy execution engine, so that a strategy execution module deployed in the strategy execution engine adjusts the running state of each server according to the cluster dynamic adjustment strategy.

[0085] In this embodiment, the strategy execution module includes a load migration control model, a power consumption coordination model, and a strategy execution model; accordingly, step S205 specifically includes:

[0086] S2051: input the cluster dynamic adjustment strategy into the strategy execution engine, so that the strategy execution engine performs the following steps:

[0087] S2052: generate one or more load migration tasks corresponding to each server according to the cluster dynamic adjustment strategy through the load migration control model.

[0088] In this embodiment, each load migration task includes a bandwidth dynamic allocation task, a pre-copy improvement task, and a processor state migration task; accordingly, step S2052 is specifically: generating bandwidth dynamic allocation tasks, pre-copy improvement tasks, and processor state migration tasks corresponding to each server according to the cluster dynamic adjustment strategy.

[0089] The load migration control model is configured with a live migration tool and a virtualization layer interface.

[0090] The bandwidth dynamic allocation task is to evaluate the available bandwidth based on the TCP BBR algorithm and allocate the migration traffic in proportion, for example: when the network delay is greater than 50ms or the packet loss rate is greater than 2%, automatically reduce the speed by 20%.

[0091] The TCP BBR algorithm is a congestion control algorithm designed to optimize data transmission efficiency and network performance by accurately measuring network bottleneck bandwidth and round-trip propagation time, thereby increasing throughput and reducing delay.

[0092] The pre-copy improvement task is to dynamically adjust the dirty page transmission rate, and the migration time is reduced by 40% compared with the default algorithm.

[0093] The processor state migration task is to realize real-time transmission of NVIDIA vGPU hardware state, including video memory data, calculation context and other key parameters, support checkpoint continuous migration of AI training tasks, and interrupt time less than 100ms.

[0094] NVIDIA vGPU is a technology that virtualizes the resources of a physical GPU, allowing multiple virtual machines or users to share the same GPU while providing hardware acceleration and resource isolation.

[0095] S2053: generate one or more power consumption tasks corresponding to each slave server according to the cluster dynamic adjustment strategy through the power consumption coordination model.

[0096] In this embodiment, each power consumption task includes a deep sleep scheduling task, a cooling coordination task, and a hardware wake-up task; accordingly, step S2053 is specifically: generating deep sleep scheduling tasks, cooling coordination tasks, and hardware wake-up tasks corresponding to the slave servers according to the cluster dynamic adjustment strategy.

[0097] The deep sleep scheduling task is an example, for example, when the cluster load of three consecutive monitoring periods is lower than 30%, the adjacent servers are coordinated to enter a deep sleep state, in the deep sleep state, the central processor core, the processing external core module and most of the peripherals are completely powered off, only the wake-up pin and the memory refresh circuit are reserved, and the power consumption is close to 0W.

[0098] The monitoring period can be 5 minutes, 10 minutes or other periods.

[0099] The cooling coordination task is to turn off redundant cooling devices, such as part of the cabinet air conditioner, according to the consensus result of the consensus algorithm, and to ensure the consistency of multi-node state switching by using the consensus algorithm.

[0100] In addition, the consensus algorithm protocol is also expanded, that is, a new entry is added to record the node power consumption mode switching event, and the event is submitted after more than half of the nodes confirm.

[0101] The hardware wake-up task is to receive an external interrupt signal through the pin of the baseboard management controller, and then perform subsequent wake-up triggering, the delay from hibernation to wake-up is controlled within 10ms, and immediate response to critical events is supported.

[0102] In addition, the corresponding detection strategy is to send a heartbeat packet every 30 seconds, and if there is no response for three consecutive times, it is determined that the node is offline.

[0103] S2054: According to each load migration task and each power consumption task, a distributed task queue is generated.

[0104] Specifically, step S2054 is specifically: according to the bandwidth dynamic allocation task, the pre-copy improvement task, the processor state migration task, the deep sleep scheduling task, the cooling coordination task and the hardware wake-up task, a distributed task queue is generated.

[0105] S2055: The distributed task queue is executed through a strategy execution model, and the running state of each server is adjusted.

[0106] In conclusion, the server operation adjustment method provided in this embodiment can send data collection instructions to the corresponding baseboard management controllers of the servers, so that the baseboard management controllers collect the device operation data of the devices in the local servers; the device operation data collected by the baseboard management controllers of the servers is received in real time; every preset time period, the device operation data of the servers is packaged to obtain the corresponding operation data set; the operation data set of each server is input into the strategy analysis engine, so that the strategy generation module deployed in the strategy analysis engine generates a cluster dynamic adjustment strategy according to the operation data set of each server; the cluster dynamic adjustment strategy is input into the strategy execution engine, so that the strategy execution module deployed in the strategy execution engine adjusts the operation state of each server according to the cluster dynamic adjustment strategy. By determining the operation data set of each server and determining the cluster dynamic adjustment strategy according to each operation data set, the operation state of the server cluster is adjusted, so that the server resource scheduling suitable for the whole global can cope with the scheduling requirements of complex scenarios.

[0107] In addition, the server operation adjustment method provided in this embodiment improves the communication rate by setting the preset hardware management interface and the two-way authentication transmission key control instruction of each server.

[0108] In addition, the server operation adjustment method provided in this embodiment improves the security of data transmission by dynamically tokenizing and hardware-level root-of-trust authenticating the transmission key control instruction to obtain the authenticated transmission key control instruction.

[0109] In addition, the server operation adjustment method provided in this embodiment reduces the bandwidth occupancy rate by compressing and encoding the device operation data of each server to obtain the processed device operation data.

[0110] In addition, the server operation adjustment method provided in this embodiment improves the cooling efficiency by evaluating the cooling efficiency according to the cluster predicted load to obtain the cooling efficiency evaluation value.

[0111] Figure 3 The server operation adjustment method provided in this embodiment Figure 2 In this embodiment, in Figure 2 Based on the provided embodiments, the specific implementation method of the adjustment report generation process of the server operation adjustment after step S205 is described in detail. As shown in Figure 3 The method comprises:

[0112] S301: After completing the operation state adjustment of each server, the corresponding adjustment result is generated.

[0113] S302: Determine whether the adjustment result meets the preset adjustment expectation value.

[0114] In addition, if it is determined that the adjustment result does not meet the preset adjustment expectation value, the adjustment result is fed back to the load migration control model to optimize the load migration control model.

[0115] S303: If it is determined that the adjustment result meets the preset adjustment expectation value, the adjustment result is input to the energy efficiency report engine to perform the following steps:

[0116] S304: Extract one or more key features from the adjustment result.

[0117] In this embodiment, each key feature is a time series energy consumption, an energy consumption history record, a knowledge graph, and a strategy log key indicator.

[0118] In addition, the retrieval of the knowledge graph includes generating a rationalization suggestion in combination with an industry standard, for example, if the rack cooling efficiency is low, it is recommended to adjust the airflow organization or add a flow guide plate.

[0119] In addition, the key features can also be the energy consumption saved by migration, the temperature change trend, and the cooling efficiency value.

[0120] S305: Input each key feature into the energy efficiency report model for processing to obtain an adjustment report, wherein the energy efficiency report model is deployed to the energy efficiency report engine.

[0121] Specifically, each key feature is input into the energy efficiency report model and processed by a deep learning model with a self-attention mechanism, a knowledge graph model, and a report generation model to obtain an adjustment report.

[0122] The adjustment report can be a text report, a visual chart, or other forms of report.

[0123] S306: Detect and process the adjustment report to determine whether the adjustment report meets the preset expected report.

[0124] S307: If it is determined that the adjustment report meets the preset expected report, output the adjustment report.

[0125] To sum up, the server operation adjustment method provided in the embodiment generates a corresponding adjustment result after completing the operation state adjustment of each server, judges whether the adjustment result meets a preset adjustment expectation value, inputs the adjustment result into an energy efficiency report engine to perform the following steps if it is determined that the adjustment result meets the preset adjustment expectation value: extracts one or more key features from the adjustment result; inputs each key feature into an energy efficiency report model for processing to obtain an adjustment report, wherein the energy efficiency report model is deployed to the energy efficiency report engine; detects and processes the adjustment report to determine whether the adjustment report meets a preset expected report; and outputs the adjustment report if it is determined that the adjustment report meets the preset expected report, so that the operation adjustment of the server can be intuitively understood.

[0126] In addition, the server operation adjustment method provided in the embodiment extracts one or more key features from the adjustment result and inputs each key feature into a preset energy efficiency report model for processing to obtain an adjustment report, so that the adjustment report is more accurate.

[0127] It should be noted that the server operation adjustment method provided in the embodiment also includes redundancy design, abnormal recovery, node disconnection, and policy conflict.

[0128] The redundancy design is that a dual-policy engine hot standby has a state synchronization delay of less than 50 ms and a switching time of less than 200 ms; and the communication link has multiple paths and can automatically switch according to the packet loss rate.

[0129] The node disconnection is that any hardware device continuously fails to respond to a heartbeat for three times, is marked as unavailable, and triggers load migration.

[0130] The policy conflict is that the latest policy priority is based on a timestamp, and a conflict event is recorded to a blockchain.

[0131] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software and a necessary general hardware platform, of course, it can also be implemented by hardware, but in many cases the former is a better implementation.

[0132] Figure 4 The structure of the server operation adjustment device provided in the embodiment of the application is shown in FIG. 4. Figure 4 As shown in FIG. 4, the embodiment of the application further provides a server operation adjustment device, which includes a sending module 401, a receiving module 402, a packaging module 403, a first input module 404, and a second input module 405.

[0133] The sending module 401 is configured to send a data collection instruction to a baseboard management controller corresponding to each server, so that the baseboard management controller collects device operation data of a plurality of devices in the local server.

[0134] The receiving module 402 is configured to receive, in real time, the device running data collected by the baseboard management controller of each server.

[0135] The packaging module 403 is configured to package the device running data of each server every preset time period to obtain a corresponding running data set.

[0136] The first input module 404 is configured to input the running data set corresponding to each server into the strategy analysis engine, so that the strategy generation module deployed in the strategy analysis engine generates a cluster dynamic adjustment strategy according to the running data set corresponding to each server.

[0137] The second input module 405 is configured to input the cluster dynamic adjustment strategy into the strategy execution engine, so that the strategy execution module deployed in the strategy execution engine adjusts the running state of each server according to the cluster dynamic adjustment strategy.

[0138] In a possible implementation, the strategy generation module includes a load prediction model and a dynamic strategy generation model. Correspondingly, the first input module 404 specifically includes:

[0139] The first input unit is configured to input the running data set corresponding to each server into the strategy analysis engine, so that the strategy analysis engine performs the following steps:

[0140] The generating unit is configured to generate a global resource state map according to the running data set corresponding to each server.

[0141] The second input unit is configured to input the running data set corresponding to each server into the load prediction model for processing to obtain a cluster predicted load.

[0142] The evaluation unit is configured to perform cooling efficiency evaluation according to the cluster predicted load to obtain a cooling efficiency evaluation value.

[0143] The third input unit is configured to input the global resource state map, the cluster predicted load, and the cooling efficiency evaluation value into the dynamic strategy generation model for processing to generate a cluster dynamic adjustment strategy.

[0144] In a possible implementation, the strategy execution module includes a load migration control model, a power consumption coordination model, and a strategy execution model. Correspondingly, the second input module 405 specifically includes:

[0145] The input unit is configured to input the cluster dynamic adjustment strategy into the strategy execution engine, so that the strategy execution engine performs the following steps:

[0146] The first generating unit is configured to generate one or more load migration tasks corresponding to each server according to the cluster dynamic adjustment strategy through a load migration control model;

[0147] The second generating unit is configured to generate one or more power consumption tasks corresponding to each slave server according to the cluster dynamic adjustment strategy through a power consumption coordination model;

[0148] The third generating unit is configured to generate a distributed task queue according to the load migration tasks and the power consumption tasks.

[0149] The executing unit is configured to execute the distributed task queue through a strategy execution model to adjust the running state of each server.

[0150] In a possible implementation, the load migration tasks include a bandwidth dynamic allocation task, a pre-copy improvement task and a processor state migration task; the power consumption tasks include a deep sleep scheduling task, a cooling coordination task and a hardware wake-up task; accordingly, the first generating unit is specifically configured to generate the bandwidth dynamic allocation task, the pre-copy improvement task and the processor state migration task corresponding to each server according to the cluster dynamic adjustment strategy; the second generating unit is specifically configured to generate the deep sleep scheduling task, the cooling coordination task and the hardware wake-up task corresponding to the slave server according to the cluster dynamic adjustment strategy; and the third generating unit is specifically configured to generate the distributed task queue according to the bandwidth dynamic allocation task, the pre-copy improvement task, the processor state migration task, the deep sleep scheduling task, the cooling coordination task and the hardware wake-up task.

[0151] In a possible implementation, the receiving module 402 specifically includes:

[0152] The setting unit is configured to set a preset hardware management interface and a bidirectional authentication transmission key control instruction of each server.

[0153] The authentication unit is configured to authenticate the transmission key control instruction to obtain an authenticated transmission key control instruction.

[0154] The receiving unit is configured to receive, according to the authenticated transmission key control instruction, device running data collected by a baseboard management controller of each server in real time through the preset hardware management interface.

[0155] In a possible implementation, the packaging module 403 specifically includes:

[0156] The processing unit is configured to compress and encode the device running data of each server through a preset data compression and encoding tool every preset time period to obtain processed device running data.

[0157] The packaging unit is configured to package each processed device operation data into a corresponding operation data set.

[0158] In a possible implementation, the apparatus further includes:

[0159] The generating module is configured to generate a corresponding adjustment result after completing the adjustment of the operation state of each server.

[0160] The first judging module is configured to judge whether the adjustment result meets a preset adjustment expectation value.

[0161] The third input module is configured to input the adjustment result to the energy efficiency report engine if it is judged that the adjustment result meets the preset adjustment expectation value, so that the energy efficiency report engine performs the following steps:

[0162] The extracting module is configured to extract one or more key features from the adjustment result.

[0163] The processing module is configured to input each key feature into the energy efficiency report model for processing to obtain an adjustment report, wherein the energy efficiency report model is deployed to the energy efficiency report engine.

[0164] The second judging module is configured to detect and process the adjustment report to judge whether the adjustment report meets a preset expected report.

[0165] The output module is configured to output the adjustment report if it is judged that the adjustment report meets the preset expected report.

[0166] The features of the embodiments of the server operation adjustment apparatus can be referred to the related descriptions of the embodiments of the server operation adjustment method, which will not be repeated here.

[0167] Figure 5 The structural schematic diagram of the electronic device provided in the present application is shown in FIG. 1. Figure 5 As shown in FIG. 1, the electronic device provided in the present embodiment includes at least one processor 501 and a memory 502. Optionally, the electronic device further includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected through a bus.

[0168] In the specific implementation process, the at least one processor 501 executes the computer execution instructions stored in the memory 502, so that the at least one processor 501 executes the above-mentioned embodiments of the server operation adjustment method.

[0169] The specific implementation process of the processor 501 can be referred to the above-mentioned method embodiments, which has similar implementation principles and technical effects, and will not be repeated here in the present embodiment.

[0170] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as hardware processor execution, or executed by a combination of hardware and software modules in the processor.

[0171] The memory can include a random access memory (RAM), and can also include a non-volatile memory (NVM), such as at least one disk memory.

[0172] The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application does not limit to only one bus or one type of bus.

[0173] Embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above server operation adjustment method embodiments when running.

[0174] In an example embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0175] Embodiments of the present application also provide a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the steps in any of the above server operation adjustment method embodiments.

[0176] The embodiment of the present application further provides another computer program product, comprising a nonvolatile computer readable storage medium, the nonvolatile computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps in any of the server operation adjustment method embodiments.

[0177] Those skilled in the art can further understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in the above description in general terms. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0178] The above describes in detail the server operation adjustment method, device, equipment and storage medium provided by the present application. The principles and implementation modes of the present application are described by applying specific examples. The above example is only used to help understand the method and its core idea of the present application. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for adjusting the operation of a server, characterized in that: include: Sending data collection instructions to the baseboard management controller corresponding to each server, so that each baseboard management controller collects device operation data of multiple devices in the local server; Receive equipment operation data collected by the baseboard management controller of each server in real time; Every preset time period, the device operation data of each server is packaged and processed to obtain the corresponding operation data set; Input the running data set corresponding to each server into the policy analysis engine, so that the policy analysis engine performs the following steps: Generate a global resource status map based on the running data sets corresponding to each server; Input the running data set corresponding to each server into the load prediction model for processing to obtain the cluster predicted load; performing cooling efficiency evaluation according to the cluster predicted load to obtain a cooling efficiency evaluation value; Inputting the global resource state map, the cluster predicted load and the cooling efficiency evaluation value into a dynamic strategy generation model for processing to generate a cluster dynamic adjustment strategy; The cluster dynamic adjustment policy is input into the policy execution engine, so that the policy execution engine performs the following steps: Generate one or more load migration tasks corresponding to each server according to the cluster dynamic adjustment strategy through the load migration control model; Generate one or more power consumption tasks corresponding to each slave server according to the cluster dynamic adjustment strategy through the power consumption coordination model; Generate a distributed task queue based on each load migration task and each power consumption task; The distributed task queue is executed through a policy execution model to adjust the operating status of each server.

2. The server operation adjustment method according to claim 1, characterized in that: The load migration tasks include bandwidth dynamic allocation tasks, pre-copy improvement tasks and processor state migration tasks; the power consumption tasks include deep sleep scheduling tasks, cooling coordination tasks and hardware wake-up tasks; Accordingly, generating one or more load migration tasks corresponding to each server according to the cluster dynamic adjustment strategy includes: Generate the bandwidth dynamic allocation task, the pre-copy improvement task, and the processor state migration task corresponding to each server according to the cluster dynamic adjustment strategy; Accordingly, generating one or more power consumption tasks corresponding to each slave server according to the cluster dynamic adjustment strategy includes: Generate the deep sleep scheduling task, the cooling coordination task and the hardware wake-up task corresponding to the slave server according to the cluster dynamic adjustment strategy; Accordingly, generating a distributed task queue according to each load migration task and each power consumption task includes: A distributed task queue is generated according to the bandwidth dynamic allocation task, the pre-copy improvement task, the processor state migration task, the deep sleep scheduling task, the cooling coordination task and the hardware wake-up task.

3. The server operation adjustment method according to claim 1, characterized in that: The real-time reception of device operation data collected by the baseboard management controller of each server includes: Set up key control instructions for transmitting two-way authentication between the preset hardware management interface and each server; authenticating the transmission key control instruction to obtain an authenticated transmission key control instruction; According to the authenticated transmission key control instruction, the device operation data collected by the baseboard management controller of each server is received in real time through the preset hardware management interface.

4. The server operation adjustment method according to claim 1, characterized in that: The device operation data of each server is packaged and processed at every preset time period to obtain a corresponding operation data set, including: At preset time intervals, the device operation data of each server is compressed and encoded using a preset data compression and encoding tool to obtain processed device operation data; The processed device operation data are packaged into corresponding operation data sets.

5. The server operation adjustment method according to any one of claims 1 to 4, characterized in that: After inputting the cluster dynamic adjustment policy into the policy execution engine so that the policy adjustment module deployed in the policy execution engine adjusts the operating status of each server according to the cluster dynamic adjustment policy, the method further includes: After completing the adjustment of the operating status of each server, generate the corresponding adjustment results; Determining whether the adjustment result meets a preset adjustment expected value; If it is determined that the adjustment result meets the preset adjustment expected value, the adjustment result is input into the energy efficiency reporting engine, so that the energy efficiency reporting engine executes the following steps: extracting one or more key features from the adjustment result; Inputting each key feature into an energy efficiency reporting model for processing to obtain an adjustment report, wherein the energy efficiency reporting model is deployed to the energy efficiency reporting engine; Performing detection and processing on the adjustment report to determine whether the adjustment report meets the preset expected report; If it is determined that the adjustment report meets the preset expected report, the adjustment report is output.

6. A server operation adjustment device, characterized in that: include: A sending module is used to send data collection instructions to the baseboard management controller corresponding to each server, so that each baseboard management controller collects device operation data of multiple devices in the local server; A receiving module is used to receive the equipment operation data collected by the baseboard management controller of each server in real time; The packaging module is used to package the device operation data of each server at preset time periods to obtain the corresponding operation data set; The first input module is configured to input the operating data set corresponding to each server into the policy analysis engine, so that the policy analysis engine performs the following steps: generating a global resource status map based on the operating data set corresponding to each server; inputting the operating data set corresponding to each server into a load prediction model for processing to obtain a cluster predicted load; performing a cooling efficiency evaluation based on the cluster predicted load to obtain a cooling efficiency evaluation value; inputting the global resource status map, the cluster predicted load, and the cooling efficiency evaluation value into a dynamic policy generation model for processing to generate a cluster dynamic adjustment policy; The second input module is used to input the cluster dynamic adjustment strategy into the policy execution engine, so that the policy execution engine performs the following steps: through the load migration control model, one or more load migration tasks corresponding to each server are generated according to the cluster dynamic adjustment strategy; through the power consumption coordination model, one or more power consumption tasks corresponding to each slave server are generated according to the cluster dynamic adjustment strategy; based on each load migration task and each power consumption task, a distributed task queue is generated; through the policy execution model, the distributed task queue is executed to adjust the operating status of each server.

7. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the server operation adjustment method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the server operation adjustment method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Multi-cloud storage cluster management method and device, equipment and storage medium

    CN119376931A