Server operation adjustment method and device, equipment and storage medium
By receiving and processing the server's device operation data to generate a cluster dynamic adjustment strategy, the problem that the single-node mode cannot cope with complex scenario scheduling is solved, and global resource scheduling and efficiency improvement is achieved.
Patent Information
- Application Number
- CN202510954819.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-07-11
AI Technical Summary
In the prior art, the operating status of the local server is monitored through a single node mode, which only adapts to local server resource scheduling and cannot cope with the scheduling needs of complex scenarios.
By receiving the substrate management controller of each server, collecting device operation data, packaging and inputting it to the policy analysis engine to generate a cluster dynamic adjustment policy, and the policy execution engine performs adjustments to achieve global server resource scheduling.
It realizes server resource scheduling suitable for global, can cope with the scheduling needs of complex scenarios, and improves communication rate, data transmission security and cooling efficiency.
Smart Images

Figure CN120492173A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of server technology, and in particular to a method, apparatus, device, and storage medium for adjusting server operation. Background Art
[0002] Server resource scheduling refers to the rational allocation and scheduling of various resources within a server or computer system, based on the resource requirements of the task and the system's resource availability, to achieve efficient task execution and resource management. Server resource scheduling primarily achieves resource allocation and scheduling by adjusting the server's operation. Therefore, adjusting the server's operation is particularly important in server resource scheduling.
[0003] Current server operation adjustment methods in the related art primarily rely on the server's local baseboard management controller to monitor the local server's operating status and adjust the local server's operating status based on the status to implement server resource scheduling. This is a single-node control model. However, this single-node approach to monitoring the local server's operating status in the related art is only suitable for local server resource scheduling and cannot meet the scheduling requirements of complex scenarios. Summary of the Invention
[0004] The present application provides a server operation adjustment method, device, equipment and storage medium to at least solve the problem in the related art that the method of monitoring the operating status of the local server through a single-node mode is only suitable for local server resource scheduling and cannot meet the scheduling requirements of complex scenarios.
[0005] This application provides a method for adjusting the operation of a server, including:
[0006] Sending data collection instructions to the baseboard management controller corresponding to each server, so that each baseboard management controller collects device operation data of multiple devices in the local server;
[0007] Receive equipment operation data collected by the baseboard management controller of each server in real time;
[0008] Every preset time period, the device operation data of each server is packaged and processed to obtain the corresponding operation data set;
[0009] Inputting the running data set corresponding to each server into the policy analysis engine, so that the policy generation module deployed in the policy analysis engine generates a cluster dynamic adjustment policy based on the running data set corresponding to each server;
[0010] The cluster dynamic adjustment policy is input into the policy execution engine, so that the policy execution module deployed in the policy execution engine adjusts the operating status of each server according to the cluster dynamic adjustment policy.
[0011] The present application also provides a server operation adjustment device, comprising:
[0012] A sending module is used to send data collection instructions to the baseboard management controller corresponding to each server, so that each baseboard management controller collects device operation data of multiple devices in the local server;
[0013] A receiving module is used to receive the equipment operation data collected by the baseboard management controller of each server in real time;
[0014] The packaging module is used to package the device operation data of each server at preset time periods to obtain the corresponding operation data set;
[0015] A first input module is used to input the operating data set corresponding to each server into the policy analysis engine, so that the policy generation module deployed in the policy analysis engine generates a cluster dynamic adjustment policy based on the operating data set corresponding to each server;
[0016] The second input module is used to input the cluster dynamic adjustment policy into the policy execution engine, so that the policy execution module deployed in the policy execution engine adjusts the operating status of each server according to the cluster dynamic adjustment policy.
[0017] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server operation adjustment methods when executing the computer program.
[0018] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server operation adjustment methods are implemented.
[0019] The server operation adjustment method, apparatus, device and storage medium provided in the embodiments of the present application receive the device operation data collected by the baseboard management controller of each server; package and process the device operation data of each server to obtain a corresponding operation data set; input the operation data set corresponding to each server into the policy analysis engine, so that the policy generation module deployed in the policy analysis engine generates a cluster dynamic adjustment policy according to the operation data set corresponding to each server; input the cluster dynamic adjustment policy into the policy execution engine, so that the policy execution module deployed in the policy execution engine adjusts the operation status of each server according to the cluster dynamic adjustment policy, and adjusts the operation status of the server cluster by determining the operation data set of each server and determining the cluster dynamic adjustment policy according to each operation data set, so as to be suitable for global server resource scheduling and can cope with the scheduling requirements of complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A schematic diagram of an application scenario of the server operation adjustment method provided in an embodiment of the present application;
[0022] Figure 2 Schematic diagram of the process of the server operation adjustment method provided in the embodiment of the present application Figure 1 ;
[0023] Figure 3 Schematic diagram of the process of the server operation adjustment method provided in the embodiment of the present application Figure 2 ;
[0024] Figure 4 A schematic diagram of the structure of the server operation adjustment device provided in an embodiment of the present application;
[0025] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0027] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0028] Server resource scheduling refers to the reasonable allocation and scheduling of various resources in a server system or computer system based on the resource requirements of the task and the resource situation of the system, so as to achieve efficient task execution and resource management. The resource scheduling of the server is mainly achieved by adjusting the operation of the server to achieve resource allocation and scheduling. Therefore, the operation adjustment of the server is particularly important in server resource scheduling. In the related art, the current server operation adjustment method mainly monitors the operation status of the local server through the local baseboard management controller of the server, and adjusts the operation status of the local server according to the operation status to achieve server resource scheduling, which belongs to a single-node control mode. However, in the related art, the method of monitoring the operation status of the local server through the single-node mode is only suitable for local server resource scheduling and cannot meet the scheduling needs of complex scenarios.
[0029] In order to solve the above technical problems, the embodiments of the present application propose the following technical concepts: the inventor takes into account the device operation data received from each server, packages and processes the device operation data of each server, obtains the corresponding operation data set, uses the policy generation module in the policy analysis engine to process the operation data set of each server, generates a cluster dynamic adjustment strategy, uses the policy execution module in the policy execution engine to adjust the operation status of each server according to the cluster dynamic adjustment strategy, determines the operation data set of each server, and determines the cluster dynamic adjustment strategy according to each operation data set to adjust the operation status of the server cluster, so that it is suitable for global server resource scheduling and can meet the scheduling requirements of complex scenarios.
[0030] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0031] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the server operation adjustment method depends, the specific application environment architecture or specific hardware architecture is described here. Figure 1 , Figure 1 Schematic diagram of an application scenario of the server operation adjustment method provided in an embodiment of the present application.
[0032] like Figure 1 As shown, the scenario includes: an electronic device 10 and servers 20.
[0033] The electronic device 10 may be a data center terminal or a main server.
[0034] Each server 20 is a cluster composed of multiple servers, wherein any server 20 includes a corresponding baseboard management controller 201 and multiple devices 202 .
[0035] The electronic device 10 sends a data collection instruction to the baseboard management controller 201 corresponding to each server 20, so that each baseboard management controller 201 collects the device operation data of multiple devices 202 in the local server; receives the device operation data collected by the baseboard management controller 201 of each server 20 in real time; packages and processes the device operation data of each server 20 at preset time intervals to obtain a corresponding operation data set; inputs the operation data set corresponding to each server 20 into a policy analysis engine, so that the policy generation module deployed in the policy analysis engine generates a cluster dynamic adjustment policy based on the operation data set corresponding to each server 20; inputs the cluster dynamic adjustment policy into a policy execution engine, so that the policy execution module deployed in the policy execution engine adjusts the operation status of each server 20 according to the cluster dynamic adjustment policy. A detailed embodiment is used below to explain the details.
[0036] Figure 2 Schematic diagram of the process of the server operation adjustment method provided in the embodiment of the present application Figure 1 ,like Figure 2 As shown, an embodiment of the present application provides a method for adjusting the operation of a server, and the method is described in detail as follows:
[0037] S201: Sending a data collection instruction to a baseboard management controller corresponding to each server, so that each baseboard management controller collects device operation data of multiple devices in the local server.
[0038] In this embodiment, the data collection instruction may be an energy efficiency query request or other instructions.
[0039] Exemplarily, the energy efficiency query request is: obtaining the interface endpoint of the cluster power data, the operating parameters of the central processing unit, the operating parameters of the graphics processing unit, the operating parameters of the memory bar, the total power consumption of the cluster, the average power utilization efficiency and the cooling efficiency details array.
[0040] Each object in the cooling efficiency details array includes: a rack identifier and a cooling efficiency score.
[0041] The baseboard management controller may be integrated into the non-volatile memory.
[0042] Exemplarily, the device operation data of multiple devices include the power consumption of the central processing unit, the power consumption of the graphics processing unit, the power consumption of the memory bar, the power consumption of the entire machine, the core temperature of the central processing unit, the inlet and outlet temperatures of the rack, the utilization rate of the central processing unit, the utilization rate of the graphics processing unit and the network bandwidth.
[0043] For example, a high-frequency collection method is used for the power consumption, temperature and utilization rate of the central processing unit, and a low-frequency collection method is used for other equipment operation data.
[0044] The high-frequency acquisition method may be any frequency among 1 second / time, 2 seconds / time or 3 seconds / time, or other frequencies.
[0045] The low-frequency acquisition method may be any frequency among 5 seconds / time, 10 seconds / time or 15 seconds / time, or other frequencies.
[0046] In addition, it also includes high-frequency monitoring data, such as: using the UDP protocol to transmit data, and transmitting data fields with a change greater than five percent.
[0047] Among them, the UDP protocol is the User Datagram Protocol, which is a connectionless transport layer protocol.
[0048] S202: Receive device operation data collected by the baseboard management controller of each server in real time.
[0049] Specifically, step S202 includes:
[0050] S2021: Setting key control instructions for transmitting two-way authentication between the preset hardware management interface and each server.
[0051] In this embodiment, the preset hardware management interface may be a Redfish API or other management interfaces.
[0052] Among them, Redfish API refers to the open standard interface developed by the Distributed Management Task Force, which realizes a unified remote management interface for hardware devices such as servers and storage through the baseboard management controller.
[0053] In this embodiment, the two-way authentication is HTTPS two-way authentication.
[0054] In this embodiment, the transmission of key control instructions refers to a set of preset parameters used to achieve precise control of a device, system, or communication link.
[0055] S2022: Authenticate the transmission key control instruction to obtain the authenticated transmission key control instruction.
[0056] Specifically, dynamic token and hardware-level root of trust authentication are performed on the transmission key control instruction to obtain the authenticated transmission key control instruction.
[0057] For example, dynamic token authentication generates a one-time token with a validity period of 60 seconds to prevent replay attacks.
[0058] Exemplarily, hardware-level root of trust authentication is used to ensure data integrity by hardware signing the transmitted data.
[0059] S2023: Based on the authenticated transmission key control instructions, the device operation data collected by the baseboard management controller of each server is received in real time through the preset hardware management interface.
[0060] S203: Packaging the device operation data of each server at every preset time period to obtain a corresponding operation data set.
[0061] Specifically, step S203 includes:
[0062] S2031: At predetermined time intervals, the device operation data of each server is compressed and encoded using a predetermined data compression and encoding tool to obtain processed device operation data.
[0063] In this embodiment, the preset time period may be any one of 1 hour, 5 hours, or 10 hours, or other time periods.
[0064] Exemplarily, every hour, the CPU utilization, image processor utilization and memory bar occupancy are compressed at a preset ratio using a preset data compression and encoding tool to obtain compressed CPU utilization, image processor utilization and memory bar occupancy; the compressed CPU utilization, image processor utilization and memory bar occupancy are encoded and merged to obtain an 8-bit composite field.
[0065] The preset ratio may be 1 / 3, 1 / 2 or other ratios.
[0066] S2032: Packaging the processed device operation data into corresponding operation data sets.
[0067] S204: Inputting the operating data set corresponding to each server into the policy analysis engine, so that the policy generation module deployed in the policy analysis engine generates a cluster dynamic adjustment policy according to the operating data set corresponding to each server.
[0068] In this embodiment, the strategy generation module includes a load prediction model and a dynamic strategy generation model; accordingly, step S204 specifically includes:
[0069] S2041: Input the operating data set corresponding to each server into the policy analysis engine, so that the policy analysis engine performs the following steps:
[0070] S2042: Generate a global resource status map based on the operating data sets corresponding to each server.
[0071] Specifically, step S2042 includes:
[0072] S20421: The running data sets corresponding to each server are stored and processed in a graph database to obtain one or more graph information.
[0073] In this embodiment, each graph information may include node topology relationship, rack heat distribution coefficient, real-time global status, and other graph information.
[0074] S20422: Generate a global resource status map based on each map information.
[0075] S2043: Inputting the operating data set corresponding to each server into the load prediction model for processing to obtain the cluster predicted load.
[0076] In this embodiment, the load prediction model may be a long short-term memory network model or other models.
[0077] In addition, historical load data, temperature trends, and business cycles can be used as auxiliary data and input into the load forecasting model to assist in obtaining cluster predicted load.
[0078] The cluster predicted load may be a load predicted within 5 minutes, 10 minutes, or other time periods.
[0079] S2044: Evaluate cooling efficiency based on the cluster predicted load to obtain a cooling efficiency evaluation value.
[0080] Specifically, the cooling efficiency is evaluated based on thermodynamic parameters and cluster predicted load to obtain a cooling efficiency evaluation value.
[0081] The thermodynamic parameters may include inlet and outlet air temperatures, air flow organization, or other parameters.
[0082] S2045: Inputting the global resource status map, cluster predicted load, and cooling efficiency evaluation value into the dynamic strategy generation model for processing to generate a cluster dynamic adjustment strategy.
[0083] In addition, the temperature gradient and electricity price time-sharing signal can be used as auxiliary data and input into the dynamic strategy generation model to assist in obtaining the cluster dynamic adjustment strategy.
[0084] S205: Input the cluster dynamic adjustment policy into the policy execution engine, so that the policy execution module deployed in the policy execution engine adjusts the operating status of each server according to the cluster dynamic adjustment policy.
[0085] In this embodiment, the policy execution module includes a load migration control model, a power consumption coordination model, and a policy execution model; accordingly, step S205 specifically includes:
[0086] S2051: Input the cluster dynamic adjustment policy into the policy execution engine, causing the policy execution engine to execute the following steps:
[0087] S2052: Generate one or more load migration tasks corresponding to each server according to the load migration control model and the cluster dynamic adjustment strategy.
[0088] In this embodiment, each load migration task includes a bandwidth dynamic allocation task, a pre-copy improvement task and a processor state migration task; accordingly, step S2052 is specifically: generating a bandwidth dynamic allocation task, a pre-copy improvement task and a processor state migration task corresponding to each server according to the cluster dynamic adjustment strategy.
[0089] The load migration control model is configured with a hot migration tool and a virtualization layer interface.
[0090] The dynamic bandwidth allocation task is to evaluate available bandwidth based on the TCP BBR algorithm and allocate migration traffic proportionally. For example, when the network delay is greater than 50ms or the packet loss rate is greater than 2%, the speed is automatically reduced by 20%.
[0091] Among them, the TCP BBR algorithm is a congestion control algorithm that aims to optimize data transmission efficiency and network performance by accurately measuring network bottleneck bandwidth and round-trip propagation time, thereby improving throughput and reducing latency.
[0092] The pre-copy improvement task is to dynamically adjust the dirty page transfer rate, reducing migration time by 40% compared to the default algorithm.
[0093] Among them, the processor state migration task is to achieve real-time transparent transmission of NVIDIA vGPU hardware status, including key parameters such as video memory data and computing context, and support continuous migration of checkpoints for AI training tasks with an interruption time of less than 100ms.
[0094] Among them, NVIDIA vGPU is a technology that virtualizes the resources of a physical GPU, allowing multiple virtual machines or users to share the same GPU while providing hardware acceleration and resource isolation.
[0095] S2053: Generate one or more power consumption tasks corresponding to each slave server according to the cluster dynamic adjustment strategy through the power consumption coordination model.
[0096] In this embodiment, each power consumption task includes a deep sleep scheduling task, a cooling collaboration task, and a hardware wake-up task; accordingly, step S2053 is specifically: generating a deep sleep scheduling task, a cooling collaboration task, and a hardware wake-up task corresponding to the slave server according to the cluster dynamic adjustment strategy.
[0097] Among them, the deep sleep scheduling task, for example, is: when the cluster load is less than 30% for three consecutive monitoring cycles, the adjacent servers are coordinated to enter the deep sleep state. In the deep sleep state, the central processing unit core, processing core modules and most peripherals are completely powered off, leaving only the wake-up pin and memory refresh circuit, and the power consumption is close to 0W.
[0098] The monitoring period may be 5 minutes, 10 minutes or other periods.
[0099] Among them, the cooling collaboration task is: according to the consensus result of the consensus algorithm, shut down redundant cooling equipment, such as some cabinet air conditioners; use the consensus algorithm to ensure the consistency of multi-node state switching.
[0100] In addition, it also includes the extension of the consensus algorithm protocol: a new entry is added to record the node power consumption mode switching event, which is submitted after confirmation by more than half of the nodes.
[0101] Among them, the hardware wake-up task is: after receiving the external interrupt signal through the pin of the baseboard management controller, subsequent wake-up triggering is performed, and the delay from sleep to wake-up is controlled within 10ms, supporting immediate response to key events.
[0102] In addition, the corresponding detection strategy is to send a heartbeat packet every 30 seconds, and if there is no response for three consecutive times, the node is considered offline.
[0103] S2054: Generate a distributed task queue according to each load migration task and each power consumption task.
[0104] Specifically, step S2054 is specifically: generating a distributed task queue according to the bandwidth dynamic allocation task, the pre-copy improvement task, the processor state migration task, the deep sleep scheduling task, the cooling coordination task and the hardware wake-up task.
[0105] S2055: Execute the distributed task queue through the policy execution model to adjust the operating status of each server.
[0106] In summary, the server operation adjustment method provided in this embodiment sends data collection instructions to the baseboard management controller corresponding to each server, so that each baseboard management controller collects the device operation data of multiple devices in the local server; receives the device operation data collected by the baseboard management controller of each server in real time; packages and processes the device operation data of each server at preset time periods to obtain a corresponding operation data set; inputs the operation data set corresponding to each server into the policy analysis engine, so that the policy generation module deployed in the policy analysis engine generates a cluster dynamic adjustment policy according to the operation data set corresponding to each server; inputs the cluster dynamic adjustment policy into the policy execution engine, so that the policy execution module deployed in the policy execution engine adjusts the operation status of each server according to the cluster dynamic adjustment policy, and adjusts the operation status of the server cluster by determining the operation data set of each server and determining the cluster dynamic adjustment policy according to each operation data set, so as to make it suitable for global server resource scheduling and can cope with the scheduling requirements of complex scenarios.
[0107] In addition, the server operation adjustment method provided in this embodiment improves the communication rate by setting a preset hardware management interface and transmitting key control instructions for bidirectional authentication between each server.
[0108] In addition, the server operation adjustment method provided in this embodiment improves the security of data transmission by performing dynamic token and hardware-level trust root authentication on the transmission key control instructions to obtain authenticated transmission key control instructions.
[0109] In addition, the server operation adjustment method provided in this embodiment reduces bandwidth occupancy by compressing and encoding the device operation data of each server to obtain the processed device operation data.
[0110] In addition, the server operation adjustment method provided in this embodiment evaluates the cooling efficiency according to the cluster predicted load to obtain a cooling efficiency evaluation value, thereby improving the cooling efficiency.
[0111] Figure 3 Schematic diagram of the server operation adjustment method provided in the embodiment of the present application Figure 2 In the embodiment of the present application, Figure 2 Based on the embodiment provided, a detailed description is given of the specific implementation method of the process of generating the adjustment report on the operation adjustment of the server after step S205. Figure 3 As shown, the method includes:
[0112] S301: After completing the adjustment of the operating status of each server, a corresponding adjustment result is generated.
[0113] S302: Determine whether the adjustment result meets the preset adjustment expected value.
[0114] In addition, if it is determined that the adjustment result does not meet the preset adjustment expected value, the adjustment result is fed back to the load migration control model to optimize the load migration control model.
[0115] S303: If it is determined that the adjustment result meets the preset adjustment expected value, the adjustment result is input into the energy efficiency reporting engine to execute the following steps:
[0116] S304: Extract one or more key features from the adjustment result.
[0117] In this embodiment, the key features are time series energy consumption, energy consumption history records, knowledge graphs, and policy log key indicators.
[0118] In addition, the retrieval of the knowledge graph includes: generating rational suggestions based on industry standards, for example: if the rack cooling efficiency is low, it is recommended to adjust the airflow organization or add guide plates.
[0119] Additionally, key features may be energy savings from migration, temperature trends, and cooling efficiency values.
[0120] S305: Inputting each key feature into the energy efficiency report model for processing to obtain an adjustment report, wherein the energy efficiency report model is deployed to the energy efficiency report engine.
[0121] Specifically, each key feature is input into the energy efficiency report model, and processed through the deep learning model of the self-attention mechanism, the knowledge graph model and the report generation model to obtain an adjustment report.
[0122] The adjustment report may be a text report, a visual chart, or a report in other forms.
[0123] S306: Detect and process the adjustment report to determine whether the adjustment report complies with the preset expected report.
[0124] S307: If it is determined that the adjustment report meets the preset expected report, the adjustment report is output.
[0125] In summary, the server operation adjustment method provided in this embodiment generates a corresponding adjustment result after completing the adjustment of the operation status of each server; determines whether the adjustment result meets the preset adjustment expected value; if it is determined that the adjustment result meets the preset adjustment expected value, the adjustment result is input into the energy efficiency reporting engine to perform the following steps: extract one or more key features from the adjustment result; input each key feature into the energy efficiency reporting model for processing to obtain an adjustment report, wherein the energy efficiency reporting model is deployed to the energy efficiency reporting engine; detects and processes the adjustment report to determine whether the adjustment report meets the preset expected report; if it is determined that the adjustment report meets the preset expected report, the adjustment report is output, so that the operation adjustment status of the server can be intuitively understood.
[0126] In addition, the server operation adjustment method provided in this embodiment extracts one or more key features from the adjustment results, and inputs each key feature into a preset energy efficiency report model for processing to obtain an adjustment report, thereby making the adjustment report more accurate.
[0127] It should be noted that the server operation adjustment method provided in this embodiment also includes redundancy design, abnormal recovery, node loss and policy conflict.
[0128] Among them, the redundancy design is: dual-strategy engine hot standby, state synchronization delay less than 50ms, and switching time less than 200ms; the communication link has multiple paths and can automatically switch according to the packet loss rate.
[0129] Node disconnection occurs when any hardware device fails to respond to heartbeats for three consecutive times, marking it as unavailable and triggering load migration.
[0130] Policy conflicts are: the latest policy priority based on the timestamp, and conflict events are recorded in the blockchain.
[0131] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0132] Figure 4 This is a schematic diagram of the structure of the server operation adjustment device provided in the embodiment of the present application. Figure 4 As shown, an embodiment of the present application further provides a server operation adjustment device, including: a sending module 401, a receiving module 402, a packaging module 403, a first input module 404 and a second input module 405.
[0133] The sending module 401 is used to send a data collection instruction to the baseboard management controller corresponding to each server, so that each baseboard management controller collects the device operation data of multiple devices in the local server;
[0134] Receiving module 402, for receiving in real time the equipment operation data collected by the baseboard management controller of each server;
[0135] The packaging module 403 is used to package the device operation data of each server at every preset time period to obtain a corresponding operation data set;
[0136] The first input module 404 is used to input the operating data set corresponding to each server into the policy analysis engine, so that the policy generation module deployed in the policy analysis engine generates a cluster dynamic adjustment policy based on the operating data set corresponding to each server;
[0137] The second input module 405 is used to input the cluster dynamic adjustment policy into the policy execution engine, so that the policy execution module deployed in the policy execution engine adjusts the operating status of each server according to the cluster dynamic adjustment policy.
[0138] In a possible implementation, the strategy generation module includes a load prediction model and a dynamic strategy generation model; accordingly, the first input module 404 specifically includes:
[0139] The first input unit is configured to input the operating data set corresponding to each server into the policy analysis engine, so that the policy analysis engine performs the following steps:
[0140] A generating unit, configured to generate a global resource status map based on the running data sets corresponding to each server;
[0141] The second input unit is used to input the operating data set corresponding to each server into the load prediction model for processing to obtain the cluster predicted load;
[0142] An evaluation unit, configured to evaluate cooling efficiency according to cluster predicted load and obtain a cooling efficiency evaluation value;
[0143] The third input unit is used to input the global resource status map, cluster predicted load and cooling efficiency evaluation value into the dynamic strategy generation model for processing to generate a cluster dynamic adjustment strategy.
[0144] In a possible implementation, the policy execution module includes a load migration control model, a power consumption coordination model, and a policy execution model; accordingly, the second input module 405 specifically includes:
[0145] The input unit is used to input the cluster dynamic adjustment policy into the policy execution engine, so that the policy execution engine performs the following steps:
[0146] The first generating unit is configured to generate one or more load migration tasks corresponding to each server according to the cluster dynamic adjustment strategy through the load migration control model;
[0147] The second generating unit is configured to generate one or more power consumption tasks corresponding to each slave server according to the cluster dynamic adjustment strategy through the power consumption coordination model;
[0148] A third generating unit is configured to generate a distributed task queue according to each load migration task and each power consumption task;
[0149] The execution unit is used to execute the distributed task queue through the policy execution model and adjust the operating status of each server.
[0150] In one possible implementation, each load migration task includes a bandwidth dynamic allocation task, a pre-copy improvement task, and a processor state migration task; each power consumption task includes a deep sleep scheduling task, a cooling collaboration task, and a hardware wake-up task; accordingly, the first generation unit is specifically used to: generate the bandwidth dynamic allocation task, pre-copy improvement task, and processor state migration task corresponding to each server according to the cluster dynamic adjustment strategy; accordingly, the second generation unit is specifically used to: generate the deep sleep scheduling task, cooling collaboration task, and hardware wake-up task corresponding to the slave server according to the cluster dynamic adjustment strategy; accordingly, the third generation unit is specifically used to: generate a distributed task queue according to the bandwidth dynamic allocation task, pre-copy improvement task, processor state migration task, deep sleep scheduling task, cooling collaboration task, and hardware wake-up task.
[0151] In a possible implementation, the receiving module 402 specifically includes:
[0152] A setting unit, used to set the transmission key control instructions for the two-way authentication between the preset hardware management interface and each server;
[0153] An authentication unit, used to authenticate the transmission key control instruction and obtain the authenticated transmission key control instruction;
[0154] The receiving unit is used to receive the equipment operation data collected by the baseboard management controller of each server in real time through the preset hardware management interface according to the authenticated transmission key control instruction.
[0155] In a possible implementation, the packaging module 403 specifically includes:
[0156] The processing unit is used to compress and encode the device operation data of each server using a preset data compression and encoding tool at each preset time period to obtain processed device operation data;
[0157] The packaging unit is used to package the processed device operation data into corresponding operation data sets.
[0158] In a possible implementation, the apparatus further includes:
[0159] A generation module is used to generate corresponding adjustment results after completing the adjustment of the operating status of each server;
[0160] The first judgment module is used to judge whether the adjustment result meets the preset adjustment expected value;
[0161] The third input module is configured to input the adjustment result to the energy efficiency reporting engine if it is determined that the adjustment result meets the preset adjustment expected value, so that the energy efficiency reporting engine executes the following steps:
[0162] an extraction module, configured to extract one or more key features from the adjustment result;
[0163] a processing module, configured to input each key feature into an energy efficiency report model for processing to obtain an adjustment report, wherein the energy efficiency report model is deployed to an energy efficiency report engine;
[0164] The second judgment module is used to detect and process the adjustment report to determine whether the adjustment report meets the preset expected report;
[0165] The output module is used to output the adjustment report if it is determined that the adjustment report meets the preset expected report.
[0166] For the description of the features in the embodiment corresponding to the server operation adjustment device, please refer to the relevant description of the embodiment corresponding to the server operation adjustment method, which will not be repeated here.
[0167] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 5 As shown, the electronic device provided by this embodiment includes: at least one processor 501 and a memory 502. Optionally, the electronic device further includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus.
[0168] During the specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502, so that the at least one processor 501 executes the above-mentioned embodiment of the server operation adjustment method.
[0169] The specific implementation process of the processor 501 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0170] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.
[0171] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.
[0172] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.
[0173] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned server operation adjustment method embodiments when running.
[0174] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0175] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned server operation adjustment method embodiments are implemented.
[0176] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, implementing the steps in any of the above-mentioned server operation adjustment method embodiments.
[0177] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0178] The above is a detailed introduction to the operation adjustment method, device, equipment and storage medium of a server provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for adjusting the operation of a server, characterized in that: include: Sending data collection instructions to the baseboard management controller corresponding to each server, so that each baseboard management controller collects device operation data of multiple devices in the local server; Receive equipment operation data collected by the baseboard management controller of each server in real time; Every preset time period, the device operation data of each server is packaged and processed to obtain the corresponding operation data set; Inputting the operating data sets corresponding to each server into the policy analysis engine, so that the policy generation module deployed in the policy analysis engine generates a cluster dynamic adjustment policy based on the operating data sets corresponding to each server; The cluster dynamic adjustment policy is input into a policy execution engine, so that a policy execution module deployed in the policy execution engine adjusts the operating status of each server according to the cluster dynamic adjustment policy.
2. The server operation adjustment method according to claim 1, characterized in that: The strategy generation module includes a load prediction model and a dynamic strategy generation model; Accordingly, the operation data set corresponding to each server is input into the policy analysis engine, so that the policy generation module deployed in the policy analysis engine generates a cluster dynamic adjustment policy based on the operation data set corresponding to each server, including: Input the running data set corresponding to each server into the policy analysis engine, so that the policy analysis engine performs the following steps: Generate a global resource status map based on the running data sets corresponding to each server; Inputting the operating data set corresponding to each server into the load prediction model for processing to obtain the cluster predicted load; performing cooling efficiency evaluation according to the cluster predicted load to obtain a cooling efficiency evaluation value; The global resource state map, the cluster predicted load and the cooling efficiency evaluation value are input into the dynamic strategy generation model for processing to generate a cluster dynamic adjustment strategy.
3. The server operation adjustment method according to claim 1, characterized in that: The policy execution module includes a load migration control model, a power consumption coordination model and a policy execution model; Accordingly, the cluster dynamic adjustment policy is input into the policy execution engine, so that the policy execution module deployed in the policy execution engine adjusts the operating status of each server according to the cluster dynamic adjustment policy, including: The cluster dynamic adjustment policy is input into the policy execution engine, so that the policy execution engine performs the following steps: Generate one or more load migration tasks corresponding to each server according to the cluster dynamic adjustment strategy through the load migration control model; Generate one or more power consumption tasks corresponding to each slave server according to the cluster dynamic adjustment strategy through the power consumption coordination model; Generate a distributed task queue based on each load migration task and each power consumption task; The distributed task queue is executed through the policy execution model to adjust the operating status of each server.
4. The server operation adjustment method according to claim 3, characterized in that: The load migration tasks include bandwidth dynamic allocation tasks, pre-copy improvement tasks and processor state migration tasks; the power consumption tasks include deep sleep scheduling tasks, cooling coordination tasks and hardware wake-up tasks; Accordingly, generating one or more load migration tasks corresponding to each server according to the cluster dynamic adjustment strategy includes: Generate the bandwidth dynamic allocation task, the pre-copy improvement task, and the processor state migration task corresponding to each server according to the cluster dynamic adjustment strategy; Accordingly, generating one or more power consumption tasks corresponding to each slave server according to the cluster dynamic adjustment strategy includes: Generate the deep sleep scheduling task, the cooling coordination task and the hardware wake-up task corresponding to the slave server according to the cluster dynamic adjustment strategy; Accordingly, generating a distributed task queue according to each load migration task and each power consumption task includes: A distributed task queue is generated according to the bandwidth dynamic allocation task, the pre-copy improvement task, the processor state migration task, the deep sleep scheduling task, the cooling coordination task and the hardware wake-up task.
5. The server operation adjustment method according to claim 1, characterized in that: The real-time reception of device operation data collected by the baseboard management controller of each server includes: Set up key control instructions for transmitting two-way authentication between the preset hardware management interface and each server; authenticating the transmission key control instruction to obtain an authenticated transmission key control instruction; According to the authenticated transmission key control instruction, the device operation data collected by the baseboard management controller of each server is received in real time through the preset hardware management interface.
6. The server operation adjustment method according to claim 1, characterized in that: The device operation data of each server is packaged and processed at every preset time period to obtain a corresponding operation data set, including: At preset time intervals, the device operation data of each server is compressed and encoded using a preset data compression and encoding tool to obtain processed device operation data; The processed device operation data are packaged into corresponding operation data sets.
7. The server operation adjustment method according to any one of claims 1 to 6, characterized in that: After inputting the cluster dynamic adjustment policy into the policy execution engine so that the policy adjustment module deployed in the policy execution engine adjusts the operating status of each server according to the cluster dynamic adjustment policy, the method further includes: After completing the adjustment of the operating status of each server, generate the corresponding adjustment results; Determining whether the adjustment result meets a preset adjustment expected value; If it is determined that the adjustment result meets the preset adjustment expected value, the adjustment result is input into the energy efficiency reporting engine, so that the energy efficiency reporting engine executes the following steps: extracting one or more key features from the adjustment result; Inputting each key feature into an energy efficiency reporting model for processing to obtain an adjustment report, wherein the energy efficiency reporting model is deployed to the energy efficiency reporting engine; Performing detection and processing on the adjustment report to determine whether the adjustment report meets the preset expected report; If it is determined that the adjustment report meets the preset expected report, the adjustment report is output.
8. A server operation adjustment device, characterized in that: include: A sending module is used to send data collection instructions to the baseboard management controller corresponding to each server, so that each baseboard management controller collects device operation data of multiple devices in the local server; A receiving module is used to receive the equipment operation data collected by the baseboard management controller of each server in real time; The packaging module is used to package the device operation data of each server at preset time periods to obtain the corresponding operation data set; A first input module is used to input the operating data set corresponding to each server into the policy analysis engine, so that the policy generation module deployed in the policy analysis engine generates a cluster dynamic adjustment policy based on the operating data set corresponding to each server; The second input module is used to input the cluster dynamic adjustment policy into the policy execution engine, so that the policy execution module deployed in the policy execution engine adjusts the operating status of each server according to the cluster dynamic adjustment policy.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the server operation adjustment method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the server operation adjustment method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Energy storage cluster optimization scheduling method based on big data analysis
CN119294784A
Multi-cloud storage cluster management method and device, equipment and storage medium
CN119376931A
Data center energy optimization method and device based on large language model, medium and product
CN119829299A
Intelligent load balancing system with artificial intelligence in distributed cloud networks
DE202024105500U1