Operation control method and apparatus for substrate management controller
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]本申请实施例提供了一种基板管理控制器的运行控制方法及装置,以至少解决相关技术中对服务器运行数据的采集效率较低的问题
[0018]通过本申请,目标决策模型用于运算各种调节操作对于将输入的变化幅度调节到目标变化幅度所达到的调节效果,在基板管理控制器按照参考频率采集服务器运行数据的过程中,通过检测采集到的目标服务器运行数据的变化幅度的稳定程度,进而在目标服务器的运行数据变化幅度不稳定的情况下,通过使用目标决策模型检测多个调节操作中每个调节操作在目标服务器运行数据的当前变化幅度下对应的调节效果,进而根据各个调节操作在目标服务器运行数据的当前变化幅下的调节效果能够在多个调节操作中筛选出满足目标收益条件的调节操作,从而对参考频率进行调节,实现了根据运行数据采集过程中的变化幅度确定对采集数据的参考频率的动态调节,使得采集到的运行数据的变化幅度保持稳定,可以解决对服务器运行数据的采集效率较低的问题,达到提高对服务器的运行数据的采集效率的效果。
Smart Images

Figure CN117742816B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and more specifically, to a method and apparatus for operating a baseboard management controller. Background Technology
[0002] Baseboard Management Controllers (BMCs) are typically used in devices such as servers, switches, and smart network interface cards (NICs). Through configuration files, they collect data from threshold sensors on these devices, such as those monitoring voltage, current, temperature, and power consumption, enabling device monitoring and maintenance. Currently, in typical BMC sensor monitoring solutions, developers usually set a fixed polling period for each monitored sensor. If the period is set too long, the accuracy of the monitored data decreases; if the period is set too short, the polling frequency is too high, putting excessive pressure on the BMC processor and causing other functions to lag. Furthermore, in practical applications, sensors such as temperature sensors directly affect server cooling strategies and frequently change, requiring high-frequency data collection; while voltage sensors typically have relatively stable readings and do not require high-frequency data collection. Therefore, setting a uniform monitoring period for all these sensors is limited, as it is detrimental to data accuracy and places additional pressure on the BMC processor. Summary of the Invention
[0003] This application provides a method and apparatus for controlling the operation of a baseboard management controller, which at least solves the problem of low efficiency in collecting server operation data in related technologies.
[0004] According to one embodiment of this application, an operation control method for a baseboard management controller is provided, comprising:
[0005] In an exemplary embodiment, during the process of the baseboard management controller acquiring server operating data using a reference frequency, a stability parameter of the acquired target server operating data is detected, wherein the stability parameter indicates the stability of the change range of the target server operating data; when the stability parameter indicates that the change range of the target server operating data is unstable, a target decision model is used to detect a target benefit parameter corresponding to each of a plurality of adjustment operations under the current change range of the target server operating data, wherein the target decision model is used to calculate the adjustment effect achieved by executing each adjustment operation on a decision objective, the decision objective including adjusting the input change range to a target change range, and the target benefit parameter indicating the adjustment effect; the reference frequency is adjusted to a target frequency by using a target adjustment operation whose target benefit parameter among the plurality of adjustment operations satisfies the target benefit condition; and the baseboard management controller is controlled to acquire the server operating data using the target frequency.
[0006] Optionally, the step of detecting the target benefit parameter corresponding to each of the multiple adjustment operations under the current change range of the target server's running data through the target decision model includes: inputting the current change range as the current system state parameter into the target decision model, wherein the system state parameter set of the target decision model records multiple change ranges, the action parameter set of the target decision model records the multiple adjustment operations, and the transition probability parameter set of the target decision model records the corresponding change ranges, the target change range, each adjustment operation, and the transition probability, wherein the transition probability is the probability of instructing each adjustment operation to convert the change range into the target change range, and the target decision model is used to calculate the target benefit parameter corresponding to each action parameter through an activation function based on the transition probability corresponding to the current system state parameter; and obtaining multiple sets of corresponding adjustment operations and target benefit parameters output by the target decision model.
[0007] Optionally, the step of calculating the target benefit parameter corresponding to each action parameter through an activation function based on the transition probability corresponding to the current system state parameter includes: calculating the initial benefit parameter corresponding to each action parameter through an activation function based on the transition probability corresponding to the current system state parameter; obtaining the reference benefit parameter corresponding to each action parameter under the current system state parameter from a decision table, wherein the decision table records system state parameters, action parameters, and benefit parameters with corresponding relationships, and the benefit parameter recorded in the decision table is used to indicate the actual adjustment effect achieved by each adjustment operation actually executed for the decision target; selecting a first adjustment operation from the plurality of adjustment operations whose initial benefit parameter is greater than a first preset benefit parameter; obtaining a candidate frequency obtained by adjusting the reference frequency using the first adjustment operation; and determining from the data acquisition cycles before the current data acquisition cycle... A reference data acquisition period for data acquisition using the candidate frequency; obtaining a second benefit parameter for each adjustment operation in the reference data acquisition period, wherein the second benefit parameter indicates the actual adjustment effect achieved by the adjustment operation under the system state in the reference data acquisition period; selecting candidate adjustment operations from the plurality of adjustment operations whose second benefit parameter is greater than a second preset benefit parameter; calculating a first product value of a target discount factor and the second benefit parameter corresponding to the candidate adjustment operation, wherein the target discount factor characterizes the importance of the second benefit parameter corresponding to the candidate adjustment operation to the target benefit parameter; calculating a target sum value of the first product value and the initial benefit parameter corresponding to the first adjustment operation; determining a second product value of a first learning rate and the target sum value, and the sum value of the second learning rate and the reference benefit parameter as the target benefit parameter.
[0008] Optionally, after controlling the baseboard management controller to collect the server operating data at the target frequency, the method further includes: detecting the reference change amplitude of the reference server operating data collected by the baseboard management controller at the target frequency; determining the actual benefit parameter corresponding to the target adjustment operation based on the reference change amplitude; and updating the decision table based on the corresponding target adjustment operation and the actual benefit parameter.
[0009] Optionally, obtaining the reference benefit parameter corresponding to each action parameter under the current system state parameter from the decision table includes: generating a target random number; comparing the target random number with a current random number threshold; if the target random number is greater than or equal to the current random number threshold, obtaining the reference benefit parameter corresponding to each action parameter under the current system state parameter from the decision table; and if the target random number is less than the current random number threshold, determining the initial benefit parameter as the reference benefit parameter.
[0010] Optionally, after updating the decision table according to the corresponding target adjustment operation and the actual benefit parameter, the method further includes: adjusting the current random number threshold to a target random number threshold as the random number threshold for the next time, wherein the target random number threshold is less than the current random number threshold.
[0011] Optionally, the stability parameter of the detected target server operating data includes: acquiring server operating data collected by the baseboard management controller within a target time period using the reference frequency as the target server operating data; determining the current variation range of the target server operating data; matching the current variation range with the target variation range; if the current variation range matches the target variation range, determining the stability parameter as a first parameter, wherein the first parameter indicates that the variation range of the target server operating data is stable; if the current variation range does not match the target variation range, determining the stability parameter as a second parameter, wherein the second parameter indicates that the variation range of the target server operating data is unstable.
[0012] According to another embodiment of this application, an operation control device for a baseboard management controller is provided, comprising: a first detection module, used to detect a stability parameter of the collected target server operation data during the process of the baseboard management controller collecting server operation data at a reference frequency, wherein the stability parameter is used to indicate the stability of the change amplitude of the target server operation data;
[0013] The second detection module is used to detect, through a target decision model, the target benefit parameter corresponding to each of the multiple adjustment operations under the current change range of the target server's operating data when the stability parameter indicates that the change range of the target server's operating data is unstable. The target decision model is used to calculate the adjustment effect achieved by executing each adjustment operation on the decision objective, the decision objective including adjusting the input change range to the target change range, and the target benefit parameter is used to indicate the adjustment effect.
[0014] The adjustment module is used to adjust the reference frequency to the target frequency by adopting the target adjustment operation that satisfies the target benefit condition corresponding to the target benefit parameter in the plurality of adjustment operations;
[0015] The control module is used to control the baseboard management controller to collect server operation data at the target frequency.
[0016] According to yet another embodiment of this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and the computer program is configured to perform the steps in any of the above method embodiments when it is run.
[0017] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0018] This application utilizes a target decision model to calculate the adjustment effect of various adjustment operations on adjusting the input variation range to the target variation range. During the process of the baseboard management controller collecting server operating data at a reference frequency, the model detects the stability of the variation range of the collected target server operating data. Furthermore, when the variation range of the target server operating data is unstable, the target decision model detects the adjustment effect of each adjustment operation under the current variation range of the target server operating data. Based on the adjustment effect of each operation under the current variation range of the target server operating data, the model can select adjustment operations that meet the target benefit conditions from multiple adjustment operations, thereby adjusting the reference frequency. This achieves dynamic adjustment of the reference frequency of the collected data based on the variation range during the data collection process, ensuring the stability of the variation range of the collected operating data. This solves the problem of low efficiency in collecting server operating data and improves the efficiency of server operating data collection. Attached Figure Description
[0019] Figure 1 This is a hardware structure block diagram of a server device for a baseboard management controller operation control method according to an embodiment of this application;
[0020] Figure 2 This is a flowchart of the operation control method of the baseboard management controller according to an embodiment of this application;
[0021] Figure 3 This is an operation control flowchart of an optional baseboard management controller according to an embodiment of this application;
[0022] Figure 4 This is a structural block diagram of the operation control device of the substrate management controller according to an embodiment of this application. Detailed Implementation
[0023] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0025] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of a server device for a baseboard management controller operation control method according to an embodiment of this application. (See diagram for example.) Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0026] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the operation control method of the baseboard management controller in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0027] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0028] This embodiment provides an operation control method for a baseboard management controller. Figure 2 This is a flowchart of the operation control method of the baseboard management controller according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:
[0029] Step S202: During the process of the baseboard management controller acquiring server operating data using a reference frequency, the stability parameter of the acquired target server operating data is detected, wherein the stability parameter is used to indicate the stability of the change amplitude of the target server operating data;
[0030] Step S204: When the stability parameter indicates that the change range of the target server's operating data is unstable, the target decision model detects the target benefit parameter corresponding to each of the multiple adjustment operations under the current change range of the target server's operating data. The target decision model is used to calculate the adjustment effect achieved by executing each adjustment operation on the decision objective. The decision objective includes adjusting the input change range to the target change range. The target benefit parameter is used to indicate the adjustment effect.
[0031] Step S206: Using the target adjustment operation in the plurality of adjustment operations that corresponds to the target benefit parameter satisfying the target benefit condition, the reference frequency is adjusted to the target frequency;
[0032] Step S208: Control the baseboard management controller to collect server operation data using the target frequency.
[0033] Through the above steps, the target decision model is used to calculate the adjustment effect of various adjustment operations on adjusting the input change range to the target change range. During the process of the baseboard management controller collecting server operating data according to the reference frequency, the stability of the change range of the collected target server operating data is detected. Then, when the change range of the target server operating data is unstable, the target decision model is used to detect the adjustment effect of each adjustment operation under the current change range of the target server operating data. Based on the adjustment effect of each adjustment operation under the current change range of the target server operating data, the adjustment operation that meets the target benefit condition can be selected from multiple adjustment operations, thereby adjusting the reference frequency. This realizes the dynamic adjustment of the reference frequency of the collected data based on the change range during the data collection process, so that the change range of the collected operating data remains stable. This can solve the problem of low data collection efficiency of the server operating data and improve the data collection efficiency of the server operating data.
[0034] The entity performing the above steps can be a baseboard management controller or a device that controls the working state of the baseboard management controller, etc. This solution does not limit this.
[0035] In the embodiment provided in step S202 above, the stability of the target server's operating data can be detected, but is not limited to, by detecting the difference between two adjacent detected server operating data. For example, if multiple operating data are collected at a reference collection frequency within a certain period of time, multiple initial differences are obtained by calculating the initial difference between two adjacent operating data, and the mean square error or the mean of the multiple initial differences is calculated to obtain the target difference, i.e., the stability parameter is the target difference. If the target difference is greater than the preset difference, it is determined that the change range of the target server's operating data is unstable; otherwise, it is determined that the change range of the operating data is stable.
[0036] Optionally, in this embodiment of the application, the stability of the target server's running data can be detected by detecting the rate of change of the difference in the running data. For example, if multiple running data are collected at a reference collection frequency within a certain period of time, multiple initial differences are obtained by calculating the initial difference between two adjacent running data, and the target rate of change of the multiple initial differences is calculated, i.e., the stability parameter is the target rate of change. If the target rate of change is greater than the preset rate of change, it is determined that the change amplitude of the target server's running data is unstable; otherwise, it is determined that the change amplitude of the running data is stable.
[0037] Optionally, in this embodiment, the BMC establishes a separate thread for each sensor that needs to be monitored. The sensor is used to detect a corresponding server operating data, and each thread runs an independent operating control method of the baseboard management controller.
[0038] In the embodiment provided in step S204 above, each adjustment operation in the multiple adjustment operations corresponds to an adjustment method for adjusting the data acquisition frequency, namely the adjustment direction and adjustment step size, such as adding 5 to the reference frequency, adding 6 to the reference frequency, subtracting 7 from the reference frequency, etc. This solution does not limit this.
[0039] Optionally, in this embodiment, the adjustment process for adjusting the acquisition frequency of server operating data is an iterative adjustment process. That is, each time the stability parameter indicates that the server operating data is unstable, the target benefit parameter of each adjustment operation needs to be detected through the target decision model, so as to select the adjustment operation to adjust the data acquisition frequency. The decision target of the target decision model is to adjust the change range of the acquired server operating data to the target change range. After the target adjustment operation is selected and the reference frequency is adjusted through the target benefit parameter output by the target decision model, it is possible that the change range of the operating data obtained by data acquisition using the adjusted frequency is still unstable. Therefore, it is necessary to repeatedly execute the operation control method of the baseboard management controller of this application, so as to continuously iteratively adjust the data acquisition frequency.
[0040] Optionally, in this embodiment, the target decision model records the optimization logic for finding the optimal adjustment operation suitable for the corresponding system state from multiple adjustment operations under different system states with varying magnitudes of change. The target decision model can be, but is not limited to, a reinforcement learning-based model. Reinforcement learning primarily addresses how an agent makes decisions in an environment to maximize its gains. The core of reinforcement learning consists of five main parts: agent, environment, state, policy, and rewards. When the agent is in a certain state in the environment, it selects a policy to execute according to the algorithm. After execution, the agent receives a reward from the environment and updates its state. Reinforcement learning seeks the optimal solution of the policy over a continuous time series, enabling the agent to maximize its gains. In the reinforcement learning model, the BMC can be viewed as an agent. The goal of the BMC is to find the most suitable monitoring cycle for each sensor under daily working conditions, in order to obtain the most accurate data change acquisition and lower system usage. BMC can use dynamically adjusting the monitoring cycle as the strategy to be executed, and the degree of change in the collected sensor data as a factor in evaluating the reward, thus serving as a reference factor for subsequent adjustments to the monitoring cycle strategy to influence the next round of execution. This analysis reveals that BMC's dynamic adjustment of the monitoring cycle satisfies the five elements of reinforcement learning. Therefore, using reinforcement learning algorithms to dynamically adjust the BMC monitoring cycle is a feasible approach.
[0041] In the embodiment provided in step S206 above, the target benefit condition is used to select the adjustment operation with the best effect on the adjustment of the change range of the collected target server running data from multiple adjustment operations. The target benefit condition may be to determine the adjustment operation corresponding to the largest benefit parameter among multiple target benefit parameters as the target adjustment operation, or to determine the adjustment operation whose target benefit parameter is greater than a preset benefit parameter among multiple adjustment operations as the target adjustment operation.
[0042] As an optional embodiment, the step of detecting the target benefit parameter corresponding to each of the multiple adjustment operations under the current change magnitude of the target server's operating data through the target decision model includes:
[0043] The current change magnitude is input into the target decision model as the current system state parameter. The target decision model's system state parameter set records multiple change magnitudes, its action parameter set records multiple adjustment operations, and its transition probability parameter set records corresponding change magnitudes, the target change magnitude, each adjustment operation, and a transition probability. The transition probability is the probability that each adjustment operation will cause the change magnitude to be converted into the target change magnitude. The target decision model is used to calculate the target benefit parameter corresponding to each action parameter using an activation function based on the transition probability corresponding to the current system state parameter.
[0044] Obtain multiple sets of corresponding adjustment operations and target benefit parameters output by the target decision model.
[0045] Optionally, in this embodiment, the target decision model is a model used to describe sequential decision problems. It formalizes the problem of decision-makers needing to make decisions in uncertain environments into a state space, a decision space, a state transition probability, and a reward function. The theoretical foundation of reinforcement learning is MDP. Therefore, it is first necessary to model the entire behavior of the BMC's dynamic adjustment sensor monitoring cycle as an MDP, satisfying the five elements of MDP, namely:
[0046] Decision cycle T: Set the decision cycle to be once every few rounds after the BMC reads the sensor data;
[0047] System parameter state S: The system state is defined as the actual value read by the sensor and the average change of the value read in each round under the current strategy. These "current sensor value - change of running data" are combined and integrated into a set S.
[0048] Action parameter set (adjustment operation set) A: The action set is set to {increase reading frequency, maintain the current reading frequency, decrease reading frequency};
[0049] Transition probability P: Transition probability describes the probability that an agent will take a certain action in a certain state, and can be described as: p t (·|s,a),∑ i∈S p t (i|s, a) = 1. In this embodiment, the transition probability will be given based on the Q-learning algorithm.
[0050] Incentive function R: The incentive function is formulated based on the magnitude of numerical changes during the previous round of decision-making. An expected range of change is set; the greater the deviation from this range, the higher the penalty. Conversely, if the change falls within the expected range, a reward is given.
[0051] By recording multiple change magnitudes in the system state parameter set, multiple adjustment operations in the action parameter set, and corresponding change magnitudes, target change magnitudes, each adjustment operation, and transition probability in the transition probability parameter set of the decision model, the target decision model can calculate the target benefit parameter corresponding to each action parameter through the activation function based on the transition probability corresponding to the current system state parameter, thereby ensuring the accuracy of the benefit parameters of each adjustment operation calculated under the current system state.
[0052] As an optional embodiment, the step of calculating the target reward parameter corresponding to each action parameter through an activation function based on the transition probability corresponding to the current system state parameter includes:
[0053] Based on the transition probability corresponding to the current system state parameters, the initial revenue parameter corresponding to each action parameter is calculated through the activation function.
[0054] Obtain the reference benefit parameter corresponding to each action parameter under the current system state parameter from the decision table. The decision table records the system state parameter, action parameter and benefit parameter with corresponding relationship. The benefit parameter recorded in the decision table is used to indicate the actual adjustment effect achieved by each adjustment operation actually executed on the decision objective.
[0055] Select the first adjustment operation from the plurality of adjustment operations, wherein the initial benefit parameter is greater than the first preset benefit parameter;
[0056] Obtain a candidate frequency obtained by adjusting the reference frequency using the first adjustment operation;
[0057] A reference data acquisition cycle for using the candidate frequency is determined from the data acquisition cycles preceding the current data acquisition cycle;
[0058] Obtain a second benefit parameter for each adjustment operation in the reference data acquisition period, wherein the second benefit parameter is used to indicate the actual adjustment effect achieved by the adjustment operation under the system state in the reference data acquisition period;
[0059] From the plurality of adjustment operations, candidate adjustment operations are selected where the second benefit parameter is greater than the second preset benefit parameter;
[0060] Calculate the first product of the target discount factor and the second revenue parameter corresponding to the candidate adjustment operation, wherein the target discount factor is used to characterize the importance of the second revenue parameter corresponding to the candidate adjustment operation to the target revenue parameter;
[0061] Calculate the target sum and value of the initial revenue parameter corresponding to the first product value and the first adjustment operation;
[0062] The target return parameter is determined by the second product of the first learning rate and the target sum, and the sum of the second learning rate and the reference return parameter.
[0063] Optionally, in this embodiment, the target decision model transforms the BMC's dynamic adjustment of the sensor monitoring cycle into an adaptive MDP model. The adaptive MDP assumes that both the state transition rate and the reward function are related to a certain unknown parameter. The unknown parameter in this model is the degree of change in the sensor values read by the BMC within a monitoring cycle. This degree of change is a parameter with a fixed value within a training cycle, and therefore can be studied in detail using an asymptotic discount optimization algorithm. The Q-learning algorithm is a value-based decision-making method under model-free reinforcement learning. It calculates how the agent should make decisions in a given state by maintaining a Q-table (Q-value table, i.e., a decision table). The size of the Q-value table is the number of states S * the number of actions A. Each value in the table is calculated using the Q-function, representing the future reward that can be obtained by taking an action in the current state. The Q-function can be described as: newQ s,a =(1-α)Q S,A +α(R S,A +γ*maxQ′(S′,A′)). The two important parameters in the Q-function above are the learning rate (α) and the discount factor (γ). The learning rate defines the proportion of a given Q-value that it will learn from a new Q-value. A value of 0 means the agent will not learn anything (old information is important), while a value of 1 means the newly discovered information is the only important information. The discount factor defines the importance of future rewards. A value of 0 means only short-term rewards are considered, while a value of 1 emphasizes long-term rewards. In practical applications, the values of these two parameters should be determined by the developers and continuously adjusted to achieve the best results in reinforcement learning. By continuously iterating and refining this Q-value table, an optimal policy can be derived based on the Q-value table during final execution. Applying the Q-learning algorithm to the dynamic adjustment of the monitoring cycle in the BMC can be understood as follows: The BMC maintains a Q-value table for each sensor monitoring thread. This Q-value table describes the reward of taking each action in each state, indirectly reflecting how the BMC will adjust in the next decision cycle. The Q-learning algorithm cannot guide the agent in decision-making because the Q-value table is initially set to all zeros.
[0064] Through the above steps, the reference return parameter indicates the actual effect of each adjustment operation on the decision objective. When calculating the target return parameter, by using the reference return parameter and the initial return parameter corresponding to each action parameter, the target return parameter contains both the influence of historical experience and the incentive of actual operation, thus ensuring the accuracy of the target return parameter.
[0065] As an optional embodiment, after controlling the baseboard management controller to collect the server operating data at the target frequency, the method further includes:
[0066] The reference change amplitude of the reference server operating data collected by the baseboard management controller at the target frequency is detected;
[0067] The actual benefit parameters corresponding to the target adjustment operation are determined based on the reference change range.
[0068] The decision table is updated based on the corresponding target adjustment operations and the actual benefit parameters.
[0069] Optionally, in this embodiment, the actual revenue parameter can be obtained by conditionally adjusting the reference revenue parameter corresponding to the target adjustment operation based on the reference change range. This is achieved by determining the reference revenue parameter adjustment value corresponding to the reference change range from the correspondence between the change range and the revenue parameter adjustment value, and then adjusting the reference revenue parameter using this adjustment value to obtain the actual revenue parameter. In other words, a desired change range is set; when the reference change range deviates from the desired range, the greater the deviation, the higher the penalty. Conversely, if it is within the desired range, a reward is given, thereby adjusting the reference revenue parameter to obtain the actual revenue parameter.
[0070] Through the above steps, during each iteration cycle, the reference change range of the running data for that cycle is detected, and the actual profit value is determined based on the reference change range. The decision table is then iteratively updated so that the profit value recorded in the iteration table conforms to the actual historical operating experience, thereby improving the accuracy of the decision table in assisting decision-making.
[0071] As an optional embodiment, obtaining the reference benefit parameter corresponding to each action parameter under the current system state parameter from the decision table includes:
[0072] Generate the target random number;
[0073] The target random number is compared with the current random number threshold;
[0074] If the target random number is greater than or equal to the current random number threshold, the reference benefit parameter corresponding to each action parameter under the current system state parameter is obtained from the decision table;
[0075] If the target random number is less than the current random number threshold, the initial profit parameter is determined as the reference profit parameter.
[0076] Optionally, in this embodiment, the target random number can be, but is not limited to, randomly generated by the ε-greedy algorithm. In the early stages of model iteration, the values recorded in the decision table may be inaccurate, failing to guide the agent in decision-making. Therefore, the ε-greedy algorithm is needed to assist in decision-making. The ε-greedy algorithm sets a given value ε. Simultaneously, at the beginning of each iteration, a random value (the current random number threshold) is requested. When the random value is less than the given value ε (the target random number), a strategy is forcibly executed. When the random value is greater than ε, the operation suggested in the decision table is adopted. The following is the logic for determining the target return parameter according to this application:
[0077]
[0078] By following the steps above, the target random number is compared with the current random number threshold by setting the current random number threshold. Based on the comparison result, it is determined whether to obtain the reference benefit parameter according to the decision table, thereby achieving a balance between exploration and utilization strategies and ensuring that the strategy eventually converges to the optimal strategy.
[0079] As an optional embodiment, after updating the decision table according to the corresponding target adjustment operation and the actual benefit parameter, the method further includes:
[0080] The current random number threshold is adjusted to a target random number threshold as the random number threshold for the next time, wherein the target random number threshold is less than the current random number threshold.
[0081] Based on the above, after each decision, by adjusting the current random number threshold to the target random number threshold as the random number threshold for the next decision, the random number threshold decreases with each iteration of the decision table, thereby gradually increasing the influence of the decision table on the decision-making process.
[0082] As an optional embodiment, the stability parameters of the detected target server operating data include:
[0083] The server operation data collected by the baseboard management controller within the target time period using the reference frequency is used as the target server operation data;
[0084] Determine the current magnitude of change in the target server's operating data;
[0085] Match the current change magnitude with the target change magnitude;
[0086] If the current change range matches the target change range, the stability parameter is determined to be the first parameter, wherein the first parameter is used to indicate that the change range of the target server's operating data is stable;
[0087] If the current change magnitude does not match the target change magnitude, the stability parameter is determined to be a second parameter, wherein the second parameter is used to indicate that the change magnitude of the target server's operating data is unstable.
[0088] In practical applications, BMC establishes a separate thread for each sensor that needs to be monitored. Each thread runs a separate set of reinforcement learning training algorithms and sensor reading logic. Simultaneously, the Q-value table data generated during training is periodically exported and updated to the file system, eliminating the need for retraining upon subsequent BMC startups. Furthermore, different types of threshold sensors exhibit significant variations in application; for example, temperature sensors change relatively frequently, while voltage sensors show minimal and relatively stable changes. Therefore, the parameters used during training should be adjusted accordingly. Figure 3 This is a flowchart illustrating the operation control of an optional baseboard management controller according to an embodiment of this application, such as... Figure 3 As shown, it includes at least the following steps:
[0089] S301 creates a separate thread for each sensor that needs to be monitored. Each thread runs a separate set of reinforcement learning training algorithms and sensor reading logic. The monitoring thread is started and initialized.
[0090] S302, query the file system to see if there is a configuration file related to the sensor. The configuration file includes the Q-value table (decision table), the excitation function R, the decay factor γ, the learning rate α, and the ε value.
[0091] S303, if a configuration file exists, import the Q-value table, activation function R, decay factor γ, learning rate α, and ε value from the configuration file.
[0092] S304 If no configuration file exists, the above parameters will be initialized with fixed values.
[0093] S305, start executing the Q-learnig reinforcement learning process, read sensor values in one iteration cycle, repeat the reading several times, and calculate the degree of data fluctuation and the temperature at this time in the current iteration cycle.
[0094] S306, check if the iteration cycle has been reached. If the iteration cycle has been reached, proceed to step S307. If the iteration cycle has not been reached, continue to proceed to step S305.
[0095] In step S307, the reward is calculated from the data generated in step S305, and the Q-value table is updated synchronously according to the Q-function. Simultaneously, the ε-greedy algorithm is used to select the strategy to be executed in the next iteration cycle, i.e., the frequency of sensor readings.
[0096] S308 reduces the value of the epsilon parameter after each execution of the ε-greedy algorithm, allowing the agent (BMC) to gradually rely on the Q-value table to make policy choices instead of making them completely random as reinforcement learning training progresses.
[0097] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0098] This embodiment also provides an operation control device for a baseboard management controller, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0099] Figure 4 This is a structural block diagram of the operation control device of the baseboard management controller according to an embodiment of this application, as shown below. Figure 4 As shown, the device includes: a first detection module, used to detect the stability parameter of the target server operating data during the process of the baseboard management controller collecting server operating data at a reference frequency, wherein the stability parameter is used to indicate the stability of the change amplitude of the target server operating data;
[0100] The second detection module is used to detect, through a target decision model, the target benefit parameter corresponding to each of the multiple adjustment operations under the current change range of the target server's operating data when the stability parameter indicates that the change range of the target server's operating data is unstable. The target decision model is used to calculate the adjustment effect achieved by executing each adjustment operation on the decision objective, the decision objective including adjusting the input change range to the target change range, and the target benefit parameter is used to indicate the adjustment effect.
[0101] The adjustment module is used to adjust the reference frequency to the target frequency by adopting the target adjustment operation that satisfies the target benefit condition corresponding to the target benefit parameter in the plurality of adjustment operations;
[0102] The control module is used to control the baseboard management controller to collect server operation data at the target frequency.
[0103] This application utilizes a target decision model to calculate the adjustment effect of various adjustment operations on adjusting the input variation range to the target variation range. During the process of the baseboard management controller collecting server operating data at a reference frequency, the model detects the stability of the variation range of the collected target server operating data. Furthermore, when the variation range of the target server operating data is unstable, the target decision model detects the adjustment effect of each adjustment operation under the current variation range of the target server operating data. Based on the adjustment effect of each operation under the current variation range of the target server operating data, the model can select adjustment operations that meet the target benefit conditions from multiple adjustment operations, thereby adjusting the reference frequency. This achieves dynamic adjustment of the reference frequency of the collected data based on the variation range during the data collection process, ensuring the stability of the variation range of the collected operating data. This solves the problem of low efficiency in collecting server operating data and improves the efficiency of server operating data collection.
[0104] Optionally, the second detection module includes:
[0105] An input unit is used to input the current change magnitude as the current system state parameter into the target decision model. The target decision model's system state parameter set records multiple change magnitudes, its action parameter set records multiple adjustment operations, and its transition probability parameter set records corresponding change magnitudes, the target change magnitude, each adjustment operation, and a transition probability. The transition probability is the probability that each adjustment operation will cause the change magnitude to be converted into the target change magnitude. The target decision model is used to calculate the target benefit parameter corresponding to each action parameter using an activation function based on the transition probability corresponding to the current system state parameter.
[0106] The acquisition unit is used to acquire multiple sets of corresponding adjustment operations and target benefit parameters output by the target decision model.
[0107] Optionally, the detection unit is used for:
[0108] Based on the transition probability corresponding to the current system state parameters, the initial revenue parameter corresponding to each action parameter is calculated through the activation function.
[0109] Obtain the reference benefit parameter corresponding to each action parameter under the current system state parameter from the decision table. The decision table records the system state parameter, action parameter and benefit parameter with corresponding relationship. The benefit parameter recorded in the decision table is used to indicate the actual adjustment effect achieved by each adjustment operation actually executed on the decision objective.
[0110] Select the first adjustment operation from the plurality of adjustment operations, wherein the initial benefit parameter is greater than the first preset benefit parameter;
[0111] Obtain a candidate frequency obtained by adjusting the reference frequency using the first adjustment operation;
[0112] A reference data acquisition cycle for using the candidate frequency is determined from the data acquisition cycles preceding the current data acquisition cycle;
[0113] Obtain a second benefit parameter for each adjustment operation in the reference data acquisition period, wherein the second benefit parameter is used to indicate the actual adjustment effect achieved by the adjustment operation under the system state in the reference data acquisition period;
[0114] From the plurality of adjustment operations, candidate adjustment operations are selected where the second benefit parameter is greater than the second preset benefit parameter;
[0115] Calculate the first product of the target discount factor and the second revenue parameter corresponding to the candidate adjustment operation, wherein the target discount factor is used to characterize the importance of the second revenue parameter corresponding to the candidate adjustment operation to the target revenue parameter;
[0116] Calculate the target sum and value of the initial revenue parameter corresponding to the first product value and the first adjustment operation;
[0117] The target return parameter is determined by the second product of the first learning rate and the target sum, and the sum of the second learning rate and the reference return parameter.
[0118] Optionally, the device further includes:
[0119] The third detection module is used to detect the reference change amplitude of the reference server operating data collected by the baseboard management controller at the target frequency after the baseboard management controller is controlled to collect the server operating data at the target frequency.
[0120] The determination module is used to determine the actual benefit parameter corresponding to the target adjustment operation based on the reference change range;
[0121] An update module is used to update the decision table based on the corresponding target adjustment operations and the actual benefit parameters.
[0122] Optionally, the detection unit is used for:
[0123] Generate the target random number;
[0124] The target random number is compared with the current random number threshold;
[0125] If the target random number is greater than or equal to the current random number threshold, the reference benefit parameter corresponding to each action parameter under the current system state parameter is obtained from the decision table;
[0126] If the target random number is less than the current random number threshold, the initial profit parameter is determined as the reference profit parameter.
[0127] Optionally, the square device further includes:
[0128] An adjustment module is used to adjust the current random number threshold to a target random number threshold as the next random number threshold after updating the decision table according to the target adjustment operation and the actual benefit parameter with a corresponding relationship, wherein the target random number threshold is less than the current random number threshold.
[0129] Optionally, the first detection module includes:
[0130] The acquisition unit is used to acquire server operation data collected by the baseboard management controller within a target time period using the reference frequency as the target server operation data;
[0131] The first determining unit is used to determine the current change range of the target server's operating data;
[0132] A matching unit is used to match the current change magnitude with the target change magnitude;
[0133] The second determining unit is configured to determine the stability parameter as a first parameter when the current change range matches the target change range, wherein the first parameter is used to indicate that the change range of the target server's operating data is stable;
[0134] The second determining unit is used to determine the stability parameter as a second parameter when the current change range does not match the target change range, wherein the second parameter is used to indicate that the change range of the target server's operating data is unstable.
[0135] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0136] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.
[0137] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0138] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0139] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0140] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0141] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0142] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A method for controlling the operation of a baseboard management controller, characterized in that, include: During the process of acquiring server operating data using a reference frequency, the baseboard management controller detects the stability parameter of the acquired target server operating data, wherein the stability parameter is used to indicate the stability of the change amplitude of the target server operating data; When the stability parameter indicates that the change range of the target server's operating data is unstable, the target decision model detects the target benefit parameter corresponding to each of the multiple adjustment operations under the current change range of the target server's operating data. The target decision model is used to calculate the adjustment effect achieved by executing each adjustment operation on the decision objective, the decision objective including adjusting the input change range to the target change range, and the target benefit parameter is used to indicate the adjustment effect. The reference frequency is adjusted to the target frequency by using the target adjustment operation in which the target benefit parameter satisfies the target benefit condition among the plurality of adjustment operations; The baseboard management controller is controlled to collect server operating data at the target frequency; The step of detecting the target benefit parameter corresponding to each adjustment operation among multiple adjustment operations under the current change range of the target server's running data through a target decision model includes: inputting the current change range as a current system state parameter into the target decision model, wherein the system state parameter set of the target decision model records multiple change ranges, the action parameter set of the target decision model records the multiple adjustment operations, and the transition probability parameter set of the target decision model records change ranges with corresponding relationships, the target change range, each adjustment operation, and a transition probability, wherein the transition probability is the probability of instructing each adjustment operation to convert the change range into the target change range, and the target decision model is used to calculate the target benefit parameter corresponding to each action parameter through an activation function based on the transition probability corresponding to the current system state parameter; and obtaining multiple sets of adjustment operations and target benefit parameters with corresponding relationships output by the target decision model.
2. The method according to claim 1, characterized in that, The step of calculating the target reward parameter corresponding to each action parameter through an activation function based on the transition probability corresponding to the current system state parameter includes: Based on the transition probability corresponding to the current system state parameters, the initial revenue parameter corresponding to each action parameter is calculated through the activation function. Obtain the reference benefit parameter corresponding to each action parameter under the current system state parameter from the decision table. The decision table records the system state parameter, action parameter and benefit parameter with corresponding relationship. The benefit parameter recorded in the decision table is used to indicate the actual adjustment effect achieved by each adjustment operation actually executed on the decision objective. Select the first adjustment operation from the plurality of adjustment operations, wherein the initial benefit parameter is greater than the first preset benefit parameter; Obtain a candidate frequency obtained by adjusting the reference frequency using the first adjustment operation; A reference data acquisition cycle for using the candidate frequency is determined from the data acquisition cycles preceding the current data acquisition cycle; Obtain a second benefit parameter for each adjustment operation in the reference data acquisition period, wherein the second benefit parameter is used to indicate the actual adjustment effect achieved by the adjustment operation under the system state in the reference data acquisition period; From the plurality of adjustment operations, candidate adjustment operations are selected where the second benefit parameter is greater than the second preset benefit parameter; Calculate the first product of the target discount factor and the second revenue parameter corresponding to the candidate adjustment operation, wherein the target discount factor is used to characterize the importance of the second revenue parameter corresponding to the candidate adjustment operation to the target revenue parameter; Calculate the target sum and value of the initial revenue parameter corresponding to the first product value and the first adjustment operation; The target return parameter is determined by the second product of the first learning rate and the target sum, and the sum of the second learning rate and the reference return parameter.
3. The method according to claim 2, characterized in that, After controlling the baseboard management controller to collect the server operating data at the target frequency, the method further includes: The reference change amplitude of the reference server operating data collected by the baseboard management controller at the target frequency is detected; The actual benefit parameters corresponding to the target adjustment operation are determined based on the reference change range. The decision table is updated based on the corresponding target adjustment operations and the actual benefit parameters.
4. The method according to claim 3, characterized in that, The step of obtaining the reference benefit parameter corresponding to each action parameter under the current system state parameter from the decision table includes: Generate the target random number; The target random number is compared with the current random number threshold; If the target random number is greater than or equal to the current random number threshold, the reference benefit parameter corresponding to each action parameter under the current system state parameter is obtained from the decision table; If the target random number is less than the current random number threshold, the initial profit parameter is determined as the reference profit parameter.
5. The method according to claim 4, characterized in that, After updating the decision table based on the corresponding target adjustment operation and the actual benefit parameter, the method further includes: The current random number threshold is adjusted to a target random number threshold as the random number threshold for the next time, wherein the target random number threshold is less than the current random number threshold.
6. The method according to claim 1, characterized in that, The stability parameters of the target server's operational data collected during detection include: The server operation data collected by the baseboard management controller within the target time period using the reference frequency is used as the target server operation data; Determine the current magnitude of change in the target server's operating data; Match the current change magnitude with the target change magnitude; If the current change range matches the target change range, the stability parameter is determined to be the first parameter, wherein the first parameter is used to indicate that the change range of the target server's operating data is stable; If the current change magnitude does not match the target change magnitude, the stability parameter is determined to be a second parameter, wherein the second parameter is used to indicate that the change magnitude of the target server's operating data is unstable.
7. An operation control device for a substrate management controller, characterized in that, include: The first detection module is used to detect the stability parameter of the target server operating data during the process of the baseboard management controller collecting server operating data at a reference frequency. The stability parameter is used to indicate the stability of the change amplitude of the target server operating data. The second detection module is used to detect, through a target decision model, the target benefit parameter corresponding to each of the multiple adjustment operations under the current change range of the target server's operating data when the stability parameter indicates that the change range of the target server's operating data is unstable. The target decision model is used to calculate the adjustment effect achieved by executing each adjustment operation on the decision objective, the decision objective including adjusting the input change range to the target change range, and the target benefit parameter is used to indicate the adjustment effect. The adjustment module is used to adjust the reference frequency to the target frequency by adopting the target adjustment operation that satisfies the target benefit condition corresponding to the target benefit parameter in the plurality of adjustment operations; The control module is used to control the baseboard management controller to collect server operating data at the target frequency; The second detection module includes: an input unit, used to input the current change amplitude as a current system state parameter into the target decision model, wherein the system state parameter set of the target decision model records multiple change amplitudes, the action parameter set of the target decision model records multiple adjustment operations, and the transition probability parameter set of the target decision model records change amplitudes with corresponding relationships, the target change amplitude, each adjustment operation, and a transition probability, wherein the transition probability is the probability of converting the change amplitude into the target change amplitude by instructing each adjustment operation, and the target decision model is used to calculate the target benefit parameter corresponding to each action parameter through an activation function based on the transition probability corresponding to the current system state parameter; and an acquisition unit, used to acquire multiple sets of adjustment operations and target benefit parameters with corresponding relationships output by the target decision model.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 6.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Data acquisition method, device and system
CN113129473A
Running control method and device of server
CN115576408A