Parameter adjusting method of cloud equipment and electronic equipment
By obtaining the performance characteristics of cloud devices and using policy networks for dynamic prediction and policy updates, the problem of poor parameter adjustment of traditional cloud devices is solved, and real-time adaptive adjustment of cloud device parameters and performance optimization are achieved.
Patent Information
- Application Number
- CN202511244345.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Traditional cloud device parameter adjustment methods use static configuration or rule-based fixed methods, resulting in insufficient or excessive optimization of cloud device performance and poor adjustment effects.
By obtaining the performance characteristics of the cloud device where the large language model is located, the policy network is used for dynamic prediction, generating predicted execution actions, and updating the policy network based on the performance characteristics and historical execution actions. Parameter adjustment instructions are generated and executed in the registers of the cloud device.
It achieves real-time adaptive adjustment of cloud device parameters, improves the accuracy and effect of parameter adjustment, optimizes the inference performance of large language models, and overcomes the limitations of static configuration.
Smart Images

Figure CN120743352A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computers, and more specifically, to a method for adjusting parameters of a cloud device and an electronic device. Background Art
[0002] In a cloud device (cloud computing device) environment, when large language models play a core role in high-performance computing tasks, traditional parameter adjustment methods often use static configuration or fixed rule-based adjustment methods, which have poor effect on the parameter adjustment of cloud devices, resulting in insufficient or excessive performance optimization of cloud devices.
[0003] Therefore, there is a technical problem in the related art that the parameter adjustment effect of the cloud device is poor. Summary of the Invention
[0004] The embodiments of the present application provide a cloud device parameter adjustment method and electronic device to at least solve the technical problem of poor effect of cloud device parameter adjustment in related technologies.
[0005] According to one embodiment of the present application, a parameter adjustment method for a cloud device is provided, comprising: obtaining performance characteristics of a cloud device where a large language model is located, wherein the performance characteristics are used to indicate operating characteristics of the cloud device in a working state; using a policy network to make predictions based on the performance characteristics to obtain predicted execution actions; updating the policy network based on the performance characteristics and historical execution actions corresponding to the performance characteristics, wherein the historical execution actions include adjustment actions that have been performed on the performance characteristics of the cloud device; performing instruction mapping on the predicted execution action to obtain at least one parameter adjustment instruction corresponding to the predicted execution action; and executing at least one parameter adjustment instruction in a register of the cloud device.
[0006] According to another embodiment of the present application, a parameter adjustment device for a cloud device is provided, comprising: an acquisition unit for acquiring performance characteristics of a cloud device where a large language model is located, wherein the performance characteristics are used to indicate operating characteristics of the cloud device in a working state; a prediction unit for using a policy network to make predictions based on the performance characteristics to obtain predicted execution actions; an update unit for performing policy updates on the policy network based on the performance characteristics and historical execution actions corresponding to the performance characteristics, wherein the historical execution actions include adjustment actions that have been performed on the performance characteristics of the cloud device; a mapping unit for performing instruction mapping on the predicted execution action to obtain at least one parameter adjustment instruction corresponding to the predicted execution action; and an adjustment unit for executing at least one parameter adjustment instruction in a register of the cloud device.
[0007] According to another embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the steps of any of the above method embodiments when running.
[0008] According to another embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0009] The embodiments provided herein enable real-time collection of performance characteristics of the cloud device hosting the large language model and dynamic prediction using a policy network, overcoming the limitations of static configuration. Predicted execution actions can be directly converted into specific parameter adjustment instructions, such as those executed in the cloud device's registers, rapidly reflecting optimization results. The policy network not only generates predicted execution actions but also self-learns and updates policies based on current and historical states (performance characteristics and historical execution actions), ensuring continuous improvement and adaptability of the tuning strategy. This in turn improves the accuracy of predicted execution actions, thereby achieving the technical effect of improving parameter adjustment effectiveness for cloud devices and resolving the technical issue of poor parameter adjustment effectiveness for cloud devices in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is a hardware structure block diagram of a parameter adjustment method for a cloud device according to an embodiment of the present application;
[0011] Figure 2 is a flow chart of a method for adjusting parameters of a cloud device according to an embodiment of the present application;
[0012] Figure 3 This is a technical diagram of CPU parameter optimization for GPU large language model inference based on model reinforcement learning according to an embodiment of the present application;
[0013] Figure 4 This is a schematic diagram of state space definition and real-time acquisition according to an embodiment of the present application;
[0014] Figure 5 2 is a schematic diagram of a PPO strategy optimization technology for a hybrid action space according to an embodiment of the present application;
[0015] Figure 6 This is a schematic diagram of a lightweight intermediate layer instruction mapping according to an embodiment of the present application;
[0016] Figure 7 2 is a schematic diagram of dynamic adjustment of a multi-objective reward function according to an embodiment of the present application;
[0017] Figure 8This is a schematic diagram of experience playback and online learning according to an embodiment of the present application;
[0018] Figure 9 This is a structural block diagram of a parameter adjustment device for a cloud device according to an embodiment of the present application. DETAILED DESCRIPTION
[0019] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0020] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0021] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal or similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of a computer terminal of a cloud device parameter adjustment method according to an embodiment of the present application. Figure 1 As shown, the computer terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. The computer terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. For example, the computer terminal may also include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0022] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining the mapping relationship in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0023] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a computer terminal's communications provider. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0024] As an optional solution, a parameter adjustment method for cloud devices, such as Figure 2 As shown, the specific steps include:
[0025] S202, obtaining performance characteristics of the cloud device where the large language model is located, wherein the performance characteristics are used to indicate operating characteristics of the cloud device in a working state;
[0026] S204, using the policy network to make predictions based on the performance characteristics and obtain predicted execution actions;
[0027] S206, updating the policy of the policy network based on the performance characteristics and the historical execution actions corresponding to the performance characteristics, wherein the historical execution actions include adjustment actions that have been performed on the performance characteristics of the cloud device;
[0028] S208, performing instruction mapping on the predicted execution action to obtain at least one parameter adjustment instruction corresponding to the predicted execution action;
[0029] S210 , executing at least one parameter adjustment instruction in a register of the cloud device.
[0030] Optionally, in this embodiment, large language models, such as those based on the Transformer architecture, are typically deployed in cloud environments due to their complex structure and large number of parameters, utilizing powerful computing resources (such as GPU clusters) for training and inference. Cloud devices are the physical or virtual hardware facilities that host and execute these models, including computing nodes, storage devices, and network components.
[0031] Optionally, in this embodiment, performance characteristics refer to a series of metrics that can quantify and describe the system operating status and resource utilization efficiency of cloud devices when performing large language model inference or other computing tasks. These characteristics typically include, but are not limited to, GPU memory bandwidth utilization, CPU cache hit rate, average inter-process communication latency, CPU power consumption, and CPU temperature.
[0032] Optionally, in this embodiment, the performance characteristics of cloud devices are an important basis for monitoring and evaluating their current operational status, directly impacting system efficiency, computational latency, energy consumption, and stability. During inference on large language models, changes in performance characteristics can reflect issues such as the effectiveness of GPU and CPU collaboration, the rationality of prefetching strategies, appropriate cache allocation, and optimized thread scheduling.
[0033] Optionally, in this embodiment, the purpose of acquiring performance characteristics is to dynamically adjust parameters on the cloud device to adapt to changes in resource requirements during large language model inference and improve overall system performance. For example, if the CPU cache hit rate is low, it may mean that the cache allocation or prefetching strategy needs to be adjusted; if the GPU utilization rate is low, the thread scheduling strategy may need to be optimized to fully utilize GPU resources; if the CPU temperature is too high, the power consumption control strategy may need to be adjusted to prevent overheating and frequency throttling.
[0034] Optionally, in this embodiment, the policy network is a neural network used in a reinforcement learning framework to predict the optimal action based on the current environment state. In this embodiment, the policy network combines a long short-term memory (LSTM) network and a proximal policy optimization (PPO) algorithm to predict CPU parameter adjustment actions.
[0035] Optionally, in this embodiment, the action prediction is an optimization action predicted by a policy network based on current performance characteristics, including but not limited to adjustment of prefetching aggressiveness, cache allocation ratio, and thread scheduling policy.
[0036] Optionally, in this embodiment, historical execution actions refer to CPU parameter adjustment actions that the system has previously executed under the same or similar performance characteristics. This information is used for continuous learning and policy optimization by the policy network. Predicted execution actions are actions to be executed for adjusting CPU parameters, as predicted by the policy network based on the current situation. Parameter adjustment instructions are instructions converted from predicted execution actions that directly act on CPU registers and are used to dynamically adjust prefetch strategies, cache allocation, thread scheduling, and more.
[0037] Optionally, in this embodiment, hardware sensors (such as GPU power sensors and CPU temperature sensors) are used to monitor GPU memory bandwidth utilization, CPU cache hit rate and other indicators in real time, which are then summarized through a data acquisition module, and a sliding average filter is used to eliminate abnormal fluctuations to generate a state vector reflecting the current system state.
[0038] Optionally, in this embodiment, a policy network combining LSTM and PPO is used. It takes as input the current system state vector and the last action executed, and predicts the next state and reward value. The policy network generates a hybrid action that includes prefetching aggressiveness, cache allocation ratio, and thread scheduling strategy. Gaussian sampling, sigmoid normalization, or softmax processing are used to ensure that the action meets actual operating conditions.
[0039] Optionally, in this embodiment, the advantage function of the action is calculated based on the current state and its corresponding historical action, and the policy gradient is calculated using the PPO algorithm to update the policy network parameters to adapt to the new workload and model characteristics. The target network parameters are synchronized through a soft update mechanism, and a prioritized experience replay mechanism is used to ensure efficient learning.
[0040] Optionally, in this embodiment, the predicted execution action can be interpreted as a high-level action prediction result, mapped via a lightweight intermediate layer to low-level CPU register operation instructions (at least one parameter adjustment instruction), and immediately executed in the cloud device's registers, enabling rapid adjustment of CPU parameters. This mechanism ensures that the policy can be executed across different CPU architectures, maintaining the universality and low overhead of the optimization policy.
[0041] It is understandable that this embodiment constructs an optimization framework based on model reinforcement learning, which can dynamically adjust parameters such as prefetching strategy, cache allocation, and thread scheduling according to the real-time system operation status to achieve the best balance between latency reduction, GPU utilization improvement, and CPU power consumption control. By capturing the timing characteristics of the system state through the LSTM network and combining the PPO algorithm for policy optimization, this method overcomes the limitations of static optimization, realizes global dynamic adjustment of CPU parameters, and improves the efficiency and energy efficiency of large language model reasoning. In addition, through the online learning mechanism, the system can continuously adapt to new model architectures and load changes, and maintain excellent performance without human intervention.
[0042] The embodiments provided herein enable real-time collection of performance characteristics of the cloud device hosting the large language model and dynamic prediction using a policy network, overcoming the limitations of static configuration. Predicted execution actions can be directly converted into specific parameter adjustment instructions, such as those executed in the cloud device's registers, rapidly reflecting optimization results. The policy network not only generates predicted execution actions but also self-learns and updates policies based on current and historical states (performance characteristics and historical execution actions), ensuring continuous improvement and adaptability of the tuning strategy. This in turn improves the accuracy of predicted execution actions, thereby achieving the technical effect of enhancing parameter adjustment effectiveness for cloud devices.
[0043] As an optional solution, updating the policy network based on performance characteristics and historical execution actions corresponding to the performance characteristics includes:
[0044] Perform time series prediction on performance characteristics and historical execution actions to obtain predicted performance characteristics and expected reward parameters of cloud devices;
[0045] Calculate the advantage difference between the predicted performance characteristics, expected reward parameters and performance characteristics to obtain the target advantage parameters of the policy network;
[0046] Using the target advantage parameter, the policy gradient of the policy network is updated.
[0047] Optionally, in this embodiment, time series prediction refers to the process of predicting the future state and reward parameters of the system, and is an important component of policy network learning and decision-making. Leveraging the time series prediction capabilities of the LSTM network, the policy network can predict the state of the large language model and the expected reward after an action is executed based on current and historical data.
[0048] Optionally, in this embodiment, advantage difference calculation is used to assess the quality of an action relative to the average action. Specifically, the weighted sum of the immediate reward and expected future reward of the action is superior to the average reward predicted by the policy network. By calculating advantage difference, the policy network can determine which parameter adjustments are more beneficial, thereby guiding the optimization direction of the policy.
[0049] Optionally, in this embodiment, policy gradient updates are a core step in a policy network-based reinforcement learning algorithm. Their purpose is to update the policy network's parameters to maximize the expected cumulative reward. By using the calculated target advantage parameter, the policy network can adjust its action prediction strategy, achieving more efficient performance optimization for large language model inference.
[0050] Optionally, in this embodiment, an LSTM network is used to perform time-series predictions on performance characteristics and historical execution actions, yielding predicted performance characteristics and expected reward parameters for the cloud device after taking the predicted execution action. This prediction process predicts the next state and reward value based on the current state vector and the previously executed action. The LSTM network's time-series prediction capabilities can capture and model long-term temporal dependencies between system states, providing more accurate predictions for the policy network and facilitating subsequent advantage differential calculations.
[0051] After time series prediction, the predicted performance characteristics and expected reward parameters are compared with the current actual performance characteristics. The advantage function method is used to calculate the advantage difference of the action, namely the target advantage parameter. Here, the advantage difference calculation is not only based on the immediate reward, but also takes into account the long-term expected reward. By discounting future rewards, an assessment of the long-term impact of the current action is formed. The target advantage parameter reflects the additional benefit of taking a certain action compared to the average action predicted by the policy network and is an important basis for the policy network to update its parameters.
[0052] Finally, the calculated target advantage parameters are used to update the policy gradient of the policy network. This process, based on the PPO algorithm, uses gradient descent to update the weight parameters in the policy network, increasing the probability of taking actions that increase reward. By limiting the amplitude of policy updates, the PPO algorithm avoids large policy jumps and ensures the stability of the learning process. By repeatedly repeating this process, the policy network gradually optimizes its parameters, achieving continuous improvement in the inference performance of large language models.
[0053] The embodiments provided in this application overcome the limitations of static parameter configuration when handling dynamic loads, enabling real-time adaptive adjustment of CPU parameters. Through the continuous learning capabilities of the policy network, the system can adapt to changes in different large language model architectures and workload characteristics, maintaining high performance without manual reconfiguration.
[0054] As an optional solution, the advantage difference calculation is performed on the predicted performance characteristics, expected reward parameters, and performance characteristics to obtain the target advantage parameters of the policy network, including:
[0055] Performing state value calculation on the predicted performance feature and the performance feature to obtain a first state value parameter corresponding to the predicted performance feature and a second state value parameter corresponding to the performance feature;
[0056] The target advantage parameter is obtained by multiplying the expected reward parameter by the first state value parameter and subtracting the second state value parameter.
[0057] Optionally, in this embodiment, the predicted performance characteristics are obtained by the policy network through timing prediction of the current system state and historical execution actions, and are expected to be the state that the system will reach after the predicted action is executed in the future, including multiple key indicators such as GPU memory bandwidth utilization and CPU cache hit rate.
[0058] Optionally, in this embodiment, the expected reward parameter is the expected reward value after taking the predicted execution action, based on the future state predicted by the policy network. The reward parameter includes a combination of immediate reward and future discounted reward, which is used to evaluate the long-term benefits of the action.
[0059] Optionally, in this embodiment, state value calculation is used to assess the value of a state, that is, to estimate the expected cumulative reward that can be obtained by following a certain policy from that state. State value is an important concept in reinforcement learning, used to guide the policy network to make better decisions.
[0060] Optionally, in this embodiment, a target advantage parameter is used to measure the additional benefit of taking a specific action compared to the average action quality. It is a key metric used to guide policy updates in reinforcement learning algorithms. The target advantage parameter is calculated based on the difference between the expected reward and two state value parameters, effectively assessing the long-term potential value of the current action.
[0061] Optionally, in this embodiment, a state value calculation is performed for the predicted performance characteristics and the performance characteristics. This step involves calculating the state value parameters for two states: the predicted performance characteristics and the current performance characteristics. The calculation of the state value parameters relies on the value function in the policy network, which predicts the expected cumulative reward from any given state based on the current policy. The value function yields a first state value parameter corresponding to the predicted performance characteristics and a second state value parameter corresponding to the current performance characteristics.
[0062] Based on the obtained first and second state value parameters, the target advantage parameter is calculated using the expected reward parameter.
[0063] Through the embodiments provided in this application, the policy network leverages the time series prediction capabilities of the prediction network (LSTM network) to predict the cloud device state after executing future actions, that is, to predict performance characteristics. Subsequently, based on this predicted state and the immediate reward parameter, the target advantage parameter is derived through state value calculation and reward discounting. This serves as an important basis for the policy network to update its policy gradient. This approach enables the policy network to more accurately assess the long-term value of actions, thereby improving the decision-making quality of parameter adjustments, which is of great significance for achieving dynamic and adaptive inference performance optimization for large language models.
[0064] As an optional solution, using the target advantage parameter, the policy gradient of the policy network is updated by:
[0065] Obtaining a first execution probability corresponding to the predicted execution action and a second execution probability corresponding to the historical execution action, wherein the first execution probability is used to indicate the execution probability of the predicted execution action under the performance characteristic, and the second execution probability is used to indicate the execution probability of the historical execution action before the historical execution action is executed;
[0066] Determine the logarithmic value corresponding to the first execution probability as the first gradient parameter corresponding to the policy network;
[0067] Determine the clipping parameter value corresponding to the ratio of the first execution probability to the second execution probability as the second gradient parameter corresponding to the policy network;
[0068] The policy gradient of the policy network is updated using the product of the first gradient parameter, the second gradient parameter, and the target advantage parameter.
[0069] Optionally, in this embodiment, the first execution probability in the policy network refers to the probability that the policy network predicts a specific action will be taken (predicted action) under the current system state (performance characteristics). The second execution probability, on the other hand, refers to the probability that the policy network predicts that a historical action will be taken before the historical action is executed. These two probabilities are part of the policy network's output and reflect the policy network's inclination toward various possible actions.
[0070] Optionally, in this embodiment, the logarithmic value is mainly used here to calculate the first gradient parameter of the policy network. In reinforcement learning, logarithmic probability is often used to calculate policy gradients because it can simplify the mathematical expression of the gradient, making it easier to understand and calculate.
[0071] Optionally, in this embodiment, the clipping parameter is a mechanism used in reinforcement learning algorithms (such as Proximal Policy Optimization (PPO)) to limit the magnitude of policy updates. By limiting the ratio of the first execution probability to the second execution probability within a specific range, excessive policy adjustments can be avoided, thereby maintaining the stability of the learning process.
[0072] Optionally, in this embodiment, the gradient parameter is the derivative required when the policy network updates its weights (i.e., parameters), used to guide the direction of parameter adjustment. In this embodiment, the first and second gradient parameters are calculated based on the clipping parameter values of the logarithmic execution probability and the execution probability ratio, respectively, and are key elements of the policy network gradient update.
[0073] Optionally, in this embodiment, the first execution probability of the predicted execution action and the second execution probability of the historical execution action are obtained from the policy network. These two probability values reflect the policy network's tendency to take these actions and are the basis for subsequent gradient calculations.
[0074] Next, the logarithm of the first execution probability is determined as the first gradient parameter. The greater the probability that the policy network predicts to execute an action, the more significant the impact of this logarithm on the gradient, which helps the policy network take more of these high-probability actions in the future, thereby achieving performance optimization.
[0075] Then, the clipping parameter value corresponding to the ratio of the first execution probability to the second execution probability is calculated and determined as the second gradient parameter. The clipping parameter is introduced to limit the amplitude of the policy update and prevent the policy network from experiencing excessive fluctuations during the update process, which would affect the stability of the overall performance.
[0076] Finally, the policy network's gradient update is obtained by multiplying the first and second gradient parameters with the calculated target advantage parameter. This product is used to update the policy network's parameters, namely the policy gradient. Specifically, this product is multiplied by the policy network's parameter gradient and adjusted through backpropagation, making the policy network more likely to take actions that result in a higher target advantage in future decisions.
[0077] The policy gradient update mechanism proposed in this application combines information about immediate rewards and expected future rewards. The target advantage parameter measures the long-term value of actions, which, combined with the first and second gradient parameters, guides parameter adjustments in the policy network. This gradient update method ensures that the policy network can achieve fast and stable policy optimization in dynamically changing CPU parameter adjustment tasks, significantly improving the inference efficiency of large language models on GPUs, optimizing inference latency, and controlling CPU power consumption.
[0078] As an optional solution, we can perform time series prediction on the performance characteristics and historical execution actions to obtain the predicted performance characteristics and expected reward parameters of the cloud device, including:
[0079] Obtain the inference delay characteristic parameter of the policy network, and obtain the GPU utilization and CPU power consumption of the cloud device from the performance characteristics, wherein the inference delay characteristic parameter is used to indicate the difference between the ratio of the inference delay of the policy network to the maximum inference delay threshold and 1;
[0080] The expected reward parameter is obtained by adding the first weight multiplied by the inference delay characteristic parameter, adding the second weight multiplied by the GPU utilization, and subtracting the third weight multiplied by the CPU power consumption.
[0081] Optionally, in this embodiment, the policy network's inference delay characteristic parameter is the difference between the inference time delay generated by the policy network when processing the current performance characteristics and generating the predicted execution action and the maximum inference delay threshold and 1. This parameter reflects the speed and efficiency of the policy network in executing decisions.
[0082] Optionally, in this embodiment, GPU utilization and CPU power consumption are two parameters extracted directly from the performance characteristics. GPU utilization reflects the GPU's activity level when processing large language model inference tasks, while CPU power consumption indicates the energy consumption level when executing CPU parameter adjustment strategies. Both are key factors in calculating the expected reward parameter in the reinforcement learning framework.
[0083] Optionally, in this embodiment, different performance indicators are assigned different weights in the calculation of the composite reward function. These weights, namely the first weight, the second weight, and the third weight, determine the relative importance of each performance indicator in the overall reward calculation. The first weight is used to control the impact of policy network inference latency on the expected reward, the second weight is used to emphasize the contribution of GPU utilization, and the third weight is used to balance the negative impact of CPU power consumption.
[0084] Optionally, in this embodiment, a characteristic parameter of the inference delay of the policy network in the current decision cycle is obtained. This parameter is calculated based on the comparison between the inference delay generated by the policy network and its maximum allowable threshold, and is used to reflect the real-time response speed of the policy network.
[0085] When the inference delay of the policy network is close to or lower than the maximum threshold, the inference delay characteristic parameter is close to 0, indicating that the policy network performs well in making quick decisions; conversely, the value of the inference delay characteristic parameter is large, indicating that the decision-making speed of the policy network needs to be improved.
[0086] Next, we extract the GPU utilization and CPU power consumption of cloud devices from the performance characteristics. These are important indicators for evaluating the efficiency of GPU-accelerated inference and CPU energy consumption control.
[0087] Finally, the calculation of the expected reward parameter takes into account the impact of three dimensions: policy network inference latency, GPU utilization, and CPU power consumption.
[0088] Through the embodiments provided in this application, fine-tuning of the policy network is achieved, ensuring that the parameter adjustment strategy can achieve a balance between different performance indicators, and avoiding the problem of over-optimization of one aspect while neglecting other aspects.
[0089] As an optional solution, the method further includes:
[0090] In the case where it is detected that the inference delay of the policy network is greater than a preset delay threshold, increasing the first weight;
[0091] When it is detected that the CPU temperature of the cloud device is greater than a preset temperature threshold, the third weight is increased.
[0092] Optionally, in this embodiment, the preset latency threshold is a predefined maximum allowable value for the policy network's inference latency, used to monitor the policy network's decision-making speed. If the policy network's inference latency exceeds this threshold, it indicates a slow decision-making speed, potentially impacting the real-time inference performance of large language models.
[0093] Optionally, in this embodiment, the preset temperature threshold is a safe upper limit of the CPU temperature. If the CPU temperature exceeds the preset temperature threshold, it may trigger an overheating warning or even cause automatic frequency reduction to protect the hardware, thereby affecting the operating efficiency of the entire system.
[0094] Optionally, in this embodiment, when calculating the expected reward parameter, the first weight controls the influence of the policy network's inference delay characteristic parameter on the expected reward, while the third weight adjusts the importance of CPU power consumption in reward calculation. When encountering situations where inference delay is too high or CPU temperature is overheating, by increasing the first or third weight, the policy network will place greater emphasis on reducing delay or controlling power consumption to ensure stable and efficient system operation.
[0095] Through the embodiments provided in this application, the policy network can quickly adjust its strategy when facing performance bottlenecks or thermal management problems, giving priority to reducing reasoning delays or controlling CPU power consumption. For example, in a high-load environment, if the reasoning delay of the policy network exceeds a preset threshold, increasing the first weight will force the policy network to pay more attention to improving decision-making speed to reduce the impact on the overall reasoning time. Similarly, if the CPU temperature approaches or exceeds the preset temperature threshold, the policy network will give priority to adopting CPU parameter adjustment strategies that can effectively reduce power consumption to prevent hardware overheating by increasing the third weight.
[0096] As an alternative, a policy network is used to make predictions based on performance characteristics. The predicted execution actions include:
[0097] Use the policy network to predict continuous actions based on performance characteristics and obtain multiple continuous action values;
[0098] Gaussian distribution sampling is performed on multiple continuous action values to obtain a prefetch aggressiveness parameter, and nonlinear transformation is performed on multiple continuous action values to obtain a cache allocation ratio parameter;
[0099] Use the policy network to predict discrete actions based on performance characteristics and obtain multiple discrete action values;
[0100] Perform probability normalization conversion on multiple discrete action values to obtain thread scheduling policy parameters;
[0101] The prefetch aggressiveness parameter, cache allocation ratio parameter and thread scheduling strategy parameter are combined to obtain the predicted execution action.
[0102] Optionally, in this embodiment, the policy network output can be divided into two types: continuous actions and discrete actions. Continuous action prediction involves prefetch aggressiveness and cache allocation ratios, which take values continuously within a certain range. Discrete action prediction, on the other hand, focuses on selecting a thread scheduling strategy, typically one of several predefined strategies.
[0103] Optionally, in this embodiment, Gaussian distribution sampling is used to randomly select the prefetch aggressiveness parameter from the continuous action prediction output of the policy network, ensuring randomness and diversity in parameter adjustment. A nonlinear transformation, such as a sigmoid function, is used to convert the raw value output by the policy network into the valid range of the cache allocation ratio parameter, i.e., between 0.1 and 0.9.
[0104] Optionally, in this embodiment, the discrete action prediction output of the policy network is subjected to probability normalization processing to ensure that the sum of the prediction probabilities of all possible thread scheduling strategies is equal to 1, so that random sampling can be performed through the probability distribution to select the most appropriate thread scheduling strategy.
[0105] Optionally, in this embodiment, the policy network predicts multiple continuous action values for prefetch aggressiveness and cache allocation ratio based on current performance characteristics. These values form a continuous action space, providing candidate solutions for subsequent specific parameter adjustments.
[0106] For prefetching aggressiveness, the continuous action prediction values of the policy network are sampled using a Gaussian distribution to obtain a prefetching aggressiveness parameter that is closer to the actual system dynamics. For the cache allocation ratio, the prediction values of the policy network are constrained between 0.1 and 0.9 through nonlinear transformations (such as the Sigmoid function) to generate a reasonable cache allocation ratio parameter.
[0107] The policy network also performs discrete action predictions on the performance characteristics and obtains the predicted values of multiple thread scheduling policies, which constitute the discrete action space of the thread scheduling policy.
[0108] After the policy network predicts multiple discrete action values, they are probabilistically normalized to generate a probability distribution, from which thread scheduling policy parameters are randomly selected. This normalized probability distribution ensures that every discrete action has a chance of being selected, but the probability reflects the policy network's assessment of the quality of the action.
[0109] Finally, the prefetch aggressiveness parameter obtained through Gaussian distribution sampling, the cache allocation ratio parameter after nonlinear transformation, and the thread scheduling policy parameter after probability normalization are combined to form a complete speculative execution action. This action includes comprehensive CPU parameter adjustment recommendations to optimize the inference performance of large language models.
[0110] Through the above steps, the policy network predicts a series of parameter adjustment actions based on the current system state and historical execution actions. The randomness and diversity of these actions are derived from Gaussian distribution sampling and probability normalization. The prefetching aggressiveness controls the depth of the system's data pre-reading, the cache allocation ratio determines the share of GPU-related processes in the CPU cache, and the thread scheduling policy affects how CPU resources are allocated. Ultimately, these processed parameter values are combined into a predicted execution action, which is used to guide the real-time adjustment of CPU parameters to optimize the performance of large language model inference.
[0111] As an optional solution, performing instruction mapping on the speculative execution action to obtain at least one parameter adjustment instruction corresponding to the speculative execution action includes: obtaining, through a mapping function, a first instruction corresponding to a prefetch aggressiveness parameter, a second instruction corresponding to a cache allocation ratio parameter, and a third instruction corresponding to a thread scheduling policy parameter;
[0112] Executing at least one parameter adjustment instruction in a register of the cloud device includes:
[0113] In a prefetch control register of the cloud device, executing a first instruction;
[0114] executing a second instruction in a cache allocation register of the cloud device;
[0115] In the thread scheduling mode register of the cloud device, a third instruction is executed.
[0116] Optionally, in this embodiment, the mapping function is a function that converts high-level parameter adjustment suggestions output by the policy network into low-level register operation instructions. In the present invention, the mapping function converts the prefetch aggressiveness, cache allocation ratio, and thread scheduling policy parameters in the speculative execution action into corresponding first, second, and third instructions, so that they can be directly executed in the registers of the cloud device's CPU.
[0117] Optionally, in this embodiment, the prefetch control register, cache allocation register, and thread scheduling mode register are all internal CPU registers used to control the CPU's prefetch policy, cache allocation policy, and thread scheduling policy. The prefetch control register (e.g., IA32_PREFETCH_CTL in the x86 architecture) is responsible for setting the depth and mode of prefetch operations; the cache allocation register (e.g., the relevant cache management register in the x86 architecture) controls the allocation ratio of the CPU cache; and the thread scheduling mode register (e.g., SCHED_EL1 in the ARM architecture) is used to set the scheduling algorithm and priority of threads across CPU cores.
[0118] Optionally, in this embodiment, after predicting the execution action, the policy network outputs prefetch aggressiveness parameters, cache allocation ratio parameters, and thread scheduling policy parameters. To make these high-level parameter adjustment suggestions effective on physical hardware, a mapping function is needed to convert them into specific instructions that can directly act on the CPU's internal registers, namely the first instruction, the second instruction, and the third instruction.
[0119] The first instruction is sent to the prefetch control register, which adjusts the prefetch aggressiveness. Based on the predictions made by the policy network, the prefetch control register dynamically adjusts the depth and pattern of prefetch operations according to the first instruction, thereby optimizing the data prefetching strategy during inference on large language models, improving data access efficiency, and reducing latency.
[0120] The second instruction operates on the cache allocation register, adjusting the share of the CPU cache allocated to GPU-related processes. This adjustment aims to balance the allocation of CPU cache resources, ensuring fast access to data required for large language model inference on the GPU while also controlling CPU power consumption and preventing overheating.
[0121] The third instruction is sent to the thread scheduling mode register, which sets the optimal thread scheduling policy. This step adjusts the thread scheduling algorithm and priority on the CPU to ensure optimal collaboration between the GPU and CPU, improving overall computing efficiency.
[0122] The embodiments provided in this application transform model decisions into actual improvements in system performance through precise register operations. This not only emphasizes the intelligent generation of parameter adjustment solutions by the policy network, but also demonstrates the efficiency and accuracy of executing parameter adjustment instructions, ensuring closed-loop control of the CPU parameter optimization process and providing a stable, efficient, and energy-efficient operating environment for GPU-based large language model inference.
[0123] As an optional solution, obtaining the performance characteristics of a large language model includes:
[0124] Obtain the GPU memory bandwidth utilization, CPU cache hit rate, average inter-process communication latency, CPU power consumption, and CPU temperature of the cloud device.
[0125] A five-dimensional vector consisting of GPU memory bandwidth utilization, CPU cache hit rate, average inter-process communication latency, CPU power consumption, and CPU temperature is determined as the performance feature.
[0126] Optionally, in this embodiment, GPU memory bandwidth utilization measures the ratio of GPU memory read and write speed to the maximum available bandwidth, intuitively reflecting the GPU's memory access efficiency. CPU cache hit rate indicates how often the CPU retrieves data directly from the cache when accessing data. A higher cache hit rate indicates lower overall system latency.
[0127] Optionally, in this embodiment, the average inter-process communication latency is used to evaluate the average time required to exchange data between different processes during cloud device inference. Communication latency directly affects overall computing efficiency. CPU power consumption indicates the amount of power consumed by the CPU during operation. Power consumption control is crucial for high-performance computing and energy management of mobile devices. CPU temperature refers to the operating temperature of the CPU. Excessive CPU temperature may trigger the system's automatic cooling mechanism, reducing the CPU frequency and affecting performance.
[0128] Optionally, in this embodiment, the following five key performance metrics are collected from live cloud devices: GPU memory bandwidth utilization, CPU cache hit rate, average inter-process communication latency, CPU power consumption, and CPU temperature. This data provides a comprehensive view of the system's current operating status and serves as an important basis for subsequent optimization decisions.
[0129] The five types of data collected above are then integrated into a five-dimensional vector, known as the performance signature. This vector encompasses the key parameters that affect the inference performance of cloud devices when the GPU and CPU work together, providing an accurate representation of the state for model training and policy generation.
[0130] It's understandable that GPU memory bandwidth utilization is directly related to GPU computing efficiency, CPU cache hit rate affects CPU response speed and overall latency, and average inter-process communication latency determines the efficiency of data exchange between different computing units. Furthermore, monitoring CPU power consumption and CPU temperature ensures that the system maintains reasonable energy consumption and a stable operating temperature while pursuing high performance, avoiding performance loss caused by overheating.
[0131] The embodiments provided in this application provide accurate status information for subsequent timing prediction, strategy generation, and parameter adjustment. By constructing a five-dimensional vector containing GPU memory bandwidth utilization, CPU cache hit rate, average inter-process communication latency, CPU power consumption, and CPU temperature, the system can comprehensively assess the current operating status and make further decisions and optimizations based on this status, ensuring the real-time and effectiveness of CPU parameter adjustment strategies and improving the efficiency and stability of large language model reasoning.
[0132] As an optional solution, the aforementioned cloud device parameter adjustment method can be used in the CPU parameter optimization scenario for GPU-based large language model inference based on model reinforcement learning. In this scenario, with the widespread application of large language models in fields such as natural language processing and recommendation systems, optimizing their inference performance has become a key challenge. As the primary computing unit for large language model inference, the performance of the GPU is often limited by the static nature of CPU-side parameter configuration, such as the aggressiveness of the prefetching strategy, the cache allocation ratio, and the thread scheduling strategy. Traditional methods typically use heuristic rules or offline analysis tools for parameter tuning, but such methods cannot adapt to dynamically changing load characteristics and model architecture differences.
[0133] To solve the above problems, this embodiment forms a closed-loop optimization around "state acquisition-timing prediction-strategy generation-instruction execution-reward adjustment-experience learning" and proposes a schematic diagram of GPU large language model inference CPU parameter optimization based on model reinforcement learning, as shown in the following figure: Figure 3 Specifically: First, through step 1 (state collection), the state vector s containing indicators such as GPU memory bandwidth utilization and CPU cache hit rate is collected in real time. t , the result is used as the input of the prediction network in step 2 (time series prediction, such as LSTM time series prediction), combined with the previous action a t-1 Output next state s t+1 and reward value r tThe prediction result of step 2 is input into the PPO strategy network of step 3 to generate a mixed action a including prefetch aggressiveness, cache allocation ratio and thread scheduling strategy t The actions generated in step 3 (strategy generation, such as PPO strategy) are mapped to CPU register operation instructions and executed by the lightweight intermediate layer in step 4 (instruction mapping execution); after execution, step 5 (reward weight adjustment) is performed according to the real-time status s t Dynamically adjust the reward function weights and calculate the compound reward value; finally, in step 6 (experience replay learning), the experience tuples such as state, action, and reward are stored in the buffer, and the policy network is updated using priority experience replay and gradient descent, forming a complete feedback loop from data collection to policy optimization, and realizing real-time adaptive adjustment of CPU parameters.
[0134] Specifically, in this embodiment, a state space definition and real-time acquisition schematic diagram is as follows: Figure 4 As shown in the figure, hardware sensors (such as GPU power sensor, CPU temperature sensor) collect raw data in real time. After being aggregated by the data acquisition module, outliers are removed through sliding average filtering, and finally a state vector s containing 5-dimensional indicators is generated. t .
[0135] The system state vector is shown in formula (1):
[0136] (1)
[0137] in, , U GPU (t) is the GPU memory bandwidth utilization, B used (t) is the real-time bandwidth occupancy, B total (t) is the total bandwidth; , H cache (t) is the CPU cache hit rate, C hit (t) and C miss (t) are the number of hits / misses respectively; , L IPC (t) is the average delay of inter-process communication, T comm.i (t) is the i-th communication delay, n is the number of sampling times; P CPU (t) is the CPU power consumption (W) directly collected by the power sensor; T CPU (t) is the CPU core temperature, obtained by the on-chip temperature sensor.
[0138] In this embodiment, in the LSTM time series prediction network modeling step, the current state s t With the previous action a t-1The input vector is composed and sequentially extracted through two layers of LSTM network to extract the time series features. Finally, the fully connected layer outputs the next state prediction value s t+1 and reward value r t .
[0139] The LSTM unit structure is shown in formula (2), which includes the input gate i t 、Forget t , output gate o t and cell state c t .
[0140] (2)
[0141] in, is the Sigmoid activation function, tanh is the hyperbolic tangent activation function; W is , W fs , W os , W cs is the weight matrix, b is , b fs , b os , b cs is the bias vector; h t is the hidden state, used to predict the next state s t+1 and reward r t : .
[0142] In this embodiment, a schematic diagram of a PPO strategy optimization technology for a hybrid action space is shown in FIG. Figure 5 As shown, the policy network receives the current state vector s t , respectively, by continuous action sampling to generate continuous actions (prefetch aggressiveness, cache allocation ratio) and by discrete action sampling to generate discrete actions (thread scheduling strategy), where continuous action sampling is processed by Gaussian sampling and action clipping, and discrete action sampling is normalized (Softmax normalization), and finally combined into a mixed action a t .
[0143] Optionally, the policy network Output mixed action (continuous value + discrete value):
[0144] Prefetch aggressiveness P prefetch : Sampling by Gaussian distribution, .
[0145] Cache allocation ratio C ratio : Normalized to [0.1, 0.9] by Sigmoid, .
[0146] Thread scheduling strategy T policy : discrete probability distribution, .
[0147] Optionally, the advantage function (GAE) is calculated as shown in formula (3):
[0148] (3)
[0149] in, , is the reward discount factor; is the GAE parameter; is the state value function, which is output by the policy network shared layer.
[0150] Optionally, the policy gradient is updated as shown in formula (4):
[0151] (4)
[0152] in, is the policy gradient target, which is the policy network parameter The gradient of , which is used to update the policy to maximize the expected cumulative reward ; is the expectation operator, for all possible state-action pairs Taking an average, usually approximated by sampling empirical trajectories; is the logarithm of action probability, the current strategy The log probability of choosing action a in state s; is the probability ratio, the new strategy With the old strategy For the same action a, if the probability ratio is > 1, it means that the new strategy is more inclined to choose a; if the ratio is < 1, it means that the new strategy is more inclined to avoid a; is the clipping function, which limits the probability ratio to the interval To prevent the policy from updating too much; Tailor parameters for PPO to ensure policy update stability.
[0153] In this embodiment, a schematic diagram of a lightweight intermediate layer instruction mapping is shown as follows: Figure 6 As shown, mixed action (high-level action) a t The instruction mapping layer converts it into CPU register operation instructions, and through cross-platform compilation, it obtains parameter adjustment instructions (such as the MSR instruction of x86 or the MCR instruction of ARM), which are finally written into the CPU register for execution.
[0154] The action space design is shown in formula (5), which adopts a mixed discrete-continuous action space:
[0155] (5)
[0156] Among them, Pprefetch is the prefetch aggressiveness (0-1 continuous value, 0 is conservative prefetch, 1 is aggressive prefetch); C ratio Allocate a ratio for L3 cache (0.1-0.9 continuous value); T policy Thread scheduling policy (discrete value: 0=round robin, 1=priority).
[0157] The mapping function from high-level actions to low-level register operations is shown in formula (6):
[0158] (6)
[0159] in, The mapping logic is to quantize with 8-bit precision and map it to the CPU prefetch control register (such as IA32_PREFETCH_CTL of x86).
[0160] Level prefetch step size adjustment; To round down; The mapping logic is to convert it into an integer from 1 to 9, corresponding to the L3 cache allocation ratio of 10% to 90% (for example, CachePart=5 means 50% of the cache is allocated to GPU-related processes), and write it into the cache partition register; To round up; The mapping logic is to directly control the mode bit of the CPU thread scheduling register (such as ARM's SCHED_EL1), 0x00 triggers the polling algorithm, and 0x01 enables priority preemption.
[0161] In this embodiment, a schematic diagram of dynamic adjustment of a multi-objective reward function is shown in FIG. Figure 7 As shown, real-time monitoring system status s t , detect whether the delay and temperature are out of limit. If the delay exceeds the threshold, increase the delay weight w1. If the temperature exceeds 80℃, increase the power consumption weight w3. Finally, update the reward function weight.
[0162] Reward function design,The composite reward function is shown in formula (7):
[0163] (7)
[0164] Among them, L t is the current inference delay; U t is the GPU utilization; P t is the CPU power consumption; w1, w2, w3 are adjustable weights.
[0165] The weight adaptation mechanism of the composite reward function is shown in formula (8):
[0166] (8)
[0167] in, : basic weight; : Delay exceeded ( is the delay threshold, which can be adjusted according to the QoS requirements of real-time services); : Temperature exceeds limit ( , increase the power consumption weight when the temperature exceeds the limit threshold to avoid overheating and frequency reduction, which can be adjusted according to project requirements); is an indicator function, which is 1 if the condition is met and 0 otherwise.
[0168] In this embodiment, a schematic diagram of experience playback and online learning is shown in FIG. Figure 8 As shown in the figure, after each decision, the experience tuple (state-action-reward-next-state tuple) is stored in the experience buffer, and the priority experience replay mechanism is used to sample data according to the absolute value of the advantage function. After calculating the gradient, the policy network is updated, and the target network parameters are synchronized through the soft update mechanism.
[0169] Experience tuples are stored as , using Prioritized Experience Replay (PER), as shown in formula (9):
[0170] (9)
[0171] in, Prioritize experience replay weight; is the empirical priority, based on the absolute value of the advantage function; Clipping parameters for PPO to avoid zero probability.
[0172] In this embodiment, the online update strategy is shown in formula (10):
[0173] (10)
[0174] in, , is the learning rate.
[0175] A batch of data is sampled from the buffer every 100ms, and the policy network is updated by gradient descent.
[0176] It should be noted that this embodiment models CPU parameter adjustment as a partially observable Markov decision process (POMDP) problem in a mixed discrete-continuous action space, and implements joint optimization of prefetching strategies, cache allocation, and thread scheduling. The reward function weights can be dynamically adjusted based on system state. The LSTM prediction network supports fine-tuning of some parameters and enables fast migration between different LLM architectures. Fast parameters (such as prefetching) are adjusted in a 10ms cycle, and slow parameters (such as cache allocation) are adjusted in a 100ms cycle. A multi-objective weighted reward mechanism is used to dynamically balance the optimization goals of inference latency, GPU utilization, and CPU power consumption. The reward function weights are automatically adjusted based on system thermal thresholds and power consumption limits. It can model long-term dependencies, predict the impact of different CPU parameter configurations on GPU inference performance, and use the predicted results as state inputs for reinforcement learning. It supports both discrete and continuous parameter adjustment and optimization, including but not limited to: continuous adjustment of prefetch aggressiveness; dynamic partitioning of L3 cache allocation ratios; and discrete selection of thread scheduling strategies.
[0177] It should also be noted that this embodiment can be applied, but not limited to, edge computing scenarios, adapting to resource-constrained devices and supporting real-time inference in the Internet of Things. It can also be applied, but not limited to, multi-tenant cloud service scenarios, independently maintaining LSTM hidden states for each tenant and integrating with a virtualization layer to isolate CPU parameters and avoid resource preemption. It can also be applied, but not limited to, autonomous driving scenarios, dynamically optimizing the CPU configuration for sensor data preprocessing on heterogeneous in-vehicle computing platforms (e.g., CPU+GPU+TPU), thereby reducing autonomous driving latency.
[0178] The embodiments provided in this application transcend the limitations of traditional static parameter configuration and can dynamically adjust multiple interrelated CPU parameters, including prefetching strategies, cache allocation ratios, and thread scheduling strategies, based on real-time system status, to achieve globally optimal system performance. Through a carefully designed composite reward function, the system is able to strike an optimal balance between multiple, mutually constrained objectives: reducing inference latency, improving GPU utilization, and controlling CPU power consumption. The system possesses continuous learning capabilities, maintaining optimal performance without manual intervention.
[0179] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0180] In this embodiment, a parameter adjustment device for a cloud device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments. Details that have already been described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0181] Figure 9 is a structural block diagram of a parameter adjustment device for a cloud device according to an embodiment of the present application, such as Figure 9 As shown, the device includes:
[0182] An acquiring unit 902 is configured to acquire performance characteristics of the cloud device where the large language model resides, wherein the performance characteristics are used to indicate operating characteristics of the cloud device in a working state;
[0183] A prediction unit 904 is configured to use a policy network to make predictions based on the performance characteristics and obtain a predicted execution action;
[0184] An updating unit 906 is configured to update the policy of the policy network based on the performance characteristics and historical execution actions corresponding to the performance characteristics, wherein the historical execution actions include adjustment actions that have been performed on the performance characteristics of the cloud device;
[0185] A mapping unit 908 is configured to perform instruction mapping on the predicted execution action to obtain at least one parameter adjustment instruction corresponding to the predicted execution action;
[0186] The adjustment unit 910 is configured to execute at least one parameter adjustment instruction in a register of the cloud device.
[0187] As an optional solution, the updating unit 906 includes:
[0188] The first prediction module is used to perform time series prediction based on the performance characteristics and historical execution actions to obtain the predicted performance characteristics and expected reward parameters of the cloud device;
[0189] The calculation module is used to calculate the advantage difference between the predicted performance characteristics, expected reward parameters and performance characteristics to obtain the target advantage parameters of the policy network;
[0190] The update module is used to update the policy gradient of the policy network using the target advantage parameters.
[0191] As an optional solution, the calculation module includes:
[0192] A first calculation submodule is configured to perform state value calculation on the predicted performance feature and the performance feature to obtain a first state value parameter corresponding to the predicted performance feature and a second state value parameter corresponding to the performance feature;
[0193] The second calculation submodule is configured to multiply the expected reward parameter by the first state value parameter and then subtract the second state value parameter to obtain a target advantage parameter.
[0194] As an optional solution, update modules include:
[0195] a first acquisition submodule, configured to acquire a first execution probability corresponding to the predicted execution action and a second execution probability corresponding to the historical execution action, wherein the first execution probability is used to indicate the execution probability of the predicted execution action under the performance characteristic, and the second execution probability is used to indicate the execution probability of the historical execution action before the historical execution action is executed;
[0196] A first determining submodule is configured to determine a logarithmic value corresponding to the first execution probability as a first gradient parameter corresponding to the policy network;
[0197] A second determining submodule is configured to determine a clipping parameter value corresponding to the ratio of the first execution probability to the second execution probability as a second gradient parameter corresponding to the policy network;
[0198] The third determination submodule is used to update the policy gradient of the policy network using the product result of the first gradient parameter, the second gradient parameter and the target advantage parameter.
[0199] As an optional solution, the prediction module includes:
[0200] A second acquisition submodule is used to obtain an inference delay characteristic parameter of the policy network, and to obtain a GPU utilization rate and a CPU power consumption rate of the cloud device from the performance characteristics, wherein the inference delay characteristic parameter is used to indicate a difference between a ratio of an inference delay of the policy network to a maximum inference delay threshold and 1;
[0201] The third calculation submodule is configured to obtain an expected reward parameter by adding the value of the first weight multiplied by the inference delay characteristic parameter and the value of the second weight multiplied by the GPU utilization, and then subtracting the value of the third weight multiplied by the CPU power consumption.
[0202] As an optional solution, the device further includes:
[0203] A first increasing module is configured to increase the first weight when detecting that the inference delay of the policy network is greater than a preset delay threshold;
[0204] The second adding module is used to increase the third weight when it is detected that the CPU temperature of the cloud device is greater than a preset temperature threshold.
[0205] As an optional solution, the prediction unit 904 includes:
[0206] The second prediction module is used to use the policy network to perform continuous action prediction on the performance characteristics to obtain multiple continuous action values;
[0207] A sampling module is used to perform Gaussian distribution sampling on multiple continuous action values to obtain a prefetch aggressiveness parameter, and to perform nonlinear transformation on multiple continuous action values to obtain a cache allocation ratio parameter;
[0208] The third prediction module is used to use the policy network to predict discrete actions based on the performance characteristics to obtain multiple discrete action values;
[0209] The conversion module is used to perform probability normalization conversion on multiple discrete action values to obtain thread scheduling policy parameters;
[0210] The combination module is used to combine the prefetch aggressiveness parameter, the cache allocation ratio parameter and the thread scheduling strategy parameter to obtain the predicted execution action.
[0211] As an optional solution, the mapping unit 908 includes:
[0212] A first acquisition module is configured to acquire, through a mapping function, a first instruction corresponding to a prefetch aggressiveness parameter, a second instruction corresponding to a cache allocation ratio parameter, and a third instruction corresponding to a thread scheduling policy parameter;
[0213] The adjustment unit 910 includes:
[0214] A first adjustment module is configured to execute a first instruction in a prefetch control register of the cloud device;
[0215] A second adjustment module is configured to execute a second instruction in a cache allocation register of the cloud device;
[0216] The third adjustment module is configured to execute a third instruction in the thread scheduling mode register of the cloud device.
[0217] As an optional solution, the acquiring unit 902 includes:
[0218] The second acquisition module is used to obtain the GPU memory bandwidth utilization of the cloud device, the CPU cache hit rate of the cloud device, the average inter-process communication delay of the cloud device, the CPU power consumption of the cloud device, and the CPU temperature of the cloud device;
[0219] The determination module is used to determine a five-dimensional vector consisting of GPU memory bandwidth utilization, CPU cache hit rate, average inter-process communication latency, CPU power consumption, and CPU temperature as a performance feature.
[0220] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0221] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0222] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0223] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above method embodiments when run.
[0224] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0225] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0226] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0227] An embodiment of the present application further provides a computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores the computer program product, and when the computer program is executed by a processor, the steps of the method in each embodiment of the present application are implemented.
[0228] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0229] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed across a network composed of multiple computing devices, they can be implemented using program code executable by the computing device, and thus, they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be performed in a different order than herein, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0230] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for adjusting parameters of a cloud device, characterized in that: include: Obtaining performance characteristics of the cloud device where the large language model resides, wherein the performance characteristics are used to indicate operating characteristics of the cloud device in a working state; Using a policy network, making predictions based on the performance characteristics to obtain predicted execution actions; performing a policy update on the policy network based on the performance characteristic and a historical execution action corresponding to the performance characteristic, wherein the historical execution action includes an adjustment action that has been performed on the performance characteristic of the cloud device; Performing instruction mapping on the predicted execution action to obtain at least one parameter adjustment instruction corresponding to the predicted execution action; The at least one parameter adjustment instruction is executed in a register of the cloud device.
2. The method according to claim 1, characterized in that The updating of the policy network based on the performance characteristics and the historical execution actions corresponding to the performance characteristics includes: Performing time series prediction on the performance characteristics and the historical execution actions to obtain predicted performance characteristics and expected reward parameters of the cloud device; performing advantage difference calculation on the predicted performance characteristic, the expected reward parameter, and the performance characteristic to obtain a target advantage parameter of the policy network; The policy gradient of the policy network is updated using the target advantage parameter.
3. The method according to claim 2, characterized in that The performing advantage difference calculation on the predicted performance characteristics, the expected reward parameter, and the performance characteristics to obtain the target advantage parameter of the policy network includes: Performing state value calculation on the predicted performance feature and the performance feature to obtain a first state value parameter corresponding to the predicted performance feature and a second state value parameter corresponding to the performance feature; The target advantage parameter is obtained by multiplying the first state value parameter by the expected reward parameter and subtracting the second state value parameter.
4. The method according to claim 2, characterized in that Updating the policy gradient of the policy network using the target advantage parameter includes: Obtaining a first execution probability corresponding to the predicted execution action and a second execution probability corresponding to the historical execution action, wherein the first execution probability is used to indicate the execution probability of the predicted execution action under the performance characteristic, and the second execution probability is used to indicate the execution probability of the historical execution action before the historical execution action is executed; Determine the logarithmic value corresponding to the first execution probability as a first gradient parameter corresponding to the policy network; Determine a clipping parameter value corresponding to the ratio of the first execution probability to the second execution probability as a second gradient parameter corresponding to the policy network; The policy gradient of the policy network is updated using a product result of the first gradient parameter, the second gradient parameter, and the target advantage parameter.
5. The method according to claim 2, characterized in that The performing time series prediction on the performance characteristics and the historical execution actions to obtain the predicted performance characteristics and expected reward parameters of the cloud device includes: Obtaining an inference delay characteristic parameter of the policy network, and obtaining a GPU utilization and a CPU power consumption of the cloud device from the performance characteristics, wherein the inference delay characteristic parameter is used to indicate a difference between a ratio of an inference delay of the policy network to an inference delay maximum threshold and 1; The expected reward parameter is obtained by multiplying the first weight by the value of the inference delay characteristic parameter, adding the second weight by the value of the GPU utilization, and subtracting the third weight by the value of the CPU power consumption.
6. The method according to claim 5, characterized in that The method further comprises: In a case where it is detected that the inference delay of the policy network is greater than a preset delay threshold, increasing the first weight; When it is detected that the CPU temperature of the cloud device is greater than a preset temperature threshold, the third weight is increased.
7. The method according to claim 1, characterized in that The using policy network to make predictions based on the performance characteristics to obtain predicted execution actions includes: Using the policy network, performing continuous action prediction on the performance characteristics to obtain a plurality of continuous action values; Performing Gaussian distribution sampling on the plurality of continuous action values to obtain a prefetch aggressiveness parameter, and performing nonlinear transformation on the plurality of continuous action values to obtain a cache allocation ratio parameter; Using the policy network, performing discrete action prediction on the performance characteristics to obtain multiple discrete action values; Performing probability normalization conversion on the multiple discrete action values to obtain thread scheduling policy parameters; The prefetch aggressiveness parameter, the cache allocation ratio parameter, and the thread scheduling policy parameter are combined to obtain the predicted execution action.
8. The method according to claim 7, characterized in that The performing instruction mapping on the predicted execution action to obtain at least one parameter adjustment instruction corresponding to the predicted execution action includes: Obtaining, through a mapping function, a first instruction corresponding to the prefetch aggressiveness parameter, a second instruction corresponding to the cache allocation ratio parameter, and a third instruction corresponding to the thread scheduling policy parameter; Executing the at least one parameter adjustment instruction in the register of the cloud device includes: executing the first instruction in a prefetch control register of the cloud device; executing the second instruction in a cache allocation register of the cloud device; The third instruction is executed in the thread scheduling mode register of the cloud device.
9. The method according to any one of claims 1 to 8, characterized in that The performance characteristics of the cloud device where the large language model is obtained include: Obtaining the GPU memory bandwidth utilization of the cloud device, the CPU cache hit rate of the cloud device, the average inter-process communication latency of the cloud device, the CPU power consumption of the cloud device, and the CPU temperature of the cloud device; A five-dimensional vector consisting of the GPU memory bandwidth utilization, the CPU cache hit rate, the average inter-process communication delay, the CPU power consumption, and the CPU temperature is determined as the performance feature.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Automatic operator tuning method based on reinforcement learning and related device
CN116861957A
Action generation model training method and device, electronic equipment and storage medium
CN116966585A
Model parameter adjusting method
CN117763318A
Data prefetching method and device, electronic equipment, electronic device and medium
CN118093020A
Predictive cloud platform resource scheduling method based on staged strategy gradient
CN118193209A