Method and system for dynamic frequency adjustment of CPU based on multi-agent reinforcement learning
By employing a multi-agent reinforcement learning method in data center servers to dynamically adjust CPU frequency, the problem of excessive CPU energy consumption is solved. This achieves energy reduction while meeting performance requirements, avoiding performance loss and power waste in traditional methods.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2023-12-13
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, CPUs in data center servers consume excessive energy. Traditional energy-saving methods result in performance loss and power waste, making it difficult to effectively reduce energy consumption while meeting performance requirements.
By employing a multi-agent reinforcement learning approach, each CPU core is treated as an agent to construct a local policy network and a value network. The CPU frequency is then dynamically adjusted to achieve global load awareness and precise frequency setting.
It achieves precise adjustment of CPU frequency while meeting server computing performance requirements, avoiding performance overkill, achieving energy saving, and reducing CPU power consumption.
Smart Images

Figure CN117687497B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer technology and data center server energy-saving technology, and in particular to a CPU dynamic frequency adjustment method and system based on multi-agent reinforcement learning. Background Technology
[0002] Data centers are crucial infrastructure for the IT infrastructure of large enterprises. With the development of industrialization and informatization in my country, domestic enterprises' investment in IT infrastructure has been continuously increasing. Simultaneously, the emergence of new internet service demands, coupled with large-scale investment in infrastructure construction driven by the development of cloud computing, have led to the rapid growth of data center scale, a key component of information and communication technology (ICT). Energy consumption has become a critical issue that must be addressed in the development of ICT. Server energy consumption accounts for a significant proportion of data center energy consumption, and the CPU is arguably the most power-intensive component in a server, with peak CPU power accounting for approximately 60% of the server's total power consumption under certain conditions. Therefore, achieving green computing and improving server energy efficiency in data centers has become a key and urgent research topic.
[0003] Currently, the main energy-saving methods for data center servers include Dynamic Power Management (DPM) and Dynamic Voltage and Frequency Scaling (DVFS). DPM is a widely used system-level low-power technology. Essentially, it reduces system energy consumption by dynamically setting system components to low-power operating states based on changes in system workload, while meeting user performance requirements. DVFS utilizes the characteristics of CMOS chips, where energy consumption is proportional to the square of the voltage and the clock frequency. DVFS reduces system energy consumption at the cost of extending task execution time, reflecting a trade-off between power consumption and performance. While reducing clock frequency can lower the power consumption of general-purpose processors, simply lowering the clock frequency does not save energy because the performance reduction increases task execution time. Adjusting voltage requires adjusting the frequency proportionally to meet signal propagation delay requirements. Both voltage and frequency adjustments result in a loss of system performance and increased system response latency. Summary of the Invention
[0004] The purpose of this invention is to provide a CPU dynamic frequency adjustment method and system based on multi-agent reinforcement learning. In the case of data centers with a large number of storage and security servers, this invention solves the problem of excessive CPU energy consumption by dynamically adjusting the CPU frequency to meet server computing performance requirements while reducing CPU power consumption.
[0005] The objective of this invention is achieved through the following technical solution:
[0006] A CPU dynamic frequency modulation method based on multi-agent reinforcement learning includes:
[0007] The construction of a multi-agent reinforcement learning framework includes: treating each CPU core as an individual agent, treating the CPU as a whole as a global policy network, and having a local policy network and a value network for each agent;
[0008] At each time step, each agent acquires CPU-related state information and uses its local policy network to decide on the corresponding action. It then calculates the reward information based on the changes in state information after executing the action. Its value network uses the CPU-related state information, action, and reward information to calculate the corresponding state-action value, and updates its local policy network and value network accordingly. After each agent repeats this process for multiple time steps, the global policy network is updated by combining the updated local policy networks of all agents. Afterward, each agent uses the updated global policy network to update its own local policy network. This process is repeated until the final global policy network is obtained.
[0009] The frequency of each CPU core is dynamically adjusted using the final global policy network.
[0010] A CPU dynamic frequency modulation system based on multi-agent reinforcement learning includes:
[0011] The learning framework building unit, used to build a multi-agent reinforcement learning framework, includes: treating each CPU core as an individual agent, treating the CPU as a whole as a global policy network, and each agent having a local policy network and a value network;
[0012] The training unit is used to train the final global policy network. The training process is as follows: At each time step, each agent acquires CPU-related state information and uses its own local policy network to decide on the corresponding action. Then, it calculates the reward information based on the changes in state information after executing the corresponding action. Its own value network uses CPU-related state information, action information, and reward information to calculate the corresponding state-action value, and updates its own local policy network and value network accordingly. After each agent repeats this process for multiple time steps, the updated local policy networks of all agents are combined to update the global policy network. After that, each agent uses the updated global policy network to update its own local policy network. This process is repeated continuously to obtain the final global policy network.
[0013] The CPU dynamic frequency adjustment unit is used to dynamically adjust the frequency of each CPU core using the final global policy network.
[0014] A processing device includes: one or more processors; and a memory for storing one or more programs;
[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0016] A readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.
[0017] As can be seen from the technical solution provided by the present invention, by using the multi-agent reinforcement learning dynamic frequency adjustment method, each core in the server CPU is assigned a corresponding agent, so that the local core of the CPU not only has the ability to perceive the system load globally, but also can actively adjust the frequency according to the load. Compared with traditional energy-saving methods, the more precise frequency setting can meet the performance requirements while avoiding performance overkill and power waste, thereby achieving the effect of server energy saving. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a CPU dynamic frequency modulation method based on multi-agent reinforcement learning, provided as an embodiment of the present invention;
[0020] Figure 2 A framework diagram of the multi-agent reinforcement learning scheme provided in an embodiment of the present invention;
[0021] Figure 3 A schematic diagram of a CPU dynamic frequency adjustment system based on multi-agent reinforcement learning provided in an embodiment of the present invention;
[0022] Figure 4 This is a schematic diagram of a processing device provided in an embodiment of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0024] First, the following explanations are provided for the terms that may be used in this article:
[0025] The terms “including,” “comprising,” “containing,” “having,” or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, “including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.)” should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.
[0026] The following is a detailed description of a CPU dynamic frequency adjustment scheme based on multi-agent reinforcement learning provided by this invention. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they shall be performed according to conventional conditions in the art or conditions recommended by the manufacturer. Reagents or instruments used in the embodiments of this invention, unless otherwise specified by the manufacturer, are all conventional products that can be purchased commercially.
[0027] Example 1
[0028] This invention provides a CPU dynamic frequency adjustment method based on multi-agent reinforcement learning, such as... Figure 1 As shown, it mainly includes:
[0029] 1. Construct a multi-agent reinforcement learning framework.
[0030] In this embodiment of the invention, the multi-agent reinforcement learning framework includes: treating each CPU core as a separate agent, treating the CPU as a whole as a global policy network, and each agent having a local policy network and a value network.
[0031] 2. Obtain the final global policy network through reinforcement learning.
[0032] In this embodiment of the invention, at each time step, each agent acquires CPU-related state information, makes a corresponding action using its local policy network, calculates reward information based on the changes in state information after executing the corresponding action, and calculates the corresponding state-action value using its own value network using the CPU-related state information, action, and reward information, and updates its own local policy network and value network accordingly. Each agent repeats multiple time steps, integrates the updated local policy networks of all agents to update the global policy network, and then each agent uses the updated global policy network to update its own local policy network. This process is repeated continuously to obtain the final global policy network.
[0033] In this embodiment of the invention, the convergence of the scene can be determined based on the collected data, thereby setting the maximum number of training iterations to obtain the final global policy network. For example, 2000 rounds can be set, with each update of the global policy network considered one round.
[0034] As an example, the global policy network can be updated by combining the local policy network updates of all agents in the following way: Each agent calculates the loss function of its own local policy network, obtains its own independent gradient, and uploads it to the global policy network. The global network then updates its parameters by performing gradient descent based on the received gradients.
[0035] In this embodiment of the invention, the CPU-related state information acquired by each agent includes: the frequency and utilization of its CPU core, as well as CPU power consumption, CPU temperature, and system load; the first two items are unique to each agent, while the latter three are shared by all agents. In this embodiment of the invention, the determined action is the CPU core frequency, that is, adjusting the CPU core frequency by executing the action.
[0036] In this embodiment of the invention, the local policy network is represented as π(s) i ;θ i ), i∈(1,n), where n is the number of CPU cores; θ i For the parameters of the local policy network in the i-th core agent, updating the local policy network is equivalent to updating the parameters of the local policy network; action a i =π(s) i ;θ i ), representing the local policy network π(s) i ;θ i Based on parameter θ i Utilizing CPU-related status information s i Output action a i .
[0037] In this embodiment of the invention, calculating the reward information based on the changes in state information after performing the corresponding action includes: calculating a first part of the reward based on the changes in CPU utilization and system load after performing the corresponding action; calculating a second part of the reward based on CPU power consumption and CPU temperature after performing the corresponding action; and combining the two parts of the reward to obtain the reward information r, expressed as:
[0038] r = w1*T(R) + w2*C(P)
[0039] Where T(R) is the first part of the reward, C(P) is the second part of the reward, and w1 and w2 are the weight coefficients corresponding to the two parts of the reward.
[0040] In this embodiment of the invention, the value network is q(s) i a i r i μ i ), i∈(1,n), μ i Let s be the parameters of the value network in the i-th core agent. i a i r i These represent the state information, action, and reward information of the i-th core agent, respectively. The higher the output state-action value (q-value), the better the corresponding action.
[0041] In this embodiment of the invention, updating the local policy network for each agent includes: first, calculating the value loss L. p And strategy loss L p The loss gradient is calculated, and then the parameters of the value network and the local policy network are updated using the loss gradient.
[0042] Among them, the value loss L p And strategy loss L p The formula for calculating ′ is:
[0043] L p =Σ(RV(s)) 2
[0044] L p =log(π(s))A(s)βH(π)
[0045] Among them, the value loss L p The parameters used to update the value network, policy loss L p ' is used to update the parameters of the local policy network; R is the discounted reward, R = r n +γr n-1 +γ 2 r n-2This is the sum of the current reward and the previous two rewards, where γ is the discount factor, r is the reward, and n is the training step number. V(s) is the state-action value output by the value network based on the current state s. π(s) is the output (i.e., action) of the local policy network π based on the current state s, A(s) = RV(s), where A is the advantage function used to calculate the policy gradient, H(π) is the entropy used to ensure that the policy is fully explored, and β is the entropy coefficient.
[0046] 3. Use the final global policy network to dynamically adjust the frequency of each CPU core.
[0047] The above-mentioned solution provided by the embodiments of the present invention uses a multi-agent reinforcement learning dynamic frequency adjustment method, treating each core in the server CPU as an agent. This enables the local cores of the CPU to not only have the ability to perceive the system load globally, but also to actively adjust the frequency according to the load. Compared with traditional energy-saving methods, the more precise frequency setting can meet the performance requirements while avoiding excessive performance that leads to wasted power consumption, thereby achieving the effect of server energy saving.
[0048] To more clearly demonstrate the technical solution and its effects provided by the present invention, the method provided by the embodiments of the present invention will be described in detail below with reference to specific examples.
[0049] I. Purpose and General Overview of the Invention
[0050] The purpose of this invention is to address the problem of excessive CPU power consumption in data centers with a large number of storage and security servers. By treating each CPU core as an agent, a dynamic CPU frequency adjustment method based on multi-agent reinforcement learning is proposed to reduce CPU power consumption while meeting server computing performance requirements.
[0051] In general, this invention addresses situations where a server CPU cannot adjust to its optimal voltage and frequency under the current system load, resulting in performance overkill. It uses the core of the server CPU as an intelligent agent to acquire environmental information such as the frequency, utilization, power consumption, temperature, and system load of each core. This allows the CPU to learn about the environment and adjust itself to the optimal frequency suitable for the current load, thereby achieving the goal of satisfying performance while reducing CPU power consumption.
[0052] II. Specific Plan Contents.
[0053] 1. Design reinforcement learning parameters suitable for the server.
[0054] State represents environmental information input by the agent application. It is represented as collectable CPU runtime information and the server's task load environment at a given moment, which can be dynamically measured from the server. State includes the server CPU's individual core frequencies (CoreFreq), utilization (CoreUtilization), power consumption (CPUConsumption), CPU temperature (CPUTemperature), and system load (SystemLoad).
[0055] State=(CoreFreq, CoreUtilization, CPUConsumption, CPUTemperature,
[0056] SystemLoad)
[0057] In the above information, frequency and utilization are specific to each core, meaning that the information obtained by each agent may be different. However, CPU power consumption, CPU temperature, and system load are information for the entire CPU as a whole. Therefore, at each time step, the information obtained by each agent is the same.
[0058] The statistical data is normalized and then provided to the agent to help it better learn the environment under different load scenarios. Actions are defined as CPU core frequencies, which must be within the variable frequency range that the server CPU can be set to.
[0059] a = {CoreFreq}
[0060] To ensure the performance of this reinforcement learning, a reward function can be defined as follows. One part is the CPU performance reward, denoted as T(R). The CPU's DVFS-based policy aims to improve processor performance during operation; performance is the policy's objective. Here, CPU utilization and system load are used as performance rewards. The other part is the reward, denoted as C(P), which represents the CPU's operating temperature and energy consumption. A small reward (positive reward value) is given when these two metrics are maintained at a good level, but a larger penalty (negative reward value) is given when the processor reaches its power consumption and temperature limits.
[0061] r = w1*T(R) + w2*C(P)
[0062] Wherein, the weight parameter w1+w2=1
[0063] 2. Obtain and calculate the parameters for reinforcement learning.
[0064] The parameters for the design described in Part 1 above are obtained or calculated in the following ways:
[0065] (1) The frequency of each core.
[0066] The CPU core frequency is read from the underlying path " / sys / devices / system / cpu / cpux / cpufreq / scaling_cur_freq". This path is an absolute path, which can directly provide the frequency information of each core.
[0067] (2) Utilization rate.
[0068] Utilization is calculated by retrieving the values of various parameters (user, nice, system, idle, iowait, irq, and softirq) from the / proc / stat file. To ensure data accuracy, the values are retrieved periodically (e.g., every 10ms) and subtracted from the corresponding previously retrieved value. The difference is used to calculate utilization, expressed as:
[0069] CoreUtilization
[0070] =[(user new -user old )+(nice new -nice old )+(system new -system old )+(irq new -irq old )+(softirq new -softirq old )] / [(user new -user old )+(nice new -nice old )+(system new -system old )+(idle new idle old )+(iowait new -iowait old )+(irq new -irq old )+(softirq new -softirq old )]
[0071] Wherein, user is the cumulative execution time of a normal process in user mode; nice is the cumulative execution time of a NICED process in user mode; system is the cumulative execution time of a process in kernel mode; idle is the cumulative idle time; iowait is the cumulative time spent waiting for I / O completion; irq is the hard interrupt time; softirq is the soft interrupt time; the subscript new indicates the most recently acquired parameter value, and the subscript old indicates the last acquired parameter value.
[0072] (3) CPU power consumption.
[0073] Obtain real-time CPU power consumption using the ipmitool tool.
[0074] (4) CPU temperature.
[0075] CPU temperature is obtained using the Python library psutil, specifically using `psutil.sensors_temperatures().get('coretemp')`.
[0076] (5) System load.
[0077] Get the server's average load over a period of time (e.g., 1 minute) using "psutil.getloadavg()[0]".
[0078] 3. Multi-agent reinforcement learning scheme.
[0079] To allow each CPU core to adjust its frequency to suit the current server operating environment, each CPU core is treated as an individual agent, employing an actor-critic architecture. Each agent has its own local policy network (actor) π(s) i ;θ i ), where i∈(1,n), n is the number of CPU cores, and θ is the parameter of the local policy network in the agent to be trained. Action a is obtained by inputting the state information s into the local policy network. i =π(s) i ;θ i The intelligent agent issues action a. i Later received a reward r i Then each agent will state s i and action a i and rewards r i Together, they transmit their own value network (critic) q(s) i ,a i r i μ i ), μ iThese are the parameters of the value network that needs to be trained. Then, after every set number of steps (e.g., 10 steps), the network parameters of the local CPU cores are updated to the global CPU network model, and the local network then pulls the parameters of the global network for updates and training.
[0080] 4. Application stage.
[0081] After the reinforcement learning process described above, the final global policy network can be obtained. This final global policy network is then applied to the CPU, and the frequency of each CPU core is dynamically adjusted according to the corresponding reinforcement learning agent.
[0082] like Figure 2 The diagram illustrates the framework of a multi-agent reinforcement learning scheme. Each CPU core acquires the current server environment state every 10ms. The collected environment state parameters are then calculated using formulas, normalized, and passed as State information to the CPU core's local policy network. Each agent receives different information: frequency and utilization are specific to each CPU core, while CPU power consumption, CPU temperature, and system load are the same. The states of each agent are denoted as observation1 to observationn. The local policy network generates actions a (action 1 to action n) based on the acquired state information and then calls "cpufreq-set-c{ith}-f{action}GHz" to distribute the actions and change the core's frequency. Simultaneously, it calculates reward information (reward 1 to rewardn). During training of a single core's local policy network, after learning an optimal policy using an exploration policy (the exploration policy here represents the policy network learning process; the optimal policy can be considered the policy result obtained after training with collected data), it calculates the value loss and policy loss, then calculates the loss gradient, and subsequently updates the parameters of the local policy network. After each agent repeats the process for multiple time steps (e.g., 10 time steps), the parameters of the global policy network are updated by combining the parameters of the local policy networks updated by all agents. Because multiple agents interact with the environment and integrate their information into the global network, the correlation between experiences is small. Furthermore, the training of the global policy network can be accelerated by parallel learning by multiple agents. Finally, the trained global policy network is run on a server CPU, allowing the CPU to select the optimal frequency based on environmental characteristics to achieve energy savings.
[0083] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0084] Example 2
[0085] This invention also provides a CPU dynamic frequency adjustment system based on multi-agent reinforcement learning, which is mainly used to implement the methods provided in the foregoing embodiments, such as... Figure 3 As shown, the system mainly includes:
[0086] The learning framework building unit, used to build a multi-agent reinforcement learning framework, includes: treating each CPU core as an individual agent, treating the CPU as a whole as a global policy network, and each agent having a local policy network and a value network;
[0087] The training unit is used to train the final global policy network. The training process is as follows: At each time step, each agent acquires CPU-related state information and uses its own local policy network to decide on the corresponding action. Then, it calculates the reward information based on the changes in state information after executing the corresponding action. Its own value network uses CPU-related state information, action information, and reward information to calculate the corresponding state-action value, and updates its own local policy network and value network accordingly. After each agent repeats this process for multiple time steps, the updated local policy networks of all agents are combined to update the global policy network. After that, each agent uses the updated global policy network to update its own local policy network. This process is repeated continuously to obtain the final global policy network.
[0088] The CPU dynamic frequency adjustment unit is used to dynamically adjust the frequency of each CPU core using the final global policy network.
[0089] The specific technical details of each unit in the above system have been described in detail in the method embodiments provided above, and therefore will not be repeated here.
[0090] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0091] Example 3
[0092] The present invention also provides a processing device, such as Figure 4 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.
[0093] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0094] In this embodiment of the invention, the specific types of the memory, input device, and output device are not limited; for example:
[0095] Input devices can be touchscreens, image acquisition devices, physical buttons, or mice, etc.
[0096] The output device can be a display terminal;
[0097] The memory can be random access memory (RAM) or non-volatile memory, such as disk storage.
[0098] Example 4
[0099] The present invention also provides a readable storage medium storing a computer program that, when executed by a processor, implements the method provided in the foregoing embodiments.
[0100] In this embodiment of the invention, the readable storage medium is a computer-readable storage medium and can be disposed in the aforementioned processing device, for example, as a memory in the processing device. Furthermore, the readable storage medium can also be any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0101] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A CPU dynamic frequency adjustment method based on multi-agent reinforcement learning, characterized in that, include: The construction of a multi-agent reinforcement learning framework includes: treating each CPU core as an individual agent, treating the CPU as a whole as a global policy network, and having a local policy network and a value network for each agent; At each time step, each agent acquires CPU-related state information and uses its local policy network to decide on the corresponding action. It then calculates the reward information based on the changes in state information after executing the action. Its value network uses the CPU-related state information, action, and reward information to calculate the corresponding state-action value, and updates its local policy network and value network accordingly. After each agent repeats this process for multiple time steps, the global policy network is updated by combining the updated local policy networks of all agents. Afterward, each agent uses the updated global policy network to update its own local policy network. This process is repeated until the final global policy network is obtained. The frequency of each CPU core is dynamically adjusted using the final global policy network.
2. The CPU dynamic frequency adjustment method based on multi-agent reinforcement learning according to claim 1, characterized in that, The CPU-related status information acquired by each agent includes: the frequency and utilization of its CPU core, as well as CPU power consumption, CPU temperature and system load.
3. The CPU dynamic frequency adjustment method based on multi-agent reinforcement learning according to claim 1, characterized in that, The decision-making action is the CPU core frequency, that is, adjusting the CPU core frequency by executing the action.
4. The CPU dynamic frequency adjustment method based on multi-agent reinforcement learning according to claim 1, characterized in that, Local policy network is represented as π(s) i ;θ i ), i∈(1,n), where n is the number of CPU cores; θ i For the parameters of the local policy network in the i-th core agent, updating the local policy network is equivalent to updating the parameters of the local policy network; action a i =π(s) i ;θ i ), representing the local policy network π(s) i ;θ i Based on parameter θ i Utilizing CPU-related status information s i Output action a i .
5. The CPU dynamic frequency adjustment method based on multi-agent reinforcement learning according to claim 1, characterized in that, The calculation of reward information based on changes in state information after performing the corresponding action includes: The first part of the reward is calculated based on the changes in CPU utilization and system load after the corresponding action is performed. The second part of the reward is calculated based on the CPU power consumption and CPU temperature after the corresponding action is performed. The reward information r is obtained by combining the two parts of the reward, as follows: r = w1*T(R) + w2*C(P) Where T(R) is the first part of the reward, C(P) is the second part of the reward, and w1 and w2 are the weight coefficients corresponding to the two parts of the reward.
6. The CPU dynamic frequency adjustment method based on multi-agent reinforcement learning according to claim 1, characterized in that, The value network is q(s) i ,a i ,r i μ i ), i∈(1,n), where n is the number of CPU cores, μ i Let s be the parameters of the value network in the i-th core agent. i ,a i ,r i These represent the state information, action, and reward information of the i-th core agent, respectively. The higher the value of the output state and action, the better the corresponding action.
7. A CPU dynamic frequency adjustment method based on multi-agent reinforcement learning according to claim 1 or 6, characterized in that, Updating the local policy network for each agent includes: First, calculate the loss L. p And strategy loss L p ′, and calculate the loss gradient, and then use the loss gradient to update the parameters of the value network and the local policy network; Among them, the value loss L p And strategy loss L p The formula for calculating ′ is: L p =∑(R-V(s)) 2 L p ′=log(π(s))A(s)βH(π) Among them, the value loss L p The parameters used to update the value network, policy loss L p ' is used to update the parameters of the local policy network; R is the discount reward, V(s) is the state-action value output by the value network based on the current state s, π(s) is the output of the local policy network π based on the current state s, A is the advantage function used to calculate the policy gradient, H(π) is the entropy used to ensure that the policy is fully explored, and β is the entropy coefficient.
8. A CPU dynamic frequency modulation system based on multi-agent reinforcement learning, characterized in that, include: The learning framework building unit, used to build a multi-agent reinforcement learning framework, includes: treating each CPU core as an individual agent, treating the CPU as a whole as a global policy network, and each agent having a local policy network and a value network; The training unit is used to train the final global policy network. The training process is as follows: At each time step, each agent acquires CPU-related state information and uses its own local policy network to decide on the corresponding action. Then, it calculates the reward information based on the changes in state information after executing the corresponding action. Its own value network uses CPU-related state information, action information, and reward information to calculate the corresponding state-action value, and updates its own local policy network and value network accordingly. After each agent repeats this process for multiple time steps, the updated local policy networks of all agents are combined to update the global policy network. After that, each agent uses the updated global policy network to update its own local policy network. This process is repeated continuously to obtain the final global policy network. The CPU dynamic frequency adjustment unit is used to dynamically adjust the frequency of each CPU core using the final global policy network.
9. A processing device, characterized in that, include: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1 to 7.
10. A readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.