Mobile equipment computing power scheduling method based on reinforcement learning
By applying a computing power scheduling method based on reinforcement learning on mobile devices, dynamically adjusting the CPU and GPU frequency, the problem of inaccurate computing power scheduling in the existing technology is solved, and more efficient system performance and energy consumption balance is achieved.
Patent Information
- Application Number
- CN202411754124.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-02
AI Technical Summary
The existing mobile device computing power scheduling mechanism cannot accurately detect excessive or insufficient frequency allocation, which makes it difficult to balance performance and energy consumption, and often causes equipment lag or high temperatures.
Using a mobile device computing power scheduling method based on reinforcement learning, the CPU and GPU frequency are dynamically adjusted according to the system characteristics of the mobile device to optimize system performance through deep deterministic policy gradient algorithm (DDPG) and reinforcement learning model.
It realizes accurate computing power scheduling decisions based on system status and operating resource requirements, reduces frame loss rate and power consumption, and avoids equipment lag and high temperature problems.
Smart Images

Figure CN119917255A_ABST
Abstract
Description
Technical field:
[0001] The present invention relates to the technical field of operating systems, and in particular to a method for scheduling computing power of a mobile device based on reinforcement learning. Background technology:
[0002] As of 2024, mobile device terminals based on the Android operating system have accounted for more than 70% of the market share. The Android operating system is an operating system derived from the Linux kernel and dedicated to human-computer interaction. It is commonly used in mobile devices such as smartphones. Mobile devices have limited battery capacity, and the usage scenarios often include a large amount of human-computer interaction. Users can clearly perceive the frame loss, power consumption and heat generation of mobile devices. Therefore, the two indicators of frame loss rate and power consumption are often used to evaluate the system performance of mobile devices.
[0003] The computing resources of mobile devices, namely the CPU and GPU frequencies, will directly affect the load capacity of the device and are directly related to the frame loss rate and power consumption of the device. Existing mobile devices usually adopt the ARM large and small core heterogeneous computing architecture. Common mobile phone processors are usually equipped with both energy-saving small cores and powerful but power-consuming large cores. The Android operating system uses this computing architecture to make full use of small core computing to reduce power consumption when the load is low; when the load is high, it uses large core computing to meet the computing power requirements of applications. In terms of computing power scheduling, the CPU performance adjustment (CPUFreq) and device frequency adjustment (devfreq) subsystems in the Linux kernel both provide frequency managers to dynamically adjust the CPU and GPU frequencies according to the application load. The Android Open Source Project (AOSP) provides a powerHint interface that allows device manufacturers to implement customized computing power scheduling policies. Device manufacturers usually provide SDK interfaces that allow applications to request the operating system to increase the processor frequency when facing sudden loads. However, existing computing power scheduling mechanisms are usually based on empirical rules or simple heuristic algorithms, and are scheduled according to predefined policies. These policies only rely on limited features (such as CPU utilization) to make passive and responsive scheduling decisions. They cannot accurately detect situations where the frequency of mobile devices is over-allocated (wasting power) or under-allocated (cannot meet user needs), and cannot dynamically perform accurate computing power scheduling based on the computing power requirements of the device or job behavior. It is difficult to achieve a balance between performance and energy consumption, which often leads to insufficient or excessive computing power resources, causing device freezes or high temperatures. The schedutil CPU frequency manager in the CPUFreq subsystem schedules frequencies by estimating the load of each CPU set and cannot access other runtime features of the operating system, such as cache misses, IPC, etc. It lacks sufficient information to identify performance issues, resulting in mobile users often encountering application UI freezes (high frame loss rates) and high power consumption in many cases. Summary of the invention:
[0004] In order to solve the problems existing in the above-mentioned prior art, the present invention proposes a mobile device computing power scheduling method based on reinforcement learning, which is used to dynamically and actively adjust the CPU and GPU frequencies according to the system characteristics when the program is running, optimize the system performance of the mobile device, and reduce the frame loss rate and power consumption. The present invention consists of the following three parts. (1) A reinforcement learning model for executing mobile device computing power scheduling decisions. (2) A mobile device system status monitoring and computing power scheduling control framework. (3) A model training framework for collaboration between the server and the mobile device. The specific steps of each part are as follows:
[0005] 1. Reinforcement learning model for executing mobile device computing power scheduling decisions
[0006] Step 1. Select the deep deterministic policy gradient algorithm, namely DDPG (Deep Deterministic Policy Gradient), as the reinforcement learning model for executing computing power scheduling decisions. For the four neural networks of Actor, Actortarget, Critic, and Critic target in this model, set each neural network to contain three hidden layers, and the number of neurons in the three hidden layers of each neural network is 32, 64, and 32 respectively. For the Actor and Actor target networks, set the number of neurons in the input layer to the number of collected features; set the number of neurons in the output layer to the number of scheduled actions, and use the tanh function as the activation function of the output layer. The other neural network layers use ReLU as the activation function. The structure of the Actor target network is exactly the same as that of the Actor network. For the Critic and Critic target networks, set the number of neurons in the input layer to the sum of the number of collected features and the number of scheduled actions; set the number of neurons in the output layer to 1. The structure of the Critic target network is exactly the same as that of the Critic network.
[0007] Step 2. Select the key system features of the mobile device as the input features of the reinforcement learning model. The specific features include the following parts. (1) Frequency configuration parameters, including: the upper and lower frequency limits and current frequency of each CPU set, the upper and lower frequency limits and current frequency of the GPU, and the current DDR frequency. (2) Operating system load and performance indicators, including: the current load of each CPU set (i.e., CPU utilization), the number of instructions per clock cycle, the last level cache miss rate, and the system available memory. (3) Foreground application load, performance indicators, and runtime status, including: the load generated by the foreground application (i.e., the CPU utilization of the foreground application process), the last level cache miss rate of the foreground application, and the CPU set number where the foreground application main thread and rendering thread are running. (4) Device status, including: whether the screen is currently being touched and the device temperature.
[0008] Step 3: Set the model action space so that the number of neurons in the output layer of the Actor network is equal to the number of frequency configuration parameters that need to be scheduled. The output of each neuron corresponds to a computing power scheduling action, which is used to adjust the corresponding frequency configuration parameters. Since the output layer of the Actor network uses tanh as the activation function, each output a of the Actor network i All meet a i ∈(-1,1). Calculate the legal scheduling range of the frequency configuration parameters and obtain the actual computing power scheduling action.
[0009] The process of calculating the legal scheduling range of the frequency configuration parameters and obtaining the actual computing power scheduling action specifically includes the following steps.
[0010] Step 31, calculate the legal scheduling range of the frequency configuration parameters. Determine the legal scheduling range of each frequency configuration parameter before each scheduling to avoid predicting illegal scheduling schemes. Examples of illegal computing power scheduling schemes include: making the frequency lower limit of a certain CPU set higher than the upper limit, or making the frequency upper / lower limit values out of range. For each frequency configuration parameter, it is necessary to determine its maximum legal scheduling step inc_len for increasing resources and its maximum legal scheduling step dec_len for decreasing resources. The maximum single scheduling step is set to 3.
[0011] Step 32: Get the actual computing power scheduling action. Taking the i-th frequency configuration parameter as an example, the i-th output of the Actor network is ai, and the legal scheduling range of the frequency configuration parameter is [dec_leni, inc_leni]. If ai<0, the actual computing power scheduling action is dec_leni×ai; if ai>0, the actual computing power scheduling action is inc_leni×ai.
[0012] Step 4: Set the segmented reward function of the reinforcement learning model, prioritize reducing the application frame drop rate, and further reduce the device power consumption on the basis of no frame drops. The reward function is designed as follows:
[0013]
[0014]
[0015] rpower = (1 - power / TDP)
[0016] The reward function includes:
[0017] (1) Frame rate reward rfram e . Where t is the average value of the rendering time of all frames during two scheduling actions. tallowe d is the allowed rendering time per frame, that is, the frame rendering can be completed within this time without frame drops. Generally, it is set to 1s / FPS - 1ms, where FPS represents the frames per second. For example, when FPS is 60, tallowed = 15.7ms. is a preset fixed value according to tallowed, and the default value is 0.9 × tallowed. is slightly less than tallowed, which can prevent the rendering time of all frames less than from bringing a higher reward value, avoiding the situation that the reinforcement learning model increases energy consumption in order to pursue a shorter frame rendering time. represents the degree to which the frame rendering time meets the expectation. If its value is less than 1, it means that no frame drops will occur; otherwise, it means that the frame rendering time exceeds tallowe d , and the larger its value, the longer the frame rendering time and the more serious the frame drop situation.
[0018] (2) Low power consumption reward rpower. The ratio of the average energy consumption per unit time Power of all frames during two scheduling actions to the CPU thermal design power (TDP) is used as a metric for efficiency.
[0019] The reward function is set as a piecewise function. When t ≥ tallowe d , at this time, the average frame rendering time exceeds the allowed frame rendering duration, and frame drops occur. At this time, the reward function only includes the frame rate reward rframe. When t < tallowed, the reward function includes both rframe and rpower, taking into account the two optimization goals of frame rate and power consumption.
[0020] (II) Mobile device system status monitoring and computing power scheduling control framework
[0021] Step 5: The mobile device system status monitoring and computing power scheduling control framework detects the system status in each frame to determine whether to call reinforcement learning to schedule CPU and GPU frequencies, which specifically includes the following steps.
[0022] Step 51: determine whether three frame drops have occurred cumulatively, that is, the frame rendering time t of three consecutive frames exceeds the allowed frame rendering time tallowe d If there have been three cumulative frame drops, computing power scheduling will be enabled.
[0023] Step 52: Determine whether there is a CPU set whose current frequency is equal to the upper or lower frequency limit of the CPU set. If so, enable computing power scheduling.
[0024] 3. Model training framework for collaboration between server and mobile devices
[0025] Step 6. Collect model training data from multiple mobile devices on the server and train a reinforcement learning model that is common to different mobile devices. Deploy a frequency scheduling model for mobile devices on the server and use Android DebugBridge (ADB) to connect to the mobile device. The server reads the system feature data collected in real time by the mobile device, calls the DDPG reinforcement learning model to predict the computing power scheduling action, executes it on the mobile device through ADB, and calculates the reward value based on the frame loss and power consumption feedback from the mobile device. Perform data collection and model training for multiple different models of devices in turn. The training data of the reinforcement learning model comes from multiple different models of devices and is common to the computing power scheduling of each type of device.
[0026] Step 7: Fine-tune the reinforcement learning model for a specific model of mobile device. When deploying a reinforcement learning model to a new device for which no training data has been collected before, use a small amount of data from the new device to fine-tune the reinforcement learning model. That is, load the trained general reinforcement learning model parameters and retrain the model using the training data from the new device. The fine-tuned model has higher accuracy in scheduling new devices.
[0027] The advantages of the present invention are: for the performance tuning problem of Android mobile device system, an intelligent computing power scheduling method based on reinforcement learning is designed. The method makes computing power scheduling decisions based on a data-driven reinforcement learning model, eliminating the sampling and exploration overhead of traditional computing power scheduling strategies; it can make accurate computing power scheduling decisions based on system status and job resource requirements, avoiding the problems of low flexibility and low accuracy in previous work; it can use the exploration ability of the reinforcement learning model to automatically explore more efficient computing power scheduling solutions. Description of the drawings:
[0028] Figure 1The figure is a flow chart of the method for scheduling computing power of mobile devices based on reinforcement learning. Specific implementation method:
[0029] In order to make the above features and effects of the present invention more clearly understood, embodiments are given below for detailed description.
[0030] 1. Operating environment. The terminal of this embodiment uses the Rockchip RK3588 development board. The terminal uses the Android 12 operating system. The terminal processor has two CPU sets: (1) CPU set 0, ARM Cortex A78, i.e., four large cores; (2) CPU set 1, ARM Cortex A55, i.e., four small cores. All cores in each CPU set have the same frequency band. The CPU set 0 includes CPU cores 0 to 3, and the CPU set 1 includes CPU cores 4 to 7. The running memory is 8GB LPDDR4, and the fixed frame rate is 60 frames per second.
[0031] 2. Specific steps. Figure 1 A flow chart of an embodiment of the present invention is shown. The present invention mainly includes the following steps:
[0032] 1. Reinforcement learning model for executing mobile device computing power scheduling decisions
[0033] According to step 1 of the present invention, a deep deterministic policy gradient algorithm, namely DDPG (Deep Deterministic Policy Gradient), is selected as a reinforcement learning model for executing computing power scheduling decisions, and a neural network structure is configured.
[0034] According to step 2 of the present invention, key system features of mobile devices are selected as input features of the reinforcement learning model. There are 21 features in total: CPU set 0 load, CPU set 1 load, CPU set 0 current frequency, CPU set 1 current frequency, DDR frequency, GPU frequency, whether the screen is currently being touched, current device temperature, current CPU set 0 frequency upper limit, current CPU set 0 frequency lower limit, current CPU set 1 frequency upper limit, current CPU set 1 frequency lower limit, current GPU frequency upper limit, current GPU frequency lower limit, number of instructions per clock cycle, cache miss rate, system available memory, foreground application cache miss rate, CPU set ID where the foreground application main thread runs, CPU set ID where the foreground application rendering thread runs, and the load generated by the foreground application.
[0035] According to step 3 of the present invention, the model action space is set. There are 6 frequency configuration parameters that need to be controlled, including: CPU set 0 frequency upper limit, CPU set 0 frequency lower limit, CPU set 1 frequency upper limit, CPU set 1 frequency lower limit, GPU frequency upper limit, GPU frequency lower limit.
[0036] According to step 31 of the present invention, the legal scheduling range of the frequency configuration parameters is calculated. Take the upper frequency limit of CPU set 0 as an example. The upper and lower frequency limits of CPU set 0 have a total of 16 orders. If the current upper frequency limit is the 15th order, the lower frequency limit is the 12th order. For example, for the upper frequency limit, the maximum legal scheduling step inc_len for increasing resources is calculated as the minimum of the following two items: (1) the order of the highest value of the upper frequency limit - the order of the current value of the upper frequency limit, (2) the maximum scheduling step 3; the maximum legal scheduling step dec_len for decreasing resources is calculated as the minimum of the following three items: (1) the order of the current value of the upper frequency limit - the order of the lowest value of the upper frequency limit, (2) the order of the current value of the upper frequency limit - floor((the order of the current value of the lower frequency limit + the order of the current value of the upper frequency limit) / 2)(3) the maximum scheduling step 3. For the lower frequency limit, its legal scheduling range is also calculated.
[0037] According to step 32 of the present invention, the actual computing power scheduling action is obtained. The action of the reinforcement learning model includes 6 outputs: i |-1 i <1,0≤i≤5}. i >0 means to increase the frequency configuration parameter of item i, and the actual scheduling step number is round(a i *inc_len i );a i <0 means reducing the i-th frequency configuration parameter. The actual scheduling step is: round(a i *dec_len i );a i =0 means that the i-th frequency configuration parameter is not adjusted.
[0038] According to step 4 of the present invention, the piecewise reward function of the reinforcement learning model is set.
[0039] 2. Mobile device system status monitoring and computing power scheduling control framework
[0040] According to step 5 of the present invention, the system state is detected in each frame to determine whether to call reinforcement learning to schedule the CPU and GPU frequencies. If scheduling is required, the 21 reinforcement learning model input features are collected and normalized using the maximum and minimum value normalization method, that is, feature normalized =(max-feature) / (max-min). Feature refers to the feature value. normalized Refers to the normalized feature. max and min refer to the maximum and minimum values of the feature in the data set, respectively. The normalized feature is input into the reinforcement learning model to obtain the scheduling action, and the changed value of each frequency configuration parameter is written into the control file corresponding to the parameter to realize computing power scheduling.
[0041] 3. Computing power scheduling framework for coordinating operating systems and reinforcement learning models
[0042] According to step 6 of the present invention, the server collects model training data from development boards RK3588, RK3399, and smartphone Google pixel 8, and conducts reinforcement learning model training that is common to all types of devices.
[0043] According to step 7 of the present invention, based on the trained general reinforcement learning model, the model is retrained based on a small amount of new data of the RK3588 development board to improve the prediction accuracy of the model on the RK3588 platform.
[0044] The present invention also proposes a storage medium for storing a program for executing any one of the reinforcement learning-based mobile device computing power scheduling methods.
[0045] The present invention also proposes a mobile device equipped with an operating system that integrates any one of the above-mentioned reinforcement learning-based mobile device computing power scheduling methods.
[0046] The present invention provides a method for scheduling computing power of mobile devices based on reinforcement learning, including: a reinforcement learning model for executing mobile device computing power scheduling decisions, a mobile device system status monitoring and computing power scheduling control framework, and a model training framework for collaboration between a server and a mobile device. The computing power scheduling method makes computing power scheduling decisions based on a data-driven reinforcement learning model, eliminating the sampling and exploration overhead of traditional computing power scheduling strategies; it can make accurate computing power scheduling decisions based on system status and job resource requirements, avoiding the problems of low flexibility and low accuracy in previous work.
Claims
1. A method for scheduling computing power of mobile devices based on reinforcement learning, characterized in that: include: (1) A reinforcement learning model for executing mobile device computing power scheduling decisions; (2) A mobile device system status monitoring and computing power scheduling control framework, which is used to detect the system status and call the reinforcement learning model to schedule computing power under specific conditions; (3) A model training framework that collaborates between the server and mobile devices to train general reinforcement learning models and fine-tune the models for specific devices to improve accuracy.
2. A method for scheduling computing power of mobile devices based on reinforcement learning as claimed in claim 1, characterized in that: The reinforcement learning model used to perform mobile device computing scheduling decisions includes: Step 1: Select the deep deterministic policy gradient algorithm, namely DDPG, as the reinforcement learning model for executing computing power scheduling decisions; for the four neural networks of Actor, Actor target, Critic, and Critic target in the model, set each neural network to contain three hidden layers, and the number of neurons in the three hidden layers of each neural network is 32, 64, and 32 respectively; for the Actor and Actor target neural networks, set the number of neurons in the input layer to the number of collected features, set the number of neurons in the output layer to the number of scheduled actions, and use the tanh function as the activation function of the output layer. The other neural network layers use ReLU as the activation function. The structure of the Actor target network is exactly the same as that of the Actor network; for the Critic and Critic target networks, set the number of neurons in the input layer to the sum of the number of collected features and the number of scheduled actions, and set the number of neurons in the output layer to 1; the structure of the Critic target network is exactly the same as that of the Critic network; Step 2: Select the key system features of the mobile device as the input features of the reinforcement learning model. The specific features include the following: (1) Frequency configuration parameters, including: the upper frequency limit, lower frequency limit, and current frequency of each CPU set, the upper frequency limit, lower frequency limit, and current frequency of the GPU, and the current DDR frequency; (2) Operating system load and performance indicators, including: the current load of each CPU set, i.e., CPU utilization, number of instructions per clock cycle, last-level cache miss rate, and system available memory; (3) Foreground application load, performance indicators, and runtime status, including: the load generated by the foreground application, i.e., the CPU utilization rate of the foreground application process, the last-level cache miss rate of the foreground application, and the CPU set number where the foreground application main thread and rendering thread are running; (4) Device status, including: whether the screen is currently being touched and device temperature. Step 3: Set the model action space so that the number of neurons in the output layer of the Actor network is equal to the number of frequency configuration parameters that need to be scheduled. The output of each neuron corresponds to a computing power scheduling action, which is used to adjust the corresponding frequency configuration parameters. Since the output layer of the Actor network uses tanh as the activation function, each output a of the Actor network i All meet a i ∈(-1,1); calculate the legal scheduling range of the frequency configuration parameters and obtain the actual computing power scheduling action; Step 4: Set the segmented reward function of the reinforcement learning model, give priority to reducing the application frame loss rate, and further reduce the device power consumption on the basis of no frame loss. The reward function is designed as follows: rpower=(1-power / TDP) 3. A method for scheduling computing power of mobile devices based on reinforcement learning as claimed in claim 2, characterized in that: This step 3 includes: Step 31, calculate the legal scheduling range of frequency configuration parameters: determine the legal scheduling range of each frequency configuration parameter before each scheduling to avoid predicting illegal scheduling schemes; illegal computing power scheduling schemes include: computing power scheduling schemes that make the frequency lower limit of a certain CPU set higher than the upper limit, or make the frequency upper / lower limit values exceed the range; for each frequency configuration parameter, it is necessary to determine its maximum legal scheduling step length inc_len for increasing resources and its maximum legal scheduling step length dec_len for decreasing resources; the maximum single scheduling step length is set to 3; Step 32, get the actual computing power scheduling action: For the i-th frequency configuration parameter, the i-th output of the Actor network is a i , the legal scheduling range of the frequency configuration parameter is [dec_len i , inc_len i ], if a i <0, the actual computing power scheduling action is dec_len i ×a i , if a i >0, the actual computing power scheduling action is inc_len i ×a i .
4. The method for scheduling computing power of mobile devices based on reinforcement learning as claimed in claim 1, characterized in that: The mobile device system status monitoring and computing power scheduling control framework includes: Step 5: The mobile device system status monitoring and computing power scheduling control framework detects the system status in each frame to determine whether to call reinforcement learning to schedule CPU and GPU frequencies.
5. A method for scheduling computing power of mobile devices based on reinforcement learning as claimed in claim 4, characterized in that: This step 5 includes: Step 51: determine whether three frame drops have occurred cumulatively, that is, the frame rendering time t of three consecutive frames exceeds the allowed frame rendering time tallowe d , if there have been three cumulative frame drops, computing power scheduling is enabled; Step 52: Determine whether there is a CPU set whose current frequency is equal to the upper or lower frequency limit of the CPU set. If so, enable computing power scheduling.
6. A method for scheduling computing power of mobile devices based on reinforcement learning as claimed in claim 1, characterized in that: The model training framework that the server and mobile devices collaborate on includes: Step 6: Collect model training data from multiple mobile devices on the server, train a reinforcement learning model that is common to different mobile devices, deploy a frequency scheduling model for mobile devices on the server, and use Android DebugBridge (ADB) to connect to the mobile device; the server reads the system feature data collected in real time by the mobile device, calls the DDPG reinforcement learning model to predict the computing power scheduling action, executes it on the mobile device through ADB, and calculates the reward value based on the frame loss and power consumption feedback from the mobile device; perform data collection and model training for multiple different models of devices in turn; the training data of the reinforcement learning model comes from multiple different models of devices and is common to the computing power scheduling of each type of device; Step 7, fine-tuning the reinforcement learning model for a specific mobile device model: When deploying a reinforcement learning model to a new device for which no training data has been collected before, use a small amount of data from the new device to fine-tune the reinforcement learning model, that is, load the trained general reinforcement learning model parameters and retrain the model using the training data from the new device. The fine-tuned model has higher accuracy in scheduling new devices.
7. A storage medium for storing a program for executing any one of the reinforcement learning-based mobile device computing power scheduling methods as described in claims 1 to 6.
8. An operating system equipped with an operating system integrating any one of the reinforcement learning-based mobile device computing power scheduling methods as described in claims 1 to 6.