Processor Low-Power State Switching Method Based on RISC-V Server Platform

By configuring multi-stage low-power state information in the ACPI firmware of the RISC-V server platform and using reinforcement learning SARSA algorithm to select the best state, the problem of difficulty in reducing power consumption during low load or idle by RISC-V server platform is solved, and high-efficiency energy consumption management and system stability are improved.

CN119576108BActive Publication Date: 2025-06-10SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510134195.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-10
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

The existing RISC-V server platform is difficult to effectively reduce power consumption when low load or idle, affecting energy efficiency and performance management.

Method used

By configuring the multi-stage low-power state information of the RISC-V processor in the DSDT table of the ACPI firmware, and using the control strategy of the reinforcement learning SARSA algorithm to select the best low-power state for the processor, the operating system calls the SBI firmware service to allow the processor to enter a specified low-power state.

Benefits of technology

It significantly reduces the power consumption of the server platform at low load or idle time, improves system stability and energy efficiency management, and improves adaptability and portability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576108B_ABST
    Figure CN119576108B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for switching the low-power state of a processor based on a RISC-V server platform, belonging to the field of computer technology. The steps include: (1) configuring multi-level low-power state information of the RISC-V processor; (2) initializing the CPU idle state management module using the multi-level low-power state information; (3) the CPU idle state management module selects the optimal low-power state for the RISC-V processor based on the control strategy of the reinforcement learning SARSA algorithm; (4) the RISC-V processor enters the specified low-power state. The present invention realizes efficient power management of the RISC-V processor on the server platform, can significantly reduce energy consumption while ensuring system performance, effectively improves the energy efficiency ratio of the server, and is particularly suitable for the power management of large-scale server clusters and the construction of green data centers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for switching the low-power state of a processor based on a RISC-V server platform, belonging to the field of computer technology. Background Art

[0002] In the RISC-V architecture, ACPI (Advanced Configuration and Power Interface) is an important standard interface for implementing power management. In traditional x86 and ARM processors, ACPI is widely used to define and manage the power states of the system. As the application of RISC-V in servers and embedded systems has been gradually promoted, the RISC-V ecosystem has also started to support ACPI to better be compatible with existing operating systems and hardware management software.

[0003] Processor states (C-States) are part of ACPI power management and are used to represent different energy-saving states of the processor when it is idle. The levels of C-States range from C0 to Cn, where C0 represents that the processor is working, and C1 to Cn represent different depths of sleep states. The deeper the C-state, the lower the power consumption, but the higher the latency to resume to the working state may be. Through ACPI, the operating system can dynamically monitor and manage the C-States of the processor, and then adjust the power consumption and performance of the processor according to the system load to achieve power consumption optimization.

[0004] Implementing the conversion of processor states (C-States) in the RISC-V architecture is crucial because it directly affects the energy efficiency and performance management of the system. Processor state conversion allows the RISC-V processor to enter a deeper energy-saving state when it is idle or under low load, thus significantly reducing power consumption, which is particularly critical in servers, mobile devices, and embedded systems. By efficiently managing C-States, the RISC-V architecture can dynamically adjust the working state of the processor to balance performance requirements and power consumption limits, improve power utilization efficiency, extend battery life, and reduce heat dissipation and operating costs in scenarios such as data centers. Therefore, implementing processor state conversion is a key factor for the expansion of RISC-V applications, especially for obtaining wide applications in the fields of low power consumption and high-performance computing. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a method for switching the low-power state of a processor on a RISC-V server platform, aiming to reduce the power consumption of the RISC-V server platform under low load or idle conditions and achieve efficient energy consumption management.

[0006] Term Explanation:

[0007] 1. RISC-V: RISC-V is an open-source instruction set architecture that adopts the concepts of reduced instruction set, modular design, and scalability. This makes RISC-V processor designs simple, low-cost, and capable of being flexibly extended according to requirements. This design flexibility enables RISC-V to be applicable to various different application scenarios, including embedded systems, servers, high-performance computing, and Internet of Things devices, etc.

[0008] 2. UEFI (Unified Extensible Firmware Interface): UEFI is a firmware interface specification designed to replace the traditional BIOS (Basic Input / Output System). It provides a more flexible and standardized boot solution, can load operating systems and applications, and offers more powerful hardware initialization and configuration functions.

[0009] 3. ACPI (Advanced Configuration and Power Interface): ACPI is the Advanced Configuration and Power Interface, a standard interface for implementing advanced configuration and power management, aiming to provide a unified interface between the operating system and computer hardware for energy management and hardware configuration.

[0010] 4. OSPM (Open Source Power Management): OSPM is an open-source energy management framework aiming to provide effective energy management and optimization solutions for computer systems. OSPM is committed to developing portable, flexible, and extensible software tools to support various operating systems and hardware platforms.

[0011] 5. Differentiated System Description Table (DSDT): DSDT contains information related to system-specific hardware. By defining devices, resources, methods, and events in the ACPI namespace, etc., it provides the key information for the operating system and firmware to dynamically configure hardware and manage power at runtime.

[0012] 6. CPU Idle State Management (CPUIdle): A subsystem of the Linux kernel responsible for managing the transition of the CPU idle state. It monitors the CPU load situation and automatically selects different idle states according to the length of idle time, thus saving power when the processor is not working. The CPUIdle framework determines which low-power state to enter when the CPU enters the idle state based on factors such as the current system load and task scheduling, and quickly resumes to the working state when a task arrives. This mechanism is crucial for saving electrical energy, improving battery life, and optimizing system performance.

[0013] 7. Governor (Control Strategy): In the Linux kernel, a Governor is a scheduling strategy module responsible for dynamically managing the CPU's frequency and power state to achieve a balance between performance and power consumption. Different Governors adjust the CPU's operating frequency and idle state in real time based on factors such as system load, temperature, and user requirements.

[0014] 8. _LPI (Low Power Idle) - Low Power State Object: _LPI is an object defined in the ACPI specification used to describe the low-power states that system components can enter when idle. Through the _LPI object, the system can identify and manage the low-power modes of different hardware components to reduce energy consumption when idle.

[0015] 9. SBI (Supervisor Binary Interface) Firmware Service: The SBI firmware service is a standard interface provided for the operating system kernel in the RISC-V architecture to enable access to low-level system functions between the operating system and the hardware. The SBI service is responsible for handling some hardware-related operations such as power management, timer configuration, and interrupt control, and acts as a bridge especially between different privilege levels.

[0016] 10. Reinforcement Learning: Reinforcement learning is a machine learning method where an agent learns decision-making strategies by interacting with the environment to achieve goals in specific tasks. The agent tries different actions to obtain feedback (rewards or punishments) from the environment and uses this feedback to adjust its strategy to maximize long-term rewards. It is particularly suitable for solving problems in dynamic, complex, and partially uncertain environments, such as robot control, game strategies, and resource allocation scenarios.

[0017] 11. SARSA (State-Action-Reward-State-Action): SARSA is a reinforcement learning algorithm for an agent to learn an optimal strategy by interacting with the environment in a dynamic environment. It updates the strategy based on the current state and action, as well as the subsequent reward and new state, making the agent more efficient in future decision-making. By associating the current state-action pair with the next state-action pair, SARSA can gradually improve decision-making in a dynamic environment and is suitable for scenarios that require online learning.

[0018] The technical solution of the present invention is as follows:

[0019] A method for switching the low-power state of a processor in a RISC-V server platform, the steps are as follows:

[0020] (1)Configure the multi-level low-power state information of the RISC-V processor in the DSDT table of the ACPI firmware;

[0021] (2)The ACPI driver module in the operating system initializes the CPU idle state management module using the multi-level low-power state information of the RISC-V processor in the DSDT table;

[0022] (3)The CPU idle state management module selects the best low-power state for the RISC-V processor based on the control strategy of the reinforcement learning SARSA algorithm;

[0023] (4)The operating system calls the SBI firmware service to let the RISC-V processor enter the specified low-power state.

[0024] Preferably according to the present invention, in step (1), the multi-level low-power state information is:

[0025] Define a low-power state object in the DSDT table of the ACPI firmware. This object provides complete information about the low-power state, including the minimum residency time, worst-case wake-up latency, residency counter frequency, residency counter register, entry method, and state name. The _LPI object uses a Resource Descriptor to specify the entry method for the processor's low-power state. The processor state transition in the RISC-V system is implemented using the SBI HSM extension.

[0026] Preferably according to the present invention, in step (2), specifically: The CPU idle state management module includes a core component, a control strategy component, and a driver component. The core component is the center of the entire idle management system, responsible for coordinating the workflow of CPU idle state management. Through the interface provided by the core component, other modules of the operating system can request the CPU to enter the idle state. The core component will select a suitable idle state according to the instructions of the control strategy, thereby achieving power consumption optimization. The core component is also responsible for managing the state transitions of different processors to ensure reasonable allocation of idle states in a multi-core system;

[0027] The control strategy component determines which idle state the system should select under different loads. By introducing the idea of reinforcement learning, the control strategy component dynamically optimizes the idle state selection process using the SARSA algorithm. While considering the balance between power consumption and performance, it continuously adjusts and improves the correction factor for idle state selection, enabling the system to adapt to complex load changes and achieve more accurate and efficient decision-making;

[0028] The driving component is the part that interacts with specific hardware and is responsible for implementing the actual state transitions. It places the CPU in the corresponding C-State according to the selection of the control strategy and manages the hardware details during the state transition process. The driving component also provides the core component with a list of available C-States and feedback on the execution of the state transition.

[0029] The initialization process includes the initialization process of the core component, the initialization process of the control strategy component, and the initialization process of the driving component.

[0030] According to a further preference of the present invention, the initialization process of the core component is to initialize the key variables and functions required for idle state management. First, the key variables are initialized, that is, the variables storing the registered device nodes. Then, the functions initialize the entry into the idle state, the selection of the idle state, and the recording of the idle state one by one, providing the core functions for idle state management.

[0031] According to a preference of the present invention, the initialization process of the control strategy component is as follows: First, the scheduler structure is defined and initialized, including the name, score, and two key function pointers of the scheduler, which are used to select the idle state and reflect the state change respectively. Then, the registration function is called to register the scheduler structure. This process ensures that the scheduler is correctly loaded at system startup and can perform state management according to system requirements during runtime. The registration operation of this scheduler is usually completed in the post-initialization phase of the core of the kernel, ensuring that it is initialized and ready to provide services in the early stage of the system.

[0032] According to a preference of the present invention, the initialization process of the driving component is to register the device driver and perform the initialization operation when the device starts.

[0033] The drive component extracts the low-power state description information related to the RISC-V processor by parsing the DSDT table. The specific low-power state objects are determined by the architecture and hardware resources of the processor and usually include multiple low-power states, such as "RISC-V WFI", "RISC-V RET_DEFAULT", and "RISC-V NONRET_DEFAULT". Each state object has different parameters and characteristics, such as the minimum residence time, the worst wake-up latency, the counting frequency, and the dependency relationship with the upper-level state. Specifically, the "RISC-V WFI" state has a short minimum residence time and wake-up latency and is suitable for quickly entering the low-power mode, while the "RISC-V NONRET_DEFAULT" state has a longer residence time and wake-up latency and is suitable for a deeper low-power mode. The drive component classifies and manages according to the characteristics of different low-power states, generates corresponding management entries, and maintains a low-power state table during system operation. These table entries include key parameters such as the exit latency, the target residence time, and the power consumption, and are uniformly described through the "CPU idle state" structure to ensure that each CPU can enter different low-power modes according to specific hardware requirements. Then, the drive component sets corresponding callback functions to dynamically adjust the power consumption state of the CPU when needed.

[0034] Preferably according to the present invention, in step (3), specifically:

[0035] The control strategies include the Ladder strategy and the Menu strategy. The Ladder strategy is a simple fixed-order strategy, that is, to reach a higher level, it must go up step by step from a lower level. In the Ladder strategy, the ladder governor will first enter the shallowest Idle state (idle state), and then if it stays long enough, it will enter a deeper Idle state, and so on until it reaches the deepest Idle state. When awakened, it will restart the CPU as quickly as possible; when idle next time, it will start from the shallowest Idle state again. In this strategy, the system may not enter the deepest Idle state for a long time, resulting in some loss of low power consumption. The Ladder strategy is not suitable for server platforms with large load changes.

[0036] The Menu strategy is an intelligent dynamic strategy, which mainly selects the best C-State idle state based on the predicted time (desired residence time) and the system's latency tolerance to achieve a more balanced optimization between performance and power consumption. The Idle Menu governor decides which idle state to enter according to the CPU load, the time of the next timer event, the number of active I / O threads, etc.

[0037] During the process of calculating the predicted idle time, the Menu controller gives a prediction value based on the time of the next timer event. However, since the time point of the next timer event may be interfered by other events, resulting in a deviation in the actual idle time, a set of correction factors needs to be introduced to further adjust this prediction value. The correction factors are calculated based on past empirical data, especially the ratio of past predicted time to actual event time, through a dynamic averaging algorithm. To improve the accuracy of the prediction, the Menu controller uses 12 different correction factors, corresponding to different time periods respectively. The value of each correction factor is adjusted according to the fluctuations of historical data to adapt to different system behaviors. Especially in different io waiting scenarios, the Menu controller checks the consistency of the time intervals through the last 8 residence times. If the standard deviation of these time intervals is less than the threshold, the average value of the 8 residence times is selected as a supplementary prediction value, and the minimum value of the above two prediction values is taken as the final predicted time, so as to more accurately adjust the idle state of the CPU to optimize energy efficiency;

[0038] The Menu controller uses the performance multiplier, the expected residence time, and the system latency requirement to calculate the system latency tolerance. The system latency requirement is used as the first system tolerance, and another system tolerance is calculated through the formula: expected residence time / (1 + 10 * the number of IO waiting tasks on the current CPU). The minimum value of the two system tolerances is taken as the system latency tolerance;

[0039] Select the idle state according to the predicted time (desired residence time) and the system latency tolerance. Compare the predicted time with the residence times of all Idle states. The condition for selecting a specific Idle state is that the corresponding residence time is less than the predicted time. In addition, compare the exit latency of the state with the system latency tolerance, and the exit latency of the state needs to be less than the system latency tolerance. Select the idle state with the minimum power consumption that meets the two conditional factors as the best low-power state.

[0040] The Menu controller is responsible for selecting the appropriate idle state. Although the deep idle state can significantly save power consumption, it requires additional entry and exit overheads and will increase latency. Therefore, in order to achieve a better balance between power consumption and performance, the present invention uses the idea of reinforcement learning to dynamically adjust and improve the idle state selection process. Reinforcement learning is a method in which an agent obtains rewards through interacting with the environment and updates its decision-making strategy. Initially, the agent understands the environment and takes actions according to the initial strategy. Subsequently, the environment feedbacks new state information and reward values, and the agent dynamically adjusts its strategy according to these feedbacks. In this process, the agent transfers from the current state to the next state and improves its decision-making through reward and punishment feedback.

[0041] Preferably according to the present invention, the SARSA algorithm is used to optimize the idle state selection strategy. The SARSA algorithm is a policy-based reinforcement learning algorithm suitable for online decision-making problems. Based on state-action pairs, the Q value is updated through the following formula: Q(S t ,a t )←Q(S t ,a t )+α[r t+1 +γQ(S t+1 ,a t+1 )−Q(S t+1 ,a t+1 )], where S t represents the state at the current moment, a t represents the action taken in the current state, r t+1 is the immediate reward obtained when entering the next state after executing the action, S t+1 represents the next state transferred to after executing the action, a t+1 represents the action selected in the next state, Q(S t ,a t ) is the Q value of the current state-action pair, used to represent the expected cumulative reward that can be obtained by taking action a t in state S t . α is the learning rate, used to balance the weights of new information and historical information, and γ is the discount factor, used to measure the influence degree of future rewards. By adjusting these parameters, the SARSA algorithm can effectively balance between optimizing short-term energy consumption and improving long-term performance.

[0042] In the present invention, the SARSA model learns the value of state-action pairs by considering the transition from one state-action pair to another, and adjusts the policy according to the current state "S", the selected action "A", and its reward "R", until entering the new state "S1" and selecting the next action "A1". This mechanism dynamically evaluates and adjusts the correction factor for idle state selection, enabling the Menu controller to adapt to the changing environment, thereby achieving more accurate idle state selection.

[0043] Preferably according to the present invention, specifically in step (4), it involves the cooperation between the operating system and the SBI firmware. The operating system issues a request to specify the target low-power state of the RISC-V processor. After receiving the ecall call request, the SBI firmware parses the call parameters to determine the target low-power state and related configuration parameters, and controls the RISC-V processor to quickly transition from the running state to the specified low-power state.

[0044] The beneficial effects of the present invention are as follows:

[0045] 1. The present invention reduces the power consumption of the server platform system. By intelligently switching the processor state, it significantly reduces the power consumption of the server platform under low load or idle conditions, thereby effectively reducing the overall energy consumption.

[0046] 2. The present invention improves the stability of the server system. By dynamically adjusting the processor state, it avoids system overheating caused by long-term high-power operation, enhancing the reliability and stability of the system.

[0047] 3. The present invention has strong adaptability and portability. Through the SBI interface, the operating system can request and manage the low-power mode without directly accessing the hardware, ensuring good compatibility and consistency of the RISC-V platform among different hardware implementations. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a schematic flowchart of the method of the present invention;

[0049] Figure 2 is a schematic diagram of the energy consumption management architecture of the operating system and UEFI involved in the embodiment of the present invention;

[0050] Figure 3 is a schematic flowchart of the initialization process of the CPU idle state management module involved in the embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The present invention will be further described below through embodiments in conjunction with the drawings, but is not limited thereto.

[0052] Embodiment 1:

[0053] As Figure 1 shown, this embodiment provides a method for switching the low-power state of a processor in a RISC-V server platform, and the steps are as follows:

[0054] Step 1: Based on the startup method of the UEFI firmware, configure the multi-level low-power state information of the RISC-V processor in the ACPI DSDT table to achieve more refined power management.

[0055] The DSDT table configuration will contain information about different low-power states, including the minimum residency time, worst-case wake-up latency, residency counter frequency, residency counter register, entry method, and state name. The _LPI object uses a resource descriptor to specify the entry method for the processor's low-power state. When the RISC-V processor enters the core WFI state, it executes the wfi instruction. The entry methods for other low-power states are implemented using the SBI HSM extension. The HSM extension defines the SUSPENDED, SUSPEND_PENDING, RESUMING_PENDING states and the sbi_hart_suspend() interface to allow the processor to enter the platform-level low-power state.

[0056] Step 2: By obtaining the low-power state parameters defined in the DSDT table, the Linux operating system can initialize the CPU idle state management module, thereby accurately identifying and controlling different idle states of the processor and dynamically managing the power consumption level of the system.

[0057] The CPU idle state management module of the Linux operating system includes a core component, a control policy component, and a driver component. See Figure 2 As shown, the initialization process of the CPU idle state management module in this embodiment includes the following steps:

[0058] Step 201, Initialize the CPUIdle core component;

[0059] The core component is responsible for implementing the management framework for the idle state and providing support for the selection and management of the processor's idle state to the system. The initialization process of the core component includes the initialization of key variables and functions required for idle state management. First, the system initializes the variable used to store the registered device nodes. Then, the functions initialize entering the idle state, selecting the idle state, and recording the idle state one by one to provide the core functions of idle state management. Specifically, it provides a unified Idle interface for other modules related to Idle, including the cpuidle_select, cpuidle_reflect, cpuIdle_enter, and cpuIdle_enter_s2idle functions.

[0060] Step 202, Initialize the CPUIdle control policy component;

[0061] The control strategy component first defines and initializes a scheduler structure, which includes the name, score of the scheduler, and two key function pointers for selecting the idle state and reflecting state changes respectively. Then, it calls the registration function to register the scheduler structure. Specifically, it calls the cpuidle_register_governor() interface to register the governor. This process ensures that the scheduler is correctly loaded at system startup and can manage the state according to system requirements during runtime.

[0062] Step 203, initialize the CPUIdle driver of the specific CPU;

[0063] During the initialization process, the driver component registers a device driver and performs relevant initialization operations when the device starts. The driver component extracts the low-power state description information related to the RISC-V processor by parsing the DSDT table. The driver component classifies and manages according to the characteristics of different low-power states, generates corresponding management entries, and maintains a low-power state table during system runtime. These table entries include key parameters such as exit latency, target stay time, and power consumption, and are uniformly described through the "CPU idle state" structure to ensure that each CPU can enter different low-power modes according to specific hardware requirements. Specifically, in cpuidle_driver, it calls sbi_cpuidle_init_cpu() to initialize and register the cpuidle driver for all CPUs, mainly initializing the attributes of the Idle States and the Idle functions. Then, the driver component sets the corresponding callback functions so that the power consumption state of the CPU can be dynamically adjusted when needed. Specifically, it calls the sbi_idle_init_cpuhp() function to set the callback function for the hot-pluggable CPUHP_AP_CPU_PM_STARTING state. The sbi_cpuidle_init_cpu() function finally calls the cpuidle_register() function to register the above driver.

[0064] Step 204, schedule to enter the Idle loop and select to enter the Idle State according to the Governor;

[0065] After the Idle thread is scheduled, it will repeatedly execute the do_idle() function, execute the interface provided by the CPUIdle framework for scheduling, call the cpuidle_select() function for the Governor to select the corresponding Idle State, and then call the call_cpuidle() function to execute the entry function of this Idle State. The sbi_cpuidle_enter_state() function in the sbi_cpuidle driver calls the sbi_hart_suspend() interface extended by the HSM to put the processor into the low-power state. Execute the cpuidle_reflect() function to notify the governor to exit the current Idle State.

[0066] Step 3: The CPU idle state management module selects a suitable low-power state for the RISC-V processor based on the control strategy of the reinforcement learning SARSA algorithm.

[0067] The control strategy adopts the Ladder and Menu strategies to adapt to different system loads and power consumption requirements. The Ladder strategy is a simple fixed-order strategy that selects idle states in a gradually increasing order and is suitable for environments with little load change. In contrast, the Menu strategy is a more intelligent dynamic strategy that mainly selects the best C-State idle state based on the predicted time (desired residence time) and the system's latency tolerance, thus achieving a more balanced optimization between performance and power consumption. Specifically, the Menu governor decides which idle state to enter based on the CPU load, the time of the next timer event, the number of active I / O threads, etc.

[0068] During the process of calculating the predicted idle time, the Menu controller initially calculates a predicted time based on the time of the next timer event. However, since the time point of the next timer event may be disturbed by other events, resulting in a deviation in the actual idle time, a set of correction factors needs to be introduced to further adjust this predicted value. The correction factors are calculated based on past empirical data, especially the ratio of past predicted time to actual event time, through a dynamic averaging algorithm. To improve the accuracy of the prediction, the Menu controller uses 12 different correction factors corresponding to different time periods. The value of each correction factor is adjusted according to the fluctuations of historical data to adapt to different system behaviors, especially in different IO waiting scenarios. The Menu controller checks the consistency of the time intervals through the most recent 8 residence times. If the standard deviation of these intervals is less than the threshold, the system will select the average value of these times as a supplementary predicted value. Finally, the system takes the minimum value of the above two predicted values as the final prediction of the idle time, so as to more accurately adjust the idle state of the CPU to optimize energy efficiency.

[0069] The Menu controller uses the performance multiplier, the expected residence time, and the system latency requirement to calculate the system latency tolerance. The system latency requirement is used as the first system tolerance, and another system tolerance is calculated through the formula: expected residence time / (1 + 10 * the number of IO waiting tasks on the current CPU). The minimum value of the previous two system tolerances is taken as the system latency tolerance.

[0070] Finally, the appropriate idle state is selected based on two factors: the calculated predicted time (expected residence time) and the system latency tolerance. The calculated expected residence time is compared with the residence times of all Idle states. The condition for selecting a specific Idle state is that the corresponding residence time should be less than the expected predicted time. In addition, the exit latency of the state is compared with the system latency tolerance, and the exit latency of the state needs to be less than the system latency tolerance. The best low-power state is the idle state that meets the two condition factors and has the minimum power consumption.

[0071] The Menu controller is responsible for selecting the appropriate idle state. Although the deep idle state can significantly save power consumption, it requires additional entry and exit overhead and will increase latency. Therefore, in order to achieve a better balance between power consumption and performance, the present invention uses the idea of reinforcement learning to dynamically adjust and improve the idle state selection process. Reinforcement learning is a method in which an agent interacts with the environment to obtain rewards and updates its decision-making strategy. Initially, the agent understands the environment and takes actions according to the initial strategy. Subsequently, the environment feedbacks new state information and reward values, and the agent dynamically adjusts its strategy based on these feedbacks. In this process, the agent transfers from the current state to the next state and improves its decision-making through reward and punishment feedback.

[0072] The SARSA algorithm is used to optimize the idle state selection strategy. SARSA is a policy-based reinforcement learning algorithm suitable for online decision-making problems. Based on state-action pairs, the algorithm updates the Q value through the following formula: Q(S t ,a t ) ← Q(S t ,a t ) + α[r t+1 + γQ(S t+1 ,a t+1 ) − Q(S t+1 ,a t+1 )], where S t represents the state at the current moment, a t represents the action taken in the current state, r t+1 is the immediate reward obtained when entering the next state after executing the action, S t+1 represents the next state transferred to after executing the action, a t+1 represents the action selected in the next state, Q(S t ,a t ) is the Q value of the current state-action pair, used to represent the expected cumulative reward that can be obtained by taking action a t in state S t . α is the learning rate, used to balance the weights of new information and historical information, and γ is the discount factor, used to measure the influence degree of future rewards. By adjusting these parameters, the SARSA algorithm can effectively balance between optimizing short-term energy consumption and improving long-term performance.

[0073] In the present invention, the SARSA model learns the values of state-action pairs by considering the transition from one state-action pair to another, and adjusts the policy according to the current state "S", the selected action "A", and its reward "R" until entering a new state "S1" and selecting the next action "A1". This mechanism dynamically evaluates and adjusts the correction factor for idle state selection, enabling the Menu controller to adapt to the changing environment, thereby achieving more accurate idle state selection.

[0074] Step 4: The operating system calls the SBI firmware service to let the RISC-V processor enter the specified low-power state.

[0075] This process involves the cooperation between the operating system and the SBI firmware. The operating system sends a request to specify the target low-power state of the RISC-V processor. After receiving the request, the SBI firmware parses the call parameters to determine the target low-power state and related configuration parameters, and controls the RISC-V processor to quickly transfer from the running state to the specified low-power state.

[0076] Figure 3 Schematic diagram of the power consumption management architecture of the operating system and UEFI used in this embodiment. Refer to Figure 3 , the modules in the operating system related to the CPU idle state component include the ACPI driver module and the CPU idle state management module. Among them, the ACPI driver module can perform operations based on the ACPI system description table provided by UEFI, and the CPU idle state management module can be used to monitor the CPU load situation and automatically select different C-States according to the length of the idle time.

[0077] In a specific application, after the system is powered on and enters the operating system, the operating system initialization can be performed. If the firmware running on the system is UEFI, it will jump to the ACPI driver module under the Linux kernel driver. The ACPI driver module parses the ACPI system description table of UEFI, parses and obtains the resource description related to the low-power state of the processor from the ACPI system description table, enumerates the Idle states in cpuidle_driver and initializes its entry function, and finally registers through cpuidle_register().

[0078] In the actual CPU operating environment, different CPUs have differences in the requirements for the Idle state and the entry / exit methods. Among them, power consumption and exit latency have become an irreconcilable contradiction in the Idle scheduling process. How to save power as much as possible while meeting the performance requirements has become an important part of the CPUIdle subsystem. The control strategy component is responsible for providing strategies such as selecting the idle state. There are two strategies provided in the Linux kernel: Menu and Ladder. The Menu governor has an array of correction factors consisting of 12 indexes for the selection of the idle state. After the idle state exits, the ratio of the actual idle time to the predicted idle time is taken, and the correction coefficient of the last transition is calculated. The calculated correction factor will be updated in the corresponding index of the one-dimensional correction factor array. The existing logic does not consider the fact that the current transition occurs from which index to which index or from the idle state to the idle state. It only considers the final result.

[0079] In this embodiment, a circular linked list containing 3 members is used to track the "state - action - reward - state - action" situation. Each structure contains a state index, an action index, and the reward information obtained. A reward_matrix[number_of_idle_states x 12] is designed to index state - action pairs and store the reward factors according to the actions taken during the state transition. After the idle state exits, the index of reward_matrix [state][action] will be updated to the new reward factor calculated based on SARSA. For example, when the CPU is in state 0 and executes action 1 on instance T, it enters state 1 on instance T + 1 and obtains some rewards. In this case, the index of reward_matrix[state - 0][action - 1] will update the reward by referring to the structure pointer.

[0080] The specific steps of SARSA modeling under the CPUIdle subsystem include:

[0081] Step 301, enable the Menu controller;

[0082] Allocate a circular linked list consisting of 3 structure pointers, which contains information such as idle state index, action index, reward value, and the next pointer, and assign the idle_sarsa_ptr of the idle data structure to the head of the linked list.

[0083] Idle_sarsa_ptr = ptr0;

[0084] Allocate a reward matrix [number_of_idle_states x 12] to index state - action pairs and maintain the respective state transition reward values, and initialize the matrix with the default unit reward value.

[0085] Step 302, select the idle state;

[0086] When predicting the next timer event, consider the previous reward factors obtained during the respective state transitions - based on the environment, consider the next action metric;

[0087] Action_index = get_bucket_index(next_timer_event, number_of_iotask);

[0088] Predicted_event_time = scheduler_next_timer_event x reward_matrix[state_index][action_index];

[0089] Idle_sarsa_ptr ->action_index = action_index;

[0090] Idle_sarsa_ptr = Idle_sarsa_ptr ->next;

[0091] Step 303, exit the idle state and update the reward;

[0092] idle_sarsa_ptr->state_index = Obtain the current state index.

[0093] Idle_sarsa_ptr ->reward= state_action_reward;

[0094] Update the reward matrix with the reward by referencing the linked list entry.

[0095] Idx_ptr_local = idle_sarsa_ptr->next;

[0096] s t =idx_ptr_local->state_index;

[0097] a t =idx_ptr_local->action_index;

[0098] s t1 =idx_ptr_local->next->state_index;

[0099] a t1 =idx_ptr_local->next->action_index;

[0100] Reward_matrix [s t [a t += learning_rate x (idx_ptr_local->next->reward)

[0101] + discounting_factor x Reward_matrix [s t1 [a t1 - Reward_matrix [s t [a t );

[0102] In this embodiment, the RISC-V processor can enter a deeper idle state when using the Menu controller optimized by the reinforcement learning SARSA algorithm than when using the conventional Menu controller. However, the average error of the idle time for all states and all RISC-V processors is still very low, indicating better idle state selection and more power savings.

[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0104] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention. All of these are within the protection scope of the present invention.

Claims

1. A processor low power consumption state switching method for a RISC-V server platform, characterized in that: Here are the steps: (1) Configuring multi-level low power state information of the RISC-V processor, the multi-level low power state information is: defining a low power state object in the DSDT table of the ACPI firmware, the object provides complete information of the low power state, including minimum dwell time, worst case wake-up delay, dwell counter frequency, dwell counter register, entry method and state name; (2) Initializing the CPU idle state management module using multi-level low power state information; (3) The CPU idle state management module selects the optimal low-power state for the RISC-V processor based on the control strategy of the reinforcement learning SARSA algorithm. Specifically: The control strategies include trapezoidal strategy and vegetable strategy; The Menu controller gives a prediction value based on the time of the next timer event. The Menu controller checks the consistency of the time intervals through the latest 8 dwell times. If the standard deviation of these time intervals is less than the threshold, the average of the 8 dwell times is selected as the supplementary prediction value, and the minimum of the above two prediction values ​​is taken as the final prediction time. The Menu controller uses the performance multiplier, the expected residence time, and the system delay requirement to calculate the system delay tolerance. The system delay requirement is used as the first system tolerance. The other system tolerance is calculated using the formula: expected residence time / (1+10*number of IO waiting tasks on the current CPU). The minimum value of the two system tolerances is taken as the system delay tolerance. Select an idle state based on the predicted time and system delay tolerance, compare the predicted time with the stay time of all idle states, and select a specific idle state if the corresponding stay time is less than the predicted time. In addition, compare the exit delay of the state with the system delay tolerance. The exit delay of the state must be less than the system delay tolerance. Select the idle state that meets the two conditions and has the lowest power consumption as the best low-power state. (4)The RISC-V processor enters a specified low-power state.

2. The processor low power consumption state switching method of the RISC-V server platform according to claim 1, characterized in that: In step (2), specifically: the CPU idle state management module includes a core component, a control strategy component and a drive component, and the initialization process includes a core component initialization process, a control strategy component initialization process and a drive component initialization process.

3. The processor low power consumption state switching method of the RISC-V server platform according to claim 2, characterized in that: The core component initialization process is to initialize the key variables and functions required for idle state management. First, the key variables are initialized, that is, the variables that store registered device nodes. Then, the functions are initialized one by one for entering the idle state, selecting the idle state, and recording the idle state.

4. The processor low power consumption state switching method of the RISC-V server platform as claimed in claim 3, characterized in that: The initialization process of the control strategy component is as follows: first, define and initialize the scheduler structure, which includes the scheduler's name, score, and two key function pointers, which are used to select the idle state and reflect the state change respectively. Then, call the registration function to register the scheduler structure.

5. The processor low power consumption state switching method of the RISC-V server platform according to claim 4, characterized in that: The driver component initialization process is to register the device driver and perform initialization operations when the device starts.

6. The processor low power consumption state switching method of the RISC-V server platform according to claim 5, characterized in that: In step (3), the SARSA algorithm is used to optimize the idle state selection strategy. Based on the state-action pair, the algorithm updates the Q value through the following formula: Q(S t ,a t )←Q(S t ,a t )+α[r t+1 +γQ(S t+1 ,a t+1 )-Q(S t+1 ,a t+1 )], where S t Indicates the current state, a t Indicates the action taken in the current state, r t+1 is the immediate reward obtained when entering the next state after executing the action, S t+1 Indicates the next state to be transferred to after executing the action, a t+1 Indicates the action selected in the next state, Q(S t ,a t ) is the Q value of the current state-action pair, which is used to represent the state S t Take action a t The expected cumulative reward that can be obtained, α is the learning rate, which is used to balance the weight of new information and historical information, and γ is the discount factor, which is used to measure the impact of future rewards.

7. The processor low power consumption state switching method of the RISC-V server platform according to claim 6, characterized in that: Step (4) specifically involves the collaboration between the operating system and the SBI firmware. The operating system sends a request to specify the target low-power state of the RISC-V processor. After receiving the ecall call request, the SBI firmware parses the call parameters to determine the target low-power state and related configuration parameters, and controls the RISC-V processor to quickly switch from the running state to the specified low-power state.

Citation Information

Patent Citations

  • Energy consumption management method for inserting system

    CN101067758A

  • Computer system and computer system power management method

    CN101145080A