Intelligent control method and system for frequency modulation and voltage regulation of server CPU
By evaluating the jitter patterns and impact of server CPUs, an adjustment and control scheme is generated, which solves the problems of untimely latency and fluctuation in existing frequency and voltage regulation technologies. It realizes intelligent frequency and voltage regulation in suitable scenarios, thereby improving CPU performance and stability.
Patent Information
- Application Number
- CN202512006303.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-24
AI Technical Summary
In existing technologies, the dynamic frequency adjustment method for server CPUs fails to take into account real-time operating scenarios, resulting in frequent or inappropriate frequency and voltage adjustments, introducing additional latency and performance fluctuations, making it difficult to achieve accurate and timely adjustments while ensuring that performance is not negatively affected.
By collecting CPU operation data, it is determined whether jitter exists and its pattern is evaluated. Based on the jitter frequency, transmitted information and amplified information, the degree of impact is assessed, and an adjustment and control scheme is generated. Frequency and voltage adjustment is only performed in suitable scenarios.
Reduce resource waste, minimize the negative impact of frequency and voltage regulation on CPU tasks, improve the intelligent control adaptability and accuracy of frequency and voltage regulation, and optimize energy efficiency and stability.
Smart Images

Figure CN121918995A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of frequency and voltage regulation control, and in particular to an intelligent control method and system for frequency and voltage regulation of a server CPU. Background Technology
[0002] As the core infrastructure of cloud computing, big data, and various network services, the energy efficiency and stability of server CPUs are crucial. Dynamic frequency and voltage regulation technology for server CPUs can optimize energy efficiency, control temperature, and manage heat dissipation under different load conditions, thereby improving overall performance. Among related technologies, traditional CPU dynamic frequency regulation methods mostly rely on preset fixed strategies, such as frequency adjustment based on CPU utilization or by user-space programs. However, these existing technologies do not deeply consider whether the server's current real-time operating scenario is truly suitable for frequency and voltage regulation. Frequent or inappropriate frequency and voltage adjustments may introduce additional latency or even exacerbate performance fluctuations. Therefore, existing technologies struggle to achieve precise and timely adjustments to CPU frequency and voltage without negatively impacting actual server performance, resulting in unmaximized efficiency and significant room for improvement. Summary of the Invention
[0003] The purpose of this invention is to provide an intelligent control method and system for frequency and voltage regulation of server CPUs, so as to solve the problems mentioned in the background art.
[0004] Firstly, this application provides an intelligent control method for frequency and voltage regulation of a server CPU, employing the following technical solution: Collect server CPU operating data and determine whether the server CPU is experiencing CPU jitter based on the operating data; If CPU jitter exists, collect the jitter data and determine whether the CPU jitter has a regularity based on the jitter data; If the CPU jitter is regular, then collect the CPU jitter frequency, jitter transmission information, and jitter amplification information; The impact of CPU jitter is assessed based on the CPU jitter frequency, jitter transmission information, and jitter amplification information. The decision to adjust the frequency and voltage of the server CPU is based on the degree of jitter. If the frequency and voltage of the server CPU are adjusted, a corresponding adjustment and control scheme is generated based on the operating data.
[0005] Preferably, the step of collecting CPU jitter data and determining whether the CPU jitter has a regularity based on the jitter data is as follows: The jitter data is used to determine whether the CPU jitter is caused by the user. If it is caused by the user, user information is collected, including task information and operation information. Based on the task information, confirm the real-time running task and determine whether the real-time running task is fully automatic control. If it is not fully automatic control, then obtain the percentage of user control over the real-time running task; The degree of operational change of the real-time running task is evaluated based on the operation information, and the degree of task change of the real-time running task is obtained by combining the user control ratio. Determine if CPU jitter has a pattern based on the degree of task variability.
[0006] Preferably, the step of determining whether CPU jitter is caused by the user based on jitter data is as follows: Collect factors that cause CPU jitter, filter out factors that are irrelevant to the running task, and record them as jitter factors; Based on the operational data, determine whether the server CPU operation has jitter factors. If jitter factors exist, collect the jitter factors that exist in real time and record them as real-time factors. Retrieve historical jitter data corresponding to real-time factors, collect real-time jitter data, and compare the jitter similarity between historical jitter data and real-time jitter data; Set a jitter similarity standard. If the jitter similarity reaches the jitter similarity standard, it is determined that the CPU jitter is not caused by the user. If the jitter similarity does not meet the jitter similarity standard, then the CPU jitter is determined to be caused by the user.
[0007] Preferably, the step of obtaining the user's control percentage over real-time running tasks is as follows: Real-time running tasks are divided into user operation tasks and automatic operation tasks. The task ratio of user operation tasks is collected and recorded as the operation volume ratio. Determine whether the user's operation task has operation rules. If it does, obtain the corresponding rule operation range based on the operation rules. Obtain the total operation range of the real-time running task, and compare the rule operation range with the total operation range to obtain the operation range percentage; If there are no operation rules, the average control range of users over user operation tasks is statistically analyzed based on operation information and used as the rule operation range. The percentage of user control over real-time running tasks is obtained by combining the percentage of operation volume and the percentage of operation scope.
[0008] Preferably, the step of evaluating the degree of change in user operations during real-time task execution based on operation information is as follows: Determine whether the operation of the real-time running task has a theoretical basis. If it does not have a theoretical basis, then obtain the average operation frequency of the user on all tasks based on the operation information and record it as the user operation frequency. The average operation frequency of all users for all tasks is calculated and recorded as the total operation frequency. The ratio of user operation frequency to total operation frequency is then calculated and recorded as the task variability. The average operation frequency of users on real-time running tasks is calculated and recorded as the task operation frequency. The average operation frequency of all users on real-time running tasks is calculated and recorded as the task standard frequency. The ratio of the standard frequency of a task to the frequency of its operation is used as the task's importance, and the degree of operational variability is obtained by combining the task's variability. If there is a theoretical basis, the frequency of occurrence of the theoretical basis is statistically analyzed and used as the degree of operational variability.
[0009] Preferably, the step of evaluating the degree of CPU jitter impact based on the CPU jitter frequency, jitter transmission information, and jitter amplification information specifically includes: The system determines whether the CPU jitter propagation is fixed based on the jitter propagation information. If it is not fixed, the system collects the call chain information. The propagation amplification value and propagation variation value of CPU jitter are estimated based on the call chain information and jitter amplification information, and the propagation impact is obtained by combining the propagation amplification value and propagation variation value. The task situation is assessed based on the jitter amplification information to evaluate the impact of CPU jitter, and the degree of task impact is determined based on the task situation. The impact of CPU jitter is obtained by superimposing the degree of transmission impact, the degree of task impact, and the CPU jitter frequency.
[0010] Preferably, the step of estimating the propagation amplification value and propagation variation value of CPU jitter based on call chain information and jitter amplification information, and obtaining the degree of propagation impact based on the combined propagation amplification value and propagation variation value, specifically includes: Collect jitter dependencies, determine whether the dependencies are serial synchronous, and if they are serial synchronous, collect all call chains in the real-time running task; The jitter amplification coefficient of all nodes in the call chain is collected based on the jitter amplification information, and the final node is obtained based on the call chain. Collect the starting point and starting value of CPU jitter. Based on the jitter amplification coefficient of all nodes in the starting point and the final node, accumulate and amplify the starting value to obtain the call chain amplification value. Calculate the average value of all call chain amplification values as the propagation amplification value. If it is not serial synchronous, then determine whether the dependency relationship is parallel asynchronous. If it is parallel asynchronous, then obtain the largest jitter amplification factor among all nodes in the call chain based on the jitter amplification factor, calculate the average of the largest jitter amplification factors in all call chains as the average amplification factor, and calculate the transmission amplification value based on the average amplification factor and the initial value. If it is not parallel and asynchronous, then determine whether the dependency relationship is an asynchronous callback. If it is an asynchronous callback, then calculate the average amplification value of all call chains as the propagation amplification value. The transmission level of the call chain is collected, the transmission change value is evaluated based on the transmission level, and the transmission amplification value is superimposed to obtain the degree of transmission impact.
[0011] Preferably, the step of collecting the call chain's transmission level and evaluating the transmission change value based on the transmission level is as follows: Group the call chains with the same passing level into one category to obtain multiple call chains of different levels. Count the number of call chains of different levels and record it as the number of levels. Statistically analyze the frequency of call chain changes for real-time running tasks, and count the total number of call chains as the call chain count; Obtain task branch information for real-time running tasks, and count the number of branches for real-time running tasks based on the task branch information; The propagation change value is obtained by combining the number of levels, the frequency of call chain changes, the number of call chains, and the number of branches.
[0012] Preferably, the steps for assessing the impact of CPU jitter on tasks based on jitter amplification information, and for determining the degree of task impact based on the task information, are as follows: Get the task information of CPU jitter, find the jitter starting task based on the task information, collect the call chain nodes affected by jitter and record them as jitter nodes; Count the number of nodes used by non-jitter-initiated tasks, calculate the ratio of jitter-initiated nodes to the number of nodes used, and record it as the jitter-node ratio. Collect the jitter amplification coefficient corresponding to the jitter node and the usage frequency of the jitter node in the non-jitter-starting task, and sum the jitter amplification coefficient and usage frequency to obtain the jitter interference degree of the non-jitter-starting task; The percentage of task requirements for all non-jitter-initiated tasks is calculated, and the impact on the task is obtained by summing the corresponding jitter interference levels.
[0013] Secondly, the intelligent control system for frequency and voltage regulation of a server CPU provided in this application adopts the following technical solution: An intelligent control system for frequency and voltage regulation of a server CPU includes: The jitter detection module collects the server CPU's operating data and determines whether the server CPU is experiencing jitter based on the operating data. The pattern detection module collects CPU jitter data if CPU jitter exists, and determines whether the CPU jitter has a pattern based on the jitter data. The jitter acquisition module collects the CPU jitter frequency, jitter transmission information, and jitter amplification information if the CPU jitter is regular. The jitter impact module evaluates the degree of CPU jitter impact based on the CPU jitter frequency, jitter transmission information, and jitter amplification information. The adjustment and control module determines whether to adjust the frequency and voltage of the server CPU based on the degree of jitter. If the server CPU needs to be adjusted in frequency and voltage, a corresponding adjustment and control scheme is generated based on the operating data.
[0014] In summary, this application includes at least one of the following beneficial technical effects: 1. Based on the server CPU's operating data, determine if CPU jitter exists. Analyze the jitter data to determine if the jitter is patterned. If jitter is present and patterned, assess the impact of the jitter based on its frequency, transmitted information, and amplified information. Determine whether to adjust the server CPU's frequency and voltage based on the impact. If so, generate a corresponding adjustment and control scheme based on the operating data. Determine if the real-time scenario is suitable for adjusting the server CPU's frequency and voltage, thereby controlling the adjustment. This reduces unnecessary resource waste from frequency and voltage adjustments and minimizes the negative impact on server CPU tasks caused by them, improving the scenario adaptability of intelligent control for server CPU frequency and voltage adjustments.
[0015] 2. Based on the real-time factors causing CPU jitter, determine whether the CPU jitter is user-induced. If it is user-induced and the running task is user-controlled, calculate the user's control percentage over the real-time running task based on the proportion of task volume, the proportion of operation range, and the user's average control range. Calculate the task variability of the real-time running task based on theoretically based operation frequency, the user's operation frequency variation in different scenarios, and the user control percentage. Finally, determine whether the CPU jitter exhibits a pattern based on the task variability. By evaluating whether the CPU jitter exhibits a pattern, further select whether to control the server CPU frequency and voltage adjustment, reducing situations where poor frequency and voltage adjustment effects or even negative consequences are caused by a lack of understanding of CPU jitter, thus improving the actual effectiveness of frequency and voltage adjustment and enhancing the implementation effect of intelligent control of server CPU frequency and voltage adjustment.
[0016] 3. If the propagation of CPU jitter is not fixed, the amplification value during the propagation process of jitter with different dependencies is evaluated separately, and the propagation amplification value is obtained by averaging the amplification of all call chains. The propagation change value is obtained by statistically analyzing the number of call chains at different levels, the frequency of call chain changes, the number of call chains, and the number of branches in the running task. The propagation impact is then obtained by summing the propagation amplification value. The jitter interference degree is obtained by combining the proportion of jitter nodes in the task, the jitter amplification coefficient, and the usage frequency of jitter nodes. The task demand proportion of all non-jitter-initiated tasks is statistically analyzed, and the corresponding jitter interference degree is summed to obtain the task impact degree. Finally, the propagation impact degree, task impact degree, and CPU jitter frequency are summed to obtain the CPU jitter impact degree. Evaluating the propagation, amplification, and impact on tasks of CPU jitter helps determine whether the current scenario is suitable for frequency and voltage adjustment of the server CPU, reducing the inconvenience caused by frequency and voltage adjustment to tasks and improving the accuracy of intelligent control of server CPU frequency and voltage adjustment. Attached Figure Description
[0017] Figure 1 This is a schematic diagram illustrating the specific steps of an embodiment of the intelligent control method for frequency and voltage regulation of a server CPU according to the present invention.
[0018] Figure 2 This is a schematic diagram of the module connections of an embodiment of an intelligent control system for frequency and voltage regulation of a server CPU according to the present invention. Detailed Implementation
[0019] The following examples and... Figures 1-2 The present invention will be described in further detail, but the embodiments of the present invention are not limited thereto.
[0020] This invention discloses an intelligent control method for frequency and voltage regulation of a server CPU, specifically including the following steps: Step S1: Collect the server CPU's operating data and determine whether the server CPU is experiencing CPU jitter based on the operating data.
[0021] To collect server CPU usage data and determine if CPU jitter exists, a script can be written to periodically read the Linux system's ` / proc / stat` file. This file records the cumulative values of various CPU times since system startup, including user mode, system mode, and idle time—in other words, runtime data. After plotting the collected usage data into a curve in chronological order, if frequent and dramatic alternating peaks and troughs are observed in the usage rate over a short period (e.g., seconds to minutes), rather than a relatively stable curve, then CPU jitter can be identified.
[0022] Step S2: If CPU jitter exists, collect the jitter data and determine whether the CPU jitter is regular based on the jitter data.
[0023] If the CPU jitter is irregular, the frequency of the CPU jitter will determine whether to adjust the frequency and voltage of the server CPU. If the CPU frequency reaches a preset threshold, it will be determined not to adjust the frequency and voltage.
[0024] Step S3: If the CPU jitter is regular, then collect the CPU jitter frequency, jitter transmission information, and jitter amplification information.
[0025] Step S4: Evaluate the degree of CPU jitter impact based on the CPU jitter frequency, jitter transmission information, and jitter amplification information.
[0026] Step S5: Determine whether to adjust the frequency and voltage of the server CPU based on the degree of jitter. If the frequency and voltage of the server CPU are adjusted, generate a corresponding adjustment and control scheme based on the operating data.
[0027] In scenarios where frequency and voltage regulation of the server CPU is required, a corresponding adjustment scheme is generated based on the operating data and existing frequency and voltage regulation technologies (such as DVFS technology).
[0028] In practical applications, frequency and voltage regulation of server CPUs can optimize energy efficiency, control temperature, and manage heat dissipation, thereby improving overall performance. However, not all scenarios are suitable for frequency and voltage regulation of server CPUs. For example, extremely high frequency jitter is unsuitable for frequency and voltage regulation of server CPUs, mainly because its control cycle is extremely short, contradicting the stabilization time required for CPU frequency and voltage adjustments. Frequency and voltage switching is not instantaneous and requires a certain stabilization time to avoid signal distortion or system crashes. Extremely high frequency jitter will cause frequency and voltage regulation commands to be triggered again before stabilization, resulting in the system being in an unstable transitional state. This not only fails to achieve performance optimization but also exacerbates latency and power consumption, and may even lead to serious hardware errors or data loss. Therefore, server CPU frequency and voltage regulation strategies need to be based on a relatively stable and predictable cycle to ensure system stability and reliability. Therefore, in cases where server CPUs exhibit extremely high frequency jitter, frequency and voltage regulation may actually degrade server CPU performance. Thus, confirming whether a scenario is suitable for frequency and voltage regulation and then controlling the server CPU accordingly is beneficial for improving the performance of frequency and voltage regulation.
[0029] The steps for collecting CPU jitter data and determining whether the CPU jitter has a regularity based on the jitter data are as follows: Step S21: Determine whether the CPU jitter is caused by the user based on the jitter data. If it is caused by the user, collect user information, which includes task information and operation information.
[0030] Jitter data includes periodic jitter (the deviation of each CPU clock cycle from the ideal cycle), adjacent cycle jitter (the change between two consecutive cycles), and time interval error (TIE, i.e., the cumulative deviation of the signal edge from the ideal position), etc. If it is not caused by the user, then the factors causing CPU jitter should be used to determine whether it has a pattern.
[0031] Step S22: Based on the task information, confirm the real-time running task and determine whether the real-time running task is fully automatic.
[0032] The task information includes the type of task currently running, the task duration, etc. If there is no user operation control for the task, the task is determined to be fully automatic.
[0033] Step S23: If it is not fully automatic control, then obtain the user's control percentage over the real-time running task.
[0034] Step S24: Evaluate the degree of change in the user's operation of the real-time running task based on the operation information, and obtain the degree of change in the real-time running task by combining the user's control ratio.
[0035] Multiply the degree of operational variability by the proportion of user control to obtain the degree of task variability for real-time running tasks.
[0036] Step S25: Determine whether CPU jitter has a regularity based on the degree of task variability.
[0037] In practical applications, a threshold for task variability can be set to determine if there is a pattern to the jitter. If the task variability reaches the threshold, the jitter is considered irregular because the task is constantly changing, and therefore the resulting jitter will also be constantly changing. Server CPU frequency and voltage regulation relies on predicting short-term future loads to adjust voltage and frequency in advance. If CPU jitter is completely irregular, it means that load changes are random and unpredictable, making it difficult for the system to accurately determine when to increase or decrease frequency and voltage. This blindly following chaotic signals leads to a misalignment between the voltage and frequency switching settling time and the actual load demand. This not only fails to optimize energy efficiency but also introduces additional latency and power consumption due to frequent and ineffective state transitions, potentially exacerbating system instability.
[0038] The steps to determine whether CPU jitter is caused by the user based on jitter data are as follows: Step S211: Collect the factors that cause CPU jitter, filter out the factors that are not related to the running task and record them as jitter factors.
[0039] Many factors can cause CPU jitter, such as thread contention and resource contention. When multiple threads execute concurrently and compete for CPU resources, the operating system's process scheduling algorithm may fail to allocate time slices evenly, causing fluctuations in CPU utilization. System activities such as interrupt handling are also important factors, as are speculative execution failures at the CPU microarchitecture level. Some of these factors are inherent to the system's technology and operation, while others are related to the user's scheduling and control of tasks.
[0040] Step S212: Determine whether the server CPU operation has jitter factors based on the running data. If jitter factors exist, collect the jitter factors that exist in real time and record them as real-time factors.
[0041] Step S213: Retrieve historical jitter data corresponding to real-time factors, collect real-time jitter data, and compare the jitter similarity between historical jitter data and real-time jitter data.
[0042] Jitter data refers to quantitative metrics describing fluctuations in CPU operating time, including Time Interval Error (TIE), which is the deviation between the actual arrival time of a signal edge and its ideal position. Periodic jitter represents the difference between a single clock cycle and the ideal cycle. Adjacent cycle jitter, long-term jitter, and other metrics reflect CPU jitter characteristics.
[0043] Step S214: Set the jitter similarity standard. If the jitter similarity reaches the jitter similarity standard, it is determined that the CPU jitter is not caused by the user.
[0044] Step S215: If the jitter similarity does not meet the jitter similarity standard, then it is determined that the CPU jitter is caused by the user.
[0045] In practical applications, similarity can be obtained through cosine similarity comparison. If the jitter similarity reaches the similarity standard, it can be determined that the jitter is caused by the server CPU itself, such as its manufacturing process or operation. If it is caused by the server itself, historical jitter data corresponding to that factor is obtained to determine if it has a pattern. If a pattern can be found in the historical jitter data, then the CPU jitter is determined to be regular.
[0046] The specific steps for obtaining the user's percentage of control over real-time running tasks are as follows: Step S231: Divide the real-time running tasks into user operation tasks and automatic operation tasks, collect the task ratio of user operation tasks and record it as the operation volume ratio.
[0047] During server CPU operation, multiple tasks typically exist. Some tasks are fully automated, while others involve user operations. Tasks involving user operations are categorized as user-operated tasks; otherwise, they are classified as automated tasks. Fully automated tasks are usually triggered by preset scripts or management systems, requiring no manual intervention. For example, they may use shell or Python scripts to monitor system resources (such as CPU usage, memory, and disk space) and automatically execute alerts or resource cleanup operations when thresholds are exceeded. User-operated tasks, on the other hand, directly rely on real-time user interactions, such as high-concurrency requests (e.g., flash sales on e-commerce platforms or player actions in online games). These requests cause the CPU to frequently process dynamic content generation, database queries, and real-time calculations.
[0048] Step S232: Determine whether the user's operation task has operation rules. If it does, obtain the corresponding rule operation range based on the operation rules.
[0049] Different tasks have different operation rules, and some are entirely dependent on the user's wishes. For example, in game artificial intelligence (AI) or robot control, behavior trees are often used to define task rules, which contain various node types: action tasks directly change the game state or perform specific actions (such as making the character move or attack), but the actions are predefined and can only be selected from the rules.
[0050] Step S233: Obtain the total operation range of the real-time running task, and compare the rule operation range with the total operation range to obtain the operation range percentage.
[0051] Some tasks allow users to operate within a certain range, while those beyond the range are automatically controlled. This results in a limited range of user operation in some semi-automatic tasks, so the operation range does not account for 100%.
[0052] Step S234: If there are no operation rules, the average control range of the user's operation tasks is statistically calculated based on the operation information and used as the rule operation range.
[0053] Step S235: Combine the percentage of operation volume and the percentage of operation scope to obtain the percentage of user control over real-time running tasks.
[0054] In practical applications, if no operational rules are available, the scope of operation is entirely controlled by the user and implemented according to their wishes. Different users have different operational habits; therefore, their average operational range for a task is used as the rule's operational range. The ratio of this average operational range to the total operational range is then used to obtain the operational range percentage. Multiplying the operational volume percentage by the operational range percentage yields the user's control percentage over the real-time running task.
[0055] The steps for evaluating the degree of change in user operations during real-time task execution based on operation information are as follows: Step S241: Determine whether the operation of the real-time running task has a theoretical basis. If it does not have a theoretical basis, then obtain the average operation frequency of the user on all tasks based on the operation information and record it as the user operation frequency.
[0056] Whether there is a theoretical basis refers to whether a user needs to meet a certain condition or scenario before they can perform an operation on a task. Some tasks can be performed by users, but only when a certain condition is triggered. Such tasks have a theoretical basis, meaning that users cannot operate arbitrarily. Operation information includes data such as the operation frequency of all users for all tasks.
[0057] Step S242: Calculate the average operation frequency of all users for all tasks and record it as the total operation frequency. Compare the ratio of user operation frequency to total operation frequency and record it as the task variability.
[0058] Different users operate on tasks at different frequencies. For example, when given a choice of actions, some users are accustomed to performing them, while others are not. Therefore, the higher the user's operation frequency, the more likely the user is to prefer variety, and the greater the task's variability. This is because users may change tasks at any time according to their own wishes.
[0059] Step S243: Calculate the average operation frequency of users on real-time running tasks and record it as the task operation frequency; calculate the average operation frequency of all users on real-time running tasks and record it as the task standard frequency.
[0060] Step S244: The ratio of the task standard frequency to the task operation frequency is compared and used as the task importance. The operation variability is obtained by combining the task variability.
[0061] Different tasks are perceived as having varying degrees of importance by different users. For more important tasks, the frequency of operation can easily lead to task changes and increase the risk of the task. Therefore, if a user operates on the same task more frequently than other users, it indicates that the user does not value the task as much.
[0062] Step S245: If there is a theoretical basis, the frequency of occurrence of the theoretical basis is statistically analyzed and used as the degree of operational variability.
[0063] In practical applications, weighted coefficients are set for task importance and task variability, and the operation variability is calculated by weighted summation. A higher task importance results in a lower operation variability because users value the task more and therefore operate more cautiously, reducing the frequency of operations. Conversely, a higher task variability indicates that users typically prefer to modify tasks, making them more likely to cause CPU fluctuations, thus resulting in a higher task operation variability.
[0064] The steps for assessing the impact of CPU jitter based on CPU jitter frequency, jitter transmission information, and jitter amplification information are as follows: Step S41: Determine whether the CPU jitter transmission is fixed based on the jitter transmission information. If it is not fixed, collect the call chain information.
[0065] The information involved in jitter propagation is multi-layered and multi-dimensional. Jitter propagation information mainly includes the jitter's start timestamp, duration (e.g., the stabilization time required for CPU frequency switching), jitter interval (the period length of a periodic jitter), and propagation delay (the time difference between the jitter source and its perception by the target process), etc. While CPU jitter propagation is not inherently fixed, under extremely specific and strictly controlled environments, jitter propagation can exhibit high statistical stability. Therefore, unless under particularly stringent conditions, it cannot be concluded that CPU jitter propagation is not fixed.
[0066] Step S42: Estimate the propagation amplification value and propagation variation value of CPU jitter based on the call chain information and jitter amplification information, and obtain the degree of propagation impact based on the combined propagation amplification value and propagation variation value.
[0067] Step S43: Assess the task status of CPU jitter impact based on jitter amplification information, and obtain the degree of task impact based on the task status.
[0068] Step S44: Superimpose the degree of transmission impact, the degree of task impact, and the CPU jitter frequency to obtain the degree of CPU jitter impact.
[0069] In practical applications, as mentioned earlier, extremely high-frequency CPU jitter is unsuitable for frequency and voltage adjustment of server CPUs. Whether frequency and voltage adjustment are suitable for low-frequency, regular CPU jitter requires further discussion. If the jitter's propagation impact is low and has almost no effect on other tasks, frequency and voltage adjustment is feasible because the jitter is controlled within a highly isolated and predictable range. Its changes mainly stem from the CPU's own power state switching (P-States), rather than uncontrollable external interference (such as interrupts or resource contention). In this case, fine-tuning of voltage and frequency can optimize the energy efficiency of a specific task or slightly improve its performance without causing a chain reaction to other parts of the system. For example, in a deeply tuned server environment, a background task responsible for periodic log compression is bound to a single CPU core isolated by isolcpus and nohz_full. The minor jitter generated during the task's execution (such as cache miss delays due to memory access) primarily stems from the data access patterns of its own algorithm. Furthermore, due to core isolation, this jitter will not propagate to other critical tasks processing user requests. In this situation, one could try slowing down the background task slightly by reducing the operating frequency and voltage of the isolated core. The aim is to reduce power consumption and heat generation.
[0070] The steps for estimating the propagation amplification and propagation variation of CPU jitter based on call chain information and jitter amplification information, and for obtaining the degree of propagation impact by combining the propagation amplification and propagation variation, are as follows: Step S421: Collect the jitter dependency relationship and determine whether the dependency relationship is serial synchronous. If it is serial synchronous, collect all call chains in the real-time running task.
[0071] Here, the "dependency" of jitter specifically refers to the causal relationship and timing constraints that exist when jitter is passed between system components. That is, the jitter behavior of one component will directly affect or trigger the jitter response of another component.
[0072] Step S422: Collect the jitter amplification coefficient of all nodes in the call chain based on the jitter amplification information, and obtain the final node based on the call chain.
[0073] Jitter amplification information includes the average coefficient of CPU jitter amplification at each node, i.e., the jitter amplification coefficient, and the final node refers to the last node where jitter no longer propagates. A call chain is the path formed by the call relationships between functions, methods, or services during program execution, used to describe the complete process of request processing.
[0074] Step S423: Collect the starting point and starting value of CPU jitter. Based on the jitter amplification coefficient of all nodes in the starting point and final node, accumulate and amplify the starting value to obtain the call chain amplification value. Calculate the average value of all call chain amplification values as the transmission amplification value.
[0075] In step S424, if it is not serial synchronous, determine whether the dependency relationship is parallel asynchronous. If it is parallel asynchronous, obtain the largest jitter amplification factor among all nodes in the call chain based on the jitter amplification factor, calculate the average of the largest jitter amplification factors in all call chains as the average amplification factor, and calculate the transmission amplification value based on the average amplification factor and the initial value.
[0076] In step S425, if it is not parallel asynchronous, then determine whether the dependency relationship is an asynchronous callback. If it is an asynchronous callback, then calculate the average amplification value of all call chains as the propagation amplification value.
[0077] Step S426: Collect the transmission level of the call chain, evaluate the transmission change value based on the transmission level, and superimpose the transmission amplification value to obtain the degree of transmission impact.
[0078] The smaller the impact of jitter propagation, the better it is for frequency and voltage regulation. However, jitter propagation is related to the actual call chain, and different call chains will lead to different degrees of propagation impact. In actual task execution, the call chain is usually dynamically changing, so analyzing a single call chain can easily lead to biased results. Applying the average amplification value of all call chains is beneficial for obtaining a more comprehensive and accurate assessment of the propagation impact.
[0079] In practical applications, different dependencies lead to different propagation amplification scenarios. Serial synchronous calls execute strictly sequentially; each dependency must wait for the previous one to return before proceeding, resulting in strictly cascading jitter. Upstream jitter directly increases downstream waiting time, with the overall response time approximately equal to the sum of the times of each stage (including jitter). Therefore, the propagation amplification value accumulates, ultimately resulting in a larger amplification value. Parallel asynchronous calls, on the other hand, initiate multiple downstream calls simultaneously, waiting for all or some results to return before continuing. Jitter is "taken advantage of," with the overall response time depending on the response time (including jitter) of the slowest dependency. Therefore, under this dependency relationship, jitter does not continuously accumulate; instead, the maximum amplification value is used as the amplification value for that call chain. As for asynchronous callbacks, the main call is not blocked, and the result is notified via messages, events, or callback functions, making the jitter propagation path uncertain. Jitter can cause delayed responses or out-of-order arrivals, potentially leading to inconsistent states or logical errors; therefore, its average amplification value is used as the propagation amplification value.
[0080] The steps for collecting the call chain's transmission hierarchy and evaluating the transmission change value based on the transmission hierarchy are as follows: Step S4261: Divide the call chains with the same passing level into one category to obtain multiple level call chains, count the number of level call chains and record it as the level number.
[0081] The propagation and impact of jitter vary significantly across different system layers, primarily depending on the design goals, resource management mechanisms, and interaction methods with the external environment at each layer. For example, at the hardware level, especially in the CPU cache subsystem, jitter propagation is heavily influenced by memory access patterns. When the CPU frequently accesses a rapidly changing dataset that is larger than the L1 cache but smaller than the L2 cache, the L1 cache, due to its limited capacity, will continuously swap in and out data rows (i.e., cache jitter), causing access latency fluctuations of ±10 to 20 cycles. This jitter propagates to the L2 cache, but because the L2 cache has a larger capacity, its fluctuations are partially absorbed (potentially reduced to ±50 cycles).
[0082] Step S4262: Calculate the frequency of call chain changes for real-time running tasks, and count the total number of call chains as the number of call chains.
[0083] Step S4263: Obtain task branch information for real-time running tasks, and count the number of branches for real-time running tasks based on the task branch information.
[0084] Branch information includes data such as the number of branches in the running task.
[0085] Step S4264: Combine the number of levels, the frequency of call chain changes, the number of call chains, and the number of branches to obtain the propagation change value.
[0086] In practical applications, nodes in the call chain are prone to change, while the hierarchical structure (abstract business logic and data flow hierarchy) is relatively stable. Therefore, from a hierarchical perspective, the probability of a call chain changing is low. Thus, the larger the number of layers, the greater the potential for layer changes, leading to changes in the transmission method and consequently, a larger transmission change value. The call chain is indeed dynamic, and the larger the number of call chains corresponding to a task, the greater the degree of change, and therefore, the larger the transmission change value. However, even with multiple call chains, it doesn't mean the task will frequently change call chains. Therefore, the higher the frequency of call chain changes, the larger the transmission change value. Similarly, the more branches there are, the more likely the task is to change, which can also lead to jitter in transmission, resulting in a larger transmission change value. Combining the entropy weight method to calculate the objective weights of the number of layers, call chain change frequency, and the number of branches, the transmission change value is obtained through information entropy weighting.
[0087] The steps for assessing the impact of CPU jitter on tasks based on jitter amplification information, and for determining the degree of impact on tasks based on task information, are as follows: Step S431: Obtain the task status of CPU jitter, obtain the jitter initiation task based on the task status, collect the call chain nodes affected by jitter and record them as jitter nodes.
[0088] Step S432: Count the number of nodes used by non-jittering starting tasks, calculate the ratio of jittering nodes to the number of nodes used, and record it as the jittering node ratio.
[0089] The jitter node ratio refers to the proportion of jitter nodes present in a non-jitter-initiated task.
[0090] Step S433: Collect the jitter amplification coefficient corresponding to the jitter node and the usage frequency of the jitter node in the non-jitter starting task, and sum the jitter interference degree of the non-jitter starting task based on the jitter amplification coefficient and the usage frequency.
[0091] The jitter interference level is obtained by multiplying the jitter amplification factor of the jitter node by the usage frequency and then summing the results.
[0092] Step S434: Calculate the percentage of task requirements for all non-jitter-initiated tasks, and sum the percentages based on the corresponding jitter interference to obtain the degree of task impact.
[0093] In practical applications, the impact on tasks is obtained by multiplying the jitter interference level by the corresponding task demand ratio and then summing the results. The task demand ratio can be set by the user or determined by the proportion of resources used by the task to the total resources. The fewer tasks affected by CPU jitter, the smaller the overall impact of frequency and voltage regulation caused by CPU jitter during the frequency and voltage regulation process. In other words, jitter may affect the current frequency and voltage regulation, but it only affects a very small portion of the tasks, and the impact on the overall performance of the server CPU is still relatively small.
[0094] An intelligent control system for frequency and voltage regulation of a server CPU, comprising the following steps, using an intelligent control method for frequency and voltage regulation of a server CPU as described above: The jitter detection module collects the server CPU's operating data and determines whether the server CPU is experiencing jitter based on the operating data.
[0095] The pattern detection module collects CPU jitter data if CPU jitter exists, and determines whether the CPU jitter has a pattern based on the jitter data.
[0096] The jitter acquisition module collects the CPU jitter frequency, jitter transmission information, and jitter amplification information if the CPU jitter is regular.
[0097] The jitter impact module assesses the degree of CPU jitter impact based on the CPU jitter frequency, jitter transmission information, and jitter amplification information.
[0098] The adjustment and control module determines whether to adjust the frequency and voltage of the server CPU based on the degree of jitter. If the server CPU needs to be adjusted in frequency and voltage, a corresponding adjustment and control scheme is generated based on the operating data.
[0099] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A smart control method for frequency and voltage regulation of a server CPU, characterized in that, Includes the following steps: Collect server CPU operating data and determine whether the server CPU is experiencing CPU jitter based on the operating data; If CPU jitter exists, collect the jitter data and determine whether the CPU jitter has a regularity based on the jitter data; If the CPU jitter is regular, then collect the CPU jitter frequency, jitter transmission information, and jitter amplification information; The impact of CPU jitter is assessed based on the CPU jitter frequency, jitter transmission information, and jitter amplification information. The decision to adjust the frequency and voltage of the server CPU is based on the degree of jitter. If the frequency and voltage of the server CPU are adjusted, a corresponding adjustment and control scheme is generated based on the operating data.
2. The intelligent control method for frequency and voltage regulation of a server CPU according to claim 1, characterized in that, The steps for collecting CPU jitter data and determining whether the CPU jitter has a regularity based on the jitter data are as follows: The jitter data is used to determine whether the CPU jitter is caused by the user. If it is caused by the user, user information is collected, including task information and operation information. Based on the task information, confirm the real-time running task and determine whether the real-time running task is fully automatic control. If it is not fully automatic control, then obtain the percentage of user control over the real-time running task; The degree of operational change of the real-time running task is evaluated based on the operation information, and the degree of task change of the real-time running task is obtained by combining the user control ratio. Determine if CPU jitter has a pattern based on the degree of task variability.
3. The intelligent control method for frequency and voltage regulation of a server CPU according to claim 2, characterized in that, The steps to determine whether CPU jitter is caused by the user based on jitter data are as follows: Collect factors that cause CPU jitter, filter out factors that are irrelevant to the running task, and record them as jitter factors; Based on the operational data, determine whether the server CPU operation has jitter factors. If jitter factors exist, collect the jitter factors that exist in real time and record them as real-time factors. Retrieve historical jitter data corresponding to real-time factors, collect real-time jitter data, and compare the jitter similarity between historical jitter data and real-time jitter data; Set a jitter similarity standard. If the jitter similarity reaches the jitter similarity standard, it is determined that the CPU jitter is not caused by the user. If the jitter similarity does not meet the jitter similarity standard, then the CPU jitter is determined to be caused by the user.
4. The intelligent control method for frequency and voltage regulation of a server CPU according to claim 2, characterized in that, The specific steps for obtaining the user's percentage of control over real-time running tasks are as follows: Real-time running tasks are divided into user operation tasks and automatic operation tasks. The task ratio of user operation tasks is collected and recorded as the operation volume ratio. Determine whether the user's operation task has operation rules. If it does, obtain the corresponding rule operation range based on the operation rules. Obtain the total operation range of the real-time running task, and compare the rule operation range with the total operation range to obtain the operation range percentage; If there are no operation rules, the average control range of users over user operation tasks is statistically analyzed based on operation information and used as the rule operation range. The percentage of user control over real-time running tasks is obtained by combining the percentage of operation volume and the percentage of operation scope.
5. The intelligent control method for frequency and voltage regulation of a server CPU according to claim 2, characterized in that, The steps for evaluating the degree of change in user operations during real-time task execution based on operation information are as follows: Determine whether the operation of the real-time running task has a theoretical basis. If it does not have a theoretical basis, then obtain the average operation frequency of the user on all tasks based on the operation information and record it as the user operation frequency. The average operation frequency of all users for all tasks is calculated and recorded as the total operation frequency. The ratio of user operation frequency to total operation frequency is then calculated and recorded as the task variability. The average operation frequency of users on real-time running tasks is calculated and recorded as the task operation frequency. The average operation frequency of all users on real-time running tasks is calculated and recorded as the task standard frequency. The ratio of the task standard frequency to the task operation frequency is used as the task importance, and the operation variability is obtained by combining the task variability. If there is a theoretical basis, the frequency of occurrence of the theoretical basis is statistically analyzed and used as the degree of operational variability.
6. The intelligent control method for frequency and voltage regulation of a server CPU according to claim 1, characterized in that, The steps for assessing the impact of CPU jitter based on CPU jitter frequency, jitter transmission information, and jitter amplification information are as follows: The system determines whether the CPU jitter propagation is fixed based on the jitter propagation information. If it is not fixed, the system collects the call chain information. The propagation amplification value and propagation variation value of CPU jitter are estimated based on the call chain information and jitter amplification information, and the propagation impact is obtained by combining the propagation amplification value and propagation variation value. The task situation is assessed based on the jitter amplification information to evaluate the impact of CPU jitter, and the degree of task impact is determined based on the task situation. The impact of CPU jitter is obtained by superimposing the degree of transmission impact, the degree of task impact, and the CPU jitter frequency.
7. The intelligent control method for frequency and voltage regulation of a server CPU according to claim 6, characterized in that, The steps for estimating the propagation amplification and propagation variation of CPU jitter based on call chain information and jitter amplification information, and for obtaining the degree of propagation impact by combining the propagation amplification and propagation variation, are as follows: Collect jitter dependencies, determine whether the dependencies are serial synchronous, and if they are serial synchronous, collect all call chains in the real-time running task; The jitter amplification coefficient of all nodes in the call chain is collected based on the jitter amplification information, and the final node is obtained based on the call chain. Collect the starting point and starting value of CPU jitter. Based on the jitter amplification coefficient of all nodes in the starting point and the final node, accumulate and amplify the starting value to obtain the call chain amplification value. Calculate the average value of all call chain amplification values as the propagation amplification value. If it is not serial synchronous, then determine whether the dependency relationship is parallel asynchronous. If it is parallel asynchronous, then obtain the largest jitter amplification factor among all nodes in the call chain based on the jitter amplification factor, calculate the average of the largest jitter amplification factors in all call chains as the average amplification factor, and calculate the transmission amplification value based on the average amplification factor and the initial value. If it is not parallel and asynchronous, then determine whether the dependency relationship is an asynchronous callback. If it is an asynchronous callback, then calculate the average amplification value of all call chains as the propagation amplification value. The transmission level of the call chain is collected, the transmission change value is evaluated based on the transmission level, and the transmission amplification value is superimposed to obtain the degree of transmission impact.
8. The intelligent control method for frequency and voltage regulation of a server CPU according to claim 7, characterized in that, The steps for collecting the call chain's transmission hierarchy and evaluating the transmission change value based on the transmission hierarchy are as follows: Group the call chains with the same passing level into one category to obtain multiple call chains of different levels. Count the number of call chains of different levels and record it as the number of levels. Statistically analyze the frequency of call chain changes for real-time running tasks, and count the total number of call chains as the call chain count; Obtain task branch information for real-time running tasks, and count the number of branches for real-time running tasks based on the task branch information; The propagation change value is obtained by combining the number of levels, the frequency of call chain changes, the number of call chains, and the number of branches.
9. The intelligent control method for frequency and voltage regulation of a server CPU according to claim 7, characterized in that, The steps for assessing the impact of CPU jitter on tasks based on jitter amplification information, and for determining the degree of impact on tasks based on task information, are as follows: Get the task information of CPU jitter, find the jitter starting task based on the task information, collect the call chain nodes affected by jitter and record them as jitter nodes; Count the number of nodes used by non-jitter-initiated tasks, calculate the ratio of jitter-initiated nodes to the number of nodes used, and record it as the jitter-node ratio. Collect the jitter amplification coefficient corresponding to the jitter node and the usage frequency of the jitter node in the non-jitter-starting task, and sum the jitter amplification coefficient and usage frequency to obtain the jitter interference degree of the non-jitter-starting task; The percentage of task requirements for all non-jitter-initiated tasks is calculated, and the impact on the task is obtained by summing the corresponding jitter interference levels.
10. An intelligent control system for frequency and voltage regulation of a server CPU, characterized in that, The application of the intelligent control method for frequency and voltage regulation of a server CPU as described in any one of claims 1-9 includes: The jitter detection module collects the server CPU's operating data and determines whether the server CPU is experiencing jitter based on the operating data. The pattern detection module collects CPU jitter data if CPU jitter exists, and determines whether the CPU jitter has a pattern based on the jitter data. The jitter acquisition module collects the CPU jitter frequency, jitter transmission information, and jitter amplification information if the CPU jitter is regular. The jitter impact module evaluates the degree of CPU jitter impact based on the CPU jitter frequency, jitter transmission information, and jitter amplification information. The adjustment and control module determines whether to adjust the frequency and voltage of the server CPU based on the degree of jitter. If the server CPU needs to be adjusted in frequency and voltage, a corresponding adjustment and control scheme is generated based on the operating data.