Intelligent computing resource energy-saving scheduling system based on load prediction
By monitoring cache access data and changes in CPU/GPU frequency, combining load classification identification and intelligent regulation strategies, resource allocation is dynamically adjusted, and the problems of insufficient load prediction accuracy and resource scheduling lag in the existing technology are solved, achieving more efficient load response and energy consumption management.
Patent Information
- Application Number
- CN202510481065.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The prior art lacks accuracy in load prediction and resource scheduling, resulting in lagging resource allocation, untimely response to burst loads, and lack of comprehensive evaluation of load dynamic characteristics, affecting computing performance and energy consumption optimization.
By monitoring cache access data, analyzing access frequency fluctuations and frequency changes, classifying loads in combination with the load classification identification module, the intelligent regulation strategy module dynamically adjusts CPU/GPU resource allocation and memory allocation to optimize resource matching and energy consumption management.
It improves the accuracy of load characteristics judgment, enhances burst load recognition capabilities, optimizes computing resource allocation, reduces energy consumption waste caused by frequency adjustment, ensures stable task execution, and reduces idle and unbalanced resources.
Smart Images

Figure CN120045332A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power management, and in particular to an intelligent computing resource energy-saving scheduling system based on load prediction. Background Art
[0002] The field of power management technology includes the control and optimization of energy consumption in environments such as computer systems, server clusters and data centers. Its core content is to reduce energy consumption and improve energy efficiency by monitoring, analyzing and regulating the power consumption of computing resources. This technical field involves dynamic power management, voltage regulation, frequency adjustment and power allocation strategies to adapt to different workload requirements. The power management of computing devices usually relies on hardware-based energy consumption control mechanisms and software-based task scheduling strategies, such as dynamic voltage and frequency adjustment, load balancing, task migration and other means. In addition, in cloud computing and distributed computing environments, power management also includes the optimization of virtual machine resource allocation and intelligent control of cooling systems to reduce unnecessary energy waste. With the popularization of artificial intelligence and large-scale computing, the application scope of power management technology has gradually expanded, not only focusing on the energy consumption of a single computing device, but also involving the energy optimization of the entire computing infrastructure.
[0003] Among them, the intelligent computing resource energy-saving scheduling system refers to the dynamic scheduling of computing resources based on load prediction to achieve energy efficiency optimization of computing tasks. The system uses a load prediction model to estimate future resource requirements based on the real-time load of computing tasks, and schedules tasks based on the power management mechanism of the computing device. Specifically, the system uses historical load data to train the prediction model, predicts future computing needs through regression analysis or time series modeling, and then performs computing task allocation based on the energy consumption characteristics of the computing nodes. Scheduling methods include task redistribution based on load balancing, task migration based on energy efficiency ratio, and computing resource management based on dynamic frequency adjustment. The system also uses a multi-objective optimization method to schedule computing resources based on the temperature, power usage, and task execution characteristics of the computing nodes, so that computing tasks can reduce power consumption while meeting performance requirements.
[0004] Existing technologies rely on historical statistical data, which makes it difficult to accurately predict real-time load changes, resulting in delayed resource allocation and untimely response to sudden loads. Load detection relies on static threshold judgment, which has limited accuracy in identifying sudden loads and may lead to mismatched computing resource adjustments. CPU / GPU frequency adjustment is based on a fixed power threshold and lacks a comprehensive assessment of the dynamic characteristics of the load, which may affect computing performance or lead to resource waste. Task scheduling uses static load balancing, which does not fully consider the characteristics of computing tasks. Some nodes run at high power consumption for a long time, and energy consumption optimization is insufficient. The computing resource allocation method is rigid and difficult to adapt to complex load changes, which affects the overall energy consumption control efficiency. The cooling system has limited energy consumption control capabilities, further increasing energy waste. Summary of the invention
[0005] The purpose of the present invention is to solve the shortcomings existing in the prior art and to propose an intelligent computing resource energy-saving scheduling system based on load prediction.
[0006] In order to achieve the above object, the present invention adopts the following technical solution: the intelligent computing resource energy-saving scheduling system based on load prediction includes: The cache access monitoring module obtains cache access data, monitors cache access patterns, calculates access frequency changes, evaluates access change ranges, counts hit rates and migration times, and obtains cache access feature data; The burst load determination module calculates the access frequency fluctuation range based on the cache access characteristic data, determines whether the growth rate is abnormal, compares the drop in the hit rate, evaluates the continuity of data fluctuation, and obtains the burst load status result; The frequency fluctuation analysis module obtains the CPU / GPU operating frequency, monitors frequency changes, calculates the fluctuation rate at adjacent moments, determines abnormal fluctuation amplitude, evaluates fluctuation rate changes, and obtains CPU / GPU frequency fluctuation status results; The load classification identification module calculates the difference between the two data based on the burst load status result and the CPU / GPU frequency fluctuation status result, compares the number of data anomalies, classifies them into burst, medium or steady-state loads, calculates the time proportion of multiple categories, and screens the load types in combination with the computing task type to obtain the load classification result; The intelligent control strategy module determines the load category based on the load classification result, compares resource allocation and data transmission, adjusts CPU / GPU resource allocation and memory allocation, and obtains a dynamic load control solution.
[0007] As a further solution of the present invention, the cache access feature data includes access frequency changes, access change amplitude, hit rate, and migration times; the burst load status results include access frequency fluctuation range, abnormal growth rate, hit rate decrease amplitude, and data fluctuation continuity; the CPU / GPU frequency fluctuation status results include CPU / GPU operating frequency, frequency change, fluctuation rate, abnormal fluctuation amplitude, and fluctuation rate change; the load classification results include load category, multi-category time proportion, and computing task type to filter load types; the load dynamic control scheme includes CPU / GPU resource allocation adjustment, memory allocation adjustment, and resource allocation and data transmission comparison.
[0008] As a further solution of the present invention, the cache access monitoring module includes: The cache access data collection submodule obtains cache access records, collects access time, access address and access interval, extracts access frequency change trends, calculates the incremental interval of multiple accesses in the access time series, sets access density indicators based on the time interval difference, classifies and summarizes the access density indicators, matches the corresponding access patterns based on the classification and summary results, and obtains cache access time series characteristics; The access pattern calculation submodule calculates the change value of multi-address access frequency based on the cache access time series characteristics, counts the access increment in a short period of time, determines the access growth rate and its change trend, and uses the formula: ; Calculate and obtain the cache access frequency change coefficient, and combine the access pattern classification calculation to obtain the cache access pattern characteristics; in, Represents the cache access frequency variation coefficient, which measures the degree to which the cache access frequency changes over time. Representative The address number of the access is used to record the specific address information when the cache is accessed. Representative The address number of the access is used as a comparison benchmark for calculating the changes in adjacent accesses. Representative The timestamp of the visit, which records the specific time when the visit occurred. Representative The timestamp of the visit is used as a comparison benchmark for calculating time changes. Represents the total number of visits, indicating the number of visits within the statistical time window; The hit rate and migration statistics submodule calculates the cache access hit rate based on the cache access pattern characteristics, counts the number of hits according to the number of accesses, obtains the hit ratio, counts the number of cache migrations, calculates the migration ratio, determines the cache load status based on the migration ratio, and obtains cache access characteristic data.
[0009] As a further solution of the present invention, the burst load determination module includes: The access frequency fluctuation calculation submodule calculates the access frequency variation range based on the cache access feature data, counts the variation range of the number of accesses per unit time, calculates the difference between the highest and lowest number of accesses, summarizes the variation range, and calculates the access fluctuation mean using the formula: ; Calculate the access frequency fluctuation range, calculate and obtain the access frequency fluctuation range, and combine the fluctuation trend to obtain the access frequency fluctuation characteristics; in, Represents the access frequency fluctuation range, which is used to measure the fluctuation of access frequency per unit time. Representative The number of visits, indicating the The number of visits counted during visits. Representative The number of visits is used as a comparison benchmark for calculating the changes in the frequency of adjacent visits. Representative The time interval between visits is the time interval between Visit to The length of time between visits, Represents the number of visits counted, indicating the total number of visits included in the calculation process; The hit rate drop evaluation submodule calculates the change range of cache hit rate based on the access frequency fluctuation characteristics, calls the access record to count the number of hits, calculates the current hit rate and compares it with the historical data, and calculates the proportion of the hit rate drop. According to the trend of the hit rate drop, the stability of the cache resource is judged, and the drop range of the hit rate is obtained; The data fluctuation continuity judgment submodule calculates the fluctuation of access data in a continuous period based on the drop in hit rate, analyzes the stability of changes in access frequency, calls access data in multiple time periods, calculates the duration of fluctuations, determines the load status according to the duration of burst access, and obtains the burst load status result.
[0010] As a further solution of the present invention, the frequency fluctuation analysis module includes: The operating frequency monitoring submodule obtains the CPU / GPU operating frequency, continuously collects frequency data according to the time series, calls the system monitoring module to record the frequency value at each moment, counts the frequency change trend at adjacent moments, calculates the difference between the maximum and minimum frequency values, and obtains the CPU / GPU operating frequency data; The fluctuation rate calculation submodule calculates the frequency change rate at adjacent moments based on the CPU / GPU operating frequency data, counts the rate fluctuation amplitude at all moments, and analyzes the fluctuation trend using the formula: ; Obtain the CPU / GPU frequency fluctuation rate through calculation, and combine it with the rate fluctuation trend to obtain the CPU / GPU fluctuation rate characteristics; in, Represents the CPU / GPU frequency fluctuation rate, represent The CPU / GPU frequency value at a time point, represent The cumulative rate value at each time point is Represents the number of statistical time points; The frequency fluctuation status assessment submodule analyzes the rate change trend based on the CPU / GPU fluctuation rate characteristics, calculates the fluctuation amplitude abnormal value, compares the rate fluctuation range of each time period, and obtains the CPU / GPU frequency fluctuation status result in combination with the rate change interval.
[0011] As a further solution of the present invention, the load classification identification module includes: The data difference calculation submodule extracts the time series of the two data based on the burst load status result and the CPU / GPU frequency fluctuation status result, calculates the mean, variance and change rate respectively, obtains the absolute difference between the two and performs normalization processing, and calculates the difference in the time series using the formula: ; Combined with the change trend, the data difference coefficient is obtained by calculation; in, Represents the data difference coefficient, which is used to measure the numerical difference between the burst load state and the CPU / GPU frequency change. Representative moments The burst load state value indicates that at time The burst load level detected by the system, Representative moments The CPU / GPU frequency value, indicating the time Calculate the actual CPU or GPU frequency the device is running at, Represents the mean of the burst load state, which indicates the average level of all burst load state values within the statistical time series range. Represents the mean of CPU / GPU frequency, which indicates the average level of all CPU / GPU frequency values within the statistical time series range. Represents the total length of the time series, indicating the total number of time points included in the statistics during the calculation process; The abnormality times statistics submodule calls the data difference coefficient, sets the sudden load abnormality threshold, performs traversal based on the time series data, calculates the difference value for each data point and compares it with the abnormality threshold, records the number of abnormal points exceeding the threshold, calculates the frequency of abnormal occurrence and its proportion in the overall data, and obtains the load abnormality times ratio; The load classification submodule sets the classification rules for burst load, medium load, and steady-state load based on the load anomaly ratio and the data difference coefficient, divides the time series data into intervals, calculates the time proportion of differentiated categories, and filters the load types in combination with the computing task type to obtain the load classification results.
[0012] As a further solution of the present invention, the intelligent control strategy module includes: The load category determination submodule calls CPU utilization, GPU utilization, memory occupancy, and data transmission rate based on the load classification result, compares multiple parameters with the load type benchmark value, screens the corresponding threshold interval, determines the load category, and obtains the load category determination result; The resource allocation comparison submodule uses the CPU / GPU resource allocation, memory usage, and data transmission rate based on the load category determination result to calculate multiple resource matching degrees using the formula: ; Obtain resource matching value through calculation, compare with resource allocation benchmark, and obtain resource allocation deviation value; in, Represents the resource matching value, which is used to measure the matching degree between the actual allocated resources and the load required resources. Representative The actual amount of resources allocated to the system in the The specific values of the CPU computing power, GPU computing power, memory usage, or data transfer rate allocated to the resource. Representative The load demand resource of the resource indicates that the system The specific values of CPU computing power, GPU computing power, memory usage or data transfer rate required on the resource. Representative The load base resource of the resource is expressed in A preset benchmark reference value or standard value for a resource category. Represents the number of resource types, indicating the total number of resource types involved in the system, including CPU computing power, GPU computing power, memory usage, data transmission rate and other different categories; The dynamic control scheme generation submodule calls the current CPU / GPU load threshold, memory usage and data transmission rate adjustment range based on the resource allocation deviation value, compares the adjustable space, adjusts the resource allocation ratio, and obtains a load dynamic control scheme.
[0013] Compared with the prior art, the advantages and positive effects of the present invention are: In the present invention, by monitoring cache access data, analyzing access frequency fluctuations, identifying access pattern changes, and improving the accuracy of load characteristic judgment. The access frequency fluctuation range calculation is combined with the hit rate reduction comparison to enhance the ability to identify sudden loads and make resource scheduling more targeted. CPU / GPU operating frequency monitoring is combined with fluctuation rate calculation to identify abnormal fluctuations, optimize computing resource allocation, and reduce energy waste caused by frequency adjustment. Load classification is based on multi-category time proportions and computing task characteristics to optimize resource matching and improve computing efficiency. Dynamically adjust CPU / GPU computing power and memory allocation, coordinate computing resources and data transmission, optimize energy consumption management, ensure stable task execution, and reduce resource idleness and imbalance problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a system flow chart of the present invention; Figure 2 This is a flow chart of the cache access monitoring module of the present invention; Figure 3 This is a flow chart of the burst load determination module of the present invention; Figure 4 This is a flow chart of the frequency fluctuation analysis module of the present invention; Figure 5 This is a flow chart of the load classification and identification module of the present invention; Figure 6 This is a flow chart of the intelligent control strategy module of the present invention. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0016] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, in the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.
[0017] Embodiment 1: See also Figure 1 The intelligent computing resource energy-saving scheduling system based on load prediction includes: The cache access monitoring module obtains cache access data, monitors cache access patterns, calculates access frequency changes, evaluates access change ranges, counts hit rates and migration times, and obtains cache access feature data; The burst load determination module calculates the access frequency fluctuation range based on cache access feature data, determines the abnormal growth rate, compares the drop in hit rate, evaluates the continuity of data fluctuation, and obtains the burst load status result; The frequency fluctuation analysis module obtains the CPU / GPU operating frequency, monitors frequency changes, calculates the fluctuation rate at adjacent moments, determines abnormal fluctuation amplitude, evaluates fluctuation rate changes, and obtains CPU / GPU frequency fluctuation status results; The load classification and identification module calculates the difference between the burst load status results and the CPU / GPU frequency fluctuation status results, compares the number of data anomalies, and classifies them into burst, medium or steady-state loads. It calculates the time proportion of multiple categories, filters the load types based on the computing task type, and obtains the load classification results. The intelligent control strategy module determines the load category based on the load classification results, compares resource allocation and data transmission, adjusts CPU / GPU resource allocation and memory allocation, and obtains a dynamic load control solution.
[0018] Cache access feature data includes access frequency changes, access change range, hit rate, and migration times. Burst load status results include access frequency fluctuation range, abnormal growth rate, hit rate decrease range, and data fluctuation continuity. CPU / GPU frequency fluctuation status results include CPU / GPU operating frequency, frequency change, fluctuation rate, abnormal fluctuation range, and fluctuation rate change. Load classification results include load category, multi-category time proportion, and computing task type to filter load types. Load dynamic control solutions include CPU / GPU resource allocation adjustment, memory allocation adjustment, and resource allocation and data transmission comparison.
[0019] See also Figure 2 ,The cache access monitoring module includes: The cache access data collection submodule obtains cache access records, collects access time, access address and access interval, extracts access frequency change trends, calculates the incremental interval of multiple accesses in the access time series, sets access density indicators based on the time interval difference, classifies and summarizes the access density indicators, matches the corresponding access patterns based on the classification and summary results, and obtains cache access time series characteristics; The cache access data collection submodule is used to obtain cache access records and collect access time, access address and access interval. First, by parsing the cache log data, extract the timestamp of each access request, the corresponding cache address number and the time interval between two adjacent accesses, and store these data in a structured data table for subsequent processing, as shown in Table 1. Then, the access time series is analyzed, and the fluctuation trend of the access frequency is obtained by calculating the increment of the continuous access interval change. In the analysis of the time series data, a benchmark for the time interval difference is set. The benchmark value can be determined by counting the mean and standard deviation of all access intervals. Assume that the cache access data is as follows: Table 1 Cache access time series data ; As shown in Table 1, the reference value of the access interval can be set to the mean of all access intervals plus a standard deviation, for example: seconds, if the standard deviation is calculated seconds, the reference value is set to For the classification of access density indicators, it is set that if the access interval is less than 5 seconds, it is considered as high-density access, 5 to 10 seconds is medium-density access, and more than 10 seconds is low-density access. Based on this classification, the access mode classification can be obtained. For example, the access interval of access A1 drops from 15 seconds to 2 seconds, indicating that the access frequency increases suddenly. This mode can be classified as a burst access mode, while the access interval of access A2 is relatively stable, which is classified as a regular access mode. Through the classification of these modes, the cache access time series characteristics are obtained.
[0020] The access pattern calculation submodule calculates the change value of multi-address access frequency based on the cache access time series characteristics, counts the access increment in a short period of time, and determines the access growth rate and its change trend. The formula is: ; Calculate and obtain the cache access frequency change coefficient, and combine the access pattern classification calculation to obtain the cache access pattern characteristics; in, Represents the cache access frequency variation coefficient, which measures the degree to which the cache access frequency changes over time. Representative The address number of the access is used to record the specific address information when the cache is accessed. Representative The address number of the access is used as a comparison benchmark for calculating the changes in adjacent accesses. Representative The timestamp of the visit, which records the specific time when the visit occurred. Representative The timestamp of the visit is used as a comparison benchmark for calculating time changes. Represents the total number of visits, indicating the number of visits within the statistical time window; The cache access pattern calculation submodule needs to calculate the multi-address access frequency change value based on the cache access time series characteristics. First, for a certain cache system, the access status of all cache addresses is recorded within a specific time window. Each access operation is accompanied by the access address number. and timestamp , the module first needs to traverse all access records, and arrange the data in ascending order according to the timestamp to build the access sequence , then calculate the absolute difference between adjacent access address numbers And sum and average. In addition, it is necessary to calculate the time-weighted access rate change, that is, the access frequency of each time Frequency compared to the previous visit The square difference between the two is summed and then squared to get the rate change. To illustrate this process, we can assume a real scenario. For example, in a server cache, the access address number may represent the data block ID. Assume that 6 accesses occur within 5 seconds, and the access address number sequence is , the access timestamp sequence is , then calculate: 1. Calculate the mean of the access address change: ; Calculate the time-weighted access rate change: ; Calculate separately: ; ; ; ; ; ; Finally get : ; This value reflects the degree of change in cache access frequency. If the value is much higher than the set benchmark value (for example, 50), it means that the current cache access pattern has changed significantly and its access trend needs to be further analyzed. See Table 1, which shows the calculation results of the access frequency change coefficient in different time windows.
[0021] The hit rate and migration statistics submodule calculates the cache access hit rate based on the cache access pattern characteristics, counts the number of hits according to the number of accesses, obtains the hit ratio, counts the number of cache migrations, calculates the migration ratio, determines the cache load status based on the migration ratio, and obtains cache access characteristic data.
[0022] The hit rate and migration statistics submodule calculates the cache access hit rate based on the cache access pattern characteristics. First, the total number of accesses is counted. and hit count , calculate the hit rate: ; In Table 1, assuming that all access data of cache A1 hits, then , total number of visits , then the hit rate is Then, count the number of cache migrations. If an address appears in different cache intervals, it is counted as a migration. For example, if A1 is moved from the main cache to the L2 cache, the number of migrations is Increase, calculate the migration ratio: ; Assuming the number of migrations is 1, the migration ratio ,Finally, the cache load status is judged based on the migration ratio.,If the migration ratio exceeds 50%, the cache load is high.,If it is less than 20%, the cache load is low.,In this case, the migration ratio is 20%, which belongs to the low load state.,Finally, the cache access characteristic data is obtained.
[0023] See also Figure 3 , the burst load determination module includes: The access frequency fluctuation calculation submodule calculates the access frequency variation range based on the cache access feature data, counts the variation range of the number of accesses per unit time, calculates the difference between the highest and lowest number of accesses, summarizes the variation range, and calculates the access fluctuation mean using the formula: ; Calculate the access frequency fluctuation range, calculate and obtain the access frequency fluctuation range, and combine the fluctuation trend to obtain the access frequency fluctuation characteristics; in, Represents the access frequency fluctuation range, which is used to measure the fluctuation of access frequency per unit time. Representative The number of visits, indicating the The number of visits counted during visits. Representative The number of visits is used as a comparison benchmark for calculating the changes in the frequency of adjacent visits. Representative The time interval between visits is the time interval between Visit to The length of time between visits, Represents the number of visits counted, indicating the total number of visits included in the calculation process; The access frequency fluctuation calculation submodule calculates the range of access frequency based on cache access feature data. First, the access number sequence and the corresponding access time interval in the cache access record are extracted, and the access number data are arranged in chronological order to count the access number change amplitude within a unit time. For the setting of the unit time window, a fixed time length can be used, such as 10 seconds, 30 seconds or 1 minute, and the maximum and minimum values of each access number within the time window are counted, and the difference is calculated to obtain the access number fluctuation amplitude. For example, assuming that the access number data of a cache address is as follows: Table 2 Visit frequency fluctuation statistics ; As shown in Table 2, within the time window of 0-40 seconds, the maximum number of visits is 8 and the minimum number of visits is 3, so the range of the number of visits is , further calculate the access fluctuation mean, using the formula: ; in, Representative The number of visits, Representative The time interval between visits, Represents the number of visits counted, brought into the data calculation: ; ; ; Calculate the access frequency fluctuation range After calculating the access frequency fluctuation range, the value is compared with the fluctuation data in different time windows. If the fluctuation trend shows an increasing or sudden increase, it indicates that the cache access has a drastic fluctuation characteristic. If the fluctuation tends to be stable, the access change is stable. Finally, combined with the fluctuation trend, the access frequency fluctuation characteristics are obtained.
[0024] The hit rate drop evaluation submodule calculates the change range of cache hit rate based on the access frequency fluctuation characteristics, calls the access record to count the number of hits, calculates the current hit rate and compares it with the historical data, and calculates the proportion of the hit rate drop. According to the trend of the hit rate drop, the stability of the cache resource is judged, and the drop range of the hit rate is obtained; The hit rate drop evaluation submodule calculates the change range of the cache hit rate based on the access frequency fluctuation characteristics. First, the historical hit count data is extracted from the cache access record, and the hit rate in the current time window is calculated. The statistical formula is as follows ; in, is the number of hits, is the total number of visits. For example, assuming the number of visits in the current time window is , hit count , the current hit rate is calculated as follows: ; Then, compare with historical data. Assuming that the hit rate in the previous time window is 85%, the drop in hit rate is calculated as follows: ; Further statistics on the percentage of hit rate decrease are calculated as follows: ; If the decrease rate exceeds 20%, it is determined that there is a problem with the stability of the cache resource. If the decrease rate is less than 10%, the cache resource is determined to be stable. In this example, the decrease rate is 17.65%, which is between 10% and 20%, and is determined to be a slight fluctuation state. Finally, the decrease in hit rate is obtained.
[0025] The data fluctuation continuity judgment submodule calculates the fluctuation of access data in a continuous period based on the drop in hit rate, analyzes the stability of changes in access frequency, calls access data in multiple time periods, calculates the duration of fluctuations, determines the load status according to the duration of burst access, and obtains the burst load status result.
[0026] The data fluctuation continuity judgment submodule calculates the fluctuation of access data in continuous time periods based on the drop in hit rate. First, the access data in multiple time periods are called to calculate the duration of fluctuations, set time windows, and analyze the access frequency change trends in multiple time windows. For example, assuming that the access frequencies in the past five time windows are [50, 40, 55, 35, 60], to calculate the duration of fluctuations, first calculate the change range of the access frequency in each time window: ; The mean of the fluctuation range is calculated as follows: ; If the fluctuation mean is greater than the set threshold (such as 15), the fluctuation is determined to be continuous. Next, the load state is determined based on the duration of the burst access. The duration threshold of the burst access is set. For example, if the burst access lasts for more than 3 time windows and the fluctuation mean is greater than 15, it is determined to be a high load state. In this example, the fluctuation amplitude exceeds 15 in three consecutive windows, so it is determined to be a burst load state, and finally the burst load state result is obtained.
[0027] See also Figure 4 , the frequency fluctuation analysis module includes: The operating frequency monitoring submodule obtains the CPU / GPU operating frequency, continuously collects frequency data according to the time series, calls the system monitoring to record the frequency value at each moment, counts the frequency change trend at adjacent moments, calculates the difference between the maximum and minimum frequency values, and obtains the CPU / GPU operating frequency data; The operating frequency monitoring submodule obtains the CPU / GPU operating frequency and continuously collects frequency data according to the time series. First, the system monitoring module is called to collect the current CPU / GPU operating frequency at fixed time intervals (such as 100ms), and the timestamp and the corresponding frequency value are recorded and stored in the cache monitoring log. Subsequently, the collected time series data is structured and stored, as shown in Table 3. The frequency data is arranged in time sequence and the frequency change trend of adjacent moments is calculated. The specific steps include obtaining the frequency values of two adjacent time points and calculating their changes. A threshold is set to judge the change trend. For example, the threshold is set to 50MHz. If the frequency increases by more than 50MHz, it is judged as an upward trend. If it decreases by more than 50MHz, it is judged as a downward trend. If the change is within the threshold, it is judged as stable. As shown in Table 3, the difference between the maximum and minimum frequency values is counted to obtain the CPU / GPU operating frequency data.
[0028] Table 3 CPU / GPU operating frequency data record table ; As shown in Table 3, the maximum CPU frequency value is 3300MHz, and the minimum CPU frequency value is 3100MHz, so the frequency variation range is MHz, similarly, the GPU frequency range is MHz, obtain the CPU / GPU operating frequency data through calculation.
[0029] The fluctuation rate calculation submodule calculates the frequency change rate at adjacent moments based on the CPU / GPU operating frequency data, counts the rate fluctuation amplitude at all moments, and analyzes the fluctuation trend using the formula: ; Obtain the CPU / GPU frequency fluctuation rate through calculation, and combine it with the rate fluctuation trend to obtain the CPU / GPU fluctuation rate characteristics; in, Represents the CPU / GPU frequency fluctuation rate, represent The CPU / GPU frequency value at a time point, represent The cumulative rate value at each time point is Represents the number of statistical time points; The fluctuation rate calculation submodule calculates the frequency change rate of adjacent moments based on the CPU / GPU running frequency data. First, the frequency data of adjacent moments are extracted from Table 3, the frequency change amount of two adjacent time points is calculated, and then divided by the time interval to obtain the frequency change rate. For example, set the time interval ms, the CPU frequency change rate is calculated as follows: ; Also calculate the GPU frequency change rate: ; ; Then, use the formula: ; in, For the The CPU / GPU frequency value at a time point, For the The cumulative rate value at each time point is To calculate the number of time points to be counted, bring in the CPU data: ; ; ; ; Calculate the CPU fluctuation rate MHz and GPU are calculated in the same way, and finally combined with the rate fluctuation trend, the CPU / GPU fluctuation rate characteristics are obtained.
[0030] The frequency fluctuation status assessment submodule analyzes the rate change trend based on the CPU / GPU fluctuation rate characteristics, calculates the fluctuation amplitude anomaly, compares the rate fluctuation range in each time period, and obtains the CPU / GPU frequency fluctuation status results based on the rate change interval.
[0031] The frequency fluctuation status assessment submodule analyzes the rate change trend based on the CPU / GPU fluctuation rate characteristics. First, the CPU / GPU frequency fluctuation rate in different time periods is obtained, and the difference between the maximum and minimum rates is calculated. Assuming that the maximum rate of the GPU is 0.4MHz / ms and the minimum rate is 0.1MHz / ms, the fluctuation range is calculated as follows: ; Then, set the abnormal fluctuation threshold, for example, set the threshold to 0.35MHz / ms. If the fluctuation range is greater than the threshold, it is judged as abnormal fluctuation, otherwise it is normal fluctuation. In this example, the GPU fluctuation range is 0.3MHz / ms, which is less than the set threshold and is judged as a normal fluctuation state. Subsequently, compare the rate fluctuation range of each time period, calculate the fluctuation mean, and finally combine the rate change interval to obtain the CPU / GPU frequency fluctuation state result.
[0032] See also Figure 5 , the load classification identification module includes: The data difference calculation submodule extracts the time series of the two data based on the burst load status results and the CPU / GPU frequency fluctuation status results, calculates the mean, variance and change rate respectively, obtains the absolute difference between the two and performs normalization, and calculates the difference in the time series using the formula: ; Combined with the change trend, the data difference coefficient is obtained by calculation; in, Represents the data difference coefficient, which is used to measure the numerical difference between the burst load state and the CPU / GPU frequency change. Representative moments The burst load state value indicates that at time The burst load level detected by the system, Representative moments The CPU / GPU frequency value, indicating the time Calculate the actual CPU or GPU frequency the device is running at, Represents the mean of the burst load state, which indicates the average level of all burst load state values within the statistical time series range. Represents the mean of CPU / GPU frequency, which indicates the average level of all CPU / GPU frequency values within the statistical time series range. Represents the total length of the time series, indicating the total number of time points included in the statistics during the calculation process; The data difference calculation submodule extracts the time series of the two data based on the burst load status results and the CPU / GPU frequency fluctuation status results. First, the time series data is obtained from the burst load monitoring record. At the same time, extract the same time series from the CPU / GPU frequency fluctuation monitoring records , align the two according to the timestamp to ensure data synchronization, and then calculate the average of the burst load status and CPU / GPU frequency. The average calculation formula is as follows: ; ; Assume that the time series length , burst load status data , CPU / GPU frequency data , then the mean is calculated as follows: ; ; Then, the variance of the burst load state and CPU / GPU frequency is calculated. The variance calculation formula is as follows: ; ; Bring in data calculation: ; ; ; ; Then, the absolute difference between the burst load state and the CPU / GPU frequency is normalized using the formula: ; Bring in data calculation: ; ; ; Finally, the data difference coefficient is calculated , combined with the change trend for analysis, and calculation to obtain the data difference coefficient.
[0033] The abnormality count submodule calls the data difference coefficient, sets the burst load abnormality threshold, traverses based on the time series data, calculates the difference value for each data point and compares it with the abnormality threshold, records the number of abnormal points exceeding the threshold, calculates the frequency of abnormal occurrence and its proportion in the overall data, and obtains the load abnormality ratio; The abnormal number statistics submodule calls the data difference coefficient and sets the abnormal threshold of the burst load. First, the abnormal threshold of the data difference coefficient is determined, for example, the threshold is set Then, based on the time series data, the data difference value is calculated for each time point. And compare it with the abnormal threshold, record the number of abnormal points exceeding the threshold, and assume that the calculated time series data difference coefficient is , the number of abnormal points exceeding the threshold is calculated as follows: ; Then, calculate the frequency of anomalies and their proportion in the overall data: ; The final load abnormality ratio is 60%.
[0034] The load classification submodule sets the classification rules for burst load, medium load, and steady-state load based on the load anomaly ratio and data difference coefficient, divides the time series data into intervals, calculates the time proportion of differentiated categories, and filters the load types based on the computing task type to obtain the load classification results.
[0035] The load classification submodule sets the classification rules for burst load, medium load, and steady-state load based on the load anomaly ratio and data difference coefficient. First, the classification criteria are defined, as shown in Table 4.
[0036] Table 4 Load classification rules ; Then, based on the calculated abnormal ratio of 60% and the data difference coefficient of 3.15, according to the classification standard in Table 4, the data series is determined to belong to burst load. Then, the proportion of different load types in the time series is calculated. Assuming that the classification results calculated in different time periods are [burst load, medium load, medium load, burst load, burst load], the burst load proportion is calculated: ; Similarly calculate the medium load ratio: ; Finally, the load type is filtered based on the computing task type to obtain the load classification result, that is, the burst load accounts for the highest proportion, and it is determined that the system is in a burst load state.
[0037] See also Figure 6 , the intelligent control strategy module includes: Based on the load classification results, the load category determination submodule calls CPU utilization, GPU utilization, memory occupancy, and data transmission rate, compares multiple parameters with the load type benchmark value, selects the corresponding threshold interval, determines the load category, and obtains the load category determination result; The load category determination submodule calls CPU utilization, GPU utilization, memory occupancy, and data transmission rate based on the load classification results. First, the CPU utilization, GPU utilization, memory occupancy, and data transmission rate are obtained from the system resource monitoring data and matched with the load classification results. For example, the benchmark values of different load types are set, as shown in Table 5.
[0038] Table 5 Load type reference values ; The current load parameters are extracted from the system monitoring data, such as CPU utilization of 65%, GPU utilization of 50%, memory occupancy of 75%, and data transmission rate of 180MB / s. The data range of Table 5 is matched and it is found that all parameters fall within the range of medium load. Therefore, the current load category is determined to be medium load, and the load category determination result is finally obtained.
[0039] The resource allocation comparison submodule uses the CPU / GPU resource allocation, memory usage, and data transfer rate based on the load category determination results to calculate the matching degree of multiple resources using the formula: ; Obtain resource matching value through calculation, compare with resource allocation benchmark, and obtain resource allocation deviation value; in, Represents the resource matching value, which is used to measure the matching degree between the actual allocated resources and the load required resources. Representative The actual amount of resources allocated to the system in the The specific values of the CPU computing power, GPU computing power, memory usage, or data transfer rate allocated to the resource. Representative The load demand resource of the resource indicates that the system The specific values of CPU computing power, GPU computing power, memory usage or data transfer rate required on the resource. Representative The load base resource of the resource is expressed in A preset benchmark reference value or standard value for a resource category. Represents the number of resource types, indicating the total number of resource types involved in the system, including CPU computing power, GPU computing power, memory usage, data transmission rate and other different categories; The resource allocation comparison submodule calls the CPU / GPU resource allocation, memory usage, and data transfer rate based on the load category determination results. First, it extracts the current resource allocation of the system, including the CPU computing resource allocation. , GPU computing resource allocation , memory usage and data transfer rates , and at the same time, determine the amount of resources corresponding to the load demand , and the amount of load benchmark resources , then use the formula: ; Assume that the current CPU resource allocation is 70%, GPU resource allocation is 55%, memory usage is 75%, and data transfer rate is 190MB / s, and the corresponding load requirements are 65%, 50%, 70%, and 180MB / s, respectively, and the load benchmark resource amounts are 80%, 60%, 85%, and 220MB / s, respectively. Calculate the resource matching degree: ; ; ; ; ; Finally, the resource matching degree is calculated , and then compare the resource allocation benchmark, set the matching threshold to 5.0, if If the threshold is exceeded, it is determined that there is a deviation in resource allocation. In this example, , is greater than 5.0, it is determined that the current resource allocation deviation is large, and finally the resource allocation deviation value is obtained.
[0040] The dynamic control scheme generation submodule calls the current CPU / GPU load threshold, memory usage, and data transmission rate adjustment range based on the resource allocation deviation value, compares the adjustable space, adjusts the resource allocation ratio, and obtains the load dynamic control scheme.
[0041] The dynamic control scheme generation submodule calls the CPU / GPU current load threshold, memory usage and data transmission rate adjustment range based on the resource allocation deviation value. First, set the adjustable resource range, for example, the CPU adjustable range is ±10%, the GPU adjustable range is ±15%, the memory occupancy adjustable range is ±10%, and the data transmission rate adjustable range is ±20%, and then compare the current resource allocation deviation value. It exceeds the threshold of 5.0, so adjustments need to be made. Assuming that the CPU resource usage is 5% higher, the CPU resource allocation is adjusted to 65%, the GPU resource allocation is adjusted to 50%, the memory usage is adjusted to 70%, and the data transmission rate is adjusted to 180MB / s. The final adjusted resource allocation is as follows: Table 6 Adjusted resource allocation plan ; According to Table 6, check whether the adjusted parameters meet the load requirements and calculate the adjusted resource matching degree. : ; ; After adjustment, the matching degree is reduced to , indicating that resource regulation has reached the optimal state, and finally a dynamic load regulation solution is obtained.
[0042] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.
Claims
1. Intelligent computing resource energy-saving scheduling system based on load prediction, characterized by: The system comprises: The cache access monitoring module obtains cache access data, monitors cache access patterns, calculates access frequency changes, evaluates access change ranges, counts hit rates and migration times, and obtains cache access feature data; The burst load determination module calculates the access frequency fluctuation range based on the cache access characteristic data, determines whether the growth rate is abnormal, compares the drop in the hit rate, evaluates the continuity of data fluctuation, and obtains the burst load status result; The frequency fluctuation analysis module obtains the CPU / GPU operating frequency, monitors frequency changes, calculates the fluctuation rate at adjacent moments, determines abnormal fluctuation amplitude, evaluates fluctuation rate changes, and obtains CPU / GPU frequency fluctuation status results; The load classification identification module calculates the difference between the two data based on the burst load status result and the CPU / GPU frequency fluctuation status result, compares the number of data anomalies, classifies them into burst, medium or steady-state loads, calculates the time proportion of multiple categories, and screens the load types in combination with the computing task type to obtain the load classification result; The intelligent control strategy module determines the load category based on the load classification result, compares resource allocation and data transmission, adjusts CPU / GPU resource allocation and memory allocation, and obtains a dynamic load control solution.
2. The intelligent computing resource energy-saving scheduling system based on load prediction according to claim 1 is characterized in that: The cache access feature data includes access frequency changes, access change amplitude, hit rate, and migration times; the burst load status results include access frequency fluctuation range, abnormal growth rate, hit rate decrease amplitude, and data fluctuation continuity; the CPU / GPU frequency fluctuation status results include CPU / GPU operating frequency, frequency change, fluctuation rate, abnormal fluctuation amplitude, and fluctuation rate change; the load classification results include load category, multi-category time proportion, and computing task type to filter load types; the load dynamic control scheme includes CPU / GPU resource allocation adjustment, memory allocation adjustment, and resource allocation and data transmission comparison.
3. The intelligent computing resource energy-saving scheduling system based on load prediction according to claim 2 is characterized in that: The cache access monitoring module includes: The cache access data collection submodule obtains cache access records, collects access time, access address and access interval, extracts access frequency change trends, calculates the incremental interval of multiple accesses in the access time series, sets access density indicators based on the time interval difference, classifies and summarizes the access density indicators, matches the corresponding access patterns based on the classification and summary results, and obtains cache access time series characteristics; The access pattern calculation submodule calculates the change value of multi-address access frequency based on the cache access time series characteristics, counts the access increment in a short period of time, determines the access growth rate and its change trend, and uses the formula: ; Calculate and obtain the cache access frequency change coefficient, and combine the access pattern classification calculation to obtain the cache access pattern characteristics; in, Represents the cache access frequency variation coefficient, which measures the degree to which the cache access frequency changes over time. Representative The address number of the access is used to record the specific address information when the cache is accessed. Representative The address number of the access is used as a comparison benchmark for calculating the changes in adjacent accesses. Representative The timestamp of the visit, which records the specific time when the visit occurred. Representative The timestamp of the visit is used as a comparison benchmark for calculating time changes. Represents the total number of visits, indicating the number of visits within the statistical time window; The hit rate and migration statistics submodule calculates the cache access hit rate based on the cache access pattern characteristics, counts the number of hits according to the number of accesses, obtains the hit ratio, counts the number of cache migrations, calculates the migration ratio, determines the cache load status based on the migration ratio, and obtains cache access characteristic data.
4. The intelligent computing resource energy-saving scheduling system based on load prediction according to claim 3 is characterized in that: The burst load determination module comprises: The access frequency fluctuation calculation submodule calculates the access frequency variation range based on the cache access feature data, counts the variation range of the access times per unit time, calculates the difference between the highest and lowest access times, summarizes the variation range, and calculates the access fluctuation mean using the formula: ; Calculate the access frequency fluctuation range, calculate and obtain the access frequency fluctuation range, and combine the fluctuation trend to obtain the access frequency fluctuation characteristics; in, Represents the access frequency fluctuation range, which is used to measure the fluctuation of access frequency per unit time. Representative The number of visits, indicating the The number of visits counted during visits. Representative The number of visits is used as a comparison benchmark for calculating the changes in the frequency of adjacent visits. Representative The time interval between visits is the time interval between Visit to The length of time between visits, Represents the number of visits counted, indicating the total number of visits included in the calculation process; The hit rate drop evaluation submodule calculates the change range of the cache hit rate based on the access frequency fluctuation characteristics, calls the access record to count the number of hits, calculates the current hit rate and compares it with the historical data, counts the proportion of the hit rate drop, judges the stability of the cache resource according to the hit rate drop trend, and obtains the hit rate drop range; The data fluctuation continuity judgment submodule calculates the fluctuation of access data in a continuous period based on the drop in hit rate, analyzes the stability of changes in access frequency, calls access data in multiple time periods, calculates the duration of fluctuations, determines the load status according to the duration of burst access, and obtains the burst load status result.
5. The intelligent computing resource energy-saving scheduling system based on load prediction according to claim 4 is characterized in that: The frequency fluctuation analysis module comprises: The operating frequency monitoring submodule obtains the CPU / GPU operating frequency, continuously collects frequency data according to the time series, calls the system monitoring module to record the frequency value at each moment, counts the frequency change trend at adjacent moments, calculates the difference between the maximum and minimum frequency values, and obtains the CPU / GPU operating frequency data; The fluctuation rate calculation submodule calculates the frequency change rate at adjacent moments based on the CPU / GPU operating frequency data, counts the rate fluctuation amplitude at all moments, and analyzes the fluctuation trend using the formula: ; Obtain CPU / GPU frequency fluctuation rate through calculation, and obtain CPU / GPU fluctuation rate characteristics based on rate fluctuation trend; in, Represents the CPU / GPU frequency fluctuation rate, represent The CPU / GPU frequency value at a time point, represent The cumulative rate value at each time point is Represents the number of statistical time points; The frequency fluctuation status assessment submodule analyzes the rate change trend based on the CPU / GPU fluctuation rate characteristics, calculates the fluctuation amplitude abnormal value, compares the rate fluctuation range of each time period, and obtains the CPU / GPU frequency fluctuation status result in combination with the rate change interval.
6. The intelligent computing resource energy-saving scheduling system based on load prediction according to claim 5 is characterized in that: The load classification identification module includes: The data difference calculation submodule extracts the time series of the two data based on the burst load status results and the CPU / GPU frequency fluctuation status results, calculates the mean, variance and change rate respectively, obtains the absolute difference between the two and performs normalization, and calculates the difference in the time series using the formula: ; Combined with the change trend, the data difference coefficient is obtained by calculation; in, Represents the data difference coefficient, which is used to measure the numerical difference between the burst load state and the CPU / GPU frequency change. Representative moments The burst load state value indicates that at time The burst load level detected by the system, Representative moments The CPU / GPU frequency value, which means at time Calculate the actual CPU or GPU frequency the device is running at, Represents the mean of the burst load state, which indicates the average level of all burst load state values within the statistical time series range. Represents the mean of CPU / GPU frequency, which indicates the average level of all CPU / GPU frequency values within the statistical time series range. Represents the total length of the time series, indicating the total number of time points included in the statistics during the calculation process; The abnormality times statistics submodule calls the data difference coefficient, sets the sudden load abnormality threshold, performs traversal based on the time series data, calculates the difference value for each data point and compares it with the abnormality threshold, records the number of abnormal points exceeding the threshold, calculates the frequency of abnormal occurrence and its proportion in the overall data, and obtains the load abnormality times ratio; The load classification submodule sets the classification rules for burst load, medium load, and steady-state load based on the load anomaly ratio and the data difference coefficient, divides the time series data into intervals, calculates the time proportion of differentiated categories, and filters the load types in combination with the computing task type to obtain the load classification results.
7. The intelligent computing resource energy-saving scheduling system based on load prediction according to claim 6 is characterized in that: The intelligent control strategy module includes: The load category determination submodule calls CPU utilization, GPU utilization, memory occupancy, and data transmission rate based on the load classification result, compares multiple parameters with the load type benchmark value, screens the corresponding threshold interval, determines the load category, and obtains the load category determination result; The resource allocation comparison submodule uses the CPU / GPU resource allocation, memory usage, and data transmission rate based on the load category determination result to calculate multiple resource matching degrees using the formula: ; Obtain resource matching value through calculation, compare with resource allocation benchmark, and obtain resource allocation deviation value; in, Represents the resource matching value, which is used to measure the matching degree between the actual allocated resources and the load required resources. Representative The actual amount of resources allocated to the system in the The specific values of the CPU computing power, GPU computing power, memory usage, or data transfer rate allocated to the resource. Representative The load demand resource of the resource indicates that the system The specific values of CPU computing power, GPU computing power, memory usage or data transfer rate required on the resource. Representative The load base resource of the resource is expressed in A preset benchmark reference value or standard value for a resource category. Represents the number of resource types, indicating the total number of resource types involved in the system, including CPU computing power, GPU computing power, memory usage, data transmission rate and other different categories; The dynamic control scheme generation submodule calls the current CPU / GPU load threshold, memory usage and data transmission rate adjustment range based on the resource allocation deviation value, compares the adjustable space, adjusts the resource allocation ratio, and obtains a load dynamic control scheme.
Citation Information
Patent Citations
Operation frequency adjustment method and device of processor and storage medium
CN114442792A
Resource scheduling method and device, readable storage medium and chip system
CN117891617A
Node computing resource allocation method, network switching subsystem and intelligent computing platform
CN117997906A
Frequency adjustment method, device, equipment, program medium and chip
CN119396270A
Server resource scheduling system integrating AI and edge computing
CN119718682A
Cited By
Power supply state monitoring method and system of protocol signal processing module
CN120566711A
Power status monitoring method and system for protocol signal processing module
CN120566711B
GPU (Graphics Processing Unit) server resource dynamic allocation management method and device, equipment and medium
CN120723472A
Memory access method and device, storage medium and program product
CN120803976A
Rack-mounted server system and method
CN120812347A