Edge calculation scheduling optimization method based on actor-commentator model
The Actor-Critic model in edge computing dynamically adjusts feature priorities and learning rates to address resource fluctuations, enhancing load prediction and scheduling efficiency in edge environments.
Patent Information
- Application Number
- CN202510804313.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The prior art lacks a continuity feature extraction mechanism for changes in node resource states in edge computing, resulting in delay in load trend identification and insufficient accuracy of scheduling strategies, especially in scenarios with high concurrency and rapid load change.
Through the edge computing scheduling optimization method based on the actor-criticist model, changes in node CPU utilization, memory occupancy and bandwidth occupancy are recorded synchronously, standard deviation and change amplitude are calculated, feature sensitivity prediction results are generated, instantaneous resource utilization trends are constructed, learning rate amplitude is dynamically adjusted, and scheduling strategies are optimized.
Improve the accuracy of load change prediction, enhance scheduling response capabilities, improve resource utilization efficiency, and reduce policy update hysteresis and service response delay.
Smart Images

Figure CN120315902A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of edge computing, and particularly to an edge computing scheduling optimization method based on an actor-critic model. Background Art
[0002] The technical field of edge computing includes a technical system that migrates computing tasks and data processing functions from a centralized data center to network edge nodes. The core content of this technical field lies in reducing data transmission latency, improving system response speed and service quality, and effectively alleviating the load pressure on the central node by deploying distributed computing resources. The edge computing architecture usually includes three levels: terminal devices, edge nodes, and cloud centers. By processing information close to the data source, optimizing resource allocation and energy consumption management, it is widely applied in intelligent manufacturing, vehicle networking, smart cities, the Internet of Things, and real-time interaction application scenarios. The overall technical field systematically covers directions such as computing resource management of edge nodes, task distribution and scheduling strategies, data synchronization and consistency guarantee, edge intelligent model deployment and inference optimization.
[0003] Among them, the edge computing scheduling optimization method based on the actor-critic model refers to a technical solution that, in an edge computing environment, for the problem of computing resource allocation and task scheduling optimization, uses the actor-critic algorithm structure in the field of reinforcement learning to achieve parallel update of decision-making and evaluation. The technical matters targeted by this patent theme cover the formulation of optimal scheduling decisions under the condition of limited resources of edge nodes in the case of multi-task concurrency. By establishing an environmental state representation, designing an action generation mechanism based on policy gradients, and a performance evaluation module based on a value function, the resource scheduling strategy of edge computing nodes is optimized in a joint training manner. Specifically, the state observation is input into the neural network, and the probability distribution of the scheduling action is output. Combining the design of the reward function to guide the synchronous update of the policy and the value function, thereby completing the optimization process of task scheduling in the edge environment.
[0004] Existing technologies rely on static value analysis in the perception of node resource status changes, lacking a mechanism for extracting the continuity characteristics of the resource fluctuation process, resulting in delayed feature recognition during the node resource fluctuation process, affecting the precise formulation of scheduling strategies. Load trend extraction is usually based on a single point in time load value, which cannot effectively reflect the directionality and continuity of load evolution, resulting in insufficient accuracy in predicting resource utilization change trends. Load status interval division relies on fixed threshold judgment, lacks a dynamic adjustment process for the change rate, and is difficult to adapt to scenarios where node load fluctuates violently, which can easily lead to slow adjustment or overreaction. The action selection method mainly relies on static priority setting, fails to dynamically select the optimal action based on the real-time load status, and has insufficient scheduling response capabilities. A unified step size adjustment is usually used in local policy updates, and the learning rate amplitude change cannot be refined according to the local gradient change trend, which reduces the adaptability of the policy in a load fluctuation environment, and is prone to cause policy update hysteresis or failure problems. In high concurrency and rapid load change scenarios, it is easy to cause a decrease in node resource utilization efficiency and an increase in service response delay. Summary of the invention
[0005] The purpose of the present invention is to solve the shortcomings existing in the prior art and propose an edge computing scheduling optimization method based on the actor-critic model.
[0006] In order to achieve the above object, the present invention adopts the following technical solution: an edge computing scheduling optimization method based on an actor-critic model, comprising the following steps: S1: Synchronously record the changes in node CPU utilization, memory occupancy, and bandwidth occupancy, calculate the standard deviation and change range, calculate the weighted ratio based on the response delay and overflow rate changes, adjust the feature priority according to the sensitivity threshold, and generate the feature sensitivity prediction result; S2: Based on the feature sensitivity prediction results, call the real-time CPU and memory occupancy rates, set the weighted average of resource weights, extract the load fluctuation trend, construct the instantaneous resource curve and generate the instantaneous resource utilization trend; S3: collecting time slice load values according to the instantaneous resource utilization trend, extracting the change rate data set, dividing the load state interval, adjusting and comparing with the node basic load threshold, and generating a dynamic load adjustment benchmark; S4: judging the instantaneous load interval according to the dynamic load adjustment benchmark, screening the sensitivity response action set, screening the high-latency actions based on the delay label, and forming a scheduling action screening list; S5: Extract the policy parameter update batch according to the scheduling action screening list, obtain the local gradient change, perform normalized difference processing, dynamically adjust the learning rate amplitude based on the difference mean and the threshold, and generate a scheduling optimization plan.
[0007] As a further solution of the present invention, the feature sensitivity prediction results include the change range of CPU utilization rate, the change range of memory occupancy rate, the change range of bandwidth occupancy rate, the change amount of task response latency, the change amount of overflow rate, and the feature priority adjustment coefficient. The instantaneous resource utilization trend includes the load fluctuation trend, the change of resource utilization direction, and the instantaneous resource utilization curve. The dynamic load regulation benchmark includes the load change rate distribution, the division of load state intervals, and the basis for adjusting the basic load threshold. The scheduling action screening list includes the sensitivity response action set, the action execution delay statistics, and the action execution priority ranking. The scheduling optimization plan includes the policy parameter update batch, the local gradient change vector, the normalized difference set mean, and the learning rate adjustment range.
[0008] As a further solution of the present invention, the specific steps of S1 are as follows: S101: Obtain the node CPU utilization rate, node memory occupancy rate, and node bandwidth occupancy rate. After continuous sampling, based on the cycle parameter value, calculate the dynamic weight change rate, and call the ratio of the mean square deviation of the change rate to the number of nodes to generate the node parameter change standard deviation. S102: Based on the node parameter change standard deviation, normalize the change rates of CPU utilization rate, memory occupancy rate, and bandwidth occupancy rate. Call the sum of the product of the delay change amount and the overflow rate change amount and the normalization amplitude respectively, and obtain the node parameter sensitivity result based on the node number ratio. S103: According to the node parameter sensitivity result, call the node feature sensitivity threshold, compare the sensitivity coefficient with the threshold to adjust the node feature priority status, establish the mapping relationship between the standard deviation change range and the feature priority, and obtain the feature sensitivity prediction result.
[0009] As a further solution of the present invention, the specific formula for the dynamic weight change rate is as follows: ; Wherein, represents the dynamic weight change rate of node at time , represents the sampling value of the CPU utilization rate, memory occupancy rate, or bandwidth occupancy rate of node at time , represents the time interval of continuous sampling, represents the total time length of the cycle parameter, represents the number of nodes, represents node 's effective sampling interval within the cycle.
[0010] As a further solution of the present invention, the specific steps of S2 are as follows: S201: Based on the feature sensitivity prediction result, call the real-time CPU utilization rate and memory occupancy rate of the node, set the node resource weight coefficient, perform weighted averaging on the CPU utilization rate and memory occupancy rate respectively, calculate the weighted weight coefficient of the node resource utilization rate, and generate the node load fluctuation trend; S202: Based on the node load fluctuation trend, call the change amount of the node resource utilization rate per unit time, compare the upper and lower bounds according to the set load change interval reference value, judge the resource utilization direction, extract the positive and negative change signs and establish a time series to obtain the node resource direction change amount; S203: According to the node load fluctuation trend and the node resource direction change amount, call the node time series order, integrate the change amplitude and direction information, generate instantaneous resource data points node by node, draw a resource change curve based on the time series, and obtain the instantaneous resource utilization trend.
[0011] As a further solution of the present invention, the calculation formula of the weighted weight coefficient of the node resource utilization rate is specifically: ; Wherein, represents the weighted weight coefficient of the node resource utilization rate, represents the weighted coefficient of the node CPU utilization rate, represents the weighted coefficient of the node memory occupancy rate, represents the value of the real-time CPU utilization rate of the node at the current moment, represents the value of the real-time memory occupancy rate of the node at the current moment, represents the total number of nodes currently participating in the statistics.
[0012] As a further solution of the present invention, the specific steps of S3 are as follows, S301: According to the instantaneous resource utilization trend, collect the time slice load values of the node within a time period, arrange the time slice load values, extract the load change amount per unit time based on the difference between adjacent time slices, integrate the change amount set, and generate a node load change rate data set; S302: Based on the node load change rate data set, set a load status interval according to the change rate range, call the load change rate data to compare with the interval boundary, classify the node load status according to the result, and obtain a node load status interval set; S303: According to the node load status interval set, call the node basic load threshold reference, calculate the load deviation value according to the comparison between the node load status and the reference, integrate the node deviation data to generate an adjustment set, and obtain a dynamic load adjustment reference.
[0013] As a further solution of the present invention, the calculation formula of the load deviation value is specifically ; wherein, represents the node load deviation value, represents the actual node load data, represents the node basic load threshold data, represents the average of the basic load thresholds of all nodes, represents all the sum of the absolute values of the differences between the actual loads and the reference loads of all nodes.
[0014] As a further solution of the present invention, the specific steps of S4 are as follows. S401: Invoke the dynamic load adjustment reference, compare the node instantaneous load value with the upper and lower bounds of the adjustment reference interval, determine the node load position, extract the node interval flag information, integrate the node position flag data set, and generate the node instantaneous load position marking value; S402: Based on the node instantaneous load position marking value, filter the actions corresponding to the sensitivity response flags in the action pool, call the action delay tag to accumulate the delay times, filter the set of actions with the top-ranked delay times, and integrate the sorting information to obtain the preferred set of sensitive actions; S403: According to the preferred set of sensitive actions, call the original execution priority of the actions, re-adjust the priority order according to the delay time ranking, assign action priority weights, integrate the node action priority change information, and obtain the scheduling action screening list.
[0015] As a further solution of the present invention, the specific steps of S5 are as follows. S501: According to the scheduling action screening list, extract the local policy parameter update batches in the feedback data, arrange the batch node policy parameter change values, extract the local gradient change amount, integrate the node change data, and generate the local gradient change amount vector set; S502: Based on the local gradient change amount vector set, extract the batch change vector data, perform the normalized difference operation between the local gradient change amount vector and the batch change vector, integrate the node normalized difference values, form the node normalized difference sequence, and obtain the normalized difference set; S503: According to the normalized difference set, compare the normalized difference mean value with the upper and lower threshold values, determine the deviation direction of the mean value, adjust the corresponding node learning rate amplitude, integrate the node learning rate change information, and obtain the scheduling optimization plan.
[0016] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In the present invention, the feature priority is adjusted by weighted delay and the change amount of the overflow rate, which enhances bottleneck recognition, synchronously extracts the load trend and resource direction, constructs an instantaneous resource curve, improves the accuracy of change prediction, divides the state according to the rate and adjusts the threshold, strengthens the load regulation, screens actions and rearranges the priority, improves the scheduling response, and dynamically adjusts the learning rate by normalizing the gradient difference, thereby accelerating the policy convergence. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 It is a schematic diagram of the step flow of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The following will describe the technical solutions in the present invention with reference to the drawings.
[0020] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "example" in the present invention should not be construed as more preferred or more advantageous than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two can be selected.
[0021] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.
[0022] In the embodiments of the present invention, sometimes the subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.
[0023] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the drawings and specific embodiments.
[0024] Please refer to Figure 1 , the edge computing scheduling optimization method based on the actor-critic model includes the following steps: S1: Obtain the CPU utilization rate of the node, the memory occupancy rate of the node, and the bandwidth occupancy rate of the node. Record the parameter values in two consecutive sampling periods through time synchronization. Calculate the standard deviation based on the rate of change of the values and extract the change amplitude. After normalizing the change amplitude of each standard deviation, calculate the sensitivity coefficient. Call the change amount of the node task response delay and the change amount of the overflow rate to calculate the weighted ratio with the change amplitude of the standard deviation. According to the set sensitivity threshold, adjust the node feature priority to generate the feature sensitivity prediction result; S2: Based on the feature sensitivity prediction result, call the real-time CPU utilization rate and memory occupancy rate of the node, set the node resource weight coefficient for weighted average, extract and generate the load fluctuation trend per unit time. Determine the node resource utilization direction according to the load change interval, integrate the load trend and direction change information to construct the node instantaneous resource curve, and generate the instantaneous resource utilization trend; S3: According to the instantaneous resource utilization trend, collect the time slice load values within the node time period, extract the load change rate data set, divide the load state interval according to the change rate set, and adjust and compare with the node basic load threshold benchmark for different state intervals to generate the dynamic load adjustment benchmark; S4: Call the dynamic load adjustment benchmark, determine the position of the node instantaneous load value within the adjustment benchmark interval, screen the action set with sensitivity response in the action pool, select the action with more delay times based on the action execution delay label, and reassign the action execution priority to generate the scheduling action screening list; S5: According to the scheduling action screening list, extract the local policy parameter update batch in the action execution feedback data, obtain the vector of the local gradient change amount within the batch, perform the normalized difference operation between the gradient change vector and the batch change vector, compare and judge according to the mean value and upper and lower limit thresholds of the normalized difference set, and adjust the learning rate amplitude to generate the scheduling optimization plan.
[0025] The feature sensitivity prediction result includes the change amplitude of CPU utilization rate, the change amplitude of memory occupancy rate, the change amplitude of bandwidth occupancy rate, the change amount of task response delay, the change amount of overflow rate, and the feature priority adjustment coefficient. The instantaneous resource utilization trend includes the load fluctuation trend, the change of resource utilization direction, and the instantaneous resource utilization curve. The dynamic load adjustment benchmark includes the load change rate distribution, the division of load state intervals, and the basis for adjusting the basic load threshold. The scheduling action screening list includes the action set with sensitivity response, the statistics of action execution delay, and the sorting of action execution priority. The scheduling optimization plan includes the local policy parameter update batch, the local gradient change vector, the mean value of the normalized difference set, and the adjustment amplitude of the learning rate.
[0026] The specific steps of S1 are as follows: S101: Obtain the CPU utilization rate of the node, the memory occupancy rate of the node, and the bandwidth occupancy rate of the node. After continuous sampling, take values based on the cycle parameters, calculate the dynamic weight change rate, and call the ratio of the mean of the squared deviations of the change rate to the number of nodes to generate the standard deviation of the node parameter change; The specific calculation formula for the dynamic weight change rate is: ; Among them, represents the dynamic weight change rate of node at time , represents the CPU utilization rate, memory occupancy rate, or bandwidth occupancy rate sampling value of node at time , represents the time interval of continuous sampling, represents the total time length of the cycle parameter, represents the number of nodes, represents node 's effective sampling interval within the cycle; For node at time , its CPU utilization rate is collected in real time through the system monitoring tool, and the sampling period is 5 seconds. The time interval seconds, is the CPU utilization rate monitoring value of node at seconds. For example, , ; The cycle parameter seconds, the number of nodes , and the effective sampling interval of each node is obtained by statistical analysis of the system log. For example, seconds, seconds, seconds, and the total sum seconds; Substituting into the formula gives: ; Basis for parameter setting: seconds is defined by the default sampling frequency of the monitoring system, is the total time of the actual effective sampling times of all nodes within the cycle, seconds is set by the cycle parameter configured in the system; and are obtained in real time through the performance monitoring interface, with the unit being percentage; The calculation result represents the weighted change rate of node , and its value is used as a component of the node change rate set for subsequent calculation of the mean of the squared deviations; Explanation of numerical results: Reflecting node At time The dynamic change amplitude relative to the previous time point, which combines the difference between the current and historical values, the normalized adjustment of the historical value, and the dynamic weights of all node sampling intervals within the period. This result is an element of the node change rate set and participates in the subsequent standard deviation calculation; Quantification of non - numerical data: CPU utilization is directly obtained through the system performance counter without additional quantification; effective sampling interval Calculated by the time difference between two consecutive valid sampling timestamps recorded in the system log. For example, timestamp seconds, seconds, then seconds; Verification of reasonable parameter range: CPU utilization The realistic reasonable range is from 0% to 100%, and the time interval In the monitoring system, it is usually from 1 to 60 seconds, and the sum Does not exceed (3600 seconds × 3 = 10800 seconds), and the actual value of 14 seconds meets the constraints; Formula derivation process: First calculate the absolute difference of the numerator , denominator , to obtain the basic change rate ; Then calculate the dynamic weight term , and the final weighted change rate is ; This result indicates that the change rate of node Is jointly affected by its own historical value and the global sampling interval distribution within the period. Its value is used as an input item for the node change rate set to calculate the ratio of the mean of squared deviations to the number of nodes, and finally generate the standard deviation of node parameter changes.
[0027] S102: Based on the standard deviation of node parameter changes, normalize the change rates of CPU utilization, memory occupancy, and bandwidth occupancy. Multiply and accumulate the call latency change amount and the overflow rate change amount with the normalization amplitude respectively, and obtain the node parameter sensitivity result based on the node number ratio; Based on the standard deviation of node parameter changes, the change rates of CPU utilization, memory occupancy, and bandwidth occupancy for each node are normalized respectively. Subtract the mean change rate of all nodes from the change rate value obtained from each sampling, and then normalize it with the standard deviation of node parameter changes. Suppose the CPU change rate of a certain node in a sampling is +2%, the mean CPU change rate is +1.5%, and the CPU standard deviation is 0.2%, then the normalization result for this time is (2% - 1.5%) / 0.2% = 2.5. Complete the normalization process for each change rate according to this method. After completion, read the change amount of node latency and the change amount of overflow rate. For example, the change amount of node latency is +10ms, and the change amount of overflow rate is +0.5%. Multiply them by the corresponding normalization amplitude respectively. For example, the change amount of latency is multiplied by the normalized CPU change of 2.5, and the result is 25. The change amount of overflow rate is multiplied by the normalized bandwidth change of 2.0, and the result is 1. Sum up the latency products and overflow rate products of each node. For example, the total latency of 10 nodes is 250, and the total overflow rate is 10. Then divide them by the number of nodes 10 respectively to obtain the latency sensitivity of 25 and the overflow rate sensitivity of 1 as the node parameter sensitivity results. Among them, the setting of the standard deviation of node parameter changes refers to the stability of node historical data. Usually, stable nodes are set with a change less than 1%, and nodes with obvious fluctuations are set with a change above 2% as the boundary. When the change rate is abnormal during normalization, if the normalization amplitude exceeds ±3, it is recognized as a high-fluctuation node.
[0028] S103: According to the node parameter sensitivity results, call the node feature sensitivity threshold, compare the sensitivity coefficient with the threshold to adjust the priority status of node features, establish the mapping relationship between the standard deviation change amplitude and the feature priority, and obtain the feature sensitivity prediction result; According to the node parameter sensitivity results, set the node feature sensitivity threshold. The sensitivity threshold for the delay change amount is set to ±20, and the sensitivity threshold for the overflow rate change amount is set to ±2. The threshold setting is based on the node service level. For example, the threshold for important service nodes is set lower to ±15, the threshold for ordinary nodes is set to ±20, and the threshold for test nodes is set to ±25. The specific setting is determined based on the node's past operation data and the system service quality requirements. For example, for important nodes, a delay change within 5 ms is considered normal, and a change above 10 ms is regarded as sensitive and abnormal. Therefore, the sensitivity threshold is set to ±15. Ordinary nodes are allowed a slightly larger change, so the threshold is set to ±20. When the node sensitivity exceeds the positive and negative thresholds, adjust the node feature priority status. If the delay sensitivity of 25 exceeds the ordinary node threshold of 20, the node priority is increased by one level. For example, if the original priority is 2, it is increased to 1. At the same time, record the change range of the node standard deviation. For example, if the standard deviation of a certain node is 0.15%, file it according to the mapping relationship: σ < 0.2% corresponds to priority 1, σ ∈ [0.2%, 0.5%) corresponds to priority 2, σ ≥ 0.5% corresponds to priority 3. For the node with σ = 0.15%, the final priority is set to 1. Finally, based on the node feature sensitivity and the change range of the standard deviation, output the node feature sensitivity prediction result, indicating that the sensitivity of the node will maintain a high-level fluctuation state in the future cycle.
[0029] The specific steps of S2 are as follows: S201: Based on the feature sensitivity prediction result, call the real-time CPU utilization rate and memory occupancy rate of the node, set the node resource weight coefficient, perform weighted averaging on the CPU utilization rate and memory occupancy rate respectively, calculate the weighted coefficient of the node resource utilization rate, and generate the node load fluctuation trend. The specific calculation formula for the weighted coefficient of the node resource utilization rate is as follows: ; Among them, represents the weighted coefficient of the node resource utilization rate, represents the weighted coefficient of the node CPU utilization rate, represents the weighted coefficient of the node memory occupancy rate, represents the value of the real-time CPU utilization rate of the node at the current moment, represents the value of the real-time memory occupancy rate of the node at the current moment, represents the total number of nodes currently participating in the statistics; Parameter definition and data source: : The weight coefficient of the CPU utilization rate. This coefficient is obtained based on historical data analysis and is used to balance the impact of the CPU utilization rate on the overall node resource weight. The specific value is set to 0.6, calculated based on the average level data and considering the importance of the CPU to the node performance.
[0030] : Weight coefficient of memory occupancy rate. This coefficient reflects the impact of memory occupancy on node performance, and the value is set to 0.4, which is also obtained based on historical data analysis, considering that the relative impact of memory occupancy on performance is slightly lower than that of CPU.
[0031] : Real-time CPU utilization rate of the node at the current moment. The data is obtained through a real-time monitoring system. For example, the measured value is 70%.
[0032] : Real-time memory occupancy rate of the node at the current moment. It is also obtained through a real-time monitoring system, and the measured value is set to 55%.
[0033] : Total number of nodes currently participating in the statistics. This value is calculated in real time. For example, there are currently 10 nodes running.
[0034] Formula derivation process and example: Calculate the weighted average CPU and memory occupancy rate: ; ; Calculate the sum of squares of absolute values: ; ; Substitute the above values into the formula of W for calculation: ; Explanation of numerical results: This result indicates that the average resource load weight coefficient of the current node group after weighing the utilization of CPU and memory is 14.35. This value reflects the average resource load situation of the nodes at a given moment. A higher value indicates that the resource usage is relatively concentrated, and load balancing operations may be required.
[0035] S202: Based on the node load fluctuation trend, call the change amount of the node's resource utilization rate per unit time, compare the upper and lower bounds according to the set load change interval reference value, judge the resource utilization direction, extract the positive and negative signs of the change, and establish a time series to obtain the node resource direction change amount; Based on the node load fluctuation trend, call the change amount of the node's resource utilization rate per unit time, and compare the upper and lower bounds according to the set benchmark value of the load change interval. The load change benchmark value is set according to the actual business carrying capacity of the node. For example, for ordinary business nodes, the benchmark value is set at ±5%, for important business nodes, the benchmark value is set at ±3%, and for test nodes, the benchmark value is set at ±10%. Node A is an important computing node, and the benchmark value is set at ±3%. Compare each change amount in the change amount set of node A {1.5%, -2.5%, 1.5%, 2%} one by one. If the change amount is greater than +3%, it is marked as a significant positive increase. If the change amount is less than -3%, it is marked as a significant negative decrease. If the change amount is within the [-3%, +3%] interval, it is marked as stable. In actual comparison, it is found that 1.5%, -2.5%, 1.5%, and 2% all fall within the [-3%, +3%] interval and are all marked as stable. At the same time, record the positive and negative signs of the change. If the change amount is positive, it is marked as +1, if it is negative, it is marked as -1, and a stable change is marked as 0. The change flag sequence of node A is {+1, -1, +1, +1}. Subsequently, establish a node time series in the order of sampling time, associate each change flag with the time stamp, and form a node resource direction change amount sequence. For example, the corresponding time stamps are T1, T2, T3, T4 respectively, then {(T1, +1), (T2, -1), (T3, +1), (T4, +1)} is obtained, and the node resource direction change amount is generated.
[0036] S203: According to the node load fluctuation trend and the node resource direction change amount, call the node time series order, integrate the change amplitude and direction information, generate instantaneous resource data points for each node one by one, draw a resource change curve based on the time series, and obtain the instantaneous resource utilization trend; According to the node load fluctuation trend and the node resource direction change amount, read the order of the node sampling time series. For each moment, integrate the change amplitude and change direction information. For example, for node A, the change amount at time T1 is 1.5% in the +1 direction, at T2 it is -2.5% in the -1 direction, at T3 it is 1.5% in the +1 direction, and at T4 it is 2% in the +1 direction. Establish instantaneous resource data points one by one. The content of the data points includes three parts: timestamp, change amplitude, and change direction, such as (T1, 1.5%, +1), (T2, 2.5%, -1), (T3, 1.5%, +1), (T4, 2%, +1). Sort all the instantaneous resource data points in chronological order to form a complete time series set. Based on this time series data, draw a resource change curve. The abscissa is the time axis, and the ordinate is the resource change amplitude. The change direction is distinguished by positive and negative. For example, positive change amounts are drawn in the positive direction, and negative change amounts are drawn in the negative direction. The change curve formed by node A from T1 to T4 rises at T1, falls at T2, rises again at T3, and continues to rise at T4, obtaining the instantaneous resource utilization trend of node A, which is a mainly upward growth trend. Among them, if there are more than three consecutive positive changes, the node trend is determined to be continuously growing; if there are more than three negative changes, it is determined to be continuously declining; otherwise, it is determined to be in a fluctuating state. Node A has three consecutive positive changes, and according to the above rules, its instantaneous resource utilization trend is determined to be continuously growing.
[0037] The specific steps of S3 are as follows: S301: According to the instantaneous resource utilization trend, collect the time slice load values within the node time period, arrange the time slice load values, extract the unit time load change amount based on the difference between adjacent time slices, and integrate the change amount set to generate a node load change rate data set; According to the instantaneous resource utilization trend, first select the corresponding node set. For each node, collect the time slice load values within the node time period at a set time interval. For example, collect the load data every 10 seconds. Record the load values continuously collected by node A within a time period as 50%, 52%, 49%, 53%, 55%, 54%. Then arrange all the collected time slice load values in chronological order to form a time series {T1: 50%, T2: 52%, T3: 49%, T4: 53%, T5: 55%, T6: 54%}. Based on the adjacent time slice load values, extract the differences, that is, subtract the previous time slice load value from the current time slice load value to obtain the load change per unit time. For example, T2 - T1 = 52% - 50% = 2%, T3 - T2 = 49% - 52% = -3%, T4 - T3 = 53% - 49% = 4%, T5 - T4 = 55% - 53% = 2%, T6 - T5 = 54% - 55% = -1%. Organize all the load changes to form a change set of node A {+2%, -3%, +4%, +2%, -1%}. Then integrate all the change data sets to establish a data set of node load change rates, where each change rate data includes a time slice identifier, a change value, and a change direction flag. For example, (T2, 2%, +1), (T3, -3%, -1), (T4, 4%, +1), (T5, 2%, +1), (T6, -1%, -1). Finally, complete the generation of the data set of node load change rates.
[0038] S302: Based on the data set of node load change rates, set the load status intervals according to the change rate range. Call the load change rate data to compare with the interval boundaries, and classify the node load status according to the results to obtain a set of node load status intervals; Based on the node load change rate data set, for each node, a set of change rates is collected, and a load status interval is set. For example, the change rate range is divided as follows: [-5%, -2%) is the decreasing interval, [-2%, +2%] is the stable interval, (2%, 5%] is the increasing interval, greater than 5% is the sharply increasing interval, and less than -5% is the sharply decreasing interval. The interval division is set according to the node service stability requirements. For example, nodes with high real-time requirements use a narrower interval setting, and ordinary nodes use a conventional interval setting. Compare the set of change amounts of node A {+2%, -3%, +4%, +2%, -1%} respectively. For example, +2% falls within the increasing interval (2%, 5%], -3% falls within the decreasing interval [-5%, -2%), +4% falls within the increasing interval (2%, 5%], +2% falls within the increasing interval (2%, 5%], and -1% falls within the stable interval [-2%, +2%]. After comparing each change rate with the interval boundaries, classify the corresponding node load status, and mark them as {increasing, decreasing, increasing, increasing, stable} respectively. Finally, organize the load status interval set of node A {increasing, decreasing, increasing, increasing, stable} for use as a basis for subsequent load adjustment actions.
[0039] S303: According to the node load status interval set, call the node basic load threshold benchmark, calculate the load deviation value based on the comparison between the node load status and the benchmark, integrate the node deviation data to generate an adjustment set, and obtain the dynamic load adjustment benchmark; The specific formula for the load deviation value is ; Among them, represents the node load deviation value, represents the node actual load data, represents the node basic load threshold data, represents the average of all node basic load thresholds, represents all the sum of the absolute values of the differences between the actual loads of all nodes and the reference loads; Parameter definition : The actual load data of node . The monitoring data shows GB.
[0040] : The basic load threshold of node . Obtained from historical data and node load analysis GB.
[0041] : The average of all node basic load thresholds. Calculated to be GB, based on the average of the load thresholds of all nodes in the network.
[0042] : Node The actual load and the base load threshold of the node. This data comes from the monitoring system of nodes in the network.
[0043] : The total number of nodes, set as .
[0044] Specific values and calculation methods of parameters Calculate : ; Calculate and take the square root: ; ; Calculate the adjustment coefficient : ; Sum : Assume that in this network, the statistical mean of the deviation is about GB, and the total number of nodes , then: ; The denominator of the calculation formula: ; Substitute the values into the main formula: ; Result interpretation This result indicates that the load deviation value of node is about , which means that the actual load of node is relatively close to the network average load, but slightly higher. This deviation value is obtained by calculating the difference between node and the average benchmark and considering the load deviation of the entire network. This value is used for further network load adjustment, which can help maintain the balance and efficiency of the network.
[0045] The specific steps of S4 are as follows: S401: Call the dynamic load adjustment benchmark, compare the instantaneous load value of the node with the upper and lower bounds of the adjustment benchmark interval, judge the node load position, extract the node interval flag information, integrate the node position flag data set, and generate the instantaneous load position mark value of the node; Call the dynamic load regulation benchmark. Based on the comparison between the instantaneous load value of the node and the upper and lower bounds of the regulation benchmark interval, first obtain the current instantaneous load value of the node. Assume that the current instantaneous load value of node A is 53%, and the dynamic load regulation benchmark of node A is [50%, 55%]. Therefore, compare the current instantaneous load value with the regulation benchmark interval. The load value of node A is between the upper and lower limits, so it can be determined that its load position is in the normal interval. Extract the load position flag information of node A as "normal interval", and then compare this flag information with other nodes to integrate the load position flag data set of all nodes. For example, the load value of node B is 56%, exceeding the upper limit of 55%, so it is marked as the "high load" interval, and the load value of node C is 48%, lower than the lower limit of 50%, marked as the "low load" interval. Finally, obtain the data set of node load position marker values, such as {node A: normal interval, node B: high load, node C: low load}, providing a basis for subsequent decision-making processes.
[0046] S402: Based on the instantaneous load position marker value of the node, screen the actions corresponding to the sensitivity response flags in the action pool, call the action delay label to accumulate the delay times, screen the set of actions with the top-ranked delay times, and integrate the sorting information to obtain the preferred set of sensitive actions; Based on the instantaneous load position marker value of the node, first read the action pool corresponding to the sensitivity response flag. The action pool contains various action response types, and each action type has a related delay label. The delay label is used to mark the delay situation when the action response is executed. For example, the delay label of an action "adjust load" is 3 times, indicating that the action has been delayed 3 times. The current load position of node A is marked as the "normal interval". Screen the actions in the action pool, and according to the accumulated delay times of the delay label, select the actions with more delay times for preference. For example, the accumulated delay times of action A is 5 times, the accumulated delay times of action B is 2 times, and the accumulated delay times of action C is 8 times. When screening, first select the actions with the top few accumulated delay times. For example, action C is 8 times, action A is 5 times, and action B is 2 times. Integrate the sorting information of the delay times to obtain a preferred set of sensitive actions, such as {action C: 8 times, action A: 5 times, action B: 2 times}. This set indicates that in the current load state, the actions with higher delay times are preferentially processed to ensure that resource regulation can be more rapid and effective.
[0047] S403: According to the preferred set of sensitive actions, call the original execution priority of the actions, re-adjust the priority order based on the delay times ranking, assign action priority weights, integrate the node action priority change information, and obtain the screening list of scheduled actions; According to the preferred set of sensitive actions, first read the original execution priority of each action. The action priority is set according to the impact of the action on the system load regulation. For example, the original execution priority of action A is 2, the original execution priority of action B is 3, and the original execution priority of action C is 1. The priority order is readjusted according to the ranking of the number of delays. Action C has the most delays and ranks first, so its priority is adjusted to 1, action A is second and is adjusted to 2, and action B has the least delays and is adjusted to 3. The adjusted priority is {action C: priority 1, action A: priority 2, action B: priority 3}. Next, assign priority weights to each action, and the weight assignment decreases from high to low priority. For example, the weight of action C is 0.6, the weight of action A is 0.3, and the weight of action B is 0.1. The node action priority change information is integrated to finally generate a scheduling action screening list, such as {action C: weight 0.6, action A: weight 0.3, action B: weight 0.1}. This list will be used by the scheduling system to decide which actions should be executed first, thereby improving the efficiency of load regulation.
[0048] The specific steps of S5 are: S501: Filter the list of scheduling actions, extract the local policy parameter update batch in the returned data, arrange the batch node policy parameter change values, extract the local gradient change, integrate the node change data, and generate a local gradient change vector set; According to the scheduling action screening list, extract the local policy parameter update batches in the backhaul data. First, determine the local policy parameter update records corresponding to each node in the current node set. For example, node A is updated in batches 1, 3, and 5, and node B is updated in batches 2, 4, and 6. When arranging the policy parameter change values of the batch nodes, arrange the policy parameter change amounts of each batch node in chronological order. For example, the change value of node A in batch 1 is +0.05, in batch 3 is +0.02, and in batch 5 is -0.01. The change value of node B in batch 2 is +0.03, in batch 4 is -0.02, and in batch 6 is +0.04. Extract the local gradient change amounts, that is, take the difference in policy change amounts between adjacent two batches. For example, the local gradient change amount of node A from batch 1 to batch 3 is +0.02 - +0.05 = -0.03, and from batch 3 to batch 5 is -0.01 - +0.02 = -0.03. Similarly, for node B, the local gradient change amount from batch 2 to batch 4 is -0.02 - +0.03 = -0.05, and from batch 4 to batch 6 is +0.04 - (-0.02) = +0.06. Integrate all node change data, and form a vector set with the local gradient change amounts extracted from each node. For example, the local gradient change amount vector of node A is {-0.03, -0.03}, and that of node B is {-0.05, +0.06}, to generate a local gradient change amount vector set for subsequent normalization processing.
[0049] S502: Based on the local gradient change amount vector set, extract the batch change vector data, perform the normalization difference operation between the local gradient change amount vector and the batch change vector, integrate the node normalization difference values, form a node normalization difference sequence, and obtain the normalization difference set; Based on the set of local gradient change amount vectors, extract batch change vector data. For example, the batch change vector of node A is {+0.05, +0.02, -0.01}, and that of node B is {+0.03, -0.02, +0.04}. For each node, perform the normalized difference operation between the local gradient change amount vector and the batch change vector. The normalization method is to subtract the mean of the corresponding batch change amount from each local gradient change amount and then divide by the maximum absolute value range. For example, the mean of the batch change of node A is (+0.05 + 0.02 - 0.01) / 3 = +0.02, and the maximum absolute value is 0.05. Therefore, the normalized difference of the first local gradient change amount of node A is (-0.03 - 0.02) / 0.05 = -1, and the normalized difference of the second local gradient change amount is (-0.03 - 0.02) / 0.05 = -1. The same applies to node B. The mean of the batch change is (+0.03 - 0.02 + 0.04) / 3 = +0.0167, and the maximum absolute value is 0.04. The normalized difference of the first local gradient change amount is (-0.05 - 0.0167) / 0.04 = -1.6667, and the second is (+0.06 - 0.0167) / 0.04 = 1.0833. Integrate the normalized difference values of each node to form a normalized difference sequence for each node. For example, the normalized difference sequence of node A is {-1, -1}, and that of node B is {-1.6667, 1.0833}. Finally, summarize to obtain the overall normalized difference set for subsequent decision-making analysis.
[0050] S503: According to the normalized difference set, compare the normalized difference mean with the upper and lower threshold values, determine the deviation direction of the mean, adjust the learning rate amplitude of the corresponding node, integrate the node learning rate change information, and obtain the scheduling optimization plan; According to the normalized difference set, first calculate the mean value of the normalized difference sequence of each node. For example, the normalized difference sequence of node A is {-1, -1}, and the mean value is -1. The normalized difference sequence of node B is {-1.6667, 1.0833}, and the mean value is (-1.6667 + 1.0833) / 2 = -0.2917. Set the upper and lower threshold values of the normalized difference mean. For example, ±0.5. The threshold values are set according to the system scheduling stability standard. If the node deviation is within the interval [-0.5, 0.5], it is considered that the load adjustment is stable; otherwise, it is considered that adjustment is needed. Compare the node mean value with the upper and lower threshold values. The mean value of node A, -1, is lower than the lower limit -0.5, and the mean value of node B, -0.2917, is within the interval. It is determined that the deviation direction of the mean value of node A is too negative, and node B is normal and does not need adjustment. Adjust the learning rate amplitude for node A. Increase the original set value of the learning rate, 0.01, by 20%. After adjustment, the learning rate is 0.012. The learning rate of node B remains unchanged. Integrate the learning rate change information of all nodes to finally obtain the scheduling optimization plan. For example, the learning rate of node A is adjusted to 0.012, and node B remains at 0.01, forming an updated learning rate scheduling list.
[0051] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An edge computing scheduling optimization method based on an actor-critic model, characterized in that It includes the following steps: S1: Synchronously record the changes in the CPU utilization rate, memory occupancy rate, and bandwidth occupancy rate of the node, calculate the standard deviation and the change range, perform weighted ratio calculation in combination with the response delay and the change in the overflow rate, adjust the feature priority according to the sensitivity threshold, and generate the feature sensitivity prediction result; S2: Based on the feature sensitivity prediction result, call the real-time CPU and memory occupancy rates, set the weighted average of the resource weights, extract the load fluctuation trend, construct the instantaneous resource curve and generate the instantaneous resource utilization trend; S3: According to the instantaneous resource utilization trend, collect the load values of the time slices, extract the change rate data set, divide the load state interval, and adjust and compare with reference to the node base load threshold to generate the dynamic load adjustment benchmark; S4: Based on the dynamic load adjustment benchmark, judge the instantaneous load interval, screen the sensitivity response action set, and screen the high-delay actions based on the delay label to form the scheduling action screening list; S5: According to the scheduling action screening list, extract the batch of policy parameter updates, obtain the local gradient change amount, perform normalized difference processing, compare according to the difference mean and the threshold, dynamically adjust the learning rate amplitude, and generate the scheduling optimization plan.
2. The edge computing scheduling optimization method based on the actor-critic model according to claim 1, wherein The feature sensitivity prediction result includes the change range of CPU utilization rate, the change range of memory occupancy rate, the change range of bandwidth occupancy rate, the change amount of task response delay, the change amount of overflow rate, and the feature priority adjustment coefficient. The instantaneous resource utilization trend includes the load fluctuation trend, the change in resource utilization direction, and the instantaneous resource utilization curve. The dynamic load adjustment benchmark includes the load change rate distribution, the division of the load state interval, and the basis for adjusting the base load threshold of the node. The scheduling action screening list includes the sensitivity response action set, the action execution delay statistics, and the action execution priority ranking. The scheduling optimization plan includes the batch of policy parameter updates, the local gradient change vector, the normalized difference set mean, and the learning rate adjustment amplitude.
3. The edge computing scheduling optimization method based on the actor-critic model according to claim 1, characterized in that The specific steps of S1 are as follows S101: Obtain the CPU utilization rate of the node, the memory occupancy rate of the node, and the bandwidth occupancy rate of the node. After continuous sampling, based on the periodic parameter values, calculate the dynamic weight change rate, and call the ratio of the mean square of the change rate deviation to the number of nodes to generate the node parameter change standard deviation; S102: Based on the node parameter change standard deviation, perform normalization processing on the change rates of CPU utilization rate, memory occupancy rate, and bandwidth occupancy rate, call the change amount of delay and the change amount of overflow rate to be multiplied and accumulated with the normalization amplitude respectively, and obtain the node parameter sensitivity result based on the ratio of the number of nodes; S103: According to the node parameter sensitivity result, call the node feature sensitivity threshold, compare the sensitivity coefficient with the threshold to adjust the node feature priority state, establish the mapping relationship between the standard deviation change range and the feature priority, and obtain the feature sensitivity prediction result.
4. The edge computing scheduling optimization method based on the actor-critic model according to claim 3, wherein The specific formula for the dynamic weight change rate is as follows: ; Among them, represents the node at time of the dynamic weight change rate, represents the node at time of the CPU utilization rate, memory occupancy rate or bandwidth occupancy rate sampling value, represents the time interval of continuous sampling, represents the total time length of the cycle parameter, represents the number of nodes, represents the node of the effective sampling interval within the cycle.
5. The edge computing scheduling optimization method based on the actor-critic model according to claim 1, wherein, The specific steps of S2 are as follows S201: Based on the feature sensitivity prediction result, call the real-time CPU utilization rate and memory occupancy rate of the node, set the node resource weight coefficient, perform weighted averaging on the CPU utilization rate and memory occupancy rate respectively, calculate the weighted weight coefficient of the node resource utilization rate, and generate the node load fluctuation trend; S202: Based on the node load fluctuation trend, call the change amount of the node resource utilization rate per unit time, compare the upper and lower bounds according to the set load change interval reference value, judge the resource utilization direction, extract the positive and negative signs of the change and establish a time series to obtain the node resource direction change amount; S203: According to the node load fluctuation trend and the node resource direction change amount, call the node time series order, integrate the change amplitude and direction information, generate instantaneous resource data points node by node, draw a resource change curve based on the time series, and obtain the instantaneous resource utilization trend.
6. The edge computing scheduling optimization method based on the actor-critic model according to claim 5, characterized in that The specific formula for the weighted weight coefficient of the node resource utilization rate is: ; Among them, represents the weighted coefficient after the weighted node resource utilization rate, represents the weighted coefficient of the node CPU utilization rate, represents the weighted coefficient of the node memory occupancy rate, represents the value of the real-time CPU utilization rate of the node at the current moment, represents the value of the real-time memory occupancy rate of the node at the current moment, represents the total number of nodes currently participating in the statistics.
7. The edge computing scheduling optimization method based on the actor-critic model according to claim 1, characterized in that The specific steps of S3 are as follows: S301: According to the instantaneous resource utilization trend, collect the time slice load values of the node within the time period, arrange the time slice load values, extract the load change amount per unit time based on the difference between adjacent time slices, integrate the change amount set, and generate the node load change rate data set; S302: Based on the node load change rate data set, set the load status interval according to the change rate range, call the load change rate data to compare with the interval boundary, classify the node load status according to the result, and obtain the node load status interval set; S303: According to the node load status interval set, call the node basic load threshold reference, calculate the load deviation value according to the comparison between the node load status and the reference, integrate the node deviation data to generate an adjustment set, and obtain the dynamic load adjustment reference.
8. The edge computing scheduling optimization method based on the actor-critic model according to claim 7, characterized in that The specific formula for the load deviation value is ; Among them, represents the node load deviation value, represents the actual node load data, represents the node basic load threshold data, represents the average of the basic load thresholds of all nodes, represents all the sum of the absolute values of the differences between the actual loads and the reference loads of all nodes.
9. The edge computing scheduling optimization method based on the actor-critic model according to claim 1, wherein The specific steps of S4 are as follows: S401: Call the dynamic load adjustment reference, compare the node instantaneous load value with the upper and lower bounds of the adjustment reference interval, judge the node load position, extract the node interval flag information, integrate the node position flag data set, and generate the node instantaneous load position marking value; S402: Based on the node instantaneous load position marking value, screen the actions corresponding to the sensitivity response flags in the action pool, call the action delay label to accumulate the delay times, screen the action set with the top-ranked delay times, and integrate the sorting information to obtain the sensitive action preference set; S403: According to the sensitive action preference set, call the original execution priority of the action, re-adjust the priority order according to the delay times ranking, assign the action priority weight, integrate the node action priority change information, and obtain the scheduling action screening list.
10. The edge computing scheduling optimization method based on the actor-critic model according to claim 1, wherein The specific steps of S5 are as follows: S501: According to the scheduling action screening list, extract the local policy parameter update batches in the feedback data, arrange the batch node policy parameter change values, extract the local gradient change amount, integrate the node change data, and generate the local gradient change amount vector set; S502: Based on the set of local gradient change amount vectors, extract batch change vector data, perform a normalized difference operation between the local gradient change amount vector and the batch change vector, integrate the node normalized difference values to form a node normalized difference sequence, and obtain a normalized difference set; S503: According to the normalized difference set, compare the normalized difference mean with the upper and lower threshold values, judge the deviation direction of the mean, adjust the corresponding node learning rate amplitude, integrate the node learning rate change information, and obtain a scheduling optimization scheme.
Citation Information
Patent Citations
Multi-target task scheduling method and device in edge computing environment
CN115292036A
Cluster load balancing method for edge computing, server and electronic equipment
CN118210631A
Cited By
Educational resource intelligent classification and accurate recommendation system and method based on artificial intelligence
CN120910323A
Intelligent classification and accurate recommendation system and method for education resources based on artificial intelligence
CN120910323B