Enterprise resource dynamic scheduling optimization method and system based on cloud edge collaboration

CN122195669BActive Publication Date: 2026-09-29BEIJING MINGQI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610408528.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-31
Publication Date
2026-09-29
Estimated Expiration
2046-03-31

AI Technical Summary

Technical Problem

[0003]然而,这些调度方法缺乏对任务间依赖关系的深入分析,特别是跨节点资源依赖关系的处理能力不足,导致强耦合任务被分散调度时造成大量的跨节点通信开销和协调成本,严重影响系统整体性能

Benefits of technology

[0045]本发明通过识别强连通分量提取强耦合任务簇,将任务进行合理分类处理,降低了调度系统的整体复杂度,提高了调度效率。对强耦合任务采用集中调度方案,对非强耦合任务设置约束边界并允许边缘节点自主决策,平衡了调度的集中性和分布式特性,提高了系统资源利用效率。通过边缘节点反馈偏差值、持续时长及应急执行状态信息,云端能够不断修正负载预测方法,形成闭环优化机制。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195669B_ABST
    Figure CN122195669B_ABST
Patent Text Reader

Abstract

The application provides an enterprise resource dynamic scheduling optimization method and system based on cloud edge collaboration, relates to the technical field of cloud edge collaboration, and comprises the following steps: collecting resource state data through an edge node, predicting a load and constructing a resource dependency graph in a cloud center, identifying a strongly connected component to extract a strongly coupled task cluster, generating a differentiated scheduling strategy for different types of tasks, autonomously executing non-strongly coupled tasks within a constraint boundary by the edge node, and triggering an emergency scheduling mechanism when a deviation exceeds a tolerance interval. The application realizes efficient utilization of resources and stability of task execution, and reduces communication overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to cloud-edge collaboration technology, and more particularly to a method and system for dynamic scheduling and optimization of enterprise resources based on cloud-edge collaboration. Background Technology

[0002] With the rapid development of information technology, enterprise resource management systems are gradually shifting from traditional centralized architectures to distributed architectures. Existing enterprise resource scheduling technologies mainly employ centralized scheduling or fully distributed scheduling schemes. Centralized scheduling involves global task allocation planning in the cloud, while distributed scheduling relies on autonomous decision-making by each node.

[0003] However, these scheduling methods lack in-depth analysis of inter-task dependencies, particularly in handling cross-node resource dependencies. This leads to significant cross-node communication overhead and coordination costs when tightly coupled tasks are distributed during scheduling, severely impacting overall system performance. Existing scheduling strategies typically make decisions based on static resource states, failing to effectively address dynamically changing resource loads. When the actual load differs significantly from expectations, the lack of effective adjustment mechanisms easily results in unreasonable resource allocation and low task execution efficiency. Summary of the Invention

[0004] This invention provides a method and system for dynamic scheduling and optimization of enterprise resources based on cloud-edge collaboration, which can solve the problems in the prior art.

[0005] A first aspect of this invention provides a method for dynamic scheduling and optimization of enterprise resources based on cloud-edge collaboration, comprising:

[0006] Edge nodes collect enterprise resource status data and upload it to the cloud scheduling center; the cloud scheduling center predicts the resource load of each edge node based on the resource status data;

[0007] The cloud-based scheduling center acquires tasks to be scheduled, constructs a resource dependency graph based on cross-node resource dependencies between tasks, extracts strongly coupled task clusters by identifying strongly connected components, generates a centralized scheduling scheme for strongly coupled task clusters based on resource load prediction results, generates constraint boundaries for non-strongly coupled tasks, and sets an error tolerance range for the constraint boundaries based on the resource load prediction results.

[0008] Edge nodes execute corresponding tasks according to the centralized scheduling scheme. For non-strongly coupled tasks, they autonomously solve the execution scheme based on the local measured resource status while satisfying the constraint boundary. When the deviation between the measured load and the predicted load exceeds the error tolerance range, a replanning request is sent to the cloud and the system switches to emergency scheduling mode. At the same time, the deviation value, duration, and emergency execution status are fed back to the cloud.

[0009] The cloud-based scheduling center corrects the load prediction method based on feedback information and regenerates the scheduling strategy when it receives a replanning request.

[0010] The steps by which the cloud-based scheduling center predicts the resource load of each edge node based on resource status data include:

[0011] Edge nodes collect enterprise resource status data and upload it to the cloud scheduling center. After receiving the enterprise resource status data uploaded by each edge node, the cloud scheduling center groups the enterprise resource status data according to the edge node identifier, and sorts each group of enterprise resource status data by timestamp to construct a time-series data sequence.

[0012] The periodic variation characteristics of resource load are extracted from the time-series data sequence. The prediction window length is determined based on the periodic variation characteristics. The resource load prediction results of each edge node are generated based on the time-series data sequence within the prediction window length.

[0013] The cloud-based scheduling center obtains tasks to be scheduled, constructs a resource dependency graph based on cross-node resource dependencies between tasks, and extracts strongly coupled task clusters by identifying strongly connected components. The steps include:

[0014] The cloud-based scheduling center obtains a set of tasks to be scheduled, parses and stores the resource requirements and execution node constraints of each task;

[0015] Traverse the set of tasks to be scheduled, identify cross-node resource dependencies between tasks, treat each task to be scheduled as a node, and treat cross-node resource dependencies as directed edges to construct a resource dependency graph;

[0016] A strong connected component identification operation is performed on the resource dependency graph. The set of nodes in the resource dependency graph that have a bidirectional reachable path between any two nodes is marked as a strong connected component. Strong connected components containing cross-node resource dependencies are selected from the strong connected components. The tasks to be scheduled corresponding to the nodes in the selected strong connected components are extracted as strongly coupled task clusters.

[0017] The tasks in the set of tasks to be scheduled, excluding strongly coupled task clusters, are marked as non-strongly coupled tasks.

[0018] The steps of generating a centralized scheduling scheme for strongly coupled task clusters based on resource load prediction results, generating constraint boundaries for non-strongly coupled tasks, and setting error tolerance intervals for the constraint boundaries based on the resource load prediction results include:

[0019] For each task in a strongly coupled task cluster, the available resources are calculated based on the resource requirements of the task and the predicted load values ​​of each edge node. Edge nodes whose available resources meet the resource requirements are selected as candidate execution nodes. The edge node with the lowest predicted load value is selected from the candidate execution nodes as the target execution node. The target execution time is determined based on the predicted load value of the target execution node. The target execution node and the target execution time are combined to generate a centralized scheduling scheme.

[0020] For loosely coupled tasks, the peak resource usage of each edge node is calculated based on the resource load prediction results as the upper limit constraint of resource usage. The task priority order constraint is determined based on the dependency relationship between loosely coupled tasks, and the constraint boundary is generated by combining them.

[0021] Extract the fluctuation range characteristics of the predicted load value from the resource load prediction results, calculate the upper and lower limits of the error tolerance interval based on the fluctuation range characteristics, and set them as the error tolerance interval of the constraint boundary; then distribute the centralized scheduling scheme, constraint boundary and error tolerance interval to the corresponding edge nodes.

[0022] Edge nodes execute corresponding tasks according to the centralized scheduling scheme. For non-strongly coupled tasks, the steps for autonomously solving the execution scheme based on the local measured resource status while satisfying the constraint boundaries include:

[0023] Edge nodes collect the measured resource status before the target execution time arrives; calculate the resource reservation amount based on the predicted resource demand in the centralized scheduling scheme and the measured resource status before execution, and establish a dedicated resource pool for strongly coupled tasks. During the execution process, the actual resource usage is continuously collected. When the actual resource usage exceeds the resource reservation amount, the difference in resources occupied by non-strongly coupled tasks is forcibly reclaimed and supplemented to the dedicated resource pool.

[0024] Edge nodes calculate the difference between their local total resources and the dedicated resource pool usage as the available resources for non-strongly coupled tasks. They then compare the available resources for non-strongly coupled tasks with the upper limit of resource usage in the constraint boundary. If the upper limit of resource usage is not exceeded, the constraint boundary is considered satisfied. The deviation between the local measured resource status and the predicted resource status in the centralized scheduling scheme is calculated. The initial priority weight is adjusted according to the magnitude and direction of the deviation to obtain the dynamic priority weight. The non-strongly coupled tasks are reordered according to the dynamic priority weight, and resource shares are allocated sequentially to generate an execution plan and start execution.

[0025] When an edge node receives a forced resource reclamation command, it selects the lowest priority task in reverse order of dynamic priority weights to suspend execution and release resources.

[0026] When the deviation between the measured load and the predicted load exceeds the error tolerance range, a replanning request is sent to the cloud and the system switches to emergency scheduling mode. Simultaneously, the deviation value, duration, and emergency execution status are fed back to the cloud. The steps include:

[0027] Edge nodes continuously collect local measured load, calculate the deviation between the measured load and the predicted load in the centralized scheduling scheme, compare the deviation with the upper and lower limits of the error tolerance range, and determine that the deviation exceeds the error tolerance range when the deviation exceeds the upper limit or falls below the lower limit.

[0028] When an edge node determines that the deviation value exceeds the error tolerance range, it generates a replanning request containing the current deviation value and the edge node identifier and sends it to the cloud scheduling center. At the same time, it starts the emergency scheduling mode, which suspends the reception of new centralized scheduling schemes and reallocates the currently executing tasks according to the local measured resource status.

[0029] Edge nodes record the start and end times when the deviation value exceeds the error tolerance range, calculate the time difference as the duration, and collect the task execution status under emergency scheduling mode as the emergency execution status. The deviation value, duration, and emergency execution status are combined to generate feedback information and uploaded to the cloud scheduling center.

[0030] The steps by which the cloud-based scheduling center corrects the load forecasting method based on feedback information and regenerates the scheduling strategy upon receiving a replanning request include:

[0031] The cloud-based dispatch center receives feedback information uploaded by edge nodes, extracts deviation values, duration, and emergency execution status from the feedback information, and archives the feedback information to the historical deviation record of the corresponding edge node according to the edge node identifier.

[0032] The cloud-based dispatch center analyzes the distribution characteristics and duration of deviation values ​​from historical deviation records to identify load fluctuation patterns that cause deviations to exceed the error tolerance range. Based on these load fluctuation patterns, it adjusts the prediction window length and the weighting parameters of periodic change characteristics in the load forecasting method.

[0033] When the cloud scheduling center receives a replanning request, it obtains the tasks currently being executed and the tasks to be scheduled by the edge node, re-predicts the resource load of the edge node according to the adjusted load prediction method, and regenerates a centralized scheduling scheme for strongly coupled task clusters based on the re-predicted resource load, and regenerates constraint boundaries and error tolerance intervals for non-strongly coupled tasks and distributes them to the edge node.

[0034] A second aspect of the present invention provides an enterprise resource dynamic scheduling and optimization system based on cloud-edge collaboration, comprising:

[0035] The data acquisition module is used by edge nodes to collect enterprise resource status data and upload it to the cloud dispatch center;

[0036] The load prediction module is used by the cloud scheduling center to predict the resource load of each edge node based on resource status data.

[0037] The task classification and scheme generation module is used by the cloud scheduling center to obtain tasks to be scheduled, construct a resource dependency graph based on the cross-node resource dependency relationship between tasks, extract strongly coupled task clusters by identifying strongly connected components, generate a centralized scheduling scheme for strongly coupled task clusters based on the resource load prediction results, generate constraint boundaries for non-strongly coupled tasks, and set an error tolerance range for the constraint boundaries based on the resource load prediction results.

[0038] The hybrid execution and anomaly detection module is used for edge nodes to execute corresponding tasks according to the centralized scheduling scheme. For non-strongly coupled tasks, it can autonomously solve the execution scheme based on the local measured resource status while satisfying the constraint boundary. When the deviation between the measured load and the predicted load exceeds the error tolerance range, it sends a replanning request to the cloud and switches to emergency scheduling mode. At the same time, it feeds back the deviation value, duration and emergency execution status to the cloud.

[0039] The adaptive optimization module is used by the cloud scheduling center to correct the load prediction method based on feedback information and regenerate the scheduling strategy when a replanning request is received.

[0040] A third aspect of the present invention provides an electronic device, comprising:

[0041] processor;

[0042] Memory used to store processor-executable instructions;

[0043] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0044] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0045] This invention extracts strongly coupled task clusters by identifying strongly connected components, and then rationally classifies and processes tasks, reducing the overall complexity of the scheduling system and improving scheduling efficiency. A centralized scheduling scheme is adopted for strongly coupled tasks, while constraints are set for non-strongly coupled tasks, allowing edge nodes to make autonomous decisions. This balances the centralized and distributed characteristics of scheduling, improving system resource utilization efficiency. By having edge nodes provide feedback on deviation values, duration, and emergency execution status information, the cloud can continuously refine the load prediction method, forming a closed-loop optimization mechanism. Attached Figure Description

[0046] Figure 1 This is a flowchart illustrating the enterprise resource dynamic scheduling and optimization method based on cloud-edge collaboration according to an embodiment of the present invention.

[0047] Figure 2 This is a flowchart of the resource scheduling execution process for edge nodes. Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0050] Figure 1 This is a flowchart illustrating the enterprise resource dynamic scheduling and optimization method based on cloud-edge collaboration according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0051] Edge nodes collect enterprise resource status data and upload it to the cloud scheduling center; the cloud scheduling center predicts the resource load of each edge node based on the resource status data;

[0052] The cloud-based scheduling center acquires tasks to be scheduled, constructs a resource dependency graph based on cross-node resource dependencies between tasks, extracts strongly coupled task clusters by identifying strongly connected components, generates a centralized scheduling scheme for strongly coupled task clusters based on resource load prediction results, generates constraint boundaries for non-strongly coupled tasks, and sets an error tolerance range for the constraint boundaries based on the resource load prediction results.

[0053] Edge nodes execute corresponding tasks according to the centralized scheduling scheme. For non-strongly coupled tasks, they autonomously solve the execution scheme based on the local measured resource status while satisfying the constraint boundary. When the deviation between the measured load and the predicted load exceeds the error tolerance range, a replanning request is sent to the cloud and the system switches to emergency scheduling mode. At the same time, the deviation value, duration, and emergency execution status are fed back to the cloud.

[0054] The cloud-based scheduling center corrects the load prediction method based on feedback information and regenerates the scheduling strategy when it receives a replanning request.

[0055] In one optional implementation, the step of the cloud scheduling center predicting the resource load of each edge node based on resource status data includes:

[0056] Edge nodes collect enterprise resource status data and upload it to the cloud scheduling center. After receiving the enterprise resource status data uploaded by each edge node, the cloud scheduling center groups the enterprise resource status data according to the edge node identifier, and sorts each group of enterprise resource status data by timestamp to construct a time-series data sequence.

[0057] The periodic variation characteristics of resource load are extracted from the time-series data sequence. The prediction window length is determined based on the periodic variation characteristics. The resource load prediction results of each edge node are generated based on the time-series data sequence within the prediction window length.

[0058] For example, edge nodes can be computing devices deployed within an enterprise, such as edge servers and gateway devices. These edge nodes periodically collect resource status data from the enterprise through built-in data acquisition modules. This resource status data includes, but is not limited to, computing resource-related metrics such as CPU utilization, memory usage, network bandwidth usage, and storage space usage, as well as business metrics such as request volume and response time of enterprise applications. The data collection frequency can be configured according to the enterprise's business characteristics, for example, once per minute or once every five minutes.

[0059] After collecting resource status data, edge nodes perform preliminary processing, including data formatting and compression, and add timestamps and edge node identification information to each data entry. Then, this data is uploaded to the cloud scheduling center via a secure network channel. The upload process employs an encrypted transmission mechanism to ensure data security.

[0060] After receiving enterprise resource status data uploaded by each edge node, the cloud-based dispatch center first verifies and decrypts the data to ensure its integrity and authenticity. Next, it groups the enterprise resource status data according to the edge node identifier, grouping data from the same edge node together. For each group, the cloud-based dispatch center sorts the data based on timestamps to construct a time-series data sequence. For example, for an edge node identified as "Edge-001," all its uploaded resource status data will be extracted and arranged in chronological order to form a time-series data sequence of that edge node's resource status.

[0061] After constructing the time-series data sequences, the cloud-based scheduling center extracts the periodic variation characteristics of resource load from these sequences. The extraction of periodic variation characteristics is achieved through time series analysis methods, mainly including the following steps: preprocessing the time-series data, including handling missing values, outlier detection and handling, and data normalization; and using methods such as autocorrelation function analysis and Fourier transform to identify periodic patterns in the time-series data. For example, by calculating the autocorrelation coefficients at different time intervals, the periodicity of the data can be discovered; through Fourier transform, the time-series data can be converted from the time domain to the frequency domain, making it easier to identify the main periodic components.

[0062] Based on the extracted periodic variation characteristics, the cloud scheduling center determines the prediction window length. The prediction window length refers to the length of historical data used to predict future resource load. Generally, the prediction window length should cover at least one complete period in order to capture the periodic changes in the data. For example, if the resource load data shows a clear daily periodic pattern (similar patterns of change every day), then the prediction window length may be set to 24 hours or an integer multiple thereof; if a weekly periodic pattern exists, the prediction window length may be set to 7 days or an integer multiple thereof.

[0063] After determining the prediction window length, the cloud scheduling center generates resource load prediction results for each edge node based on the time-series data sequence within the prediction window length. Resource load prediction employs a Long Short-Term Memory (LSTM) network algorithm. Specifically, a three-layer network structure is first constructed, consisting of an input layer, a hidden layer, and an output layer. The input layer receives the time-series data within the prediction window length, and the hidden layer contains multiple LSTM units to capture long-term dependencies and short-term fluctuations in the time-series data. Within each LSTM unit, a forget gate controls the retention of historical information, an input gate determines the amount of new information added, and an output gate controls the hidden state of the current output. During the training phase, historical time-series data is divided into training samples using a sliding window approach. Each sample contains an input sequence of the prediction window length and the corresponding future true resource load value as a label. Backpropagation and the Adam optimizer are used to update the network parameters, minimizing the mean squared error between the predicted and true values. After training, real-time data within the prediction window length before the current time is input into the network, and the output layer generates predicted resource usage values ​​for each time period in the future.

[0064] The forecast results include predicted resource usage values ​​for each edge node over a future time period, such as predicted CPU utilization and memory usage. These forecasts can be used by the cloud scheduling center for resource scheduling decisions. For example, computing resources can be allocated in advance based on the forecasts, or some tasks can be migrated to nodes with more available resources if resource shortages are predicted on a certain node, thereby optimizing overall resource utilization efficiency, improving system response speed, and reducing the risk of service interruption.

[0065] In addition, the cloud-based scheduling center regularly evaluates the accuracy of the prediction results by comparing the predicted values ​​with the actual observed values ​​to calculate the prediction error. Based on the error analysis results, the scheduling center will adjust hyperparameters such as the number of hidden layer units, learning rate, or training epochs of the LSTM network accordingly. If necessary, the network will be retrained to adapt to changes in resource load patterns in order to continuously improve the accuracy of the prediction.

[0066] This method can accurately capture the complex temporal evolution of enterprise resource load and effectively reduce the risk of resource allocation imbalance caused by prediction bias.

[0067] In one optional implementation, the cloud scheduling center obtains the tasks to be scheduled, constructs a resource dependency graph based on the cross-node resource dependencies between tasks, and extracts strongly coupled task clusters by identifying strongly connected components.

[0068] The cloud-based scheduling center obtains a set of tasks to be scheduled, parses and stores the resource requirements and execution node constraints of each task;

[0069] Traverse the set of tasks to be scheduled, identify cross-node resource dependencies between tasks, treat each task to be scheduled as a node, and treat cross-node resource dependencies as directed edges to construct a resource dependency graph;

[0070] A strong connected component identification operation is performed on the resource dependency graph. The set of nodes in the resource dependency graph that have a bidirectional reachable path between any two nodes is marked as a strong connected component. Strong connected components containing cross-node resource dependencies are selected from the strong connected components. The tasks to be scheduled corresponding to the nodes in the selected strong connected components are extracted as strongly coupled task clusters.

[0071] The tasks in the set of tasks to be scheduled, excluding strongly coupled task clusters, are marked as non-strongly coupled tasks.

[0072] For example, after obtaining the tasks to be scheduled, the cloud scheduling center first needs to parse the resource requirements and execution constraints of each task. The obtained set of tasks to be scheduled typically contains multiple tasks, each with its specific resource requirements and execution node constraints. During the parsing process, the cloud scheduling center extracts the computing resources (such as CPU, memory, and storage), network resources (such as bandwidth and latency requirements), and node limitations (such as hardware architecture and geographical location) required for each task. The parsed information is stored in a task attribute table for quick access when constructing the resource dependency graph later.

[0073] After parsing the task information, the cloud scheduling center begins to identify cross-node resource dependencies between tasks. Cross-node resource dependency refers to the execution of one task depending on data or resources generated by another task on different nodes. For example, if task A runs on node 1, and the generated data needs to be used by task B on node 2, then task B has a cross-node resource dependency on task A. The identification process is achieved by analyzing the task's input and output data streams, shared resource access patterns, and execution order constraints. Specifically, for each task pair (Ti, Tj) in the task set, the cloud scheduling center checks whether Tj needs data or resources generated by Ti, and whether Ti and Tj may be assigned to different nodes. If these conditions are met, a directed edge is established between Ti and Tj, indicating that Tj depends on Ti.

[0074] When constructing a resource dependency graph, each task to be scheduled is represented as a node in the graph, and cross-node resource dependencies between tasks are represented as directed edges from the dependent task to the dependent task. If data generated by task A is used by task B, a directed edge is established from node A to node B in the graph. The dependency graph is constructed using an adjacency list or adjacency matrix structure for efficient representation and manipulation of the graph structure. The completed resource dependency graph reflects the complex dependency network of the entire set of tasks to be scheduled.

[0075] After the resource dependency graph is constructed, the cloud scheduling center performs a strongly connected component identification operation. A strongly connected component is the largest subset of nodes in the graph that can be bidirectionally reached between any two nodes. Tarjan's algorithm or Kosaraju's algorithm is commonly used to identify strongly connected components. Taking Tarjan's algorithm as an example, this algorithm traverses all nodes in the graph using a depth-first search, assigning a discovery time and the earliest backtrackable time to each node. When these two time values ​​are equal, a strongly connected component can be identified. In the specific implementation of the algorithm, a stack is maintained to record the nodes on the currently accessed path. When a strongly connected component is identified, the nodes are popped from the stack.

[0076] From all identified strongly connected components, the cloud-based scheduling center further filters for those containing cross-node resource dependencies. Specifically, the filtering method checks whether there are edges representing cross-node resource dependencies between nodes within each strongly connected component. If a strongly connected component contains at least one edge representing a cross-node resource dependency, it is marked as a strongly connected component containing cross-node resource dependencies. In practical applications, such as enterprise order processing, an inventory query task deployed in the headquarters cloud, an order verification task deployed on regional edge nodes, and a logistics scheduling task deployed on warehouse edge nodes might form a strongly connected component because these three tasks have a cyclical data dependency: order verification requires real-time inventory information, logistics scheduling depends on order verification results, and inventory updates require logistics scheduling feedback. These tasks are distributed across the cloud and different edge nodes, requiring frequent cross-node data interactions.

[0077] Tasks within the selected strongly connected components require special handling during scheduling due to their bidirectional dependencies and cross-node resource interactions. These tasks are extracted from the original dependency graph to form independent task clusters, preparing for subsequent collaborative scheduling. Identifying these strongly coupled task clusters allows the cloud-based scheduling center to optimize deployment strategies for these tasks, such as deploying tasks within the cluster on nodes with good network connectivity or employing data localization techniques to reduce cross-node data transmission.

[0078] Tasks in the set of tasks to be scheduled that are not classified into strongly coupled task clusters are marked as weakly coupled tasks. Although these tasks may have dependencies on other tasks, there are no complex situations involving circular dependencies, and they can be handled using conventional scheduling strategies. The marking of weakly coupled tasks completes the entire task classification process and lays the foundation for subsequent differentiated scheduling.

[0079] Through the above steps, the cloud-based scheduling center successfully divides the set of tasks to be scheduled into two categories: strongly coupled task clusters and loosely coupled tasks. This allows for the adoption of appropriate scheduling strategies for different types of tasks, thereby improving overall scheduling efficiency and resource utilization.

[0080] In one optional implementation, the steps of generating a centralized scheduling scheme for strongly coupled task clusters based on resource load prediction results, generating constraint boundaries for non-strongly coupled tasks, and setting an error tolerance range for the constraint boundaries based on the resource load prediction results include:

[0081] For each task in a strongly coupled task cluster, the available resources are calculated based on the resource requirements of the task and the predicted load values ​​of each edge node. Edge nodes whose available resources meet the resource requirements are selected as candidate execution nodes. The edge node with the lowest predicted load value is selected from the candidate execution nodes as the target execution node. The target execution time is determined based on the predicted load value of the target execution node. The target execution node and the target execution time are combined to generate a centralized scheduling scheme.

[0082] For loosely coupled tasks, the peak resource usage of each edge node is calculated based on the resource load prediction results as the upper limit constraint of resource usage. The task priority order constraint is determined based on the dependency relationship between loosely coupled tasks, and the constraint boundary is generated by combining them.

[0083] Extract the fluctuation range characteristics of the predicted load value from the resource load prediction results, calculate the upper and lower limits of the error tolerance interval based on the fluctuation range characteristics, and set them as the error tolerance interval of the constraint boundary; then distribute the centralized scheduling scheme, constraint boundary and error tolerance interval to the corresponding edge nodes.

[0084] For example, a strongly coupled task cluster refers to a group of tasks that have high requirements for latency and execution resources and have close dependencies on each other. For such tasks, centralized scheduling control is required to ensure that they are executed in an optimized time sequence on the most suitable edge nodes.

[0085] When generating a centralized scheduling scheme, the resource requirements of each task are obtained, including the required computing resources, storage resources, and network bandwidth resources. Then, combined with the predicted load values ​​of each edge node, the resource availability of each edge node during the expected execution time of the task is calculated. The resource availability is the total resource amount of the edge nodes minus the resource amount occupied by the predicted load value.

[0086] Edge nodes whose available resources meet the task's resource requirements are selected as candidate execution nodes. For example, if the task requires 2 CPU cores and 4GB of memory, and edge node A has 3 CPU cores and 6GB of available memory, then node A can be selected as a candidate execution node.

[0087] From the candidate execution nodes, the edge node with the lowest predicted load value is selected as the target execution node. Selecting the node with the lowest load can reduce resource contention during task execution and improve execution efficiency. For example, if the predicted CPU load of candidate node A is 30% and the predicted CPU load of candidate node B is 20%, then node B is selected as the target execution node.

[0088] The target execution time is determined based on the predicted load of the target execution node and the resource requirements of the task. The target execution time includes the task's start time and estimated completion time. For example, if the current time is T and the task is expected to take 10 minutes to execute, then the target execution time is [T, T+10 minutes].

[0089] By combining the target execution node and the target execution time, a centralized scheduling scheme for the task is generated. For example, the centralized scheduling scheme for task A is: execution begins at time T on edge node B and is expected to be completed in T+10 minutes.

[0090] For loosely coupled tasks, a boundary constraint approach is used for scheduling control. First, based on resource load prediction results, the peak resource usage of each edge node is calculated as the upper limit constraint for resource usage. The peak resource usage refers to the highest utilization rate of various resources of the edge node within the prediction time window. For example, if the peak CPU utilization of edge node C is 80% in the next hour, its CPU resource usage upper limit constraint is set to 80%. Based on the dependencies between loosely coupled tasks, task priority constraints are determined. Dependencies are represented by a task dependency graph, where tasks are nodes and dependencies are directed edges. A topological sorting algorithm is used to determine task execution priorities. The specific process includes: first, counting the in-degree of each task node, i.e., how many other tasks depend on this task; adding tasks with an in-degree of zero to a priority queue, indicating that these tasks have no prerequisites and can be executed first; taking one task from the priority queue and setting it as the highest priority task, and decrementing the in-degree of other tasks that depend on it; repeating the above process until all tasks are sorted, the final task sequence is the priority constraint. For example, if task B depends on the execution result of task A, then task A has a higher priority than task B.

[0091] By combining resource consumption limits and task priority constraints, a constraint boundary is generated. This constraint boundary serves as a guide for decentralized task scheduling, ensuring that task scheduling does not exceed resource capacity and satisfies inter-task dependencies.

[0092] To address load fluctuations in real-world environments, the fluctuation amplitude characteristics of predicted load values ​​are extracted from resource load prediction results. Specifically, the standard deviation is used as a quantification of fluctuation amplitude. The calculation method is as follows: for the load prediction value sequence of an edge node at each moment within the prediction time window, first calculate the average value of the sequence, then calculate the square of the difference between each predicted value and the average value, sum all squared differences, divide by the number of predicted values, and finally take the square root of the result to obtain the standard deviation. A larger standard deviation indicates more severe load fluctuations. Based on the fluctuation amplitude characteristics, the upper and lower limits of the error tolerance interval are calculated using the following formulas: the upper limit equals the predicted average load plus the standard deviation, and the lower limit equals the predicted average load minus the standard deviation. For example, if the predicted average CPU load of edge node D is 50% and the standard deviation is 5%, then the error tolerance interval is set to [45%, 55%]. Setting this error tolerance interval as the constraint boundary enhances the robustness of the scheduling strategy.

[0093] The centralized scheduling scheme, constraint boundaries, and error tolerance intervals are distributed to the corresponding edge nodes. After receiving this scheduling information, the edge nodes execute tasks in tightly coupled task clusters strictly according to the centralized scheduling scheme; while the edge nodes autonomously schedule and execute tasks within the constraint boundaries and error tolerance intervals, thus achieving distributed task management.

[0094] In practical applications, such as enterprise manufacturing systems, tasks like equipment status monitoring, quality inspection, and production scheduling deployed at factory edge nodes can be identified as tightly coupled task clusters. These tasks have close data dependencies and timing constraints, requiring execution on specific edge nodes according to strict time requirements to ensure real-time production line response. Meanwhile, non-real-time tasks such as uploading equipment maintenance logs and compiling energy consumption data, deployed at workshop edge nodes, can be treated as loosely coupled tasks and flexibly scheduled as long as constraints are met. This differentiated scheduling mechanism ensures timely processing of critical production tasks while improving overall enterprise resource utilization efficiency.

[0095] This method can ensure the quality of high-priority task execution while making full use of the autonomy of edge nodes, effectively reducing the communication overhead and computing burden of the cloud scheduling center, and improving the flexibility and real-time response capability of enterprise resource scheduling.

[0096] In one optional implementation, edge nodes execute corresponding tasks according to a centralized scheduling scheme. For non-strongly coupled tasks, the steps of autonomously solving the execution scheme based on local measured resource status while satisfying the constraint boundaries include:

[0097] Edge nodes collect the measured resource status before the target execution time arrives; calculate the resource reservation amount based on the predicted resource demand in the centralized scheduling scheme and the measured resource status before execution, and establish a dedicated resource pool for strongly coupled tasks. During the execution process, the actual resource usage is continuously collected. When the actual resource usage exceeds the resource reservation amount, the difference in resources occupied by non-strongly coupled tasks is forcibly reclaimed and supplemented to the dedicated resource pool.

[0098] Edge nodes calculate the difference between their local total resources and the dedicated resource pool usage as the available resources for non-strongly coupled tasks. They then compare the available resources for non-strongly coupled tasks with the upper limit of resource usage in the constraint boundary. If the upper limit of resource usage is not exceeded, the constraint boundary is considered satisfied. The deviation between the local measured resource status and the predicted resource status in the centralized scheduling scheme is calculated. The initial priority weight is adjusted according to the magnitude and direction of the deviation to obtain the dynamic priority weight. The non-strongly coupled tasks are reordered according to the dynamic priority weight, and resource shares are allocated sequentially to generate an execution plan and start execution.

[0099] When an edge node receives a forced resource reclamation command, it selects the lowest priority task in reverse order of dynamic priority weights to suspend execution and release resources.

[0100] Combination Figure 2 The edge node resource scheduling execution flowchart illustrates this process. For example, before the target execution time arrives, the edge node collects pre-execution measured resource status data. This process is implemented through a resource monitoring module on the edge node, which periodically collects resource metrics such as CPU utilization, memory usage, network bandwidth usage, and storage space. For instance, it uses the operating system's performance monitoring API to obtain data such as the real-time utilization of each CPU core, available memory size, real-time throughput and latency of network interfaces, and disk I / O rate. This data is then aggregated to form a snapshot of the current edge node's resource status, with the collection timestamp accurate to the millisecond level to ensure data timeliness.

[0101] Edge nodes calculate resource reservations based on predicted resource requirements from the centralized scheduling scheme and actual measured resource status before execution. Specifically, if the centralized scheduling scheme predicts that a strongly coupled task requires 30% CPU resources, 4GB of memory, and 50Mbps network bandwidth, while the current measured CPU utilization is 20%, available memory is 10GB, and available bandwidth is 100Mbps, then 30% CPU resources, 4GB of memory, and 50Mbps bandwidth need to be reserved. After performing similar calculations for all resource types, the edge node establishes a dedicated resource pool for strongly coupled tasks and allocates the calculated reserved resources to this pool. The establishment of the dedicated resource pool employs an operating system-level resource isolation mechanism. Specific implementation methods include: limiting CPU core allocation and usage ratios through Linux cgroups (control groups), dividing exclusive memory regions through memory namespace technology, and setting bandwidth guarantee policies through network traffic control tools to ensure that resources in the dedicated resource pool are used only by strongly coupled tasks and are not encroached upon by other tasks. After the resource pool is established, the edge node records the reserved amount, occupancy status, and assigned task identifier for each resource in the resource allocation table for subsequent dynamic adjustment and monitoring.

[0102] During task execution, edge nodes continuously collect actual resource usage data. The resource monitoring module records the actual resource consumption of tightly coupled tasks at a preset sampling frequency (e.g., every 100 milliseconds). Monitoring metrics include underlying indicators such as the number of CPU cycles used, the number of memory pages used, and the network packet sending / receiving rate. When actual resource usage exceeds the reserved resource limit, for example, when actual CPU usage reaches 35% while the reserved limit is only 30%, the edge node triggers a resource reclamation mechanism. The specific reclamation process is as follows: First, the difference in resource usage is calculated (5% CPU resources in this example). Then, the list of currently executing non-tightly coupled tasks is traversed, sorted in reverse order according to dynamic priority weights. Starting with the lowest priority task, the reclaimable resource usage of each task is checked. The CPU quota of the task is forcibly reduced via the cgroups interface or the task execution is paused and resources are released via a signal mechanism. The reclaimed resources are then added to a dedicated resource pool to ensure the stable execution of tightly coupled tasks. The entire reclamation process is recorded in the resource scheduling log, including reclamation time, type and quantity of reclaimed resources, and the identifier of the affected task.

[0103] Edge nodes calculate the difference between their total local resources and the dedicated resource pool's usage as the available resources for non-strongly coupled tasks. Assuming an edge node has a total of 8 CPU cores, 16GB of memory, and 1Gbps of network bandwidth, with the dedicated resource pool using 3 CPU cores, 6GB of memory, and 200Mbps of bandwidth, the available resources for non-strongly coupled tasks are 5 CPU cores, 10GB of memory, and 800Mbps of bandwidth. The edge node compares these available resources with the resource usage limits in the constraint boundary, calculating the usage ratio of each resource type. If the constraint boundary stipulates that non-strongly coupled tasks can use up to 70% of the total CPU, 65% of the total memory, and 75% of the total bandwidth, and the current available resources for non-strongly coupled tasks account for 62.5% (5 / 8 CPU cores), 62.5% (10 / 16GB memory), and 80% (800 / 1000Mbps bandwidth) of the total resources respectively, then the bandwidth usage exceeds the constraint limit. The available bandwidth needs to be adjusted to 750Mbps, and then the constraint boundary is determined to be satisfied after the adjustment.

[0104] To generate the optimal execution plan, edge nodes calculate the deviation between the locally measured resource status and the predicted resource status in the centralized scheduling plan. For example, if the centralized scheduling plan predicts CPU utilization of 50%, memory usage of 8GB, and network latency of 10ms, while the measured CPU utilization is 40%, memory usage is 9GB, and network latency is 8ms, then the CPU deviation is -10%, the memory deviation is +12.5% ​​(1GB / 8GB), and the network latency deviation is -20%. Based on the magnitude and direction of these deviations, edge nodes adjust the initial priority weights of loosely coupled tasks. The initial priority weights are pre-defined task priority values ​​within the constraints issued by the cloud scheduling center, ranging from 0 to 100. Higher values ​​indicate higher priority, and are typically calculated based on the task's business importance, deadline urgency, and resource sensitivity.

[0105] The adjustment method is as follows: when the actual resource status is better than the predicted status (e.g., CPU utilization is lower than the predicted value, network latency is lower than the predicted value), the priority weight of compute-intensive and network-intensive tasks is appropriately increased; when the actual resource status is worse than the predicted status (e.g., memory usage is higher than the predicted value), the priority weight of memory-intensive tasks is correspondingly decreased. Specifically, a linear adjustment formula is used: Dynamic priority weight = Initial priority weight × (1 + Deviation percentage × Adjustment coefficient). The adjustment coefficient is set according to the task's sensitivity to different resource types. For example, the adjustment coefficient for compute-intensive tasks is set to 0.8 for CPU resources, 0.3 for memory resources, and 0.2 for network resources; the adjustment coefficient for memory-intensive tasks is set to 0.2 for CPU resources, 0.9 for memory resources, and 0.1 for network resources. Taking a computationally intensive task with an initial priority weight of 60 as an example, its dynamic priority weight = 60 × (1 + (-10%) × 0.8 + 12.5% ​​× 0.3 + (-20%) × 0.2) = 55.05.

[0106] Edge nodes reorder loosely coupled tasks based on calculated dynamic priority weights. The sorting is in descending order, with tasks of higher priority weights placed first, forming a queue of tasks to be executed. Subsequently, edge nodes allocate resource shares to loosely coupled tasks according to the sorting results. The allocation process uses a resource quota allocation algorithm. First, it checks whether the minimum resource requirements declared by the task can be met by currently available resources. If so, it allocates the minimum resource amount and attempts to add additional resources based on the task's resource elasticity coefficient until the maximum resource limit declared by the task is reached or available resources are exhausted. The allocation result is recorded in the resource allocation table. For example, a data analysis task may declare a minimum requirement of 1 CPU core and 2GB of memory, and a maximum usage of 2 CPU cores and 4GB of memory. If currently available resources are sufficient, 2 CPU cores and 4GB of memory are allocated; if available resources are scarce, only 1 CPU core and 2GB of memory are allocated. After allocation, an execution plan is generated, containing information such as the task identifier for each loosely coupled task, the number of allocated CPU cores, memory size, network bandwidth, storage I / O quota, estimated start time, and execution duration. The execution process or container instance of each task is then started by the task scheduler.

[0107] During execution, if an edge node receives a forced resource reclamation instruction (e.g., a reclamation request triggered by excessive resource usage of a tightly coupled task or an emergency resource reclamation command issued by the cloud scheduling center), it will select the lowest priority task in reverse order of dynamic priority weights to suspend execution. The suspension process includes: first, sending a pause signal to the target task process, triggering the task execution state snapshot saving mechanism, serializing the task's memory data, execution context, and intermediate results, and storing them on the local persistent storage device; then, terminating the task process and releasing the CPU quota, memory pages, and network bandwidth quota occupied by the task, and returning the released resources to the available resource pool. If the resources released by a single task are insufficient to meet the reclamation requirements, the next lowest priority task will be selected for suspension, and the above process will be repeated until the cumulative amount of reclaimed resources reaches or exceeds the amount required by the instruction. The information of the suspended task is recorded in the suspended task queue, including the task identifier, pause time, and state snapshot storage path, and execution can be resumed after the resource pressure is relieved.

[0108] Through the above mechanism, edge nodes can dynamically adjust the priority and resource allocation of non-strongly coupled tasks based on the actual local resource status, while ensuring that strongly coupled tasks receive sufficient resource guarantees. This not only guarantees the execution quality of critical tasks but also maximizes the overall utilization efficiency of edge computing resources, achieving refined dynamic scheduling and management of enterprise resources.

[0109] In one optional implementation, the step of sending a replanning request to the cloud and switching to emergency scheduling mode when the deviation between the measured load and the predicted load exceeds the error tolerance range, and simultaneously feeding back the deviation value, duration, and emergency execution status to the cloud, includes:

[0110] Edge nodes continuously collect local measured load, calculate the deviation between the measured load and the predicted load in the centralized scheduling scheme, compare the deviation with the upper and lower limits of the error tolerance range, and determine that the deviation exceeds the error tolerance range when the deviation exceeds the upper limit or falls below the lower limit.

[0111] When an edge node determines that the deviation value exceeds the error tolerance range, it generates a replanning request containing the current deviation value and the edge node identifier and sends it to the cloud scheduling center. At the same time, it starts the emergency scheduling mode, which suspends the reception of new centralized scheduling schemes and reallocates the currently executing tasks according to the local measured resource status.

[0112] Edge nodes record the start and end times when the deviation value exceeds the error tolerance range, calculate the time difference as the duration, and collect the task execution status under emergency scheduling mode as the emergency execution status. The deviation value, duration, and emergency execution status are combined to generate feedback information and uploaded to the cloud scheduling center.

[0113] For example, edge nodes continuously collect local measured load data through a resource monitoring module, with a collection frequency set to once per second. The collected metrics include CPU utilization, memory usage, network bandwidth consumption, and the current task queue length. After collection, this real-time data is compared with the predicted load at the corresponding time point in the centralized scheduling scheme distributed from the cloud. The predicted load data in the centralized scheduling scheme includes the predicted values ​​for each resource dimension within each time window.

[0114] The deviation value is calculated using a weighted comprehensive relative error method. First, the relative deviation for each resource dimension is calculated: CPU deviation rate = (actual CPU utilization - predicted CPU utilization) / predicted CPU utilization × 100%. The calculation methods for memory deviation rate, network deviation rate, and task queue deviation rate are the same. Then, the weights of each dimension are determined based on the resource consumption characteristics of the currently executing task. For example, for compute-intensive tasks, the CPU deviation weight is set to 0.5, memory deviation weight to 0.3, network deviation weight to 0.1, and task queue deviation weight to 0.1. The comprehensive deviation value is obtained by multiplying the deviation rate of each dimension by its corresponding weight and then summing the results. The calculated comprehensive deviation value is then compared with the error tolerance range issued in the centralized scheduling scheme. The error tolerance range consists of an upper limit and a lower limit, for example, an upper limit of +15% and a lower limit of -10%. When the calculated comprehensive deviation value is greater than the upper limit or less than the lower limit, it is determined that the deviation value has exceeded the error tolerance range, and an emergency mechanism needs to be activated.

[0115] When an edge node determines that its deviation value exceeds the error tolerance range, it immediately executes two parallel operations. First, it generates a replanning request data packet, containing the edge node's unique identifier, current timestamp, overall deviation value, specific deviation values ​​for each resource dimension, a list of currently executing strongly coupled tasks, a list of currently executing weakly coupled tasks and their dynamic priority weights, and a snapshot of available resource status. This request is sent to the replanning API interface of the cloud scheduling center via HTTPS. Simultaneously with sending the replanning request, the edge node immediately switches to emergency scheduling mode, setting a flag to prevent receiving and applying new centralized scheduling schemes, thus avoiding scheduling conflicts caused by using outdated prediction data.

[0116] For currently executing tasks, edge nodes handle them differently based on task type. For tightly coupled tasks, since their execution plans are centrally scheduled in the cloud and there are strict dependencies between tasks, their current execution state remains unchanged, but their resource reservation is increased by 10% to cope with load fluctuations, and the additional resources are allocated from the available resource pool of loosely coupled tasks. For loosely coupled tasks, resources are reallocated according to a dynamic priority weight mechanism. The specific steps are as follows: re-collect the current measured resource status, calculate the latest deviation from the predicted resource status in the centralized scheduling plan, and recalculate the dynamic priority weight of all loosely coupled tasks according to the dynamic priority weight calculation formula based on the deviation. The tasks are reordered according to the updated dynamic priority weight, prioritizing the resource needs of high-priority tasks. For low-priority tasks that are ranked lower and whose current resources are insufficient to support their continued execution, the execution state is saved and resources are released according to the pause mechanism.

[0117] When the emergency dispatch mode is activated, the edge node records the moment when the deviation value first exceeds the error tolerance range as the start timestamp, accurate to the millisecond. Changes in the overall deviation value are continuously monitored; when the overall deviation value returns to the error tolerance range and remains stable for more than 30 seconds, this moment is recorded as the end timestamp. The time difference between the two timestamps is calculated to obtain the duration of this abnormal state.

[0118] During the operation of the emergency dispatch mode, the edge nodes continuously collect emergency execution status information, including the percentage comparison between the execution progress of strongly coupled tasks and the expected progress of the centralized dispatch scheme, the completion rate of non-strongly coupled tasks, the average utilization rate of various resources, the average response time of tasks, the number of suspended tasks and their identification list, the number of times resources are forcibly reclaimed, and system stability indicators such as the number of task execution failures and timeouts.

[0119] After the edge nodes complete information collection, they integrate the deviation value, duration, and emergency execution status into structured feedback information. The feedback information is in JSON format and includes fields such as node identifier, anomaly start time stamp, anomaly end time stamp, duration, overall deviation value, arrays of deviation values ​​for each resource dimension, comparison of tightly coupled task execution progress, completion rate of loosely coupled tasks, average resource utilization during the emergency, average task response time, list of suspended tasks, number of resource reclamation attempts, number of failed tasks, number of timeout tasks, and error log summary. The integrated feedback information is proactively pushed to the cloud scheduling center's feedback data receiving API via an HTTPS POST request. To ensure data integrity, the request header includes an MD5 checksum. The cloud verifies the data integrity upon receipt and returns a confirmation response. If network transmission fails, the edge node caches the feedback information in a local queue and retryes every 30 seconds, up to a maximum of 5 times.

[0120] After receiving feedback, the cloud-based scheduling center adjusts the load prediction model parameters based on the actual operational data to optimize future scheduling plans. Once the cloud completes the replanning and issues a new scheduling strategy, the edge nodes verify the effectiveness of the new strategy, then exit emergency scheduling mode and resume normal scheduling processes. Through this mechanism, edge nodes can promptly detect significant deviations between load predictions and actual conditions, ensure stable system operation through local emergency scheduling during cloud replanning, and provide the cloud with detailed operational data for continuous optimization, thus realizing the adaptive and fault-tolerant capabilities of the cloud-edge collaborative scheduling system.

[0121] In one optional implementation, the steps of the cloud scheduling center correcting the load forecasting method based on feedback information and regenerating the scheduling strategy upon receiving a replanning request include:

[0122] The cloud-based dispatch center receives feedback information uploaded by edge nodes, extracts deviation values, duration, and emergency execution status from the feedback information, and archives the feedback information to the historical deviation record of the corresponding edge node according to the edge node identifier.

[0123] The cloud-based dispatch center analyzes the distribution characteristics and duration of deviation values ​​from historical deviation records to identify load fluctuation patterns that cause deviations to exceed the error tolerance range. Based on these load fluctuation patterns, it adjusts the prediction window length and the weighting parameters of periodic change characteristics in the load forecasting method.

[0124] When the cloud scheduling center receives a replanning request, it obtains the tasks currently being executed and the tasks to be scheduled by the edge node, re-predicts the resource load of the edge node according to the adjusted load prediction method, and regenerates a centralized scheduling scheme for strongly coupled task clusters based on the re-predicted resource load, and regenerates constraint boundaries and error tolerance intervals for non-strongly coupled tasks and distributes them to the edge node.

[0125] For example, the cloud-based dispatch center receives feedback data packets uploaded by edge nodes via a feedback data receiving API. Upon receiving the data, it first verifies the MD5 checksum of the data packet to ensure transmission integrity. Then, it parses the JSON-formatted feedback information and extracts key data such as edge node identifier, anomaly start timestamp, anomaly end timestamp, duration, overall deviation value, deviation values ​​for each resource dimension, execution progress of tightly coupled tasks, completion rate of loosely coupled tasks, resource utilization during emergencies, task response time, information on suspended tasks, number of resource reclamation attempts, and system stability indicators.

[0126] After extracting the data, the cloud-based dispatch center queries the historical deviation record table of the edge node in the distributed database based on the edge node identifier. The historical deviation record table is stored using a time-series database, and its structure includes fields such as record ID, node identifier, timestamp, overall deviation value, CPU deviation value, memory deviation value, network deviation value, task queue deviation value, duration, emergency status flag, progress of tightly coupled tasks, completion rate of loosely coupled tasks, and resource utilization. The system retains the historical records for each node for the most recent 90 days; data older than 90 days is archived to cold storage. New feedback information is inserted as a new record into the historical deviation record table, and the last update timestamp of that node is updated.

[0127] The cloud-based dispatch center periodically performs offline analysis of historical deviation records for each edge node. The analysis first calculates the statistical characteristics of deviation values ​​for each resource dimension, including the mean, standard deviation, skewness, and kurtosis. Simultaneously, it calculates the average, median, maximum, and distribution histogram of the duration of the deviation. Based on these statistical characteristics, load fluctuation patterns are identified. Specifically, the K-means clustering algorithm is used to categorize historical deviation records according to the magnitude of the deviation value, duration, and time period. For each cluster center, characteristics are analyzed to identify the fluctuation pattern type. Periodic fluctuations are detected using Fast Fourier Transform (FFT); when the timestamps of records in a cluster exhibit a periodic distribution, it is identified as a periodic fluctuation pattern. Sudden fluctuations are characterized by deviation values ​​significantly higher than other clusters and a short duration. Gradual fluctuations are identified through a linear regression slope significance test, showing a monotonic trend in deviation values ​​over time. Clusters that do not meet the above characteristics are classified as random fluctuation patterns.

[0128] For identified load fluctuation patterns, the load forecasting method parameters are adjusted. The forecast window length is adjusted based on the statistical characteristics of the duration of the deviation. If the average duration exceeds 30 minutes, the forecast window length is increased from 60 minutes to 90 minutes to capture longer-term trends; if the average duration is less than 10 minutes, the forecast window is shortened to 45 minutes to improve sensitivity to short-term fluctuations. After adjustment, the forecast accuracy is evaluated using a validation set; if the accuracy improves, the adjustment is retained. The periodic feature weight parameters are adjusted based on the intensity of periodic fluctuations. When a strong periodic fluctuation pattern is detected, the periodic feature weight is increased; when the periodicity is weak, the periodic feature weight is decreased, and the short-term trend feature weight is increased accordingly. The weight adjustment updates the neural network parameters using the backpropagation algorithm, and the model is fine-tuned using data from the most recent 30 days. For sudden fluctuation patterns, an anomaly detection layer is added after the forecast output, and the anomaly detection threshold is set based on the statistical characteristics of the amplitude of historical sudden fluctuations. For nodes with strong random fluctuations, exponential smoothing filtering is applied to the forecast results to reduce forecast volatility, and the smoothing factor is dynamically adjusted based on the standard deviation of the random fluctuations.

[0129] When the cloud-based scheduling center receives a replanning request from an edge node, it first verifies the request's validity and determines its urgency. It then parses the list of currently executing tasks carried in the request, which includes information such as task identifier, task type, current execution status, consumed resources, and estimated remaining execution time. The cloud-based scheduling center categorizes tasks according to their type: strongly coupled tasks (those with data dependencies or resource sharing relationships) are organized into task clusters. Weakly coupled tasks are considered remaining tasks, which can be autonomously scheduled by the edge node provided that the constraints are met.

[0130] The adjusted load forecasting method is used to re-predict the resource load trend of this edge node within future time windows. The forecast inputs include the node's recent historical load data, resource consumption characteristics of currently executing tasks, and time characteristics. The forecast model outputs predicted values ​​for CPU utilization, memory usage, network bandwidth consumption, and task queue length at each time point within the future time period, forming a predicted load time series.

[0131] Based on the re-predicted resource load time series, a new centralized scheduling scheme is generated for tightly coupled task clusters. The scheduling scheme generation process includes: constructing a directed acyclic graph (DAG) based on the dependencies between tasks in the task cluster; performing topological sorting on the graph; calculating the critical path length of each task as the basic priority; and adjusting the final priority based on the urgency of the task deadlines and the importance of the business. Based on the predicted load time series, resource-sufficient and resource-intensive periods are identified, and high-priority tasks are allocated to time windows during resource-sufficient periods. For each task, its resource requirements are predicted based on its historical execution data and the current input data size. Under the premise of satisfying dependency constraints, the task execution order is determined according to the priority and time window allocation results, generating a task execution sequence. The generated centralized scheduling scheme includes the task execution sequence, the target execution time of each task, the predicted resource requirements of each task, the dependencies between tasks, and an overall execution time estimate.

[0132] For loosely coupled tasks, the cloud scheduling center recalculates the constraint boundaries and error tolerance intervals. The constraint boundaries include resource usage limits, priority ranges, and allowable delay times. The resource usage limit is calculated based on the total resources allocated to each time period after subtracting the reserved resources for loosely coupled tasks from the total resources allocated to the predicted load time series. A certain percentage of these remaining resources is used as the resource usage limit for loosely coupled tasks, with a buffer space. The priority range is set based on the business importance and execution flexibility of the loosely coupled tasks, allowing for an initial priority weight range. The allowable delay time is calculated based on the task's deadline requirements, representing the maximum allowable delay relative to the ideal execution time. The error tolerance interval is calculated based on prediction confidence, standard deviation, and sensitivity coefficient. The prediction model calculates the prediction variance while outputting the predicted value; a smaller variance indicates higher confidence. A narrower error tolerance interval is set for high-confidence predictions. The standard deviation is calculated from the historical deviation records of that node. The sensitivity coefficient is set based on the sensitivity of the currently executing task to resource fluctuations; resource-intensive tasks have larger coefficient values, while fault-tolerant tasks have smaller coefficient values. The upper and lower limits of the tolerance interval are calculated using a combination of the coefficient, standard deviation, and confidence level.

[0133] The cloud-based scheduling center integrates the newly generated centralized scheduling scheme, constraint boundaries and error tolerance intervals, and adjusted load prediction parameters into a scheduling policy update data packet. The data packet uses a serialized format to reduce transmission size and includes the protocol version number, generation timestamp, target node identifier, centralized scheduling scheme object, constraint boundary object, error tolerance interval object, prediction model parameter update flag, and related content. The data packet is sent to the target edge node's scheduling policy receiving API via an encrypted HTTPS channel. To ensure successful delivery, a request-confirmation mechanism is used. After sending the data packet, the cloud waits for a confirmation response from the edge node. If no confirmation is received within a set time, the packet is resent, up to a maximum of several resentments. If the delivery still fails, an exception log is recorded, and the packet is delivered through a backup channel.

[0134] Upon receiving a scheduling policy update data packet, the edge node first verifies the packet's integrity and timeliness. If verification is successful, it extracts the centralized scheduling scheme, constraint boundaries, and error tolerance intervals, replacing the corresponding data structures in memory. If the data packet contains updates to prediction model parameters, these parameters are applied to the local prediction model copy. After the update is complete, the edge node exits emergency scheduling mode, resumes normal scheduling procedures, executes tightly coupled tasks according to the new centralized scheduling scheme, and autonomously schedules loosely coupled tasks according to the new constraint boundaries and error tolerance intervals.

[0135] The cloud-based scheduling center records key information about this replanning in the scheduling history log, including the triggering node, trigger time, reason, replanning time, type and value of adjusted parameters, and distribution status. These logs are used for subsequent analysis of scheduling system performance and continuous algorithm optimization.

[0136] Through the above mechanism, the cloud scheduling center can dynamically adjust the load prediction steps based on real-time feedback from edge nodes, and quickly regenerate optimized scheduling strategies when prediction deviations are detected, thereby realizing closed-loop adaptive optimization of the cloud-edge collaborative scheduling system and continuously improving resource utilization efficiency and task execution reliability.

[0137] A second aspect of the present invention provides an enterprise resource dynamic scheduling and optimization system based on cloud-edge collaboration, comprising:

[0138] The data acquisition module is used by edge nodes to collect enterprise resource status data and upload it to the cloud dispatch center;

[0139] The load prediction module is used by the cloud scheduling center to predict the resource load of each edge node based on resource status data.

[0140] The task classification and scheme generation module is used by the cloud scheduling center to obtain tasks to be scheduled, construct a resource dependency graph based on the cross-node resource dependency relationship between tasks, extract strongly coupled task clusters by identifying strongly connected components, generate a centralized scheduling scheme for strongly coupled task clusters based on the resource load prediction results, generate constraint boundaries for non-strongly coupled tasks, and set an error tolerance range for the constraint boundaries based on the resource load prediction results.

[0141] The hybrid execution and anomaly detection module is used for edge nodes to execute corresponding tasks according to the centralized scheduling scheme. For non-strongly coupled tasks, it can autonomously solve the execution scheme based on the local measured resource status while satisfying the constraint boundary. When the deviation between the measured load and the predicted load exceeds the error tolerance range, it sends a replanning request to the cloud and switches to emergency scheduling mode. At the same time, it feeds back the deviation value, duration and emergency execution status to the cloud.

[0142] The adaptive optimization module is used by the cloud scheduling center to correct the load prediction method based on feedback information and regenerate the scheduling strategy when a replanning request is received.

[0143] A third aspect of the present invention provides an electronic device, comprising:

[0144] processor;

[0145] Memory used to store processor-executable instructions;

[0146] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0147] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0148] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for dynamic scheduling and optimization of enterprise resources based on cloud-edge collaboration, characterized in that, include: Edge nodes collect enterprise resource status data and upload it to the cloud-based dispatch center; The cloud-based scheduling center predicts the resource load of each edge node based on resource status data; The cloud-based scheduling center obtains tasks to be scheduled, constructs a resource dependency graph based on cross-node resource dependencies between tasks, and extracts strongly coupled task clusters by identifying strongly connected components. Specifically, this includes: the cloud-based scheduling center obtaining a set of tasks to be scheduled, parsing and storing the resource requirements and execution node constraints of each task; traversing the set of tasks to be scheduled, identifying cross-node resource dependencies between tasks, treating each task as a node, and the cross-node resource dependencies as directed edges to construct a resource dependency graph; performing a strongly connected component identification operation on the resource dependency graph, marking a set of nodes in the resource dependency graph with a bidirectional reachable path between any two nodes as a strongly connected component, and extracting strongly coupled task clusters from the strongly connected components. Strongly connected components containing cross-node resource dependencies are selected, and the tasks to be scheduled corresponding to the nodes in the selected strongly connected components are extracted as strongly coupled task clusters. Tasks in the set of tasks to be scheduled other than strongly coupled task clusters are marked as non-strongly coupled tasks. A centralized scheduling scheme is generated for the strongly coupled task clusters based on the resource load prediction results, and constraint boundaries are generated for the non-strongly coupled tasks. An error tolerance range is set for the constraint boundaries based on the resource load prediction results. Specifically, for non-strongly coupled tasks, the peak resource occupancy of each edge node is calculated based on the resource load prediction results as the upper limit constraint of resource occupancy. The task priority order constraint is determined based on the dependency relationship between non-strongly coupled tasks, and the constraint boundaries are generated by combining them. Edge nodes execute corresponding tasks according to the centralized scheduling scheme. For non-strongly coupled tasks, they autonomously solve the execution scheme based on the local measured resource status while satisfying the constraint boundary. When the deviation between the measured load and the predicted load exceeds the error tolerance range, a replanning request is sent to the cloud and the system switches to emergency scheduling mode. At the same time, the deviation value, duration, and emergency execution status are fed back to the cloud. In emergency scheduling mode, receiving new centralized scheduling schemes is suspended, and the currently executing tasks are reallocated according to the local measured resource status. The cloud-based scheduling center corrects the load prediction method based on feedback information and regenerates the scheduling strategy when it receives a replanning request.

2. The method according to claim 1, characterized in that, The steps by which the cloud-based scheduling center predicts the resource load of each edge node based on resource status data include: Edge nodes collect enterprise resource status data and upload it to the cloud scheduling center. After receiving the enterprise resource status data uploaded by each edge node, the cloud scheduling center groups the enterprise resource status data according to the edge node identifier, and sorts each group of enterprise resource status data by timestamp to construct a time-series data sequence. The periodic variation characteristics of resource load are extracted from the time-series data sequence. The prediction window length is determined based on the periodic variation characteristics. The resource load prediction results of each edge node are generated based on the time-series data sequence within the prediction window length.

3. The method according to claim 1, characterized in that, The steps of generating a centralized scheduling scheme for strongly coupled task clusters based on resource load prediction results, generating constraint boundaries for non-strongly coupled tasks, and setting error tolerance intervals for the constraint boundaries based on the resource load prediction results include: For each task in a strongly coupled task cluster, the available resources are calculated based on the resource requirements of the task and the predicted load values ​​of each edge node. Edge nodes whose available resources meet the resource requirements are selected as candidate execution nodes. The edge node with the lowest predicted load value is selected from the candidate execution nodes as the target execution node. The target execution time is determined based on the predicted load value of the target execution node. The target execution node and the target execution time are combined to generate a centralized scheduling scheme. Extract the fluctuation range characteristics of the predicted load value from the resource load prediction results, calculate the upper and lower limits of the error tolerance interval based on the fluctuation range characteristics, and set them as the error tolerance interval of the constraint boundary.

4. The method according to claim 1, characterized in that, Edge nodes execute corresponding tasks according to the centralized scheduling scheme. For non-strongly coupled tasks, the steps for autonomously solving the execution scheme based on the local measured resource status while satisfying the constraint boundaries include: Edge nodes collect the measured resource status before the target execution time arrives; calculate the resource reservation amount based on the predicted resource demand in the centralized scheduling scheme and the measured resource status before execution, and establish a dedicated resource pool for strongly coupled tasks. During the execution process, the actual resource usage is continuously collected. When the actual resource usage exceeds the resource reservation amount, the difference in resources occupied by non-strongly coupled tasks is forcibly reclaimed and supplemented to the dedicated resource pool. Edge nodes calculate the difference between their local total resources and the dedicated resource pool usage as the available resources for non-strongly coupled tasks. They then compare the available resources for non-strongly coupled tasks with the upper limit of resource usage in the constraint boundary. If the upper limit of resource usage is not exceeded, the constraint boundary is considered satisfied. The deviation between the local measured resource status and the predicted resource status in the centralized scheduling scheme is calculated. The initial priority weight is adjusted according to the magnitude and direction of the deviation to obtain the dynamic priority weight. The non-strongly coupled tasks are reordered according to the dynamic priority weight, and resource shares are allocated sequentially to generate an execution plan and start execution. When an edge node receives a forced resource reclamation command, it selects the lowest priority task in reverse order of dynamic priority weights to suspend execution and release resources.

5. The method according to claim 1, characterized in that, When the deviation between the measured load and the predicted load exceeds the error tolerance range, a replanning request is sent to the cloud and the system switches to emergency scheduling mode. Simultaneously, the deviation value, duration, and emergency execution status are fed back to the cloud. The steps include: Edge nodes continuously collect local measured load, calculate the deviation between the measured load and the predicted load in the centralized scheduling scheme, compare the deviation with the upper and lower limits of the error tolerance range, and determine that the deviation exceeds the error tolerance range when the deviation exceeds the upper limit or falls below the lower limit. When an edge node determines that the deviation value exceeds the error tolerance range, it generates a replanning request containing the current deviation value and the edge node identifier and sends it to the cloud scheduling center. At the same time, it starts the emergency scheduling mode, which suspends the reception of new centralized scheduling schemes and reallocates the currently executing tasks according to the local measured resource status. Edge nodes record the start and end times when the deviation value exceeds the error tolerance range, calculate the time difference as the duration, and collect the task execution status under emergency scheduling mode as the emergency execution status. The deviation value, duration, and emergency execution status are combined to generate feedback information and uploaded to the cloud scheduling center.

6. The method according to claim 1, characterized in that, The steps by which the cloud-based scheduling center corrects the load forecasting method based on feedback information and regenerates the scheduling strategy upon receiving a replanning request include: The cloud-based dispatch center receives feedback information uploaded by edge nodes, extracts deviation values, duration, and emergency execution status from the feedback information, and archives the feedback information to the historical deviation record of the corresponding edge node according to the edge node identifier. The cloud-based dispatch center analyzes the distribution characteristics and duration of deviation values ​​from historical deviation records to identify load fluctuation patterns that cause deviations to exceed the error tolerance range. Based on these load fluctuation patterns, it adjusts the prediction window length and the weighting parameters of periodic change characteristics in the load forecasting method. When the cloud scheduling center receives a replanning request, it obtains the tasks currently being executed and the tasks to be scheduled by the edge node, re-predicts the resource load of the edge node according to the adjusted load prediction method, and regenerates a centralized scheduling scheme for strongly coupled task clusters based on the re-predicted resource load, and regenerates constraint boundaries and error tolerance intervals for non-strongly coupled tasks and distributes them to the edge node.

7. A cloud-edge collaborative enterprise resource dynamic scheduling and optimization system, used to implement the method of any one of claims 1-6, characterized in that, include: The data acquisition module is used by edge nodes to collect enterprise resource status data and upload it to the cloud dispatch center; The load prediction module is used by the cloud scheduling center to predict the resource load of each edge node based on resource status data. The task classification and scheme generation module is used by the cloud scheduling center to obtain tasks to be scheduled, construct a resource dependency graph based on the cross-node resource dependency relationship between tasks, extract strongly coupled task clusters by identifying strongly connected components, generate a centralized scheduling scheme for strongly coupled task clusters based on the resource load prediction results, generate constraint boundaries for non-strongly coupled tasks, and set an error tolerance range for the constraint boundaries based on the resource load prediction results. The hybrid execution and anomaly detection module is used for edge nodes to execute corresponding tasks according to the centralized scheduling scheme. For non-strongly coupled tasks, it can autonomously solve the execution scheme based on the local measured resource status while satisfying the constraint boundary. When the deviation between the measured load and the predicted load exceeds the error tolerance range, it sends a replanning request to the cloud and switches to emergency scheduling mode. At the same time, it feeds back the deviation value, duration and emergency execution status to the cloud. The adaptive optimization module is used by the cloud scheduling center to correct the load prediction method based on feedback information and regenerate the scheduling strategy when a replanning request is received.

8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Edge cloud resource collaborative scheduling method, cloud management platform and edge cloud node

    CN120675958A

  • Heterogeneous computing power scheduling optimization method based on cloud edge collaborative architecture

    CN121396990A