Cloud edge end collaborative task scheduling system for heterogeneous computing power pooling
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-16
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]为此,本发明提供了一种用于异构算力池化的云边端协同任务调度系统,用以克服现有技术中未考虑在多节点分布式协同任务调度管控场景下,因云节点、边节点及端节点间的感知与决策能力割裂、缺乏动态协同,导致全局信息利用不足,难以平衡任务执行实时性与节点可靠性,致使资源浪费、任务执行效率低的问题
[0015] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling. The system uses a collaborative scheduling module to schedule target tasks to cloud nodes, edge nodes, or device nodes for response. A data acquisition module collects decision-making collaboration characteristic information and load characteristic information of target nodes within historical periods. A data processing module analyzes the decision-making collaboration characteristic value based on the decision-making collaboration characteristic information and the load characteristic value based on the load characteristic information. A resource pool collaborative perception module determines whether the matching degree of the target node scheduling decision meets the standard based on the decision-making collaboration characteristic value. When the matching degree meets the standard, the system determines whether the stability of the target node is abnormal based on the load characteristic value. An adaptive optimization module determines the cause of the stability anomaly and the corresponding handling measures based on the stability fluctuation index. This overcomes the problems of existing technologies in multi-node distributed collaborative task scheduling and management scenarios, where the fragmented perception and decision-making capabilities and lack of dynamic collaboration among cloud nodes, edge nodes, and device nodes lead to insufficient utilization of global information, difficulty in balancing task execution real-time performance and node reliability, resulting in resource waste and low task execution efficiency.
Smart Images

Figure CN122554528A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud-edge-device collaborative scheduling technology, and in particular to a cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling. Background Technology
[0002] With the rapid development of heterogeneous computing power pooling technology and IoT technology, cloud-edge-device collaborative task scheduling systems have gradually taken shape and are widely used in fields such as multi-scenario perception data processing in smart parks. In smart park scenarios, common perception data processing tasks such as environmental monitoring, security video analysis, and equipment status inspection within the park need to support flexible execution and dynamic scheduling of the same task on any one of the three sides: the cloud big data center, the park edge computing nodes, and the terminal perception devices. Although existing technologies have adopted intelligent collaborative collection and judgment mechanisms, they still have significant shortcomings. They lack correlation verification between scheduling decision information and load information, are susceptible to instantaneous interference leading to misjudgments, and lack targeted adaptive optimization mechanisms. This can easily lead to unnecessary task migration and resource waste, and it is difficult to balance the real-time performance of task execution and node stability. There is an urgent need for a technical solution that can achieve feature correlation verification, accurate anomaly judgment, and adaptive optimization based on the real-time computing power load and network transmission status of the cloud, edge, and device sides.
[0003] Chinese Patent Publication No. CN121396990A discloses a heterogeneous computing power scheduling optimization method based on a cloud-edge collaborative architecture, comprising the following steps: First, acquiring real-time computing power status data of all available computing nodes in the cloud-edge collaborative architecture, including computing core utilization, memory usage, network bandwidth usage, and task queue length; Second, classifying computing nodes into heterogeneous types based on the real-time computing power status data to generate a three-layer computing power resource pool containing cloud computing nodes, edge computing nodes, and terminal computing nodes; Third, extracting task computing characteristics for the current set of tasks to be scheduled, including computational intensity, data dependency, and real-time requirements; Fourth, constructing an initial task allocation scheme based on the matching relationship between task computing characteristics and the three-layer computing power resource pool; Fifth, iteratively optimizing the initial task allocation scheme using a dynamic load balancing strategy to generate final task scheduling instructions; Sixth, distributing the final task scheduling instructions to the corresponding computing nodes for execution and continuously monitoring changes in computing power status during task execution. This method only obtains the real-time computing power status data of the computing nodes, extracts the task computing features, and schedules tasks based on the matching relationship between the two. It does not verify the scheduling decision with the node load, nor does it establish a feature correlation analysis and quantitative judgment system between the two. It cannot achieve the collaborative judgment of the rationality of the scheduling decision and the stability of the node load, nor can it identify the cause of the node's abnormal operation and formulate targeted adaptive optimization strategies based on the correlation between the two. Summary of the Invention
[0004] To address this, the present invention provides a cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling, which overcomes the problems in the prior art that do not consider the fragmented perception and decision-making capabilities and lack of dynamic collaboration among cloud nodes, edge nodes and device nodes in multi-node distributed collaborative task scheduling and management scenarios, resulting in insufficient utilization of global information, difficulty in balancing task execution real-time performance and node reliability, and thus resource waste and low task execution efficiency.
[0005] To achieve the above objectives, the present invention provides a cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling, comprising: The collaborative scheduling module is used to schedule target execution tasks to cloud nodes, edge nodes, or end nodes for response; The data acquisition module is used to collect decision-making collaboration characteristic information and load characteristic information of target nodes within a historical period; The data processing module is used to analyze the decision collaboration feature representation value based on the decision collaboration feature information and to analyze the load feature representation value based on the load feature information. The resource pool collaborative perception module is used to determine whether the matching degree of the target node scheduling decision meets the standard based on the decision collaborative feature characterization value, and to determine whether the stability of the target node is abnormal based on the load feature characterization value. An adaptive optimization module is used to determine the cause of stability anomalies based on the stability fluctuation index, and to match corresponding processing measures based on the cause. The corresponding processing measures include enabling a smoothing filter to filter the decision collaboration feature characterization value, and determining to issue node load optimization scheduling instructions to the collaborative scheduling module. The decision-making collaboration feature information includes the queuing time of the task and the execution timeout rate of the task; the load feature information includes the CPU utilization and the network round-trip latency of adjacent nodes.
[0006] Furthermore, the data processing module is used to analyze the decision collaboration feature representation value based on the decision collaboration feature information, wherein, The decision collaboration feature value is obtained by summing the first collaboration factor and the second collaboration factor according to a predetermined first weight ratio. The first coordination factor is determined based on the ratio of a predetermined queuing time threshold to the queuing time. The second collaboration factor is determined based on the ratio of a predetermined execution timeout rate threshold to the execution timeout rate.
[0007] Furthermore, the resource pool collaborative perception module is used to determine whether the matching degree of the target node scheduling decision meets the standard based on the decision collaborative feature representation value, including: If the value of the decision collaboration feature is less than or equal to the predetermined decision collaboration feature threshold, then the matching degree is determined to be inconsistent with the standard. If the value of the decision collaboration feature is greater than the predetermined decision collaboration feature threshold, then the matching degree is determined to meet the standard.
[0008] Furthermore, the data processing module is used to analyze load characteristic representation values based on the load characteristic information, wherein, The load characteristic value is obtained by summing the first load factor and the second load factor according to a predetermined second weighting ratio. The first load factor is determined based on the ratio of a predetermined utilization threshold to the utilization rate; The second load factor is determined based on the ratio of a predetermined network round-trip time threshold to the network round-trip time.
[0009] Furthermore, the resource pool collaborative sensing module is used to determine whether the stability of the target node is abnormal based on the load characteristic representation value, including: If the load characteristic value is less than or equal to the predetermined load characteristic threshold, then a stability anomaly is determined. If the load characteristic value is greater than the predetermined load characteristic threshold, then the stability is determined to be normal.
[0010] Furthermore, the adaptive optimization module is used to analyze the stability fluctuation index, wherein the stability fluctuation index is determined based on the difference between a predetermined load characteristic characterization threshold and the load characteristic characterization value.
[0011] Furthermore, the adaptive optimization module is used to determine the cause of stability anomalies based on the stability fluctuation index, including: If the stability fluctuation index is less than or equal to the predetermined stability fluctuation qualification threshold, it is determined that the decision collaboration feature information is subject to instantaneous interference. If the stability fluctuation index is greater than the predetermined stability fluctuation acceptable threshold, then the node load is determined to be excessive.
[0012] Furthermore, the adaptive optimization module is used to determine corresponding handling measures based on the cause of stability anomalies, including: If the cause is that the decision collaboration feature information is subject to instantaneous interference, then a smoothing filter is activated to filter the decision collaboration feature representation value. If the cause is excessive node load, then a node load optimization scheduling instruction will be issued to the collaborative scheduling module.
[0013] Furthermore, the adaptive optimization module is used to determine whether to enable a smoothing filter to filter the decision collaboration feature representation value, wherein the decision collaboration feature representation value is positively correlated with the stability fluctuation index.
[0014] Furthermore, the adaptive optimization module is used to issue node load optimization scheduling instructions to the collaborative scheduling module, wherein the execution intensity of the scheduling instructions is positively correlated with the stability fluctuation index.
[0015] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling. The system uses a collaborative scheduling module to schedule target tasks to cloud nodes, edge nodes, or device nodes for response. A data acquisition module collects decision-making collaboration characteristic information and load characteristic information of target nodes within historical periods. A data processing module analyzes the decision-making collaboration characteristic value based on the decision-making collaboration characteristic information and the load characteristic value based on the load characteristic information. A resource pool collaborative perception module determines whether the matching degree of the target node scheduling decision meets the standard based on the decision-making collaboration characteristic value. When the matching degree meets the standard, the system determines whether the stability of the target node is abnormal based on the load characteristic value. An adaptive optimization module determines the cause of the stability anomaly and the corresponding handling measures based on the stability fluctuation index. This overcomes the problems of existing technologies in multi-node distributed collaborative task scheduling and management scenarios, where the fragmented perception and decision-making capabilities and lack of dynamic collaboration among cloud nodes, edge nodes, and device nodes lead to insufficient utilization of global information, difficulty in balancing task execution real-time performance and node reliability, resulting in resource waste and low task execution efficiency.
[0016] In particular, this invention virtualizes and integrates different types of computing resources distributed across the cloud, edge, and terminal devices into a unified resource pool through a collaborative scheduling module. Then, based on the computing requirements of the target task and the real-time status of the resources, the task is scheduled to a matching cloud node, edge node, or terminal node for response. This makes computing tasks more likely to be executed on nodes near the data generation location or nodes with lower loads, reducing the resource consumption of remote data transmission and thus improving the overall throughput and global resource utilization of the heterogeneous computing pool.
[0017] In particular, this invention obtains decision-making collaboration feature information by collecting the queuing time and execution timeout rate of the target node's tasks within a historical period through the data acquisition module, and obtains load feature information by collecting the CPU utilization of the target node and the network round-trip latency of adjacent nodes within a historical period. The decision-making collaboration feature information and the load feature information are then normalized by the data processing module to obtain decision-making collaboration feature representation values and load feature representation values, thereby improving the diagnostic efficiency of the resource pool collaboration perception module for collaborative task scheduling.
[0018] In particular, the present invention uses a resource pool collaborative perception module to determine whether the matching degree of the target node scheduling decision meets the standard based on the comparison result of the decision collaborative feature representation value and the predetermined decision collaborative feature representation threshold. In response to the matching degree meeting the standard, the invention further determines whether the stability of the target node is abnormal based on the comparison result of the load feature representation value and the predetermined load feature representation threshold. This can distinguish between the instantaneous performance fluctuations of the target node and long-term load anomalies, thereby improving the system's accuracy in identifying abnormal conditions.
[0019] In particular, the present invention uses an adaptive optimization module to calculate the difference between a predetermined load characteristic representation threshold and the load characteristic representation value as a stability fluctuation index. Based on the magnitude of the stability fluctuation index, the cause of the target node's stability anomaly is determined, and differentiated processing measures are matched: a smoothing filter is enabled to filter the decision collaboration characteristic representation value, and a node load optimization scheduling instruction is issued to the collaborative scheduling module, thereby improving the long-term stability of the system. Attached Figure Description
[0020] Figure 1 This is a structural block diagram of a cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling, as described in an embodiment of the present invention. Figure 2 This invention provides a logical decision diagram for determining whether the matching degree of the target node scheduling decision conforms to the standard based on the decision collaboration feature characterization value. Figure 3 This invention provides a logical decision diagram for determining whether the stability of a target node is abnormal based on load characteristic values. Figure 4 This invention provides a logical decision diagram for determining the causes of stability anomalies and their corresponding handling measures based on the stability fluctuation index. Detailed Implementation
[0021] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0022] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0023] It should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the term "connected" should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral connection; it can refer to a mechanical connection or an electrical connection; it can refer to a direct connection or an indirect connection through an intermediate medium; it can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0024] Please see Figure 1 The diagram shown is a structural block diagram of a cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling, according to an embodiment of the present invention. The present invention provides a cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling, comprising: The collaborative scheduling module is bidirectionally connected to cloud nodes, edge nodes, and end nodes to schedule target execution tasks to cloud nodes, edge nodes, or end nodes for response. The data acquisition module is connected to the collaborative scheduling module and is also connected to the cloud node, edge node, and end node one-way data acquisition to collect decision collaboration feature information and load feature information of the target node within a historical period. The data processing module is connected to the data acquisition module and has no direct connection to the cloud node, edge node, or end node. It is used to analyze the decision collaboration feature representation value based on the decision collaboration feature information and to analyze the load feature representation value based on the load feature information. The resource pool collaborative perception module is connected to the data processing module, but has no direct connection with cloud nodes, edge nodes, or end nodes. It is used to determine whether the matching degree of the target node scheduling decision meets the standard based on the decision collaboration feature characterization value, and to determine whether the stability of the target node is abnormal based on the load feature characterization value. An adaptive optimization module is connected to the data processing module and the resource pool collaborative perception module respectively, and has no direct connection with cloud nodes, edge nodes, and end nodes. It is used to determine the cause of stability anomalies based on the stability fluctuation index, and match corresponding processing measures based on the cause. The corresponding processing measures include enabling a smoothing filter to filter the decision collaborative feature representation value, and issuing node load optimization scheduling instructions to the collaborative scheduling module. The decision-making collaboration feature information includes the queuing time of the task and the execution timeout rate of the task, while the load feature information includes the CPU utilization and the network round-trip latency of adjacent nodes.
[0025] Specifically, embodiments of the present invention provide implementation steps for a cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling, including: Step S1: Collect decision-making collaboration characteristic information of target nodes within the historical period; Step S2: Analyze the decision collaboration feature representation value based on the aforementioned decision collaboration feature information; Step S3: Determine whether the matching degree of the target node scheduling decision meets the standard based on the decision collaboration feature characterization value; Step S4: In response to the matching degree meeting the standard, collect the load characteristic information of the target node within the historical period; Step S5: Analyze the load characteristic representation value based on the load characteristic information; Step S6: Determine whether the stability of the target node is abnormal based on the load characteristic characterization value; Step S7: Calculate the stability fluctuation index in response to stability anomalies; Step S8: Determine the cause of stability anomalies and corresponding handling measures based on the stability fluctuation index: enable a smoothing filter to filter the decision collaboration feature characterization value, and issue node load optimization scheduling instructions to the collaborative scheduling module; The decision-making collaboration feature information includes the queuing time of the task and the execution timeout rate of the task, while the load feature information includes the CPU utilization and the network round-trip latency of adjacent nodes.
[0026] As is understandable, task queuing time refers to the time from when a task is assigned to the target node to when it begins execution, collected by the node's built-in task scheduling timer. It reflects the node's real-time busyness; the longer the queue, the longer the wait. The unit is milliseconds.
[0027] It is understandable that the task execution timeout rate refers to the proportion of tasks with an execution time exceeding 50 milliseconds on the target node within a single collection cycle, out of the total number of tasks. It reflects the recent performance of the node, and insufficient processing capacity will lead to task execution timeouts.
[0028] As is understandable, CPU utilization refers to the percentage of CPU usage on a target node relative to the total computing power within a unit data collection cycle, reflecting the node's load. High CPU utilization can lead to an increase in task execution timeout rates.
[0029] It is understandable that the network round-trip latency between adjacent nodes refers to the average network round-trip latency between the target node and its adjacent nodes within a unit collection period. It reflects the network congestion between nodes and is measured in milliseconds. High network round-trip latency can cause task data to be unable to be transmitted in a timely manner, leading to increased task queuing time and execution timeout rate.
[0030] In this embodiment, the single acquisition period is preset, and the preferred single acquisition period is 100 milliseconds.
[0031] This invention provides a cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling. A collaborative scheduling module schedules target tasks to cloud nodes, edge nodes, or device nodes for response. A data acquisition module collects decision-making collaboration characteristics and load characteristics of target nodes over historical periods. A data processing module analyzes decision-making collaboration characteristic values based on the decision-making collaboration characteristic information and load characteristic values based on the load characteristic information. A resource pool collaborative perception module determines whether the matching degree of the target node scheduling decision meets the standard based on the decision-making collaboration characteristic values and whether the target node's stability is abnormal based on the load characteristic values. An adaptive optimization module determines the cause of stability anomalies and corresponding handling measures based on a stability fluctuation index. This invention improves the execution efficiency of collaboratively scheduled tasks across cloud nodes, edge nodes, and device nodes.
[0032] Specifically, the data processing module is used to analyze the decision collaboration feature representation value based on the decision collaboration feature information, wherein, The decision collaboration feature value is obtained by summing the first collaboration factor and the second collaboration factor according to a predetermined first weight ratio. The first coordination factor is determined based on the ratio of a predetermined queuing time threshold to the queuing time. The second collaboration factor is determined based on the ratio of a predetermined execution timeout rate threshold to the execution timeout rate.
[0033] In this embodiment, the predetermined queuing time threshold is preset. Queuing time samples from four historical collection periods are predetermined, and the predetermined queuing time threshold is determined based on the average value of the queuing time samples. The threshold is determined within the range of [45 milliseconds, 52 milliseconds]. Based on the requirement of no task congestion under the historical normal operation of the node, an optimal value that meets 95% of the upper limit of the measured value is selected. In this embodiment, the predetermined queuing time threshold is preferably 49.4 milliseconds.
[0034] In this embodiment, the predetermined execution timeout rate threshold is preset. Specifically, execution timeout rate samples within four historical collection periods are predetermined, and the predetermined execution timeout rate threshold is determined based on the average value of the execution timeout rate samples. The threshold is determined within the range of [5.8%, 7.0%]. Based on the node performance compliance requirements, an optimal value that meets 95% of the upper limit of the measured value is selected. In this embodiment, the predetermined execution timeout rate threshold is preferably 6.7%.
[0035] In the embodiments, the predetermined first weight ratio is preset, and the preferred predetermined first weight ratio is 2:3, that is, the decision synergy feature representation value is equal to the sum of 0.4 times the first synergy factor and 0.6 times the second synergy factor.
[0036] The embodiments of the present invention calculate the precise quantitative node decision collaboration feature characterization value through a predetermined first weight ratio, and set the queuing time threshold and execution timeout rate threshold based on the average value of decision collaboration feature information in the historical collection period. This can set the judgment benchmark in line with the actual operating status of the node, thereby improving the accuracy of node scheduling decision matching degree judgment.
[0037] Please see Figure 2 As shown, this is a logical judgment diagram for determining whether the matching degree of the target node scheduling decision meets the standard based on the decision collaboration feature characterization value in an embodiment of the present invention. The resource pool collaboration perception module of the present invention is used to determine whether the matching degree of the target node scheduling decision meets the standard based on the decision collaboration feature characterization value, including: If the value of the decision collaboration feature is less than or equal to the predetermined decision collaboration feature threshold, then the matching degree is determined to be inconsistent with the standard. If the value of the decision collaboration feature is greater than the predetermined decision collaboration feature threshold, then the matching degree is determined to meet the standard.
[0038] In this embodiment, the predetermined decision collaboration feature representation threshold is preset. Specifically, the average value of the decision collaboration feature representation values over eight historical collection periods is predetermined. The predetermined decision collaboration feature representation threshold is determined based on the average value of the decision collaboration feature representation values and is determined within the range [0.96, 1.14]. Based on the actual judgment requirements of the scheduling decision matching degree of cloud nodes, edge nodes, and end nodes, an optimal value that meets 90% of the measured value is selected. In this embodiment, the predetermined decision collaboration feature representation threshold is preferably 1.03.
[0039] Understandably, when the matching degree is determined to be inconsistent with the standard, the resource pool collaborative perception module will immediately generate a matching degree anomaly judgment result and synchronize it to the collaborative scheduling module. After receiving the result, the collaborative scheduling module will perform the following scheduling adjustment actions: First, it will suspend the issuance of new execution tasks to the target node to avoid further aggravating the node scheduling mismatch problem by adding new tasks; Second, it will identify the priority of the incomplete tasks being executed on the target node, and migrate high-priority tasks to cloud nodes, edge nodes, or end nodes in the computing power pool whose scheduling decision matching degree meets the standard and whose load characteristic characterization value is less than the predetermined load characteristic characterization threshold, while low-priority tasks will be temporarily stored in the computing power pool task buffer queue; Finally, it will mark the target node as a node to be optimized for scheduling, and synchronize information such as node identifier, matching degree anomaly value, and current task execution status to the data acquisition module. The data acquisition module will increase the collection frequency of decision collaboration characteristic information for the target node from the original 100 milliseconds / time to 50 milliseconds / time, and continuously collect the queuing time and execution timeout rate data of the target node and transmit them to the data processing module.
[0040] This invention, by selecting a threshold for decision-coordination feature representation based on the average value of decision-coordination feature representation values from historical collection periods, can clearly define the qualified and unqualified intervals of node scheduling decisions, reduce the deviation of scheduling decision judgment, and thus improve the consistency of scheduling decision matching degree judgment.
[0041] Specifically, the data processing module is used to analyze load characteristic representation values based on the load characteristic information, wherein, The load characteristic value is obtained by summing the first load factor and the second load factor according to a predetermined second weighting ratio. The first load factor is determined based on the ratio of a predetermined utilization threshold to the utilization rate; The second load factor is determined based on the ratio of a predetermined network round-trip time threshold to the network round-trip time.
[0042] In this embodiment, the predetermined utilization threshold is preset. Specifically, utilization samples from four historical collection periods are predetermined, and the predetermined utilization threshold is determined based on the average value of the utilization samples. The threshold is determined within the range of [75%, 92%]. Based on the actual operating requirements of cloud nodes, edge nodes, and end nodes without CPU overload, an optimal value that meets the upper limit of the measured value of 90% is selected. In this embodiment, the predetermined utilization threshold is preferably 82.8%.
[0043] In this embodiment, the predetermined network round-trip latency threshold is preset. Specifically, network round-trip latency samples from four historical collection periods are predetermined, and the predetermined network round-trip latency threshold is determined based on the average value of the network round-trip latency samples. The threshold is determined within the range of [18 milliseconds, 23 milliseconds]. Based on the actual scheduling requirements for non-blocking data transmission between adjacent nodes of cloud nodes, edge nodes, or end nodes, an optimal value that meets 90% of the measured value is selected. In this embodiment, the predetermined network round-trip latency threshold is preferably 20.7 milliseconds.
[0044] In the embodiments, the predetermined second weighting ratio is preset, and the preferred embodiment is a predetermined second weighting ratio of 3:7, that is, the load characteristic representation value is equal to the sum of 0.3 times the first load factor and 0.7 times the second load factor.
[0045] The embodiments of the present invention calculate the load characteristic value by a predetermined second weight ratio, thereby quantifying the actual load status of cloud nodes, edge nodes and end nodes. Based on the sample mean of historical collection periods, CPU utilization threshold and network round-trip latency threshold are set, which can fit the actual operating status of node hardware and network, thereby improving the accuracy of node stability determination.
[0046] Please see Figure 3As shown, this is a logic diagram for determining whether the stability of a target node is abnormal based on load characteristic values in an embodiment of the present invention. The resource pool collaborative sensing module of the present invention is used to determine whether the stability of a target node is abnormal based on the load characteristic values, including: If the load characteristic value is less than or equal to the predetermined load characteristic threshold, then a stability anomaly is determined. If the load characteristic value is greater than the predetermined load characteristic threshold, then the stability is determined to be normal.
[0047] In this embodiment, the predetermined load characteristic representation threshold is preset. The average value of the load characteristic representation values over 8 historical collection periods is predetermined. The predetermined load characteristic representation threshold is determined based on the average value of the load characteristic representation values and is determined within the range [0.90, 1.14]. The preferred value that meets 95% of the measured value is selected according to the actual business requirements for determining the stability of cloud nodes, edge nodes, and end nodes. In this embodiment, the predetermined load characteristic representation threshold is preferably 1.08.
[0048] This invention determines node stability by setting a load characteristic representation threshold, accurately identifying the load operation status of cloud nodes, edge nodes, and end nodes, and clearly defining the boundary between node stability and abnormality. This allows the system to quickly determine the node operation status based on the load characteristic representation threshold, thereby improving the efficiency of stability determination.
[0049] Specifically, the adaptive optimization module is used to analyze the stability fluctuation index, wherein the stability fluctuation index is determined based on the difference between a predetermined load characteristic characterization threshold and the load characteristic characterization value.
[0050] Understandably, the stability volatility index is a positive value, and its magnitude reflects the gap between the actual performance and the expected standard. A larger stability volatility index indicates a greater deviation between the actual performance and the expected value, and a higher severity of anomalies; conversely, a smaller stability volatility index indicates a lower severity of anomalies.
[0051] This invention quantifies the degree of node stability anomaly by calculating the difference between the load characteristic characterization threshold and the actual load characteristic characterization value as a stability fluctuation index. By setting the stability fluctuation index to a positive value and using its magnitude to reflect the deviation between the actual effect and the expected standard, the severity of node anomaly problems can be accurately distinguished, thereby improving the accuracy of anomaly cause determination.
[0052] Please see Figure 4 As shown, this is a logic decision diagram for determining the causes of stability anomalies and corresponding handling measures based on the stability fluctuation index in an embodiment of the present invention. The adaptive optimization module of the present invention is used to determine the causes of stability anomalies based on the stability fluctuation index, including: If the stability fluctuation index is less than or equal to the predetermined stability fluctuation qualification threshold, it is determined that the decision collaboration feature information is subject to instantaneous interference. If the stability fluctuation index is greater than the predetermined stability fluctuation acceptable threshold, then the node load is determined to be excessive.
[0053] In this embodiment, the predetermined stability fluctuation qualification threshold is preset. The average value of the stability fluctuation index over 12 historical collection periods is predetermined. The predetermined stability fluctuation qualification threshold is determined based on the average value of the stability fluctuation index and is determined within the range [0.01, 0.18]. The preferred value that meets 80% of the measured value is selected according to the actual business requirements for determining the cause of stability anomalies in cloud nodes, edge nodes, and end nodes. In this embodiment, the predetermined stability fluctuation qualification threshold is preferably 0.14.
[0054] This invention determines the cause of node anomalies by setting a qualified threshold for stability fluctuations, accurately distinguishing between two types of anomalies: instantaneous interference and overload. By setting the qualified threshold for stability fluctuations based on the average value of the stability fluctuation index within the historical collection period, it can closely match the actual data of historical node anomaly determination, thereby improving the matching degree of anomaly cause determination.
[0055] Specifically, the adaptive optimization module is used to determine corresponding handling measures based on the causes of stability anomalies, including: If the cause is that the decision collaboration feature information is subject to instantaneous interference, then a smoothing filter is activated to filter the decision collaboration feature representation value. If the cause is excessive node load, then a node load optimization scheduling instruction will be issued to the collaborative scheduling module to trigger task migration and node deload operations within the computing power pool.
[0056] Understandably, by enabling the smoothing filter to filter the decision-making collaboration feature information, we can filter out the numerical fluctuations in task queuing time and execution timeout rate caused by instantaneous interference, making the calculation of the decision-making collaboration feature representation value more consistent with the actual collaborative operation status of the nodes, and making the judgment basis of the resource pool collaboration perception module more accurate.
[0057] Understandably, by issuing node load optimization scheduling instructions to the collaborative scheduling module, it is possible to adapt to changes in the operating status of nodes, reduce the risk that nodes will be directly judged as having unqualified scheduling due to short-term overload, and ensure the continuity of node participation in scheduling.
[0058] This invention, through the synergistic effect of two processing measures, can carry out differentiated processing for different causes of node stability anomalies. This reduces the risk of judgment deviation caused by instantaneous interference, while also taking into account the scheduling adaptability of nodes with excessive load. This makes the optimization strategy for the coordinated scheduling of cloud nodes, edge nodes, and end nodes more in line with the actual operating needs of heterogeneous computing power pools, thereby improving the adaptability of the scheduling system to the node operating status.
[0059] Specifically, the adaptive optimization module is used to determine whether to enable a smoothing filter to filter the decision collaboration feature representation value, wherein the decision collaboration feature representation value is positively correlated with the stability fluctuation index.
[0060] In this embodiment, the strategy for using a smoothing filter to filter the decision collaborative feature representation value is specifically as follows: ;in, The filtered collaborative feature representation value of the decision. The value represents the decision collaboration feature from the previous data collection period and is determined within the interval [0.96, 1.14]. The current collaborative feature value is defined within the interval [0.96, 1.14]. It is a stability fluctuation index and is defined within the range [0.01, 0.14]. The smoothing factor is preferably 0.36 in this embodiment.
[0061] This invention improves the accuracy of scheduling decision matching degree determination by filtering the decision collaboration feature representation value to make it fit the actual collaborative operation state of the node. The smoothing factor is calculated by using the stability fluctuation index and combined with the representation values of the previous and subsequent periods.
[0062] Specifically, the adaptive optimization module is used to issue node load optimization scheduling instructions to the collaborative scheduling module, wherein the execution intensity of the scheduling instructions is positively correlated with the stability fluctuation index.
[0063] In this embodiment, the execution intensity of the load optimization scheduling instruction is determined by an adjustment coefficient, and the specific formula for calculating the adjustment coefficient is as follows: ;in, For adjustment coefficients, It is a stability fluctuation index and is defined within the range [0.14, 0.18]. The strength adjustment coefficient is preferred in this embodiment, which is 1.03.
[0064] When the adjustment coefficient When the target node is within the interval [1.14, 1.15], a low-intensity node load optimization scheduling instruction is issued. This instruction specifies that the collaborative scheduling module will only pause issuing all low-priority non-real-time tasks to the target node exhibiting instability. Low-priority real-time tasks, all medium-priority tasks, and all high-priority tasks will maintain their original scheduling rhythm. All currently executing tasks on the target node will not be migrated and will remain in their original execution state. The low-priority tasks include low-priority real-time tasks and low-priority non-real-time tasks; the medium-priority tasks include medium-priority real-time tasks and medium-priority non-real-time tasks; and the high-priority tasks... The tasks are categorized into high-priority real-time tasks and high-priority non-real-time tasks. Real-time tasks are those with a maximum queuing time of less than or equal to 50 milliseconds, while non-real-time tasks are those with a maximum queuing time greater than 50 milliseconds. High-priority tasks are those with a time limit of less than or equal to 20 milliseconds from generation to assignment to a task execution node. Medium-priority tasks are those with a time limit of greater than 20 milliseconds and less than or equal to 30 milliseconds from generation to assignment to a task execution node. Low-priority tasks are those with a time limit of greater than 30 milliseconds from generation to assignment to a task execution node. When the adjustment coefficient When the target node is within the range [1.15, 1.17], a medium-intensity node load optimization scheduling instruction is issued. The instruction is as follows: the collaborative scheduling module first suspends the issuance of all low-priority tasks to the target node with abnormal stability. Then, it identifies and batch migrates all low-priority tasks currently being executed by the target node to cloud nodes, edge nodes, or end nodes in the computing power pool that are not abnormal in stability and whose load characteristic value is less than the predetermined load characteristic threshold. The migration process ensures the continuity and integrity of task data. All medium-priority tasks and all high-priority tasks maintain their original scheduling issuance rhythm and execution status. When the adjustment coefficient When the target node is within the range [1.17, 1.18], a high-intensity node load optimization scheduling instruction is issued. The instruction is as follows: The collaborative scheduling module first suspends all new tasks from being issued to the target node with abnormal stability. Then, it performs batch migration of all low-priority and medium-priority tasks currently being executed by the target node to cloud nodes, edge nodes, or end nodes in the computing power pool that are not abnormally stable and whose load characteristic value is less than the predetermined load characteristic value threshold. At the same time, the collaborative scheduling module schedules idle cloud nodes and edge nodes in the computing power pool to share the subdivisible sub-computation links in the high-priority tasks being executed by the target node, thereby reducing the computing pressure on the target node and reducing the CPU utilization and network round-trip latency between nodes until the load characteristic value of the target node is less than the predetermined load characteristic value threshold.
[0065] This invention calculates the execution intensity of node load optimization scheduling instructions issued to the collaborative scheduling module by associating stability fluctuation index. Based on the numerical range of the calculated adjustment coefficient, the execution intensity of scheduling instructions is divided into three gradients: low intensity, medium intensity, and high intensity. The judgment criteria can be dynamically adjusted according to the target node load exceeding the standard, thereby improving the adaptability of the scheduling decision matching degree judgment.
[0066] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A cloud-edge-end collaborative task scheduling system for heterogeneous computing power pooling, characterized in that, include: The collaborative scheduling module is used to schedule target execution tasks to cloud nodes, edge nodes, or end nodes for response; The data acquisition module is used to collect decision-making collaboration characteristic information and load characteristic information of target nodes within a historical period; The data processing module is used to analyze the decision collaboration feature representation value based on the decision collaboration feature information and to analyze the load feature representation value based on the load feature information. The resource pool collaborative perception module is used to determine whether the matching degree of the target node scheduling decision meets the standard based on the decision collaborative feature characterization value, and to determine whether the stability of the target node is abnormal based on the load feature characterization value. An adaptive optimization module is used to determine the cause of stability anomalies based on the stability fluctuation index, and to match corresponding processing measures based on the cause. The corresponding processing measures include enabling a smoothing filter to filter the decision collaboration feature characterization value, and determining to issue node load optimization scheduling instructions to the collaborative scheduling module. The decision-making collaboration feature information includes the queuing time of the task and the execution timeout rate of the task; the load feature information includes the CPU utilization and the network round-trip latency of adjacent nodes.
2. The cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling according to claim 1, characterized in that, The data processing module is used to analyze the decision collaboration feature representation value based on the decision collaboration feature information, wherein, The decision collaboration feature value is obtained by summing the first collaboration factor and the second collaboration factor according to a predetermined first weight ratio. The first coordination factor is determined based on the ratio of a predetermined queuing time threshold to the queuing time. The second collaboration factor is determined based on the ratio of a predetermined execution timeout rate threshold to the execution timeout rate.
3. The cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling according to claim 1, characterized in that, The resource pool collaborative perception module is used to determine whether the matching degree of the target node scheduling decision meets the standard based on the decision collaborative feature representation value, including: If the value of the decision collaboration feature is less than or equal to the predetermined decision collaboration feature threshold, then the matching degree is determined to be inconsistent with the standard. If the value of the decision collaboration feature is greater than the predetermined decision collaboration feature threshold, then the matching degree is determined to meet the standard.
4. The cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling according to claim 1, characterized in that, The data processing module is used to analyze load characteristic values based on the load characteristic information, wherein, The load characteristic value is obtained by summing the first load factor and the second load factor according to a predetermined second weighting ratio. The first load factor is determined based on the ratio of a predetermined utilization threshold to the utilization rate; The second load factor is determined based on the ratio of a predetermined network round-trip time threshold to the network round-trip time.
5. The cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling according to claim 1, characterized in that, The resource pool collaborative sensing module is used to determine whether the stability of the target node is abnormal based on the load characteristic characterization value, including: If the load characteristic value is less than or equal to the predetermined load characteristic threshold, then a stability anomaly is determined. If the load characteristic value is greater than the predetermined load characteristic threshold, then the stability is determined to be normal.
6. The cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling according to claim 1, characterized in that, The adaptive optimization module is used to analyze the stability fluctuation index, wherein the stability fluctuation index is determined based on the difference between a predetermined load characteristic characterization threshold and the load characteristic characterization value.
7. The cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling according to claim 6, characterized in that, The adaptive optimization module is used to determine the causes of stability anomalies based on the stability fluctuation index, including: If the stability fluctuation index is less than or equal to the predetermined stability fluctuation qualification threshold, it is determined that the decision collaboration feature information is subject to instantaneous interference. If the stability fluctuation index is greater than the predetermined stability fluctuation acceptable threshold, then the node load is determined to be excessive.
8. The cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling according to claim 7, characterized in that, The adaptive optimization module is used to determine corresponding handling measures based on the causes of stability anomalies, including: If the cause is that the decision collaboration feature information is subject to instantaneous interference, then a smoothing filter is activated to filter the decision collaboration feature representation value. If the cause is excessive node load, then a node load optimization scheduling instruction will be issued to the collaborative scheduling module.
9. The cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling according to claim 8, characterized in that, The adaptive optimization module is used to determine whether to enable a smoothing filter to filter the decision collaboration feature representation value, wherein the decision collaboration feature representation value is positively correlated with the stability fluctuation index.
10. The cloud-edge-device collaborative task scheduling system for heterogeneous computing power pooling according to claim 8, characterized in that, The adaptive optimization module is used to determine the node load optimization scheduling instructions to be issued to the collaborative scheduling module, wherein the execution intensity of the scheduling instructions is positively correlated with the stability fluctuation index.
Citation Information
Patent Citations
Heterogeneous computing power scheduling optimization method based on cloud edge collaborative architecture
CN121396990A