Cooperative task scheduling method supporting multi-region computing power hosts
By building resource perception and dynamic portraits, formulating cross-regional collaborative scheduling solutions, optimizing node allocation and data transmission, efficient and stable task scheduling of multi-regional computing power hosts is achieved, solving the problems of resource waste and delay in the existing technology, and improving system performance and user experience.
Patent Information
- Application Number
- CN202510671383.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-29
AI Technical Summary
When facing a variety of business characteristics, existing task scheduling strategies are difficult to meet the diversified resources of different tasks, resulting in waste of computing resources, task accumulation and response delays, which seriously affects system performance and user experience.
Build resource perception and dynamic portraits, formulate cross-regional collaborative scheduling solutions, optimize node allocation through adaptive weighting mechanisms and deep reinforcement learning models, perform data pre-distribution and storage optimization, manage load balancing and elastic scaling, and deploy dual live node redundancy to achieve high availability.
It improves the utilization rate of computing power resources, reduces energy consumption, shortens task response time, enhances system stability and reliability, reduces operational costs, and supports more diverse application scenarios.
Smart Images

Figure CN120560802A_ABST
Abstract
Description
Technical Field
[0001] The present invention discloses a collaborative task scheduling method supporting multi-region computing hosts, which relates to cloud computing and edge computing technologies. Background Art
[0002] With the rapid development of cloud computing, big data, and artificial intelligence technologies, the efficiency of computing hosts in handling large-scale concurrent tasks has become increasingly prominent. However, existing task scheduling strategies, when faced with diverse business characteristics, typically employ static or single optimization objectives, making it difficult to meet the diverse resource requirements of different tasks. This lack of adaptability often leads to problems such as wasted computing resources, task backlogs, and response delays, severely hampering system performance and user experience. Summary of the Invention
[0003] In response to the problems of the existing technology, the present invention provides a collaborative task scheduling method that supports multi-regional computing hosts, realizes the intelligent allocation and management of computing resources on a global scale, improves the overall concurrent processing capability of the system, optimizes the collaboration efficiency between regions, reduces delays and costs, and ensures high efficiency and stability in the task processing process. It is particularly suitable for large-scale application scenarios that require multi-regional collaboration.
[0004] The specific scheme proposed by the present invention is:
[0005] The present invention provides a collaborative task scheduling method supporting multi-region computing hosts, comprising:
[0006] Step 1: Build resource perception and dynamic portraits: Collect node resource status data in real time, and dynamically generate node resource portraits based on the resource status data.
[0007] Step 2: Make collaborative decisions on task scheduling: automatically parse tasks based on task requests submitted by users; formulate cross-region collaborative scheduling plans: select target nodes based on node resource profiles, and dynamically adjust the priority of each target node through an adaptive weight mechanism. At the same time, train a deep reinforcement learning model based on historical scheduling data to generate the optimal node allocation plan in real time.
[0008] Step 3: Data pre-distribution and storage optimization: According to the geographical location of the task target node, the input data is divided into multiple sub-blocks and stored in the storage unit closest to the execution target node first.
[0009] Build a cross-regional network quality map, evaluate the transmission delay and bandwidth utilization between target nodes, select the lowest-cost and delay-controlled data transmission path for each task, and avoid network congestion areas;
[0010] Step 4: Manage load balancing and elastic scaling: Develop a dynamic load migration mechanism and perform elastic resource expansion.
[0011] Step 5: Perform fault tolerance and high availability management: Deploy active-active node redundancy, and perform breakpoint resume and status snapshots to balance storage overhead and recovery efficiency.
[0012] Furthermore, step 1 of the collaborative task scheduling method supporting multi-region computing hosts includes:
[0013] When collecting node resource status data in real time, lightweight probes are deployed at the edge nodes in each region to collect the CPU / GPU utilization, memory usage, and network bandwidth core indicators of the computing nodes at a frequency of seconds, and the data is aggregated in real time to the central dispatch center through a distributed message queue.
[0014] When dynamically generating node resource portraits based on resource status data, a resource portrait of the computing power node is constructed based on historical data and real-time collected resource status data. The resource portrait includes computing power score, network quality level, and regional cost index.
[0015] A time series prediction model is used to predict the load change trend of each node in the next 5-10 minutes, providing a forward-looking basis for collaborative decision-making.
[0016] Furthermore, step 4 of the collaborative task scheduling method supporting multi-region computing hosts includes:
[0017] When developing a dynamic load migration mechanism, the resource utilization of each region is monitored in real time. When it is detected that the load of a target node continues to exceed the limit, some tasks are automatically migrated to the adjacent low-load target node.
[0018] Use virtual node mapping to ensure that the migration process does not affect the high-priority tasks being executed;
[0019] When performing elastic resource expansion, when resource demand in a certain area surges, it automatically connects to the public cloud resource pool, expands temporary computing power on demand, and sets expansion thresholds and recycling policies to avoid resource waste.
[0020] Furthermore, step 5 of the collaborative task scheduling method supporting multi-region computing hosts includes:
[0021] When deploying active-active node redundancy, at least one backup target node is assigned to each primary execution target node. The two nodes synchronize task status and intermediate data in real time. When the primary target node fails, the backup target node takes over the task to ensure business continuity.
[0022] When performing breakpoint resuming and status snapshots, periodically save task execution progress snapshots to distributed storage, support task recovery from any breakpoint, and set snapshot saving frequency and retention policy.
[0023] The present invention provides a collaborative task scheduling device supporting multi-region computing hosts, including a portrait management module, a decision module, a distribution storage module, and a management module.
[0024] The portrait management module builds resource perception and dynamic portraits: real-time collection of node resource status data, and dynamic generation of node resource portraits based on resource status data.
[0025] The decision-making module makes collaborative decisions on task scheduling: automatically analyzes tasks based on task requests submitted by users; formulates cross-regional collaborative scheduling plans: selects target nodes based on node resource profiles, and dynamically adjusts the priority of each target node through an adaptive weight mechanism. At the same time, it trains a deep reinforcement learning model based on historical scheduling data to generate the optimal node allocation plan in real time.
[0026] The distribution storage module performs data pre-distribution and storage optimization: according to the geographical location of the task target node, the input data is divided into multiple sub-blocks and stored in the storage unit closest to the execution target node first.
[0027] Build a cross-regional network quality map, evaluate the transmission delay and bandwidth utilization between target nodes, select the lowest-cost and delay-controlled data transmission path for each task, and avoid network congestion areas;
[0028] The management module manages load balancing and elastic scaling: formulates dynamic load migration mechanisms and performs elastic resource expansion.
[0029] The management module performs fault tolerance and high availability management: deploys active-active node redundancy, and performs breakpoint resume and status snapshots to balance storage overhead and recovery efficiency.
[0030] Furthermore, when the portrait management module of the collaborative task scheduling device supporting multi-region computing hosts collects node resource status data in real time, lightweight probes are deployed at the edge nodes of each region to collect the CPU / GPU utilization, memory usage, and network bandwidth core indicators of the computing nodes at a frequency of seconds, and the data is aggregated in real time to the central scheduling center through a distributed message queue.
[0031] When dynamically generating node resource portraits based on resource status data, a resource portrait of the computing power node is constructed based on historical data and real-time collected resource status data. The resource portrait includes computing power score, network quality level, and regional cost index.
[0032] A time series prediction model is used to predict the load change trend of each node in the next 5-10 minutes, providing a forward-looking basis for collaborative decision-making.
[0033] Furthermore, when the management module of the collaborative task scheduling device supporting multi-region computing hosts formulates a dynamic load migration mechanism, it monitors the resource utilization rate of each region in real time. When it detects that the load of a target node continues to exceed the limit, it automatically migrates some tasks to adjacent low-load target nodes.
[0034] Use virtual node mapping to ensure that the migration process does not affect the high-priority tasks being executed;
[0035] When performing elastic resource expansion, when resource demand in a certain area surges, it automatically connects to the public cloud resource pool, expands temporary computing power on demand, and sets expansion thresholds and recycling policies to avoid resource waste.
[0036] Furthermore, when the management module of the collaborative task scheduling device supporting multi-region computing hosts deploys active-active node redundancy, at least one backup target node is assigned to each main execution target node, and the two synchronize task status and intermediate data in real time. When the main target node fails, the backup target node takes over the task to ensure business continuity.
[0037] When performing breakpoint resuming and status snapshots, periodically save task execution progress snapshots to distributed storage, support task recovery from any breakpoint, and set snapshot saving frequency and retention policy.
[0038] The benefits of the present invention are:
[0039] Significantly improved the utilization of computing resources and reduced unnecessary energy consumption;
[0040] Accelerates task response time and completion time, enhancing user experience;
[0041] Enhanced the stability and reliability of the system and reduced the failure rate;
[0042] By optimizing the scheduling strategy, the operating costs are greatly reduced;
[0043] Support more diverse application scenarios to meet the needs of different users. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a schematic flow chart of the method of the present invention. DETAILED DESCRIPTION
[0045] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.
[0046] Example 1
[0047] The present invention provides a collaborative task scheduling method supporting multi-region computing hosts, which can be implemented using a three-level distributed architecture and specifically includes:
[0048] Step 1: Build resource awareness and dynamic portraits: Collect node resource status data in real time and dynamically generate node resource portraits based on the resource status data.
[0049] Wherein step 1 may specifically include:
[0050] When collecting node resource status data in real time, lightweight probes are deployed at the edge nodes in each region to collect the CPU / GPU utilization, memory usage, and network bandwidth core indicators of the computing nodes at a frequency of seconds, and the data is aggregated in real time to the central dispatch center through a distributed message queue.
[0051] When dynamically generating node resource portraits based on resource status data, a resource portrait of the computing power node is constructed based on historical data and real-time collected resource status data. The resource portrait includes computing power score, network quality level, and regional cost index.
[0052] A time series prediction model is used to predict the load change trend of each node in the next 5-10 minutes, providing a forward-looking basis for collaborative decision-making.
[0053] Step 2: Make collaborative decisions about task scheduling: Automatically parse tasks based on user-submitted task requests, such as task type (e.g., real-time computing, batch processing), data dependencies, service level agreements (SLAs), and other key parameters.
[0054] Develop a cross-regional collaborative scheduling plan: select target nodes based on node resource profiles, simultaneously optimize task execution delay, resource usage cost, and energy consumption indicators, and dynamically adjust the priority of each target node through an adaptive weight mechanism. At the same time, train a deep reinforcement learning model based on historical scheduling data to generate the optimal node allocation plan in real time.
[0055] Step 3: Data pre-distribution and storage optimization: Based on the geographic location of the task target node, the input data is divided into multiple sub-blocks and stored preferentially in the storage unit closest to the execution target node. Redundant coding technology is used to improve data availability and ensure rapid recovery in the event of a single point of failure.
[0056] Build a cross-regional network quality map, evaluate the transmission delay and bandwidth utilization between target nodes, select the lowest-cost and delay-controlled data transmission path for each task, and avoid network congestion areas;
[0057] Step 4: Manage load balancing and elastic scaling: Develop a dynamic load migration mechanism and perform elastic resource expansion.
[0058] Step 4 may specifically include:
[0059] When developing a dynamic load migration mechanism, the resource utilization of each region is monitored in real time. When it is detected that the load of a target node continues to exceed the limit, some tasks are automatically migrated to the adjacent low-load target node.
[0060] Use virtual node mapping to ensure that the migration process does not affect the high-priority tasks being executed;
[0061] When performing elastic resource expansion, when resource demand in a certain area surges, it automatically connects to the public cloud resource pool, expands temporary computing power on demand, and sets expansion thresholds and recycling policies to avoid resource waste.
[0062] Step 5: Perform fault tolerance and high availability management: Deploy active-active node redundancy, and perform breakpoint resume and status snapshots to balance storage overhead and recovery efficiency.
[0063] Step 5 may specifically include:
[0064] When deploying active-active node redundancy, at least one backup target node is assigned to each primary execution target node. The two nodes synchronize task status and intermediate data in real time. When the primary target node fails, the backup target node takes over the task to ensure business continuity.
[0065] When performing breakpoint resuming and status snapshots, periodically save task execution progress snapshots to distributed storage, support task recovery from any breakpoint, and set snapshot saving frequency and retention policy.
[0066] Example 2
[0067] The present invention provides a collaborative task scheduling device supporting multi-region computing hosts, including a portrait management module, a decision module, a distribution storage module, and a management module.
[0068] The portrait management module builds resource perception and dynamic portraits: real-time collection of node resource status data, and dynamic generation of node resource portraits based on resource status data.
[0069] The decision-making module makes collaborative decisions on task scheduling: automatically analyzes tasks based on task requests submitted by users; formulates cross-regional collaborative scheduling plans: selects target nodes based on node resource profiles, and dynamically adjusts the priority of each target node through an adaptive weight mechanism. At the same time, it trains a deep reinforcement learning model based on historical scheduling data to generate the optimal node allocation plan in real time.
[0070] The distribution storage module performs data pre-distribution and storage optimization: according to the geographical location of the task target node, the input data is divided into multiple sub-blocks and stored in the storage unit closest to the execution target node first.
[0071] Build a cross-regional network quality map, evaluate the transmission delay and bandwidth utilization between target nodes, select the lowest-cost and delay-controlled data transmission path for each task, and avoid network congestion areas;
[0072] The management module manages load balancing and elastic scaling: formulates dynamic load migration mechanisms and performs elastic resource expansion.
[0073] The management module performs fault tolerance and high availability management: deploys active-active node redundancy, and performs breakpoint resume and status snapshots to balance storage overhead and recovery efficiency.
[0074] Since the information interaction, execution process and other contents between the modules in the above-mentioned device are based on the same concept as the embodiment of the method of the present invention, the specific contents can be found in the description of the embodiment of the method of the present invention and will not be repeated here.
[0075] Likewise, the device of the present invention is beneficial in that:
[0076] Significantly improved the utilization of computing resources and reduced unnecessary energy consumption;
[0077] Accelerates task response time and completion time, enhancing user experience;
[0078] Enhanced the stability and reliability of the system and reduced the failure rate;
[0079] By optimizing the scheduling strategy, the operating costs are greatly reduced;
[0080] Support more diverse application scenarios to meet the needs of different users.
[0081] It should be noted that not all steps and modules in the above-mentioned processes and device structures are required, and certain steps or modules can be omitted according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or may be implemented by certain components in multiple independent devices.
[0082] The above embodiments are merely preferred embodiments for the purpose of fully illustrating the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are within the scope of protection of the present invention. The scope of protection of the present invention shall be subject to the claims.
Claims
1. A collaborative task scheduling method supporting multi-region computing hosts, characterized by include: Step 1: Build resource perception and dynamic portraits: Collect node resource status data in real time, and dynamically generate node resource portraits based on the resource status data. Step 2: Make collaborative decisions on task scheduling: automatically parse tasks based on task requests submitted by users; formulate cross-region collaborative scheduling plans: select target nodes based on node resource profiles, and dynamically adjust the priority of each target node through an adaptive weight mechanism. At the same time, train a deep reinforcement learning model based on historical scheduling data to generate the optimal node allocation plan in real time. Step 3: Data pre-distribution and storage optimization: According to the geographical location of the task target node, the input data is divided into multiple sub-blocks and stored in the storage unit closest to the execution target node first. Build a cross-regional network quality map, evaluate the transmission delay and bandwidth utilization between target nodes, select the lowest-cost and delay-controlled data transmission path for each task, and avoid network congestion areas; Step 4: Manage load balancing and elastic scaling: Develop a dynamic load migration mechanism and perform elastic resource expansion. Step 5: Perform fault tolerance and high availability management: Deploy active-active node redundancy, and perform breakpoint resume and status snapshots to balance storage overhead and recovery efficiency.
2. A collaborative task scheduling method supporting multi-region computing hosts according to claim 1, characterized in that Step 1 includes: When collecting node resource status data in real time, lightweight probes are deployed at the edge nodes in each region to collect the CPU / GPU utilization, memory usage, and network bandwidth core indicators of the computing nodes at a frequency of seconds, and the data is aggregated to the central dispatch center in real time through a distributed message queue. When dynamically generating node resource portraits based on resource status data, a resource portrait of the computing power node is constructed based on historical data and real-time collected resource status data. The resource portrait includes computing power score, network quality level, and regional cost index. A time series prediction model is used to predict the load change trend of each node in the next 5-10 minutes, providing a forward-looking basis for collaborative decision-making.
3. The collaborative task scheduling method supporting multi-region computing hosts according to claim 1 is characterized by: Step 4 includes: When developing a dynamic load migration mechanism, the resource utilization of each region is monitored in real time. When it is detected that the load of a target node continues to exceed the limit, some tasks are automatically migrated to the adjacent low-load target node. Use virtual node mapping to ensure that the migration process does not affect the high-priority tasks being executed; When performing elastic resource expansion, when resource demand in a certain area surges, it automatically connects to the public cloud resource pool, expands temporary computing power on demand, and sets expansion thresholds and recycling policies to avoid resource waste.
4. The collaborative task scheduling method supporting multi-region computing hosts according to claim 1 is characterized by: Step 5 includes: When deploying active-active node redundancy, at least one backup target node is assigned to each primary execution target node. The two nodes synchronize task status and intermediate data in real time. When the primary target node fails, the backup target node takes over the task to ensure business continuity. When performing breakpoint resuming and status snapshots, periodically save task execution progress snapshots to distributed storage, support task recovery from any breakpoint, and set snapshot saving frequency and retention policy.
5. A collaborative task scheduling device supporting multi-region computing hosts, characterized by Including portrait management module, decision module, distribution storage module, management module, The portrait management module builds resource perception and dynamic portraits: real-time collection of node resource status data, and dynamic generation of node resource portraits based on resource status data. The decision-making module makes collaborative decisions on task scheduling: automatically analyzes tasks based on task requests submitted by users; formulates cross-regional collaborative scheduling plans: selects target nodes based on node resource profiles, and dynamically adjusts the priority of each target node through an adaptive weight mechanism. At the same time, it trains a deep reinforcement learning model based on historical scheduling data to generate the optimal node allocation plan in real time. The distribution storage module performs data pre-distribution and storage optimization: according to the geographical location of the task target node, the input data is divided into multiple sub-blocks and stored in the storage unit closest to the execution target node first. Build a cross-regional network quality map, evaluate the transmission delay and bandwidth utilization between target nodes, select the lowest-cost and delay-controlled data transmission path for each task, and avoid network congestion areas; The management module manages load balancing and elastic scaling: formulates dynamic load migration mechanisms and performs elastic resource expansion. The management module performs fault tolerance and high availability management: deploys active-active node redundancy, and performs breakpoint resume and status snapshots to balance storage overhead and recovery efficiency.
6. The collaborative task scheduling device supporting multi-region computing hosts according to claim 5 is characterized in that When the portrait management module collects node resource status data in real time, it deploys lightweight probes at the edge nodes in each area to collect the CPU / GPU utilization, memory usage, and network bandwidth core indicators of the computing nodes at a frequency of seconds, and aggregates the data to the central dispatch center in real time through the distributed message queue. When dynamically generating node resource portraits based on resource status data, a resource portrait of the computing power node is constructed based on historical data and real-time collected resource status data. The resource portrait includes computing power score, network quality level, and regional cost index. A time series prediction model is used to predict the load change trend of each node in the next 5-10 minutes, providing a forward-looking basis for collaborative decision-making.
7. The collaborative task scheduling device supporting multi-region computing hosts according to claim 5 is characterized by: When the management module formulates a dynamic load migration mechanism, it monitors the resource utilization of each area in real time. When it detects that the load of a target node continues to exceed the limit, it automatically migrates some tasks to the adjacent low-load target node. Use virtual node mapping to ensure that the migration process does not affect the high-priority tasks being executed; When performing elastic resource expansion, when resource demand in a certain area surges, it automatically connects to the public cloud resource pool, expands temporary computing power on demand, and sets expansion thresholds and recycling policies to avoid resource waste.
8. The collaborative task scheduling device supporting multi-region computing hosts according to claim 5 is characterized in that When the management module deploys dual-active node redundancy, at least one backup target node is assigned to each main execution target node. The two synchronize task status and intermediate data in real time. When the main target node fails, the backup target node takes over the task to ensure business continuity. When performing breakpoint resuming and status snapshots, periodically save task execution progress snapshots to distributed storage, support task recovery from any breakpoint, and set snapshot saving frequency and retention policy.
Citation Information
Cited By
Resource and task cooperative scheduling method, system and equipment based on computing power alliance
CN121070623A
Resource and task cooperative scheduling method, system and device based on computing power alliance
CN121070623B