Computing power resource optimization method based on dynamic scheduling algorithm

By combining a multimodal perception layer, a dynamic priority engine, and an incremental scheduling window, the problems of resource fragmentation and insufficient load prediction in computing power scheduling are solved, achieving efficient resource utilization and improved task stability.

CN121029316APending Publication Date: 2025-11-28浪潮智慧城市科技有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511160071.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing computing power scheduling schemes lack fine-grained awareness, leading to resource fragmentation or overload, insufficient load prediction, preemptive scheduling strategies causing task restart delays, low resource utilization, and poor system stability.

Method used

A multimodal perception layer is used to build a task feature library. A dynamic priority engine is combined with elastic preemption and reinforcement learning algorithms, and an incremental scheduling window predicts the load to achieve fine-grained resource allocation and reasonable reservation.

Benefits of technology

It improves resource utilization, reduces task startup latency, increases task completion rate, reduces task interruption recovery time, and enhances system stability and load prediction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029316A_ABST
    Figure CN121029316A_ABST
Patent Text Reader

Abstract

The invention provides a computing power resource optimization method based on a dynamic scheduling algorithm, which belongs to the field of edge computing resource management, constructs a dynamic scheduling decision model by fusing multi-dimensional data such as real-time load monitoring, task feature analysis and environmental energy consumption perception, and realizes a self-adaptive resource allocation strategy in combination with a reinforcement learning algorithm. Comprising the steps of multi-mode sensing layer construction, dynamic priority engine and incremental scheduling, the computing power resource utilization rate is remarkably improved, task response delay and system energy consumption are reduced, and the method is suitable for real-time computing scenes of smart city services such as video analysis and industrial digital twinning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the fields of cloud computing, edge computing resource management, dynamic scheduling algorithm, artificial intelligence optimization, and in particular to a computing power resource optimization method based on a dynamic scheduling algorithm. BACKGROUND

[0002] In the field of computing power scheduling, existing computing power scheduling schemes mostly use static threshold or rule-driven strategies (such as round robin, priority queue), which have the following major pain points.

[0003] 1. Traditional container orchestration tools (Kubernetes) rely on preset rules to expand resources, lack fine-grained perception of task types and resource requirements, and cause problems such as resource fragmentation or overload.

[0004] 2. Existing pre-emptive scheduling strategies (such as the Pod eviction mechanism of Kubernetes) have the following fatal defects. For example, hard pre-emption causes state loss, and when forcibly terminating low-priority tasks, the intermediate calculation state is not saved, causing task restart delay to be extended from minutes to hours.

[0005] 3. Insufficient load prediction capability: existing technologies mostly use simple moving average method to predict load, which cannot capture long-period fluctuation rules. SUMMARY

[0006] To solve the above technical problems, the present application provides a computing power resource optimization method based on a dynamic scheduling algorithm, which overcomes the limitations of traditional methods, provides more fine-grained resource perception, can identify heterogeneous tasks, and provides more scientific scheduling strategies to achieve reasonable resource allocation and rolling scheduling window algorithm, improves the original prediction logic, and promotes scientific use of resources; solves the problems of low resource utilization, high task delay, and insufficient load prediction caused by load fluctuation and resource heterogeneity in cloud computing and edge computing environments.

[0007] The technical solution of the present application is:

[0008] A computing power resource optimization method based on a dynamic scheduling algorithm, comprising:

[0009] A multi-modal perception layer is constructed to collect resource states in real time, and task feature extraction is used to realize qualitative analysis of resources and support subsequent resource allocation;

[0010] A dynamic priority engine is used to realize reasonable scheduling and allocation of existing resources by using an elastic pre-emption mechanism and introducing a distributed reinforcement learning algorithm;

[0011] Incremental scheduling is used to predict the load distribution of the future scheduling period through a spatiotemporal perception prediction model, and dynamically adjust the resource reservation ratio according to the confidence interval.

[0012] Furthermore,

[0013] The construction of the multimodal sensing layer includes

[0014] Hardware awareness layer: Deploy agent programs to collect task characteristics in real time to achieve resource awareness;

[0015] Task feature extraction: parse the layer dependencies of container images, predict task startup latency, and use the lightweight performance profiler eBPF to dynamically identify task computation patterns such as compute-intensive or I / O-intensive.

[0016] Furthermore,

[0017] By deploying eBPF probes on computing nodes to collect CPU instruction set characteristics, GPU memory access patterns, and storage IOPS data in real time, a task computing feature library is established to dynamically identify tasks and determine task attributes.

[0018] in,

[0019] The construction of the task-based feature fingerprint database includes:

[0020] By statically analyzing the container image layer dependencies, the dynamic link libraries required by the task are preloaded.

[0021] Identifying characteristics of compute-intensive tasks based on instruction-level performance monitoring (IPC) data.

[0022] Furthermore,

[0023] Dynamic Priority Engine: It constructs a more reasonable allocation strategy through reinforcement learning algorithms to maximize allocation efficiency.

[0024] Dynamic priority engine adaptive scheduling, specifically including

[0025] Employ an elastic preemption mechanism: perform soft preemption on low-priority tasks, preserving their memory state to reduce restart overhead;

[0026] Introducing reinforcement learning algorithms: Task scheduling is modeled as an incomplete information game, with each task acting as an agent, using the Q-learning algorithm to select the optimal action under different states, thereby achieving reasonable competition and allocation of resources.

[0027] in,

[0028] Elastic preemption decision-making includes:

[0029] Implement incremental resource compression for low-priority tasks: initially limit CPU quota to 50%;

[0030] Establish a preemption compensation mechanism: record the intermediate state of a preempted task, and allocate high-bandwidth resources to it when its priority is increased.

[0031] Furthermore,

[0032] Incremental scheduling divides the scheduling cycle into continuous time windows. Within each window, load data is collected in real time to predict the load trend for several future windows.

[0033] You can set a window to open every 10 minutes.

[0034] Furthermore,

[0035] Incremental scheduling, specifically including

[0036] Spatiotemporal awareness prediction model: The Transformer time series model is used to capture long-term load patterns and graph convolutional networks are used to predict cross-node task diffusion trends.

[0037] Resource reservation strategy: Reserve 10% of the blank resources in the rolling window, and dynamically adjust the reservation ratio according to the prediction confidence.

[0038] Furthermore,

[0039] Resource reservations include:

[0040] 10% of blank resources are reserved when the prediction confidence level is ≥80%;

[0041] Retain 25% of resources when the prediction confidence is less than 60%.

[0042] The beneficial effects of this invention are

[0043] 1. Resource utilization optimization: By preloading container image layers through a spatiotemporal prediction model, task startup latency is reduced by 72%; fine-grained scheduling based on computational feature fingerprints reduces GPU memory fragmentation rate from 28% to below 15%.

[0044] 2. Multi-objective collaborative optimization: Under the same load, energy consumption is reduced by 35% while the task completion rate is increased by 40%.

[0045] 3. Innovative preemption mechanism: The soft preemption strategy reduces the recovery time of long-cycle training tasks from minutes to seconds, ensuring parameter consistency in distributed training.

[0046] 4. Innovative prediction mechanism: By using a rolling scheduling window for load prediction, the system's resilience is improved. Attached Figure Description

[0047] Figure 1 This is a schematic diagram illustrating the logic and implementation effect of the multimodal perception layer.

[0048] Figure 2 This is a diagram illustrating the logic and implementation of the dynamic priority engine.

[0049] Figure 3 This is a schematic diagram illustrating the logic and implementation effect of the incremental scrolling scheduling window. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0051] This invention provides a computing resource optimization method based on a dynamic scheduling algorithm, which solves the scheduling dilemma based on rule-driven methods, achieves finer-grained perception capabilities, enables reasonable scheduling of existing resources, and improves system stability by reasonably reserving future resources.

[0052] 1. Multimodal sensing input layer

[0053] This invention employs a hardware perception layer to achieve real-time acquisition of resource status, and uses task feature extraction to qualitatively assess resources and support subsequent resource allocation.

[0054] 1. Hardware Awareness Layer: Deploy agent programs to collect task characteristics in real time to achieve resource awareness.

[0055] 2. Task Feature Extraction: Parse the layer dependencies of container images, predict task startup latency, and use the lightweight performance profiler eBPF to dynamically identify task computation patterns such as compute-intensive or I / O-intensive.

[0056] 2. Dynamic Priority Engine (Adaptive Scheduling)

[0057] An elastic preemption mechanism and a distributed reinforcement learning algorithm are adopted to achieve reasonable scheduling and allocation of existing resources.

[0058] 1. Adopt an elastic preemption mechanism: Implement "soft preemption" for low-priority tasks (such as limiting CPU quotas instead of forcibly terminating them) to preserve their memory state and reduce restart overhead.

[0059] 2. Introduce reinforcement learning algorithm: Model task scheduling as an incomplete information game. Each task acts as an agent and selects the optimal action (such as allocating resources or preempting tasks) under different states (such as task load and resource utilization) through Q-learning algorithm to complete the reasonable competition and allocation of resources.

[0060] The Q-function Q(s,a) is defined as the expected reward obtained by taking action a from state s and following policy π. Its mathematical expression is:

[0061] Q(s,a)=E[R t+1 +γmax a′ Q(s t+1 ,a′)}s t =s,a t =a]

[0062] in:

[0063] ·R t+1 The reward is obtained at time t+1.

[0064] • γ is the discount factor, which determines the present value of future rewards.

[0065] ·s t+1 It is the new state at time t+1.

[0066] a′ represents the possible actions to be taken in the new state.

[0067] Q function definition

[0068] Three: Incremental Rolling Scheduling Window

[0069] The scheduling cycle is divided into continuous "time windows" (set to 10 minutes per window). Within each window, load data such as node resource utilization, task queue status, and network bandwidth are collected in real time to predict the load trend of multiple future windows, thereby ensuring system stability while improving resource utilization efficiency.

[0070] 1. Spatiotemporal Awareness Prediction Model: The Transformer time series model is used to capture long-term load patterns, and the Graph Convolutional Network (GCN) is used to predict cross-node task diffusion trends (e.g., the increase in load on edge node A may lead to an increase in latency on adjacent node B).

[0071] 2. Resource reservation strategy: Reserve 10% of the blank resources in the rolling window, and dynamically adjust the reservation ratio according to the prediction confidence (increase the reservation when the confidence is low).

[0072] The above description is merely a preferred embodiment of the present invention and is used only to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A method for optimizing computing resources based on a dynamic scheduling algorithm, characterized in that, include: A multimodal perception layer is constructed to collect resource status in real time, and task feature extraction is used to qualitatively assess resources and support subsequent resource allocation. The dynamic priority engine employs an elastic preemption mechanism and introduces a distributed reinforcement learning algorithm to achieve reasonable scheduling and allocation of existing resources. Incremental scheduling uses a spatiotemporal awareness prediction model to predict the load distribution in future scheduling cycles and dynamically adjusts the resource reservation ratio according to the confidence interval.

2. The method according to claim 1, characterized in that, The construction of the multimodal sensing layer includes Hardware awareness layer: Deploy agent programs to collect task characteristics in real time to achieve resource awareness; Task feature extraction: parse the layer dependencies of container images, predict task startup latency, and use the lightweight performance profiler eBPF to dynamically identify task computation patterns.

3. The method according to claim 2, characterized in that, By deploying eBPF probes on computing nodes to collect CPU instruction set characteristics, GPU memory access patterns, and storage IOPS data in real time, a task computing feature library is established to dynamically identify tasks and determine task attributes.

4. The method according to claim 3, characterized in that, The construction of the task-based feature fingerprint database includes: By statically analyzing the container image layer dependencies, the dynamic link libraries required by the task are preloaded. Identify characteristics of computationally intensive tasks based on instruction-level performance monitoring IPC data.

5. The method according to claim 1, characterized in that, Dynamic priority engine, adaptive scheduling, specifically includes Employ an elastic preemption mechanism: perform soft preemption on low-priority tasks, preserving their memory state to reduce restart overhead; Introducing reinforcement learning algorithms: Task scheduling is modeled as an incomplete information game, with each task acting as an agent, using the Q-learning algorithm to select the optimal action under different states, thereby achieving reasonable competition and allocation of resources.

6. The method according to claim 5, characterized in that, Elastic preemption decision-making includes: Implement incremental resource compression for low-priority tasks: initially limit CPU quota to 50%; Establish a preemption compensation mechanism: record the intermediate state of a preempted task, and allocate high-bandwidth resources to it when its priority is increased.

7. The method according to claim 1, characterized in that, Incremental scheduling divides the scheduling cycle into continuous time windows. Within each window, load data is collected in real time to predict the load trend for several future windows.

8. The method according to claim 7, characterized in that, Set it to open a window every 10 minutes.

9. The method according to claim 7, characterized in that, Incremental scheduling, specifically including Spatiotemporal awareness prediction model: The Transformer time series model is used to capture long-term load patterns and graph convolutional networks are used to predict cross-node task diffusion trends. Resource reservation strategy: Reserve 10% of the blank resources in the rolling window, and dynamically adjust the reservation ratio according to the prediction confidence.

10. The method according to claim 9, characterized in that, Resource reservations include: 10% of blank resources are reserved when the prediction confidence level is ≥80%; Retain 25% of resources when the prediction confidence is less than 60%.