Task scheduling method, system and device based on multi-domain computing power network and storage medium
Through the TabTransformer-based task classification model and multi-level strategy collaborative scheduling module, the intelligent problem of power task scheduling in the multi-domain computing power network is solved, and efficient and accurate task scheduling and resource utilization are achieved to adapt to the complex power business environment.
Patent Information
- Application Number
- CN202510822269.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional power task scheduling mechanisms are difficult to achieve high-real-time, high-precision intelligent collaborative scheduling in a multi-domain computing network environment. They face challenges such as structural heterogeneity, time sensitivity, and uneven resource distribution. Existing methods cannot effectively match task attributes and resource status, resulting in low scheduling efficiency.
A TabTransformer-based task classification model is constructed, combined with a multi-domain computing power network and a multi-level strategy collaborative scheduling module. Through task feature vectors, computing power resource status matrix, and task level labels, multi-level strategy collaborative scheduling is realized, the information link between task semantic understanding and resource scheduling is opened up, and the Meta-Q Learning mechanism is used for strategy optimization.
It improves the dispatch response capability and resource utilization of the power dispatch system, realizes intelligent self-optimization of tasks, adapts to dynamic changes in the multi-domain computing environment, and improves the accuracy and adaptability of task classification and dispatch decision-making.
Smart Images

Figure CN120653401A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of intelligent power dispatching, and in particular to a task scheduling method, system, device and storage medium based on a multi-domain computing power network. Background Art
[0002] As power systems continue to expand and their application scenarios become increasingly complex, traditional power task scheduling mechanisms are no longer able to meet the demands for high real-time, high-precision, and intelligent coordinated scheduling in a multi-domain computing resource environment. Currently, power dispatching tasks face multiple challenges, including heterogeneous structures, time-sensitive operations, and uneven resource distribution. The diverse nature of dispatching tasks and the dynamic state of resources frequently intertwine, placing higher demands on the adaptability and coordination of dispatching strategies.
[0003] However, most existing scheduling methods rely on static rules, heuristic algorithms or single-domain strategies for task distribution and resource matching, and there are significant technical bottlenecks when facing cross-domain collaborative scheduling between multi-level computing resources. Summary of the Invention
[0004] The present invention provides a task scheduling method, system, device and storage medium based on a multi-domain computing power network to solve the problem that the power task scheduling mechanism in the prior art cannot intelligently schedule power tasks when facing a multi-domain computing power network.
[0005] According to one aspect of the present invention, a task scheduling method based on a multi-domain computing power network is provided, the method comprising:
[0006] Constructing a task feature vector based on the structured attribute data of the power task to be scheduled;
[0007] Determining a task level label corresponding to the power task based on the task feature vector and the task classification model;
[0008] Get the computing power resource status matrix corresponding to the current multi-domain computing power network;
[0009] Inputting the task level label, the computing resource state matrix, and the task feature vector into a multi-level strategy collaborative scheduling model to obtain a node selection action;
[0010] According to the node selection action, the power task is added to the task queue of the corresponding target computing power node to execute the power task based on the target computing power node.
[0011] According to another aspect of the present invention, a task scheduling system based on a multi-domain computing network is provided, the system comprising:
[0012] A data acquisition and processing module is used to construct a task feature vector based on the structured attribute data of the power task to be scheduled;
[0013] A task classification module, configured to determine a task level label corresponding to the power task based on the task feature vector and the task classification model;
[0014] The multi-domain computing power network module is used to obtain the computing power resource status matrix corresponding to the current multi-domain computing power network;
[0015] A multi-level strategy collaborative scheduling module, configured to input the task level label, the computing power resource state matrix, and the task feature vector into a multi-level strategy collaborative scheduling model to obtain a node selection action;
[0016] The task scheduling module is used to add the power task to the task queue of the corresponding target computing power node according to the node selection action, so as to execute the power task based on the target computing power node.
[0017] According to another aspect of the present invention, there is provided an electronic device, comprising: at least one processor; and
[0018] a memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the task scheduling method based on the multi-domain computing power network described in any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the task scheduling method based on the multi-domain computing power network described in any embodiment of the present invention when executed.
[0021] The embodiment of the present invention provides a task scheduling method, system, device and storage medium based on a multi-domain computing power network. The method includes: constructing a task feature vector based on the structured attribute data of the power task to be scheduled; determining the task level label corresponding to the power task based on the task feature vector and the task classification model; obtaining the computing power resource state matrix corresponding to the current multi-domain computing power network; inputting the task level label, the computing power resource state matrix and the task feature vector into a multi-level strategy collaborative scheduling model to obtain a node selection action; according to the node selection action, adding the power task to the task queue of the corresponding target computing power node to execute the power task based on the target computing power node. The method classifies the power tasks to be scheduled through the task classification model and determines how to schedule the power tasks through the multi-level strategy collaborative scheduling model. It can realize the intelligent scheduling of power tasks and solve the problem that the power task scheduling mechanism in the prior art cannot intelligently schedule power tasks when facing a multi-domain computing power network.
[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 A flowchart of a task scheduling method based on a multi-domain computing network provided in Example 1 of the present invention;
[0025] Figure 2 A schematic diagram of the structure of a task scheduling system based on a multi-domain computing network provided in the second embodiment of the present invention;
[0026] Figure 3 A schematic diagram of the structure of another task scheduling system based on a multi-domain computing network provided by an embodiment of the present invention;
[0027] Figure 4 Schematic diagram of the structure of an electronic device for a task scheduling method based on a multi-domain computing power network according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only embodiments of a part of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present invention. It should be understood that the various steps described in the method implementation mode of the present invention can be performed in different orders and / or in parallel. In addition, the method implementation mode may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0030] It should be noted that the terms "first," "second," and the like in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, any variations of the terms "including" and "having" are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.
[0031] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0033] Currently, most dispatching systems still use fixed task classification standards or experience-driven dispatching processes, making it difficult to accurately identify and dynamically respond to complex power business tasks. Especially when task attributes exhibit characteristics such as divisibility, latency sensitivity, and highly volatile resource demands, traditional task identification methods often fail to accurately capture task levels and dispatch priorities, resulting in decreased computing resource allocation efficiency and delayed system dispatch responses. Even though some methods attempt to introduce machine learning models for task classification, their ability to model structured data for power business operations is limited, and the classification results are decoupled from the actual dispatching logic, making it impossible to achieve linkage between the upper and lower levels of the dispatch chain.
[0034] On the other hand, the scheduling process of computing resources lacks a systematic multi-layer policy modeling mechanism. Some studies have applied reinforcement learning to resource scheduling problems, but most still focus on single-layer policy modeling and lack a description of the differences in scheduling levels between computing domains. In a multi-domain environment, computing power is widely distributed and computing power is uneven. Without a refined policy classification and decision coordination mechanism, the scheduling system is often unable to make optimal decisions based on task attributes and resource status. In addition, existing methods generally use a unified policy network for node selection, ignoring the fact that policies should be different and synergistic between different resource domains. This leads to slow convergence of model training and weak policy generalization ability. When faced with sudden growth of tasks or resource imbalance, it exhibits poor robustness and scalability.
[0035] The feedback learning mechanism for task scheduling is also a weak link in existing scheduling technologies. Feedback data after scheduling execution (such as response time, execution efficiency, and task completion status) is often used only for statistical analysis, lacking effective mechanisms for leveraging it for model parameter updates and policy optimization. While some reinforcement learning methods have considered incorporating reward mechanisms, these common training structures fail to effectively support the joint updating of high-level domain policies and low-level node policies. This creates an optimization disconnect between the policies, making it impossible to achieve an end-to-end adaptive scheduling closed loop.
[0036] Therefore, how to build an intelligent dispatching system that can integrate intelligent task classification, multi-domain computing power topology modeling, strategy hierarchical scheduling and feedback joint optimization is a key issue that needs to be broken through in the current power dispatching field. To solve the above problems, this embodiment introduces a task classification model based on TabTransformer, a multi-domain computing power structure, and a multi-level strategy collaborative scheduling module to build a full-process intelligent closed loop from task perception to node selection to strategy update, opening up the information link between task semantic understanding and resource scheduling, significantly improving the system's dispatching response capability, resource utilization and intelligent self-optimization level, and providing a new and feasible solution for intelligent power dispatching.
[0037] Example 1
[0038] Figure 1A flowchart of a task scheduling method based on a multi-domain computing power network is provided in accordance with the first embodiment of the present invention. The method can be applied to the case of intelligent scheduling of power tasks. The method can be executed by a task scheduling system based on a multi-domain computing power network, wherein the system can be implemented by software and / or hardware and is generally integrated on an electronic device. In this embodiment, the electronic device includes but is not limited to: computers and other devices.
[0039] like Figure 1 As shown, a task scheduling method based on a multi-domain computing network provided in the first embodiment of the present invention includes the following steps:
[0040] S110 : Constructing a task feature vector based on the structured attribute data of the power task to be scheduled.
[0041] The power task may be a power business task in the power system. The task feature vector may be a mathematical representation corresponding to the structured attribute data, and the task feature vector may be obtained by converting the structured attribute data into dimensions in a vector space.
[0042] In this embodiment, structured attribute data of the power task to be scheduled may be obtained, and the corresponding task feature vector may be obtained by processing the structured attribute data.
[0043] Furthermore, structured attribute data may include task type, execution deadline, computational load estimation, fault tolerance level, communication data volume, splitability flag, and scheduling window parameters. The execution deadline can be the latest time point at which the task must be completed, and the computational load estimation can be a quantitative assessment of the computational resources required for the task (e.g., CPU time, computing power, and memory access times). The fault tolerance level can be the level of ability of the task to continue to operate or recover normally in the event of a failure (e.g., hardware failure, network interruption, or data error). The communication data volume can be the amount of data transmitted between nodes (e.g., servers or cluster nodes in a distributed system) or modules during task execution. The splitability flag can be an indicator indicating whether the task can be split into multiple subtasks for parallel execution. The scheduling window parameters can be parameters used to describe the time range or conditions within which a task is allowed to be scheduled for execution. In this embodiment, the execution deadline and scheduling window parameters can be used to describe the time sensitivity of the power task, while the computational load estimation and communication data volume can reflect the level of computing and transmission resource utilization of the power task. The fault tolerance level and splitability flag are further used to determine whether scheduling delays are allowed and whether to support task splitting, thereby enabling the system to have stronger task differentiation processing capabilities and resource matching flexibility.
[0044] By introducing multiple key fields into the task's structured attribute data, including task type, execution time limit, computational load estimation, fault tolerance level, communication data volume, splitability flag, and scheduling window parameters, the expressiveness and distinguishing capabilities of task feature vectors in the scheduling decision-making process can be significantly improved. Compared to traditional scheduling methods that rely solely on task IDs or coarse-grained priorities for scheduling control, this structured attribute data constitutes a high-dimensional representation that comprehensively characterizes task execution requirements and resource adaptation characteristics. This provides a clear input foundation for subsequent task classification, strategy stratification, and computing power tiered scheduling, effectively improving the accuracy of task scheduling and the stability of system operation.
[0045] S120 : Determine a task level label corresponding to the power task based on the task feature vector and the task classification model.
[0046] The task classification model may be a task classification model built based on TabTransformer, and the task level label may be a label indicating the level of the power task.
[0047] In this embodiment, the task feature vector may be input into a task classification model, and a task level label corresponding to the power task may be obtained through the task classification model.
[0048] In one embodiment, the task classification model includes an embedding coding submodule, a feature interaction submodule and a classification output submodule. Accordingly, the task level label corresponding to the power task is determined based on the task feature vector and the task classification model, including: embedding and mapping the discrete fields in the task feature vector through the embedding coding submodule to obtain an embedding matrix; performing linear transformation on the embedding matrix through the feature interaction submodule to obtain an interaction representation matrix, and the interaction representation matrix is used to represent the weighted output features of each discrete field; performing feature aggregation, dimensionality reduction processing and nonlinear transformation on the interaction representation matrix through the classification output submodule to obtain the task level label corresponding to the power task.
[0049] The task classification model includes an embedding encoding submodule, a feature interaction submodule, and a classification output submodule. The feature interaction submodule can be built based on the Transformer structure, including at least one multi-head self-attention mechanism layer and a feedforward neural network layer.
[0050] In this embodiment, the embedding coding submodule can be used to receive the input of discrete fields in the task feature vector, perform embedding mapping on each discrete field, generate an embedding vector, and splice all the embedding vectors in the order of the original fields to form an embedding matrix. Through the multi-head self-attention mechanism of the feature interaction submodule, a linear transformation can be performed on the embedding matrix, and the query matrix, key matrix, and value matrix can be generated in combination with the trainable parameter matrix. The attention weight matrix can be calculated based on the query matrix and the key matrix, and the interaction representation matrix can be calculated based on the attention weight matrix and the value matrix. The interaction representation matrix can represent the weighted output features of each field. The classification output submodule can be used to perform feature aggregation and dimensionality reduction on the interaction representation matrix, connect a multi-layer fully connected neural network and perform nonlinear transformation, and finally output the task level label.
[0051] For example, the task feature vector x=[x1, x2, ..., x d ]In the task classification model built based on TabTransformer, the embedding encoding submodule of the task classification model is used to encode the discrete field x in the task feature vector. i Perform embedded coding to generate the embedding matrix E∈R d×k , where d is the number of discrete fields, k is the embedding dimension, and the i-th row vector e in the embedding matrix i Represents a discrete field x i The embedding vector of . The multi-head self-attention mechanism of the feature interaction submodule performs a linear transformation on the embedding matrix E to generate the query matrix Q, key matrix K, and value matrix V. The calculation formula is:
[0052] Q=E·W Q ; K = E·W K ; V = E·W V W Q ;
[0053] Among them, W Q 、W K 、W V are the trainable weight matrices corresponding to the query matrix, key matrix, and value matrix, respectively, W Q ∈R k ×k , W K ∈R k×k , W V ∈R k×k .
[0054] Calculate the attention weight matrix A∈R d×d :
[0055]
[0056] Among them, A i,jrepresents the attention weight of the i-th field to the j-th field.
[0057] Perform weighted calculation on the attention weight matrix A and the value matrix V to generate the interaction representation matrix H∈R d×k :
[0058] H=A·V;
[0059] The interaction representation matrix H is subjected to feature aggregation and dimensionality reduction processing through the classification output submodule, and a multi-layer fully connected neural network is connected and nonlinear transformation is performed to output the task level label y∈{y1,y2,...,y c}, where c represents the number of task level categories.
[0060] This embodiment introduces a task classification model built based on TabTransformer, which can automatically learn the complex relationships between task attributes, fully explore the interactive dependencies between feature fields, and effectively enhance the modeling ability and classification accuracy of the structured features of power dispatching tasks. TabTransformer combines embedded coding with self-attention mechanism to accurately model the implicit semantic relationships between task types, perform saliency modeling on key fields in task features (such as execution time limit, computational estimation, fault tolerance level, communication volume, etc.), and can effectively distinguish task levels from scheduling priorities. This mechanism not only improves the accuracy of task classification, but also provides high-quality prior labels for the subsequent scheduling decision process, realizing a semantic closed loop from task identification to scheduling call.
[0061] Compared to traditional shallow neural networks or rule-driven classification methods, this model first embeds and maps discrete task fields through an embedding encoding submodule, preserving the semantic information between the original fields and achieving uniform dimensional alignment. This addresses the problem of discrete task attributes being difficult to directly incorporate into model training. Furthermore, the feature interaction submodule utilizes a multi-head self-attention mechanism to deeply model the interdependencies between fields. By generating a query, key, and value matrix to construct attention weights, the model automatically focuses on field combinations that have the greatest impact on the classification results. This improves the representation of task semantic features and enables adaptive discrimination across diverse task scenarios, resulting in highly accurate task-level label output and enhanced discriminative task-level label prediction. The classification output submodule aggregates and reduces the dimensionality of the interaction representation matrix, connecting it to a multi-layer nonlinear fully connected network to enhance the representation of classification boundaries. Ultimately, it outputs stable and generalizable task-level labels. This architecture not only improves the accuracy and robustness of task classification but also provides accurate and structured priors for subsequent computing domain selection and scheduling policy matching, making it a crucial step in enabling proactive awareness for intelligent scheduling.
[0062] S130. Obtain the computing power resource status matrix corresponding to the current multi-domain computing power network.
[0063] The multi-domain computing network is a distributed computing infrastructure that integrates diverse geographical regions, heterogeneous computing resources, and diverse network architectures (such as cloud computing, edge computing, fog computing, and core networks). The computing resource status matrix is a mathematical model that can quantitatively describe the real-time status of various resources in the computing network.
[0064] In this embodiment, the computing power resource status matrix corresponding to the current multi-domain computing power network can be obtained.
[0065] In one embodiment, obtaining the computing power resource status matrix corresponding to the current multi-domain computing power network includes: constructing a multi-domain computing power network structure including different categories of computing power domains; assigning unique identifiers to all computing power nodes in different computing power domains to construct a global computing power node set; collecting resource status data of each computing power node in the global computing power node set; stacking the resource status data based on the unique identifier to obtain the computing power resource status matrix corresponding to the current multi-domain computing power network.
[0066] Computing domains can include edge computing domains, regional computing domains, and central cloud computing domains, each of which contains computing nodes. The global computing node set can be the collection of computing nodes in all computing domains. Resource status data can refer to real-time or near-real-time data reflecting the current operating status, performance indicators, availability, and resource usage of computing nodes.
[0067] In this embodiment, a multi-domain computing power network structure can be constructed, which includes different categories of computing power domains. Unique identifiers can be assigned to all computing power nodes in different computing power domains, and a global computing power node set can be constructed. The resource status data of each computing power node in the global computing power node set is collected, so that the resource status data can be stacked based on the unique identifier to obtain the computing power resource status matrix corresponding to the current multi-domain computing power network.
[0068] For example, a multi-domain computing network structure is constructed, which includes three types of computing domains: edge computing domain, regional computing domain, and central cloud computing domain. Each type of computing domain includes several computing nodes. Among them, the edge computing domain can be deployed at the edge of the power grid, including a set of edge computing nodes. m e The number of edge nodes. The regional computing domain can be deployed in the regional dispatch center or the central site, including the regional computing node set m r The central cloud computing domain can be deployed on a centralized computing platform or cloud service node, including a collection of central cloud computing nodes. mc is the number of central nodes. Merge all computing power node sets to form a global computing power node set N=N e ∪N r ∪N c , assign a unique identifier i to each computing power node d And the computing domain type i , the node structure can be defined as n i =(id i , type i ), where type i ∈{edge, regional, cloud}. Collect resource status data of each computing node, which may include CPU usage Memory usage Current task queue length q i , network delay estimate l i And the historical load average h i , organize the resource status data of each computing power node into a five-dimensional resource status vector Among them, r i is the resource state vector of the i-th computing power node. All five-dimensional resource state vectors are assigned a unique identifier id i The stack generates the computing power resource state matrix R, which is used as the input of the multi-level strategy collaborative scheduling model.
[0069] This embodiment solves the problem of resource domain fragmentation and fixed scheduling paths in traditional scheduling methods by constructing a multi-domain computing power network structure, and collects the resource status of each computing power node in the edge computing power domain, regional computing power domain and central cloud computing power domain respectively to generate a computing power resource status matrix, which can effectively achieve comprehensive perception and unified representation of heterogeneous computing resources. Compared with the traditional scheduling system that is limited to a single computing power layer or relies on static resource description, the multi-domain modeling method adopted in this embodiment can dynamically reflect the resource differences and real-time status between different computing power domains, providing a globally visible scheduling basis for the policy model. The resource status matrix contains key performance indicators such as CPU utilization, memory occupancy, current load, queue length, etc., which can accurately characterize the processing capacity and schedulability of the computing power node, so that the subsequent high-level policy sub-module can make reasonable computing power domain selection based on task requirements and resource load, while providing accurate local resource input for the low-level policy and scheduling agent sub-module. This embodiment can dynamically select the optimal computing power domain based on the task level and the current system load status, realize the intelligent migration and load balancing of tasks between multiple domains, significantly improve the scheduling system's adaptability and perception granularity to dynamic changes in resources in a multi-domain computing power environment, and is the basic guarantee for achieving efficient resource matching and hierarchical collaborative scheduling.
[0070] S140: Input the task level label, the computing resource state matrix, and the task feature vector into a multi-level strategy collaborative scheduling model to obtain a node selection action.
[0071] Among them, the multi-level strategy collaborative scheduling model can be a hierarchical framework for solving resource allocation, task scheduling or decision-making coordination problems in complex systems.
[0072] In this embodiment, the task level label, computing resource state matrix and task feature vector can all be input into the multi-level strategy collaborative scheduling model to obtain the node selection action.
[0073] In one embodiment, the multi-level policy collaborative scheduling model includes a high-level policy sub-module, a low-level policy sub-module and a scheduling agent sub-module. Accordingly, the task level label, the computing power resource state matrix and the task feature vector are input into the multi-level policy collaborative scheduling model to obtain a node selection action, including: executing a policy decision process on the task level label and the computing power resource state matrix through the high-level policy sub-module to obtain the target computing power domain of the power task; obtaining the corresponding resource state vector based on the candidate nodes of the candidate computing power node set in the target computing power domain; constructing a local joint state representation based on the resource state vector and the task feature vector through the low-level policy sub-module; and performing a policy evaluation on the local joint state representation through the scheduling agent sub-module to generate a node selection action.
[0074] The multi-level policy collaborative scheduling model can include a high-level policy submodule, a low-level policy submodule, and a scheduling agent submodule. The high-level policy submodule can include an online learning component for receiving historical scheduling results and updating decision network parameters. By introducing this online learning component, historical scheduling results can be continuously absorbed and policy parameters can be dynamically updated, thereby enhancing the model's adaptability to different business scenarios and resource distribution states. The scheduling agent submodule can be constructed based on the Meta-Q Learning mechanism.
[0075] In this embodiment, the high-level policy sub-module can be used to receive the task level label and the computing power resource status matrix as the joint state input, perform feature extraction and policy evaluation on the current system state through the high-level policy network, and obtain the corresponding target computing power domain. The low-level policy sub-module can receive the resource status data of each node in the target computing power domain, and combine the task feature vector of the current power task to construct a local joint state representation as the input of the scheduling sub-module; the scheduling agent sub-module can be used to receive the local joint state representation constructed by the low-level policy sub-module and generate a node selection action.
[0076] This embodiment significantly improves the hierarchical decision-making capabilities and scheduling adaptability of smart power tasks in a multi-domain computing resource environment by constructing a multi-level policy collaborative scheduling model comprising a high-level policy submodule, a low-level policy submodule, and a scheduling agent submodule. The high-level policy submodule combines input task level labels with a computing resource state matrix, utilizes a policy network to perform global system state perception and domain-level policy evaluation, and dynamically outputs computing domain selection instructions, effectively enabling intelligent task diversion between edge, regional, and central cloud computing domains. The low-level policy submodule combines the node resource state in the target computing domain with the current task characteristics to construct a local joint state representation, ensuring that the input received by the scheduling agent submodule is targeted and contextually supported. The scheduling agent submodule utilizes a scheduling mechanism based on Meta-Q Learning, combining Q-value estimation with a fast meta-learning update strategy. This allows for continuous optimization of the policy network in complex and changing scheduling environments, enabling fast and accurate node action decisions in diverse node environments while exhibiting excellent transfer and generalization capabilities.
[0077] For example, after obtaining the target computing power domain, the candidate computing power node set in the target computing power domain determined by the high-level strategy submodule can be extracted, and the corresponding resource state vector r is extracted for the candidate nodes in the candidate computing power node set. j , the resource state vector r j With the task feature vector x t Splice the input to the low-level strategy submodule to construct a local joint state representation
[0078]
[0079] The local joint state is represented as The input is sent to the scheduling agent submodule, and the low-level policy network inside the scheduling agent submodule performs policy evaluation to generate node selection actions:
[0080]
[0081] Among them, a L is the selected target node, π L is the low-level policy network inside the scheduling agent submodule, θ L It is the parameter set of the low-level policy network inside the scheduling agent submodule.
[0082] After determining the target computing power domain, this embodiment further obtains the resource state vector of each computing power node within the domain and, combined with the task feature vector, inputs it into the low-level policy submodule of the multi-level policy collaborative scheduling model. The scheduling agent submodule is then called to generate node selection actions, thereby significantly improving the refined decision-making capabilities and resource utilization efficiency of task scheduling at the node level. By introducing a local joint state representation, it can fully consider the computing requirements and attribute characteristics of the current task and comprehensively analyze multi-dimensional indicators such as the real-time load, available resources, and operating status of each node in the target computing power domain, thereby making the node selection process context-aware. Compared to the node distribution method based on static priority or simple resource matching in traditional scheduling strategies, the policy generation mechanism is based on the decision logic of reinforcement learning and can continuously learn the optimal scheduling path in a dynamic environment, significantly improving the accuracy and adaptability of task scheduling. In addition, through the linkage between the low-level policy network and the scheduling agent submodule, the node selection action has trainability and feedback optimization capabilities, which can ensure that tasks are always assigned to the optimal or suboptimal resource nodes within the target computing power domain, thereby improving the scheduling stability, computing power utilization, and task response speed of the entire system.
[0083] In one embodiment, the high-level policy submodule executes a policy decision process on the task level label and the computing power resource status matrix to obtain the target computing power domain of the power task, including: inputting the task level label and the computing power resource status matrix as a joint state into the high-level policy submodule to construct a global joint state representation; performing feature extraction and policy evaluation on the global joint state representation based on the high-level policy network of the high-level policy submodule to generate a computing power domain selection instruction; and determining the target computing power domain corresponding to the power task based on the computing power domain selection instruction.
[0084] In this embodiment, the task registration label and the computing power resource state matrix can be input into the high-level policy sub-module as a joint state, and a global joint state representation can be constructed through the high-level policy sub-module. The global joint state representation is subjected to feature extraction and strategy evaluation based on the high-level policy network of the high-level policy sub-module to generate a computing power domain selection instruction. Based on the computing power domain selection instruction, the target computing power domain corresponding to the power task can be determined.
[0085] For example, the task level label and the computing resource state matrix are input as the joint state into the high-level strategy submodule. After the high-level strategy submodule receives the input, it constructs a global joint state representation s H , the global joint state representation can be used to represent the scheduling matching state of the current power task to be scheduled under the overall configuration of system resources. H , feature extraction and policy evaluation are performed through the high-level policy network to generate computing power domain selection instructions:
[0086] a H =π H (s H θ H );
[0087] Among them, a H Select instructions for the power domain, π H is the high-level policy network, θ H It is a set of parameters for the high-level policy network. Based on the power domain selection instruction, the target power domain corresponding to the task to be scheduled can be determined.
[0088] This embodiment inputs the task level label and computing power resource status matrix into the high-level policy submodule of the multi-level policy collaborative scheduling model, executes the policy decision process and generates computing power domain selection instructions, thereby determining the target computing power domain of the task to be scheduled. This can significantly improve the intelligent adaptability of the scheduling system in a multi-domain heterogeneous environment. Unlike traditional methods that rely on fixed rules or static matching strategies for domain selection, this high-level policy submodule can dynamically determine whether the task should be prioritized for scheduling to edge, regional, or central cloud computing power domains based on the combined characteristics of the task level and the current system resource status, realizing the transformation of the scheduling strategy from static preset to dynamic decision-making. This mechanism effectively alleviates the problem of resource concentration or imbalanced allocation, improves the load balancing capability between computing power domains, and can prioritize the allocation of delay-sensitive tasks to low-latency computing power domains and guide computing-intensive tasks to high-performance resource domains, ultimately achieving a precise match between task requirements and computing power capabilities, improving the overall efficiency of scheduling, response time, and system resource utilization.
[0089] S150. According to the node selection action, the power task is added to the task queue of the corresponding target computing power node to execute the power task based on the target computing power node.
[0090] Among them, the target computing power node can be a computing power node suitable for executing power tasks.
[0091] In this embodiment, the task to be scheduled can be added to the task queue of the target computing power node so that the power task can be executed based on the target computing power node. Before the task is executed, the task execution parameters of the target computing power node can also be configured.
[0092] A task scheduling method based on a multi-domain computing power network provided in a first embodiment of the present invention includes: constructing a task feature vector based on the structured attribute data of the power task to be scheduled; determining the task level label corresponding to the power task based on the task feature vector and the task classification model; obtaining the computing power resource state matrix corresponding to the current multi-domain computing power network; inputting the task level label, the computing power resource state matrix and the task feature vector into a multi-level strategy collaborative scheduling model to obtain a node selection action; according to the node selection action, adding the power task to the task queue of the corresponding target computing power node to execute the power task based on the target computing power node. This method classifies the power tasks to be scheduled through the task classification model, and determines how to schedule the power tasks through the multi-level strategy collaborative scheduling model, which can realize the intelligent scheduling of power tasks and solves the problem that the power task scheduling mechanism in the prior art cannot intelligently schedule power tasks when facing a multi-domain computing power network.
[0093] Based on the above embodiment, a modified embodiment of the above embodiment is proposed. It should be noted that, in order to simplify the description, only the differences from the above embodiment are described in the modified embodiment.
[0094] In one embodiment, the multi-level strategy collaborative scheduling model also includes a collaborative optimization submodule. Accordingly, after executing the power task, the method also includes: obtaining scheduling feedback information when executing the power task; constructing a state-action value function and a strategy error loss function based on the scheduling feedback information, the local joint state representation, the target computing power node and the reward signal; through the collaborative optimization submodule, updating the high-level strategy network parameters of the high-level strategy submodule based on the state-action value function, and updating the low-level strategy network parameters of the scheduling agent submodule based on the strategy error loss function.
[0095] The scheduling feedback information may include task response time, resource utilization, and task completion status identification.
[0096] In this embodiment, the multi-level policy collaborative scheduling model also includes a collaborative optimization submodule, which can be used to coordinate the parameter update process of the high-level policy submodule and the low-level policy submodule; after the task scheduling is completed, the collaborative optimization submodule can receive the reward signal generated by the scheduling agent submodule, construct a state-action value function, and optimize and adjust the parameters of the high-level policy network and the low-level policy network within the scheduling agent submodule through the Meta-Q Learning mechanism to achieve a joint training structure of multi-level scheduling strategies. After controlling the target computing power node to start the task execution process, this embodiment can collect scheduling feedback information during the task execution process and input the scheduling feedback information into the multi-level policy collaborative scheduling model to update the parameters of the high-level strategy and the low-level strategy. For example, the state-action value function and the policy error loss function can be constructed through the scheduling feedback information, the local joint state representation, the target computing power node and the reward signal, and the high-level policy network parameters of the high-level policy submodule are updated based on the state-action value function, and the low-level policy network parameters of the scheduling agent submodule are updated based on the policy error loss function.
[0097] The collaborative optimization submodule of this embodiment connects the parameters of high- and low-level policy networks, constructs a state-action value function based on task feedback reward signals, and implements joint optimization training of hierarchical policy networks, avoiding isolated policy updates and disconnected top- and bottom-level layers. Overall, this modular system strengthens the policy's hierarchical perception, dynamic decision-making, and closed-loop optimization capabilities in multi-domain scheduling, serving as the core support structure for efficient and robust intelligent power dispatch.
[0098] In one embodiment, constructing a state-action value function and a policy error loss function based on the scheduling feedback information, the local joint state representation, the target computing power node and the reward signal includes: constructing a feedback indicator set according to the scheduling feedback information; calculating a reward signal according to the feedback indicator set; constructing a state-action value function based on the local joint state representation, the target computing power node and the parameter set; and constructing a policy error loss function based on the reward signal and the state-action value function.
[0099] In this embodiment, a feedback indicator set can be constructed by scheduling feedback information, a reward signal can be calculated based on the feedback indicator set, and a state-action value function can be constructed based on the local joint state representation, the target computing power node and the parameter set. The strategy error loss function can be constructed based on the reward signal and the state-action value function.
[0100] For example, based on the scheduling feedback information, a feedback indicator set f is constructed. t ={t r ,u r , m r}, where tr is the task response time, u r is the resource utilization rate, m r It is the task completion status mark. According to the feedback indicator set f t Calculate the reward signal r t :
[0101]
[0102] Among them, α, β, and γ represent weighted coefficients, which are used to balance the effects of delay, utilization, and completion status. Target node a L and reward signal r t Forming triples Used to construct state-action value function:
[0103]
[0104] Where Q is the state-action value function and φ is the parameter set of the state-action value function. The constructed policy error loss function is:
[0105]
[0106] The Meta-Q Learning mechanism based on the loss function can be used to adjust the low-level policy network parameters θ within the scheduling agent submodule. L Perform optimization update. The reward signal r t Synchronously pass it to the collaborative optimization submodule, which coordinates the joint update of the parameters of the high-level strategy submodule and the scheduling agent submodule; according to the state-action value function and the current strategy error, the high-level strategy network parameters θ are updated respectively. H and the low-level policy network parameters θ within the scheduling agent submodule L , to achieve a joint training structure of multi-level scheduling strategies.
[0107] This embodiment inputs scheduling feedback information into a multi-level policy collaborative scheduling model and jointly updates the parameters of the high-level and low-level policy networks, significantly enhancing the scheduling system's self-learning and dynamic adaptability. By collecting multi-dimensional feedback information after task execution, including task response time, resource utilization, and task completion status, the system can construct a reward signal reflecting the actual operational performance. Based on this information, it optimizes the policy network parameters, achieving collaborative evolution between the policy networks and avoiding the policy optimization fragmentation problem in traditional multi-policy systems. This mechanism overcomes the limitations of traditional scheduling systems that rely on static rules or fixed policies, enabling the policy model to continuously learn and optimize in a real-world scheduling environment, effectively improving the model's robustness to complex working conditions, task fluctuations, and resource status changes. Furthermore, the use of a joint high- and low-level policy update structure ensures coordination and consistency between domain-level scheduling and node-level distribution policies, avoiding policy fragmentation and target drift. Overall, this feedback-driven parameter update mechanism establishes an intelligent closed-loop scheduling system, significantly improving the system's scheduling accuracy, resource matching efficiency, and long-term operational stability, and possesses significant engineering practical value.
[0108] The embodiments of the present invention provide several specific implementation methods based on the technical solutions of the above embodiments.
[0109] As a specific implementation method of this embodiment, in order to verify the feasibility of the method of this embodiment in implementation, this method is applied to an intelligent computing power scheduling platform, which covers multiple key cities and is oriented to various scenarios such as regional power grid equipment operation and maintenance, power trading regulation, and new energy access coordination. It needs to process an average of about 25,000 scheduling tasks every day. These tasks are diverse in type and have high timeliness requirements, and the computing power resources they rely on span three types of heterogeneous nodes: edge, regional, and central cloud. However, in traditional scheduling systems, task level division is rough, resource allocation lags, and the policy update cycle is long. Especially during the high-load period in summer, problems such as task backlogs, scheduling response timeouts, and resource conflicts often occur, which directly affect scheduling efficiency and system stability.
[0110] During the peak summer load period, the system first uses the data acquisition and processing module to obtain real-time structural attributes such as task type, computational load estimation, fault tolerance level, and splitability. It also simultaneously collects resource status information such as CPU, memory, and task queues for each node in the three computing domains. After standardization, the system automatically constructs task feature vectors and a computing resource status matrix, providing input for subsequent scheduling processes.
[0111] During the trial run, task feature vectors were input into a task classification module built on TabTransformer, improving the system's task classification accuracy to 94.8% to 95.1%. Task-level labels drive the high-level policy submodule to select computing domains, automatically determining whether tasks are suitable for dispatch to edge, regional, or central cloud computing domains, and achieving resource load balancing across multiple domains. Subsequently, the low-level policy submodule and the scheduling agent submodule, based on the Meta-Q Learning mechanism, combine task characteristics with the target domain resource state to construct a local joint state representation and generate the optimal node selection action. After task completion, the system collects feedback information such as task response time, node load, and execution success rate. The collaborative optimization submodule then adjusts policy parameters in real time, enabling adaptive evolution of the policy.
[0112] By selecting data from different time periods, we compared and analyzed the system's performance on multiple indicators, including task classification accuracy, scheduling response time, success rate, resource utilization, and error scheduling rate. The results are shown in Table 1 below:
[0113] Table 1 Task processing index results
[0114]
[0115]
[0116] As shown in Table 1, this embodiment achieves all-round and multi-dimensional optimization of the core performance indicators in the power intelligent dispatching system, especially in dealing with large-scale task concurrency and high resource pressure scenarios. In terms of task classification accuracy, traditional systems often rely on preset rules or static parameters to divide tasks into different levels, and the accuracy rate has long hovered around 85%. It is easy to have task level division deviations, resulting in resource misallocation or scheduling failure. After this embodiment introduces the TabTransformer-based task classification model, the system can make full use of task structured attributes such as execution time limit, computational estimation, fault tolerance level, etc., deeply model the implicit differences between tasks, and achieve more fine-grained level division. During the trial operation, the classification accuracy of the eight cities remained between 94.6% and 95.1%, greatly improving the reliability of the source of strategic decision-making.
[0117] In terms of average response time, the original system had an average response time of 2.87 seconds under high load, exceeding 3.5 seconds during peak periods. This resulted in significant delays in the scheduling process, which easily led to task backlogs and queue congestion. By coordinating high- and low-level policy calls and implementing dynamic policy learning driven by the Meta-Q Learning mechanism, this implementation reduces the average response time to 1.17-1.25 seconds, a reduction of over 55%. Regions like City B and City E maintained stable response times below 1.18 and 1.17 seconds, respectively, demonstrating exceptional timeliness control capabilities.
[0118] In terms of task execution success rate, this embodiment significantly reduces task interruption and failure rates due to insufficient resources or uneven node load. The task success rate in traditional platforms is typically maintained between 93% and 95%. However, by optimizing domain-level scheduling paths and node-level resource adaptation strategies, the success rate of this system has generally increased to over 98%, with a maximum of 98.7% (City E). This means that at the same task density, the platform can reduce the retry overhead of hundreds or even thousands of tasks, indirectly improving the overall system throughput.
[0119] The improvement in node resource utilization, especially CPU utilization, was particularly significant. Because traditional scheduling strategies lack dynamic load balancing capabilities, some edge nodes often become overloaded, leaving central cloud resources idle. However, this implementation, by constructing a local joint state vector and real-time domain-level node evaluation, achieved a more rational node allocation mechanism. During the trial run, the average CPU utilization across all regions increased by between +17.1% and +18.4%, resulting in a more balanced and efficient resource utilization structure.
[0120] Furthermore, in terms of scheduling errors, traditional scheduling systems often experience task assignment failures due to factors such as task misclassification and delayed node information, resulting in an average daily scheduling error rate exceeding 3%. However, after the deployment of this invention, this metric dropped below 2% in all test cities, with the lowest being 1.1% (in City H), demonstrating significantly enhanced fault tolerance and robustness in the scheduling chain.
[0121] Finally, in terms of computing domain selection accuracy, this system achieves continuous learning and optimization for the dynamic adaptation of tasks and domain resources. The average domain selection accuracy across all cities is above 95%, with regions like H and E reaching nearly 97%. This is particularly critical in multi-domain scheduling strategies, significantly reducing duplicate resource allocation calls and inter-domain forwarding, further improving the overall system throughput.
[0122] Compared with existing methods, this embodiment not only performs well in static indicators, but also demonstrates strong adaptability and scalability in the dynamic scheduling process, forming a closed-loop system from "high-quality input modeling" to "intelligent domain scheduling" to "task-level execution feedback", which has strong engineering practicality and promotion and application value. This embodiment integrates key technologies such as structured task perception, deep task classification, multi-domain computing power modeling and reinforcement learning scheduling optimization to construct an intelligent closed-loop scheduling process from task understanding to scheduling execution to strategy feedback. By introducing a task classification model built based on TabTransformer, the structured attributes of power business tasks are accurately identified and the task level is determined, effectively improving the hierarchical processing capability of scheduling tasks. In terms of resource modeling, a multi-domain computing power network topology including edge, regional and central clouds is constructed to collect and manage resource status information of nodes in each computing power domain to ensure the comprehensiveness and real-time nature of scheduling input.
[0123] Example 2
[0124] Figure 2 This is a structural diagram of a task scheduling system based on a multi-domain computing power network provided in Example 2 of the present invention. The device can be used in situations where power tasks are intelligently scheduled, wherein the device can be implemented by software and / or hardware and is generally integrated on an electronic device.
[0125] like Figure 2 As shown, the device includes:
[0126] The data acquisition and processing module 210 is used to construct a task feature vector based on the structured attribute data of the power task to be scheduled;
[0127] A task classification module 220 is configured to determine a task level label corresponding to the power task based on the task feature vector and the task classification model;
[0128] The multi-domain computing network module 230 is used to obtain the computing resource status matrix corresponding to the current multi-domain computing network;
[0129] A multi-level strategy collaborative scheduling module 240 is configured to input the task level label, the computing resource state matrix, and the task feature vector into a multi-level strategy collaborative scheduling model to obtain a node selection action;
[0130] The task scheduling module 250 is used to add the power task to the task queue of the corresponding target computing power node according to the node selection action, so as to execute the power task based on the target computing power node.
[0131] This embodiment provides a task scheduling system based on a multi-domain computing power network, comprising: a data acquisition and processing module for constructing a task feature vector based on the structured attribute data of the power task to be scheduled; a task classification module for determining the task level label corresponding to the power task based on the task feature vector and the task classification model; a multi-domain computing power network module for obtaining the computing power resource state matrix corresponding to the current multi-domain computing power network; a multi-level strategy collaborative scheduling module for inputting the task level label, the computing power resource state matrix, and the task feature vector into the multi-level strategy collaborative scheduling model to obtain a node selection action; and a task scheduling module for adding the power task to the task queue of the corresponding target computing power node based on the node selection action, so as to execute the power task based on the target computing power node. By classifying the power tasks to be scheduled through the task classification model and determining how to schedule the power tasks through the multi-level strategy collaborative scheduling model, intelligent scheduling of power tasks can be achieved, solving the problem that the power task scheduling mechanism in the prior art cannot intelligently schedule power tasks when facing a multi-domain computing power network.
[0132] Furthermore, the task classification model includes an embedding encoding submodule, a feature interaction submodule, and a classification output submodule. Accordingly, the task classification module 220 includes:
[0133] An embedding encoding submodule, configured to embed and map discrete fields in the task feature vector to obtain an embedding matrix;
[0134] a feature interaction submodule, configured to perform a linear transformation on the embedding matrix to obtain an interaction representation matrix, wherein the interaction representation matrix is used to represent the weighted output features of each discrete field;
[0135] The classification output submodule is used to perform feature aggregation, dimensionality reduction and nonlinear transformation on the interaction representation matrix to obtain a task level label corresponding to the power task.
[0136] Furthermore, the multi-domain computing network module 230 is specifically configured to:
[0137] Build a multi-domain computing power network structure that includes different types of computing power domains;
[0138] Assign unique identifiers to all computing nodes in different computing domains to build a global computing node set;
[0139] Collecting resource status data of each computing power node in the global computing power node set;
[0140] The resource status data is stacked based on the unique identifier to obtain a computing power resource status matrix corresponding to the current multi-domain computing power network.
[0141] Furthermore, the multi-level policy collaborative scheduling model includes a high-level policy submodule, a low-level policy submodule, and a scheduling agent submodule. Accordingly, the multi-level policy collaborative scheduling module 240 includes:
[0142] A high-level strategy submodule, configured to execute a strategy decision process on the task level label and the computing power resource state matrix to obtain a target computing power domain for the power task;
[0143] The bottom layer strategy submodule is configured to obtain a corresponding resource state vector based on a candidate node in the candidate computing power node set in the target computing power domain; and construct a local joint state representation based on the resource state vector and the task feature vector;
[0144] The scheduling agent submodule is used to perform strategy evaluation on the local joint state representation and generate node selection actions.
[0145] Furthermore, the high-level strategy submodule is specifically used to:
[0146] Inputting the task level label and the computing resource state matrix as a joint state into the high-level strategy submodule to construct a global joint state representation;
[0147] A high-level policy network based on the high-level policy submodule performs feature extraction and policy evaluation on the global joint state representation to generate a computing power domain selection instruction;
[0148] The target computing power domain corresponding to the power task is determined based on the computing power domain selection instruction.
[0149] Furthermore, the multi-level strategy collaborative scheduling model also includes a collaborative optimization submodule. Accordingly, the collaborative optimization submodule is used to:
[0150] Obtaining scheduling feedback information when executing the power task;
[0151] Constructing a state-action value function and a policy error loss function based on the scheduling feedback information, the local joint state representation, the target computing power node and the reward signal;
[0152] Through the collaborative optimization submodule, the high-level policy network parameters of the high-level policy submodule are updated based on the state-action value function, and the low-level policy network parameters of the scheduling agent submodule are updated based on the policy error loss function.
[0153] Furthermore, constructing a state-action value function and a policy error loss function based on the scheduling feedback information, the local joint state representation, the target computing power node and the reward signal includes:
[0154] Constructing a feedback indicator set according to the scheduling feedback information;
[0155] Calculating a reward signal based on the feedback indicator set;
[0156] Constructing a state-action value function based on the local joint state representation, the target computing power node and the parameter set;
[0157] A policy error loss function is constructed based on the reward signal and the state-action value function.
[0158] This embodiment constructs a multi-level policy collaborative scheduling module, consisting of a high-level policy submodule, a low-level policy submodule, a scheduling agent submodule, and a collaborative optimization submodule. It employs a policy optimization method based on the Meta-Q Learning mechanism to generate computing domain selection and node selection actions, respectively. It also drives the joint optimization and update of policy parameters through task execution feedback, forming a policy adaptive evolution capability. Ultimately, it enables efficient hierarchical scheduling of multiple types of power tasks across multi-domain computing resources, offering advantages such as fast scheduling response, high resource utilization, and sustainable policy optimization. It is suitable for various distributed power scheduling platforms and computing power control systems.
[0159] For example, Figure 3 A structural diagram of another task scheduling system based on a multi-domain computing network provided by an embodiment of the present invention is shown as follows: Figure 3 As shown, the system has a complete modular architecture, from task perception to node scheduling and then to policy self-optimization. This significantly enhances the system's intelligent dispatching capabilities and system response efficiency in complex power business environments. The system's data acquisition and processing module standardizes the integration of task structure attributes and multi-domain computing resource status, constructing a unified dispatch input and providing high-quality data support for subsequent computational steps. The introduction of a TabTransformer-based task classification module enables in-depth extraction and hierarchical judgment of multi-dimensional task attributes, enhancing the ability to drive differentiated dispatch policies. The multi-domain computing network module establishes a three-tiered heterogeneous computing topology: edge, regional, and central cloud, enabling cross-domain resource visibility and unified management of dispatch calls. The core multi-level policy collaborative scheduling module combines a hierarchical structure with the Meta-Q Learning mechanism to support joint modeling of high- and low-level policies and adaptive execution path selection. The collaborative optimization submodule, driven by scheduling feedback, also jointly updates high- and low-level policy parameters. The task scheduling module, driven by policy, distributes and executes node-level tasks, forming a closed-loop system from classification judgment to resource scheduling and feedback optimization. The overall system has the characteristics of precise task adaptation, flexible resource call, and self-driven strategy evolution. It shows stronger stability, scalability and intelligent optimization capabilities in scenarios of multi-task concurrency and dynamic fluctuations in power resources.
[0160] This embodiment organically integrates task type understanding, domain-level scheduling, node-level distribution, and strategy optimization to build an end-to-end, scalable, and evolvable intelligent task scheduling architecture. This not only improves the intelligent scheduling capabilities of the power business scheduling system in scenarios such as complex computing resource distribution, diverse task requirements, and frequent state changes, but also provides a feasible reference framework for future high-load computing power scheduling systems such as the Industrial Internet and grid-edge cloud collaborative control. Overall, this embodiment has significant advantages such as high scheduling efficiency, strong strategy self-learning, and good system scalability, and has broad prospects for practical engineering applications.
[0161] Example 3
[0162] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0163] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0164] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0165] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the task scheduling method based on the multi-domain computing network.
[0166] In some embodiments, the task scheduling method based on the multi-domain computing power network can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the task scheduling method based on the multi-domain computing power network described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the task scheduling method based on the multi-domain computing power network by any other appropriate means (for example, by means of firmware).
[0167] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0168] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0169] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0170] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0171] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0172] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0173] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0174] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A task scheduling method based on a multi-domain computing power network, characterized in that: The method comprises: Constructing a task feature vector based on the structured attribute data of the power task to be scheduled; Determining a task level label corresponding to the power task based on the task feature vector and the task classification model; Get the computing power resource status matrix corresponding to the current multi-domain computing power network; Inputting the task level label, the computing resource state matrix, and the task feature vector into a multi-level strategy collaborative scheduling model to obtain a node selection action; According to the node selection action, the power task is added to the task queue of the corresponding target computing power node to execute the power task based on the target computing power node.
2. The method according to claim 1, characterized in that The task classification model includes an embedding coding submodule, a feature interaction submodule, and a classification output submodule. Accordingly, determining the task level label corresponding to the power task based on the task feature vector and the task classification model includes: Performing embedding mapping on the discrete fields in the task feature vector by the embedding coding submodule to obtain an embedding matrix; Performing a linear transformation on the embedding matrix through the feature interaction submodule to obtain an interaction representation matrix, wherein the interaction representation matrix is used to represent the weighted output features of each discrete field; The interaction representation matrix is subjected to feature aggregation, dimensionality reduction processing and nonlinear transformation by the classification output submodule to obtain a task level label corresponding to the power task.
3. The method according to claim 1, characterized in that Obtaining the computing power resource status matrix corresponding to the current multi-domain computing power network includes: Build a multi-domain computing power network structure that includes different types of computing power domains; Assign unique identifiers to all computing nodes in different computing domains to build a global computing node set; Collecting resource status data of each computing power node in the global computing power node set; The resource status data is stacked based on the unique identifier to obtain a computing power resource status matrix corresponding to the current multi-domain computing power network.
4. The method according to claim 1, wherein The multi-level strategy collaborative scheduling model includes a high-level strategy submodule, a low-level strategy submodule, and a scheduling agent submodule. Accordingly, the task level label, the computing power resource state matrix, and the task feature vector are input into the multi-level strategy collaborative scheduling model to obtain a node selection action, including: The high-level strategy submodule performs a strategy decision process on the task level label and the computing power resource state matrix to obtain a target computing power domain for the power task; Obtaining a corresponding resource state vector based on a candidate node in the candidate computing power node set in the target computing power domain; Constructing a local joint state representation based on the resource state vector and the task feature vector through the underlying strategy submodule; The scheduling agent submodule performs a policy evaluation on the local joint state representation to generate a node selection action.
5. The method according to claim 4, characterized in that The high-level strategy submodule performs a strategy decision process on the task level label and the computing power resource state matrix to obtain a target computing power domain for the power task, including: Inputting the task level label and the computing resource state matrix as a joint state into the high-level strategy submodule to construct a global joint state representation; A high-level policy network based on the high-level policy submodule performs feature extraction and policy evaluation on the global joint state representation to generate a computing power domain selection instruction; The target computing power domain corresponding to the power task is determined based on the computing power domain selection instruction.
6. The method according to claim 4, characterized in that The multi-level strategy collaborative scheduling model further includes a collaborative optimization submodule. Accordingly, after executing the power task, the method further includes: Obtaining scheduling feedback information when executing the power task; Constructing a state-action value function and a policy error loss function based on the scheduling feedback information, the local joint state representation, the target computing power node and the reward signal; Through the collaborative optimization submodule, the high-level policy network parameters of the high-level policy submodule are updated based on the state-action value function, and the low-level policy network parameters of the scheduling agent submodule are updated based on the policy error loss function.
7. The method according to claim 6, characterized in that The constructing of a state-action value function and a policy error loss function based on the scheduling feedback information, the local joint state representation, the target computing power node, and the reward signal includes: Constructing a feedback indicator set according to the scheduling feedback information; Calculating a reward signal based on the feedback indicator set; Constructing a state-action value function based on the local joint state representation, the target computing power node and the parameter set; A policy error loss function is constructed based on the reward signal and the state-action value function.
8. A task scheduling system based on a multi-domain computing network, characterized in that: The system comprises: A data acquisition and processing module is used to construct a task feature vector based on the structured attribute data of the power task to be scheduled; A task classification module, configured to determine a task level label corresponding to the power task based on the task feature vector and the task classification model; The multi-domain computing power network module is used to obtain the computing power resource status matrix corresponding to the current multi-domain computing power network; A multi-level strategy collaborative scheduling module, configured to input the task level label, the computing power resource state matrix, and the task feature vector into a multi-level strategy collaborative scheduling model to obtain a node selection action; The task scheduling module is used to add the power task to the task queue of the corresponding target computing power node according to the node selection action, so as to execute the power task based on the target computing power node.
9. An electronic device, characterized in that: The device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the task scheduling method based on the multi-domain computing power network according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the task scheduling method based on a multi-domain computing power network according to any one of claims 1 to 7 when executed.