Electronic information dynamic scheduling and management system and method

Through the improved Hungarian algorithm and DRL-OPT dynamic reward and punishment reinforcement learning model combined with Q-Learning, the problem of poor dynamic adaptability in the existing electronic information resource scheduling methods is solved, efficient resource matching and scheduling is achieved, and system response efficiency and stability are improved.

CN120263734AActive Publication Date: 2025-07-04GUANGDONG DUOMEIDA INTELLIGENT TERMINAL CO LTD

Patent Information

Application Number
CN202510729117.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing electronic information resource scheduling methods have poor dynamic adaptability and are unable to respond to network fluctuations and node load changes in real time, have low resource coordination efficiency, high task execution delay, traditional algorithms are prone to fall into local optimality, cross-layer resource coupling relationships have not been considered, and the reinforcement learning scheduling model converges slowly in dynamic scenarios.

Method used

A dynamic scheduling and management system for electronic information is designed, including system data acquisition module, task instruction reception module, coarse-grained matching module and fine-grained optimization module. The improved Hungarian algorithm and DRL-OPT dynamic reward and punishment reinforcement learning model are used for resource matching and scheduling, combined with Q-Learning to select the optimal node, and load balancing is achieved through resource borrowing and return protocol.

Benefits of technology

It realizes efficient resource matching and scheduling, dynamically responds to task requirements, improves the overall response efficiency and service quality of the system, avoids local optimal traps, reduces computing complexity, supports elastic scaling of resources, and ensures system stability and economy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263734A_ABST
    Figure CN120263734A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information management, in particular to an electronic information dynamic scheduling and management system and method. Real-time electronic information data is collected through a system data acquisition module, a resource matrix is constructed based on the real-time electronic information data, and a task instruction receiving module receives a system task request and analyzes the system task request to generate a task description vector; a coarseness matching module traverses the initial information matrix data according to the task description vector based on an improved Hungary algorithm to perform coarseness matching; the fine-grained optimization module performs fine-grained optimization on the candidate node subset data by using a DRL-OPT dynamic reward and punishment reinforcement learning model, selects a node for maximizing long-term reward through Q-Learning and outputs an optimal node, and the system scheduling management module schedules a system according to a resource allocation instruction, can adaptively learn a system operation rule, and outputs a resource allocation instruction to the DRL-OPT dynamic reward and punishment reinforcement learning model to the DRL-OPT dynamic reward and punishment reinforcement learning model. And dynamically adjusting the strategy according to the historical distribution effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic information management, and in particular to an electronic information dynamic scheduling and management system and method. Background Art

[0002] In the prior art, the scheduling methods of electronic information resources are mostly based on static rules or single indicators, and have the following defects: poor dynamic adaptability, unable to respond in real time to network fluctuations, node load changes or sudden task requirements; low resource collaboration efficiency: heterogeneous resources are difficult to be uniformly scheduled, resulting in local resource overload or idleness; high task execution delay, traditional algorithms are prone to fall into local optimum, and the global resource matching efficiency is insufficient. There is a task allocation method based on load balancing in the prior art solutions, but the coupling relationship of cross-layer resources is not considered; the reinforcement learning scheduling model proposed by the traditional solution has a slow convergence speed in dynamic scenarios due to the lack of a multi-dimensional feedback mechanism. Summary of the Invention

[0003] The purpose of the present invention is to solve the above problems, and a kind of electronic information dynamic scheduling and management system and method are designed.

[0004] Furthermore, in the above-mentioned electronic information dynamic scheduling and management system, the electronic information dynamic scheduling and management system includes the following modules: System data acquisition module, used for deploying monitoring agents on each physical node, collecting real-time electronic information data, constructing a resource matrix based on the real-time electronic information data, and performing normalization processing to obtain initial information matrix data; Task instruction receiving module, used for receiving system task requests, parsing the system task requests to generate task description vectors, and calculating urgency factors to obtain task requirement data; Coarse-grained matching module, used for traversing the initial information matrix data based on the improved Hungarian algorithm according to the task description vector to perform coarse-grained matching to obtain candidate node subset data; Fine-grained optimization module, used for using the DRL-OPT dynamic reward and punishment reinforcement learning model to perform fine-grained optimization on the candidate node subset data, selecting the node that maximizes the long-term reward through Q-Learning, outputting the optimal node, and obtaining resource allocation instructions; System scheduling management module, used for scheduling the system according to the resource allocation instructions, monitoring the load condition of the optimal node, and starting a resource borrowing and returning protocol when the real-time load of the optimal node exceeds a preset threshold.

[0005] Furthermore, in the above-mentioned electronic information dynamic scheduling and management system, the system data acquisition module includes the following sub-modules: The data acquisition sub-module is used to collect the usage of computing resources, storage resources, network resources and energy consumption data of physical nodes to obtain real-time electronic information data; the computing resources include CPU utilization rate, GPU CUDA core occupancy rate and the number of idle FPGA computing units; the storage resources include disk IOPS, remaining storage space and SSD life wear value; the network resources include inter-node latency, remaining bandwidth and TCP retransmission rate; the energy consumption data includes current power and energy consumption per unit computing power. The matrix establishment sub-module is used to establish a resource matrix according to the 4D indicators of each physical node in the real-time electronic information data in the order of nodes. The matrix normalization sub-module is used to normalize the resources by using the min-max normalization method to make the indicators of each dimension within the same value range to obtain the initial information matrix data.

[0006] Furthermore, in the above electronic information dynamic scheduling and management system, the task instruction receiving module includes the following sub-modules: The task receiving sub-module is used to receive system task requests, obtain the computing amount, memory size, data amount and deadline in the system task requests to obtain task computing parameters. The urgency calculation sub-module is used to integrate the task computing parameters into a task description vector, obtain the current time and the deadline in the task description vector, and calculate the urgency factor of the task based on the preset weight coefficients and service quality levels ; ; Among them, and represent the preset weight coefficients, represents the current time, represents the deadline, represents the service quality level.

[0007] Furthermore, in the above electronic information dynamic scheduling and management system, the coarse-grained matching module includes the following sub-modules: The traversal sub-module is used to traverse the initial information matrix data according to the computing amount, memory requirements and data amount in the task description vector, screen out the nodes that meet the basic requirements of the task for computing amount, memory and data amount, and form the first candidate node subset with the nodes. The sorting sub-module is used to sort the first candidate node subset from more to less according to the remaining resources of the nodes to obtain the second candidate node subset. The calculation module is used to establish a bipartite graph based on the improved Hungarian algorithm, with the left part being the task description vector and the right part being the second candidate node subset. Introduce slack variables for multi-resource constraints , and establish the Lagrangian function: ; Among them, represents the remaining resource amount of node , represents the matching deviation value between the resource efficiency of the node and the task description vector; represents the resources required by the task description vector; Obtain a sub-module for outputting the Top-5 candidate nodes through row and column reduction and augmented path search, and obtain the candidate node subset data.

[0008] Furthermore, in the above-mentioned electronic information dynamic scheduling and management system, the fine-grained optimization module includes the following units: A model definition unit for defining the state space, action space, and reward function of the DRL-OPT dynamic reward and punishment reinforcement learning model; A node output unit for using the real-time state of the candidate node subset data as the input as the state space of the model, generating binary actions of selecting and not selecting each node of the candidate node subset using the action space, and outputting a single optimal node; A reward function unit for determining that the reward function includes a multi-dimensional dynamic reward and punishment mechanism, at least including positive rewards, negative rewards, and dynamic adjustments.

[0009] Furthermore, in the above-mentioned electronic information dynamic scheduling and management system, the fine-grained optimization module further includes the following units: An action selection unit for selecting the node that maximizes the long-term reward through Q-Learning, collecting the state data of the candidate nodes based on the candidate node subset data, and selecting actions through the ε-greedy strategy; A greedy strategy unit for the greedy strategy to at least include randomly exploring nodes that have not been fully evaluated with probability ε and selecting the node with the largest current Q value with probability 1-ε; A reward calculation unit for calculating the immediate reward according to the actual result, updating the Q value using the Q-Learning formula, outputting the optimal node, and obtaining the resource allocation instruction; ; Among them, represents the state-action value function, represents the learning rate, represents the discount factor, represents the next state, represents the next action, represents the maximum Q value that can be obtained among all possible actions in state .

[0010] Further, in the above-mentioned electronic information dynamic scheduling and management system, the system scheduling and management module includes the following units: The node management unit is used to encapsulate the optimal node ID and task parameters into a scheduling instruction and send it to the node manager, and use the node manager to verify the resource availability. If the verification passes, the task execution is started; otherwise, the exception handling is triggered. The threshold preset unit is used to preset multiple levels of thresholds, including at least a warning threshold and an emergency threshold. The warning threshold is that the CPU usage rate ≥ 80%, the memory occupancy rate ≥ 90%, and it lasts for 5 minutes; the emergency threshold is that the CPU usage rate ≥ 95%, the memory occupancy rate ≥ 98%, and it lasts for 1 minute. The resource borrowing and returning unit is used to start the resource borrowing and returning protocol when the real-time load of the optimal node exceeds the preset threshold.

[0011] Further, in the above-mentioned electronic information dynamic scheduling and management method, the electronic information dynamic scheduling and management method includes the following steps: Deploy a monitoring agent on each physical node, collect real-time electronic information data, construct a resource matrix based on the real-time electronic information data, and perform normalization processing to obtain initial information matrix data. Receive a system task request, parse the system task request to generate a task description vector, and calculate the urgency factor to obtain task requirement data. Based on the improved Hungarian algorithm, traverse the initial information matrix data according to the task description vector for coarse-grained matching to obtain candidate node subset data. Use the DRL-OPT dynamic reward and punishment reinforcement learning model to perform fine-grained optimization on the candidate node subset data, select the node that maximizes the long-term reward through Q-Learning, output the optimal node, and obtain a resource allocation instruction. Schedule the system according to the resource allocation instruction, monitor the load condition of the optimal node, and start the resource borrowing and returning protocol when the real-time load of the optimal node exceeds the preset threshold.

[0012] Further, in the above-mentioned electronic information dynamic scheduling and management method, the step of deploying a monitoring agent on each physical node, collecting real-time electronic information data, constructing a resource matrix based on the real-time electronic information data, and performing normalization processing to obtain initial information matrix data includes: Collect the usage of computing resources, storage resources, network resources, and energy consumption data of physical nodes to obtain real-time electronic information data; the computing resources include CPU utilization rate, GPU CUDA core occupancy rate, and the number of idle FPGA computing units; the storage resources include disk IOPS, remaining storage space, and SSD life wear value; the network resources include inter-node latency, remaining bandwidth, and TCP retransmission rate; the energy consumption data includes current power and energy consumption per unit computing power. According to the 4D indicators of each physical node in the real-time electronic information data, establish a resource matrix in the order of nodes. Use the min-max normalization method to normalize the resources, so that the indicators of each dimension are within the same value range, and obtain the initial information matrix data.

[0013] Further, in the above method for dynamic scheduling and management of electronic information, it is characterized in that the receiving system task request, parsing the system task request to generate a task description vector, and calculating the urgency factor to obtain task requirement data, including: Receive a system task request, obtain the computing volume, memory size, data volume, and deadline in the system task request to obtain task calculation parameters. Integrate the task calculation parameters into a task description vector, obtain the current time and the deadline in the task description vector, and calculate the urgency factor of the task based on the preset weight coefficient and service quality level. ; ; Among them, and represent the preset weight coefficients, represents the current time, represents the deadline, represents the service quality level.

[0014] Its beneficial effects are as follows: 1. By deploying monitoring agents at physical nodes to collect real-time electronic information data and constructing a resource matrix for normalization processing, it ensures that the basic data obtained by the system has high timeliness and standardization characteristics. The resource status information can truly reflect the node capabilities, providing a reliable basis for subsequent task matching and avoiding allocation deviations caused by data lag or format chaos. 2. The intelligent parsing and priority determination of task requirements parse the system task requests to generate task description vectors, realizing the quantitative expression of task requirements. Dynamically distinguishing the urgency of tasks and the differences in resource requirements, it preferentially guarantees the resource supply for high-priority tasks, effectively improving the overall response efficiency and service quality of the system. 3. The efficient resource screening of coarse-grained matching traverses the initial information matrix data based on the improved Hungarian algorithm for coarse-grained matching, and can quickly screen out a subset of candidate nodes that meet the basic task conditions from a large number of nodes. Compared with traditional global search, this algorithm significantly reduces the computational complexity and shortens the task allocation decision time. 4. The dynamic resource adaptation of fine-grained optimization uses the DRL-OPT dynamic reward and punishment reinforcement learning model combined with the Q-Learning algorithm to deeply optimize the subset of candidate nodes, and selects the optimal nodes with the goal of maximizing long-term rewards. It can adaptively learn the operation rules of the system, dynamically adjust the strategy according to the historical allocation effects, balance the resource utilization rate and the task completion quality, avoid local optimal traps, and realize the continuous evolution of resource allocation strategies. 5. Load balancing and elastic resource scheduling construct a closed-loop dynamic resource management system by monitoring the load of the optimal nodes and starting the resource borrowing and returning protocol. When the node load exceeds the threshold, the borrowing and returning protocol can automatically trigger resource migration or sharing, effectively preventing performance bottlenecks caused by single-point overload and ensuring the system stability. At the same time, this mechanism supports the elastic scaling of resources, reduces the hardware redundancy cost, and improves the economic efficiency of resource use. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention.

[0016] Figure 1 It is a schematic diagram of the first embodiment of an electronic information dynamic scheduling and management system in an embodiment of the present invention; Figure 2 It is a schematic diagram of the second embodiment of an electronic information dynamic scheduling and management system in an embodiment of the present invention; Figure 3 It is a schematic diagram of the third embodiment of an electronic information dynamic scheduling and management system in an embodiment of the present invention; Figure 4 It is a schematic diagram of the first embodiment of an electronic information dynamic scheduling and management method in an embodiment of the present invention. Detailed Implementation Manner

[0017] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0018] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the term "including" used in the description of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0019] The present invention will be specifically described below with reference to the accompanying drawings. As Figure 1 shown, an electronic information dynamic scheduling and management system, the electronic information dynamic scheduling and management system includes the following modules: 101. System data acquisition module, which is used to deploy monitoring agents on each physical node, collect real-time electronic information data, construct a resource matrix based on the real-time electronic information data, and perform normalization processing to obtain initial information matrix data; Specifically, this embodiment further includes a data acquisition sub-module, which is used to collect the usage of computing resources, storage resources, network resources and energy consumption data of physical nodes to obtain real-time electronic information data; the computing resources include CPU utilization rate, GPU CUDA core occupancy rate and FPGA computing unit idle number; the storage resources include disk IOPS, remaining storage space and SSD life wear value; the network resources include inter-node delay, remaining bandwidth and TCP retransmission rate; the energy consumption data includes current power and unit computing power energy consumption; Matrix establishment sub-module, which is used to establish a resource matrix according to the 4D indicators of each physical node in the real-time electronic information data in the order of nodes; Matrix normalization sub-module, which is used to normalize the resources by using the min-max normalization method to make the indicators of each dimension within the same value range to obtain initial information matrix data.

[0020] 102. Task instruction receiving module, which is used to receive system task requests, parse the system task requests to generate task description vectors, and calculate urgency factors to obtain task requirement data; Specifically, this embodiment further includes a task receiving sub-module, which is used to receive system task requests, obtain the computing amount, memory size, data amount and deadline in the system task requests to obtain task computing parameters; The urgency calculation sub-module is used to integrate task calculation parameters into a task description vector, obtain the current time and the deadline in the task description vector, and calculate the urgency factor of the task based on preset weight coefficients and service quality levels. ; ; Among them, and represent preset weight coefficients, represents the current time, represents the deadline, represents the service quality level.

[0021] The urgency factor is a dynamic value calculated by a formula, which is used to comprehensively evaluate two core dimensions of the task: 1. Static service quality requirements, that is, the service quality level, which reflects the importance level of the task; 2. Dynamic time urgency, the difference between the current time and the deadline, which reflects the remaining processing time window through preset weights and through a linear combination, and finally generate a comprehensive urgency score for the task.

[0022] The urgency factor plays a core role in establishing a multi-dimensional task priority evaluation system; 1. Intelligent priority sorting. When processing multiple tasks simultaneously, the system automatically generates an execution queue according to the value; 2. Dynamic resource allocation. High tasks trigger an elastic resource allocation mechanism, and low tasks are temporarily stored or downgraded; 3. Time limit guarantee mechanism. When the difference between the current time and the deadline approaches zero, the time item will generate exponential growth pressure and automatically trigger the "deadline" emergency processing mode.

[0023] 103. The coarse-grained matching module is used to perform coarse-grained matching on the initial information matrix data based on the improved Hungarian algorithm according to the task description vector to obtain candidate node subset data; Specifically, in this embodiment, there is also a traversal sub-module, which is used to traverse the initial information matrix data according to the computational memory requirements and data volume in the task description vector, filter out the nodes that meet the basic requirements of the task for computational power, memory, and data volume, and form the nodes into a first candidate node subset; The sorting sub-module is used to sort the first candidate node subset in descending order of the remaining resource volume of the nodes to obtain a second candidate node subset; The calculation module is used to establish a bipartite graph based on the improved Hungarian algorithm, with the left part being the task description vector and the right part being the second candidate node subset; Introduce slack variables for multi-resource constraints , establish the Lagrangian function: ; Among them, represents the remaining resource amount of node ; represents the matching deviation value between the resource efficiency of the node and the task description vector; represents the resources required by the task description vector; Obtain a sub-module, which is used to output the Top-5 candidate nodes through row and column reduction and augmented path search, and obtain the candidate node subset data.

[0024] 104. Fine-grained optimization module, which is used to perform fine-grained optimization on the candidate node subset data by using the DRL-OPT dynamic reward and punishment reinforcement learning model, select the node that maximizes the long-term reward through Q-Learning, output the optimal node, and obtain the resource allocation instruction; Specifically, this embodiment further includes a model definition unit, which is used to define the state space, action space, and reward function of the DRL-OPT dynamic reward and punishment reinforcement learning model; The node output unit is used to use the real-time state of the candidate node subset data as the input as the state space of the model, generate binary actions of selection and non-selection for each node of the candidate node subset by using the action space, and output a single optimal node; The reward function unit is used to determine that the reward function includes a multi-dimensional dynamic reward and punishment mechanism, at least including positive reward, negative reward, and dynamic adjustment.

[0025] State Space: Taking the real-time state of the candidate node subset as the input, including: Real-time data of the node resource matrix (vectors after normalizing CPU utilization, remaining memory, network bandwidth occupancy, I / O latency, etc.); Task requirement data (task description vector, urgency factor, data input and output scale); Node historical scheduling records (recent task completion rate, average response time, number of resource allocation conflicts).

[0026] Action Space: Generate "select / do not select" binary actions for each node of the candidate node subset, and finally output a single optimal node (ensuring that only one node is allocated to execute the task each time of scheduling).

[0027] Design a multi-dimensional dynamic reward and punishment mechanism to comprehensively evaluate long-term benefits: Positive rewards: tasks are completed on time (weighted based on the urgency factor), the utilization rate of node resources is improved (to avoid excessive idleness), and the load balance is improved (the difference in load between adjacent nodes is reduced); Negative punishments: tasks are overdue (inversely proportional to the urgency), nodes are overloaded (exceeding the resource threshold), and resource allocation conflicts (duplicate scheduling or insufficient resources); Dynamic adjustment: The reward weights are corrected in real time according to the system load fluctuations (for example, when the load is high, overload is punished first, and when the load is low, resource utilization is optimized first).

[0028] An action selection unit, which is used to select the node that maximizes the long-term reward through Q-Learning, collect the state data of candidate nodes based on the candidate node subset data, and select actions through the ε-greedy strategy; A greedy strategy unit, where the greedy strategy at least includes randomly exploring nodes that have not been fully evaluated with probability ε and selecting the node with the largest current Q value with probability 1-ε; A reward calculation unit, which is used to calculate the immediate reward according to the actual result, update the Q value using the Q-Learning formula, output the optimal node, and obtain the resource allocation instruction; ; Among them, represents the state-action value function, represents the learning rate, represents the discount factor, represents the next state, represents the next action, represents in the state the maximum Q value that can be obtained among all possible actions.

[0029] 105. The system scheduling and management module is used to schedule the system according to the resource allocation instruction, monitor the load condition of the optimal node, and start the resource borrowing and returning protocol when the real-time load of the optimal node exceeds the preset threshold.

[0030] Specifically, this embodiment further includes: A node management unit, which is used to encapsulate the optimal node ID and task parameters into a scheduling instruction and send it to the node manager, use the node manager to verify the resource availability, and if the verification passes, start the task execution, otherwise trigger exception handling; A threshold preset unit, which is used to preset multiple levels of thresholds, at least including a warning threshold and an emergency threshold. The warning threshold is CPU usage rate ≥ 80%, memory occupancy rate ≥ 90%, lasting for 5 minutes; The emergency threshold is CPU usage rate ≥ 95%, memory occupancy rate ≥ 98%, lasting for 1 minute; A resource borrowing and returning unit, which is used to start the resource borrowing and returning protocol when the real-time load of the optimal node exceeds the preset threshold.

[0031] 1. Resource allocation instruction execution; Scheduling execution: The optimal node ID and task parameters (input data path, computing resource requirements) are encapsulated as scheduling instructions and sent to the node manager; The node manager verifies resource availability (based on the latest resource matrix). If the verification passes, the task execution is started. Otherwise, exception handling is triggered (return to step 3 for rematching).

[0032] 2. Load monitoring and threshold judgment; Monitoring indicators: Real-time collection of multi-dimensional load data of the optimal node: Computing resources: CPU usage (average core load), memory usage (remaining physical memory / total amount); Storage resources: disk I / O throughput, remaining storage space; Network resources: inbound / outbound bandwidth utilization, network latency (RTT).

[0033] Threshold setting: Preset multiple thresholds (warning threshold, emergency threshold), for example: Warning threshold: CPU usage ≥ 80%, memory usage ≥ 90%, for 5 minutes; Emergency threshold: CPU usage ≥ 95%, memory usage ≥ 98% for 1 minute.

[0034] 3. Resource borrowing and returning protocol startup logic; Trigger conditions: When the real-time load of a node exceeds the warning threshold, resource pre-detection is triggered; when it exceeds the emergency threshold, the borrowing and returning process is immediately started.

[0035] Resource borrowing process: Candidate node screening: Screen nodes with a load of less than 50% and matching resource types from the system (given priority to nodes in the same physical machine cluster to reduce network transmission overhead); Borrowing Negotiation: Send resource borrowing request to candidate nodes (including borrowing resource type, quantity, and expected duration); The candidate node responds whether it agrees based on its own resource scheduling strategy (such as reserving minimum resources); Select the borrowing node according to the response priority (nodes with high resource surplus and high historical collaboration efficiency are given priority).

[0036] Resource transfer: If computing resources are borrowed, partial load transfer can be achieved through container migration or task offloading; If it is for borrowing storage / network resources, establish temporary data channels or bandwidth allocation rules (it is necessary to avoid affecting the original tasks of the borrowing source node).

[0037] Resource return mechanism: Set a borrowing timeout mechanism (for example, the default borrowing duration does not exceed 30 minutes), and the return process will be automatically triggered when it expires; When the load of the optimal node drops below the threshold (for example, less than 70% for 10 consecutive minutes), actively release the borrowed resources; Both the borrowing and lending parties update the resource matrix and record the borrowing and returning history (used for calculating the reward function in step 4 to encourage cooperation between nodes).

[0038] 4. Fault tolerance and compensation in the borrowing and returning process; If all candidate nodes reject the borrowing or a failure occurs during the transfer process, trigger system-level resource reallocation: Suspend some processes of the current task and release the resources of the overloaded node; Return to step 3 to regenerate a subset of candidate nodes and include more available nodes (including cross-cluster nodes).

[0039] Give compensation rewards to the nodes providing resources (such as preferentially scheduling their subsequent tasks and increasing resource quotas) to form a virtuous cooperation ecosystem.

[0040] The beneficial effects are as follows: 1. By deploying monitoring agents at physical nodes to collect real-time electronic information data and constructing a resource matrix for normalization processing, it ensures that the basic data obtained by the system has high timeliness and standardization characteristics. The resource status information can truly reflect the node capabilities, providing a reliable basis for subsequent task matching and avoiding allocation deviations caused by data lag or format chaos. 2. The intelligent parsing and priority determination of task requirements parse the system task requests to generate task description vectors, realizing the quantitative expression of task requirements. Dynamically distinguish the urgency of tasks and the differences in resource requirements, and prioritize the resource supply for high-priority tasks, effectively improving the overall response efficiency and service quality of the system. 3. The efficient resource screening of coarse-grained matching traverses the initial information matrix data based on the improved Hungarian algorithm for coarse-grained matching, and can quickly screen out a subset of candidate nodes that meet the basic task conditions from a large number of nodes. Compared with traditional global search, this algorithm significantly reduces the computational complexity and shortens the task allocation decision time. 4. The dynamic resource adaptation of fine-grained optimization uses the DRL-OPT dynamic reward and punishment reinforcement learning model combined with the Q-Learning algorithm to deeply optimize the subset of candidate nodes, and selects the optimal nodes with the goal of maximizing long-term rewards. It can adaptively learn the operation rules of the system, dynamically adjust the strategy according to the historical allocation effect, balance the resource utilization rate and the task completion quality, avoid the local optimal trap, and realize the continuous evolution of the resource allocation strategy. 5. Load balancing and elastic resource scheduling build a closed-loop dynamic resource management system by monitoring the load of the optimal nodes and starting the resource borrowing and returning protocol. When the node load exceeds the threshold, the borrowing and returning protocol can automatically trigger resource migration or sharing, effectively preventing performance bottlenecks caused by single-point overload and ensuring system stability. At the same time, this mechanism supports the elastic scaling of resources, reduces the hardware redundancy cost, and improves the economic efficiency of resource use.

[0041] In this embodiment, please refer to Figure 2 , the second embodiment of an electronic information dynamic scheduling and management system in the embodiment of the present invention. The system data acquisition module includes the following sub-modules: The data acquisition sub-module is used to collect the usage of computing resources, storage resources, network resources, and energy consumption data of physical nodes to obtain real-time electronic information data; the computing resources include CPU utilization rate, GPU CUDA core occupancy rate, and the number of idle FPGA computing units; the storage resources include disk IOPS, remaining storage space, and SSD life wear value; the network resources include inter-node delay, remaining bandwidth, and TCP retransmission rate; the energy consumption data includes the current power and the energy consumption per unit computing power; The matrix establishment sub-module is used to establish a resource matrix according to the 4D indicators of each physical node in the real-time electronic information data in the order of nodes; A matrix normalization sub-module is used to normalize resources by using the min-max normalization method, so that the indexes of each dimension are within the same value range, and the initial information matrix data is obtained.

[0042] Its beneficial effect is that by deploying monitoring agents on physical nodes to collect real-time electronic information data and constructing a resource matrix for normalization processing, it is ensured that the basic data obtained by the system has the characteristics of high timeliness and standardization. The resource status information can truly reflect the node capabilities, providing a reliable basis for subsequent task matching and avoiding allocation deviations caused by data lag or format confusion.

[0043] In this embodiment, please refer to Figure 3 , the third embodiment of an electronic information dynamic scheduling and management system in the embodiment of the present invention. The fine-grained optimization module includes the following sub-units: An action selection unit is used to select a node that maximizes the long-term reward through Q-Learning, collect the state data of candidate nodes based on the candidate node subset data, and select an action through the ε-greedy strategy; A greedy strategy unit, where the greedy strategy at least includes randomly exploring nodes that have not been fully evaluated with probability ε and selecting the node with the maximum current Q value with probability 1-ε; A reward calculation unit is used to calculate the immediate reward according to the actual result, update the Q value using the Q-Learning formula, output the optimal node, and obtain the resource allocation instruction; ; Among them, represents the state-action value function, represents the learning rate, represents the discount factor, represents the next state, represents the next action, represents the maximum Q value that can be obtained among all possible actions in state .

[0044] Its beneficial effect is to select the optimal node with the goal of maximizing the long-term reward. It can adaptively learn the operation rules of the system, dynamically adjust the strategy according to the historical allocation effect, balance the resource utilization rate and the task completion quality, avoid the local optimal trap, and realize the continuous evolution of the resource allocation strategy.

[0045] The above describes an electronic information dynamic scheduling and management system provided by the embodiment of the present invention. Next, a description of an electronic information dynamic scheduling and management method of the embodiment of the present invention will be given. Please refer to Figure 4 , an embodiment of an electronic information dynamic scheduling and management method in the embodiment of the present invention includes: Step 401: Deploy monitoring agents on each physical node, collect real-time electronic information data, construct a resource matrix based on the real-time electronic information data, and perform normalization processing to obtain initial information matrix data; Step 402: Receive a system task request, parse the system task request to generate a task description vector, and calculate the urgency factor to obtain task requirement data; Step 403: Based on the improved Hungarian algorithm, traverse the initial information matrix data according to the task description vector for coarse-grained matching to obtain candidate node subset data; Step 404: Use the DRL-OPT dynamic reward and punishment reinforcement learning model to perform fine-grained optimization on the candidate node subset data, select the node that maximizes the long-term reward through Q-Learning, output the optimal node, and obtain a resource allocation instruction; Step 405: Schedule the system according to the resource allocation instruction, monitor the load condition of the optimal node, and when the real-time load of the optimal node exceeds a pre-set threshold, initiate a resource borrowing and returning protocol.

[0046] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. An electronic information dynamic scheduling and management system, characterized in that The electronic information dynamic scheduling and management system includes the following modules: A system data acquisition module, which is used to deploy monitoring agents on each physical node, collect real-time electronic information data, construct a resource matrix based on the real-time electronic information data, and perform normalization processing to obtain initial information matrix data; A task instruction receiving module, which is used to receive system task requests, parse the system task requests to generate task description vectors, and calculate urgency factors to obtain task requirement data; A coarse-grained matching module, which is used to perform coarse-grained matching on the initial information matrix data based on the improved Hungarian algorithm according to the task description vectors to obtain candidate node subset data; A fine-grained optimization module, which is used to perform fine-grained optimization on the candidate node subset data by using the DRL-OPT dynamic reward and punishment reinforcement learning model, select the node that maximizes the long-term reward through Q-Learning, output the optimal node, and obtain a resource allocation instruction; A system scheduling and management module, which is used to schedule the system according to the resource allocation instruction, monitor the load condition of the optimal node, and start a resource borrowing and returning protocol when the real-time load of the optimal node exceeds a preset threshold.

2. The electronic information dynamic scheduling and management system according to claim 1, wherein The system data acquisition module includes the following sub-modules: A data acquisition sub-module, which is used to collect the usage conditions of computing resources, storage resources, network resources, and energy consumption data of physical nodes to obtain real-time electronic information data; the computing resources include CPU utilization rate, GPU CUDA core occupancy rate, and the number of idle FPGA computing units; the storage resources include disk IOPS, remaining storage space, and SSD life wear value; the network resources include inter-node delay, remaining bandwidth, and TCP retransmission rate; the energy consumption data includes current power and energy consumption per unit computing power; A matrix establishment sub-module, which is used to establish a resource matrix according to the 4D indicators of each physical node in the real-time electronic information data in the order of nodes; A matrix normalization sub-module, which is used to normalize the resources by using the min-max normalization method to make the indicators of each dimension within the same value range to obtain initial information matrix data.

3. An electronic information dynamic scheduling and management system as claimed in claim 1, wherein The task instruction receiving module includes the following sub-modules: A task receiving sub-module, which is used to receive system task requests, obtain the computing amount, memory size, data amount, and deadline in the system task requests to obtain task computing parameters; An urgency calculation sub-module, configured to integrate the task calculation parameters into a task description vector, obtain the current time and the deadline in the task description vector, and calculate an urgency factor of the task based on a preset weight coefficient and a service quality level ; ; Among them, and represent preset weight coefficients, represents the current time, represents the deadline, represents the service quality level.

4. An electronic information dynamic scheduling and management system according to claim 1, characterized in that, The coarse-grained matching module includes the following sub-modules: A traversal sub-module, which is used to traverse the initial information matrix data according to the computing amount, memory requirement, and data amount in the task description vector, screen out the nodes that meet the basic requirements of the task for computing amount, memory, and data amount, and form the nodes into a first candidate node subset; A sorting sub-module, which is used to sort the first candidate node subset from more to less according to the remaining resources of the nodes to obtain a second candidate node subset; A calculation module, which is used to establish a bipartite graph based on the improved Hungarian algorithm, with the left part being the task description vector and the right part being the second candidate node subset; Introduce slack variables for multi-resource constraints , and establish the Lagrangian function: ; Among them, represents the remaining resource amount of the node , represents the matching deviation value between the node resource efficiency and the task description vector; represents the resources required by the task description vector. An obtaining sub-module, which is used to output the Top-5 candidate nodes through row and column reduction and augmented path search to obtain candidate node subset data.

5. An electronic information dynamic scheduling and management system as claimed in claim 1, characterized in that, The fine-grained optimization module includes the following units: A model definition unit, which is used to define the state space, action space, and reward function of the DRL-OPT dynamic reward and punishment reinforcement learning model; A node output unit, which is used to use the real-time state of the candidate node subset data as the input of the state space of the model, generate binary actions of selection and non-selection for each node of the candidate node subset by using the action space, and output a single optimal node; A reward function unit, which is used to determine that the reward function includes a multi-dimensional dynamic reward and punishment mechanism, at least including positive rewards, negative rewards, and dynamic adjustment.

6. An electronic information dynamic scheduling and management system according to claim 1, characterized in that, The fine-grained optimization module further includes the following units: An action selection unit, which is used to select a node that maximizes the long-term reward through Q-Learning, collect the state data of the candidate nodes based on the candidate node subset data, and select actions through an ε-greedy strategy; A greedy strategy unit, where the greedy strategy at least includes randomly exploring nodes that have not been fully evaluated with probability ε and selecting the node with the largest current Q value with probability 1 - ε; A reward calculation unit, which is used to calculate the immediate reward according to the actual result, update the Q value using the Q-Learning formula, output the optimal node, and obtain a resource allocation instruction; ; Among them, represents the state-action value function, represents the learning rate, represents the discount factor, represents the next state, represents the next action, represents the maximum Q value that can be obtained among all possible actions in the state below.

7. An electronic information dynamic scheduling and management system according to claim 1, characterized in that, The system scheduling and management module includes the following units: A node management unit, which is used to encapsulate the optimal node ID and task parameters into a scheduling instruction and send it to the node manager, use the node manager to verify the resource availability, and if the verification passes, start task execution, otherwise trigger exception handling; A threshold preset unit, which is used to preset multiple levels of thresholds, at least including a warning threshold and an emergency threshold. The warning threshold is that the CPU usage rate ≥ 80% and the memory occupancy rate ≥ 90% for 5 minutes; The emergency threshold is that the CPU usage rate ≥ 95% and the memory occupancy rate ≥ 98% for 1 minute; A resource borrowing and returning unit, which is used to start a resource borrowing and returning protocol when the real-time load of the optimal node exceeds a preset threshold.

8. An electronic information dynamic scheduling and management method, characterized in that, The electronic information dynamic scheduling and management method includes the following steps: Deploy a monitoring agent on each physical node, collect real-time electronic information data, construct a resource matrix based on the real-time electronic information data, and perform normalization processing to obtain initial information matrix data; Receive a system task request, parse the system task request to generate a task description vector, and calculate a urgency factor to obtain task requirement data; Based on the improved Hungarian algorithm, traverse the initial information matrix data according to the task description vector for coarse-grained matching to obtain candidate node subset data; Use the DRL-OPT dynamic reward and punishment reinforcement learning model to perform fine-grained optimization on the candidate node subset data, select a node that maximizes the long-term reward through Q-Learning, output the optimal node, and obtain a resource allocation instruction; Schedule the system according to the resource allocation instruction, monitor the load condition of the optimal node, and start a resource borrowing and returning protocol when the real-time load of the optimal node exceeds a preset threshold.

9. The electronic information dynamic scheduling and management method according to claim 8, characterized in that, Deploy monitoring agents on each physical node, collect real-time electronic information data, construct a resource matrix based on the real-time electronic information data, and perform normalization processing to obtain initial information matrix data, including: Collect the usage of computing resources, storage resources, network resources, and energy consumption data of physical nodes to obtain real-time electronic information data; the computing resources include CPU utilization rate, GPU CUDA core occupancy rate, and the number of idle FPGA computing units; the storage resources include disk IOPS, remaining storage space, and SSD life wear value; the network resources include inter-node latency, remaining bandwidth, and TCP retransmission rate; the energy consumption data includes current power and energy consumption per unit computing power. According to the 4D indicators of each physical node in the real-time electronic information data, establish a resource matrix in the order of nodes; Normalize the resources using the min-max normalization method to make the indicators of each dimension within the same value range, and obtain the initial information matrix data.

10. An electronic information dynamic scheduling and management method according to claim 8, characterized in that Receive the system task request, parse the system task request to generate a task description vector, and calculate the urgency factor to obtain task requirement data, including: Receive the system task request, obtain the computing volume, memory size, data volume, and deadline in the system task request to obtain task calculation parameters; Integrate the task calculation parameters into a task description vector, obtain the current time and the deadline in the task description vector, and calculate the urgency factor of the task based on a preset weight coefficient and service quality level ; ; Among them, and represent preset weight coefficients, represents the current time, represents the deadline, represents the service quality level.

Citation Information

Patent Citations

  • Scheduling automation system application state management method

    CN119292745A

  • Big data dynamic allocation and optimal scheduling method based on reinforcement learning

    CN119311407A

  • Multi-access edge computing architecture for cloud-network integration

    WO2022217503A1

  • Task scheduling method and device for artificial intelligence (AI) network function service

    WO2024011376A1

Cited By

  • Self-adaptive weight task scheduling method and system based on time sequence differential learning

    CN120973493A