Intelligent monthly settlement task scheduling method and system based on reinforcement learning

Through an intelligent scheduling method based on reinforcement learning, combined with graph neural networks and proximal strategy optimization algorithms, the problems of poor scheduling adaptability and low resource utilization efficiency in the traditional financial monthly settlement process are solved, dynamic and adaptive global optimization scheduling is achieved, and the stability and efficiency of the financial monthly settlement process are improved.

CN120803681AActive Publication Date: 2025-10-17INSPUR GENERSOFT CO LTD

Patent Information

Application Number
CN202511315859.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-10-17
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

The traditional financial monthly settlement process suffers from poor adaptability to static scheduling, low resource utilization efficiency, inability to achieve global optimization, and weak ability to respond to emergencies. Especially in the complex financial monthly settlement scenarios of large group companies, the existing system has difficulty coping with dynamic changes and multiple constraints.

Method used

A dynamic scheduling system is constructed by adopting an intelligent scheduling method based on reinforcement learning, combined with graph neural networks and proximal policy optimization algorithms. Through multi-objective reward functions and hard constraint mechanisms, dynamic, adaptive, and global optimization scheduling of complex financial monthly settlement processes is achieved.

Benefits of technology

Significantly shorten the monthly settlement cycle, improve resource utilization efficiency, ensure process stability and scalability, adapt to different business strategy requirements, and achieve true global optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803681A_ABST
    Figure CN120803681A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and provides a monthly settlement task intelligent scheduling method and system based on reinforcement learning, and the technical scheme is that a time sequence state vector is constructed and obtained based on obtained monthly settlement task multi-dimensional state feature data; encoding the time sequence state vector to obtain time sequence features, constructing a task dependency graph, extracting global features of the task dependency graph, fusing the time sequence features and the global features of the task dependency graph to obtain fusion features, and generating action probability distribution based on the fusion features; based on the action probability distribution, introducing a hard constraint and a soft constraint, constructing a multi-target reward function, optimizing the reinforcement learning agent based on a near-end strategy optimization algorithm, and outputting an action strategy based on the optimized reinforcement learning agent; and converting the optimized reinforcement learning agent output action strategy into a specific monthly settlement task scheduling instruction. And dynamic, self-adaptive and global optimization scheduling of the complex financial monthly settlement process is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to a monthly task intelligent scheduling method and system based on reinforcement learning. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] With the continuous deepening of enterprise fine operation management, the financial monthly settlement process of large group companies has become a core task involving massive data, multi-system cooperation and complex dependency. Taking a communication infrastructure service enterprise as an example, its financial monthly settlement needs to handle cost accounting businesses such as rental fees, electricity fees and maintenance fees, and the process has the characteristics of extremely large data volume, numerous participants and complex business logic.

[0004] At present, this process usually relies on traditional workflow engines and static scheduling strategies based on fixed rules, but it has the following defects: Poor adaptability of static scheduling: existing systems mostly use predefined workflow templates, and the task execution order and dependency relationship are fixed before the process is started. However, the actual business environment is highly dynamic, and the task dependency relationship often changes temporarily due to local policy adjustment, system interface upgrade or data anomaly. The static scheduling model cannot respond to these changes in real time, and it is easy to cause the entire process to be delayed or interrupted due to a blockage.

[0005] Low resource utilization efficiency: the monthly settlement process is subject to strict constraints on computing resources such as servers, database performance and human resources (number and skills of financial personnel). Especially during the peak period of monthly settlement, traditional rule scheduling such as first-come-first-served cannot dynamically allocate resources according to global resource load, which is easy to form a performance bottleneck, resulting in coexistence of idle and congested resources, and it is difficult to optimize the utilization rate.

[0006] Cannot achieve global optimization: the monthly settlement optimization goal is multi-dimensional, and needs to minimize the total completion time, maximize resource utilization and ensure that key tasks are not delayed. Rule scheduling strategies based on human experience, such as simply setting priorities, can only achieve local optimization and lack the ability to learn from historical scheduling experience, so they cannot make scientific decisions under multiple constraints to achieve global goals.

[0007] Weak response to emergencies: the process must be completed within a strict time limit (such as T+5 days), and traditional systems lack effective dynamic restructuring and emergency scheduling mechanisms when tasks are delayed unexpectedly, often relying on manual intervention, which is not only inefficient but also difficult to ensure reliability.

[0008] Although reinforcement learning shows potential in handling dynamic and complex problems in task scheduling theory research, direct application still faces great challenges in this specific industrial application scenario of financial monthly closing. Existing research lacks comprehensive consideration of super-large-scale data, multi-level distributed organizational architecture, heterogeneous system integration, and strong business constraints, and has not yet formed a mature and implementable systematic solution. SUMMARY

[0009] To solve at least one technical problem in the background art, the present application provides a reinforcement learning-based monthly closing task intelligent scheduling method and system, which realizes dynamic, adaptive, and global optimization scheduling of complex financial monthly closing processes, significantly shortens the monthly closing cycle, improves resource utilization efficiency, and guarantees process stability and scalability.

[0010] To achieve the above-mentioned purpose, the present application adopts the following technical solutions: The first aspect of the present application provides a reinforcement learning-based monthly closing task intelligent scheduling method, comprising the following steps: Obtaining multi-dimensional state feature data of monthly closing tasks; Constructing a time series state vector based on the obtained multi-dimensional state feature data of monthly closing tasks; Constructing a reinforcement learning agent, taking the time series state vector as the state space, encoding the time series state vector to obtain a time series feature, constructing a task dependency graph, extracting global features of the task dependency graph, fusing the time series feature and the global features of the task dependency graph to obtain fused features, and generating an action probability distribution based on the fused features; Based on the action probability distribution, introducing hard constraints and soft constraints, constructing a multi-objective reward function, optimizing the reinforcement learning agent based on a proximal policy optimization algorithm, and outputting an action policy based on the optimized reinforcement learning agent; Converting the action policy output by the optimized reinforcement learning agent into specific monthly closing task scheduling instructions and executing the monthly closing task scheduling instructions.

[0011] Further, the extraction of the global features of the task dependency graph comprises: For each node i Sampling a fixed number of neighbor nodes; Aggregating the features of the neighbor nodes to generate neighbor messages; Concatenating the current node features with the aggregated neighbor messages, and updating them through a learnable weight matrix and a nonlinear activation function to obtain the new features of the node i In the first k layer; After K-layer convolution, the embedding vectors of all nodes are obtained, and the global features of the graph are obtained by global average pooling of all node embeddings.

[0012] Further, the hard constraint is to introduce an action masking mechanism, analyze the current task dependency graph in real time, dynamically generate an action mask, which forces the probability of all actions that violate the current dependency constraints to be zero, and then re-normalize the probability distribution of the remaining valid actions to ensure that the agent only samples from legal actions.

[0013] Further, the multi-objective reward function is: wherein, , , denotes the adjustment coefficient, denotes the time constraint, denotes the resource constraint, denotes the delay penalty, denotes the stability reward, , , and denote the weights.

[0014] Further, the loss function of the reinforcement learning agent based on the proximal policy optimization algorithm is: , , , , , , wherein, denotes the total loss, denotes the loss of updating the actor policy network, denotes the loss of updating the value function, denotes the entropy reward loss, and denote the weights, is the probability ratio of the new and old policies, is the advantage function, is a hyperparameter, γ is the discount factor, and λ is the GAE parameter, denotes the importance sampling ratio, and denote the future and current time difference errors, and denote the value functions of the state and the next state, denotes the initial state value, denotes the actual cumulative return, denotes the policy function, denotes the entropy of the policy.

[0015] Further, when encoding the time sequence state vector to obtain the time sequence feature, the time sequence state vector is encoded through a fully connected network or a small Transformer encoder to obtain the time sequence feature.

[0016] Further, the multi-dimensional state feature data of the monthly closing task includes a task state, a resource state, a dependency graph, and an environment state; the task state includes a task queue, an execution duration, a remaining task number, a priority, and a task type; the resource state includes available computing nodes, CPU / memory / disk I / O load, network bandwidth, personnel to be allocated, and a skill matrix; the dependency graph includes a dynamic analysis task DAG dependency relationship; and the environment state includes a time to deadline, a historical delay situation, and system alarms.

[0017] Further, when constructing the time sequence state vector based on the obtained multi-dimensional state feature data of the monthly closing task, the multi-dimensional state feature data is normalized, the normalized data is embedded and encoded, and a graph neural network is used to encode a complex graph structure dependency relationship into a fixed-length graph embedding vector.

[0018] The second aspect of the present application provides a monthly closing task intelligent scheduling system based on reinforcement learning, comprising: a data acquisition module configured to acquire multi-dimensional state feature data of a monthly closing task; a time sequence state vector constructed based on the acquired multi-dimensional state feature data of the monthly closing task; an action generation module configured to construct a reinforcement learning agent, encode the time sequence state vector to obtain a time sequence feature, construct a task dependency graph, extract global features of the task dependency graph, fuse the time sequence feature and the global features of the task dependency graph to obtain fused features, and generate an action probability distribution based on the fused features; an action optimization module configured to introduce hard constraints and soft constraints based on the action probability distribution, construct a multi-objective reward function, optimize the reinforcement learning agent based on a proximal policy optimization algorithm, and output an action policy based on the optimized reinforcement learning agent; a task execution module configured to convert the action policy output by the optimized reinforcement learning agent into specific monthly closing task scheduling instructions and execute the monthly closing task scheduling instructions.

[0019] The third aspect of the present application provides a computer device.

[0020] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the above-described monthly closing task intelligent scheduling method based on reinforcement learning when executing the program.

[0021] Compared with the prior art, the present application has the following advantages: The present application innovatively deepens the reinforcement learning and the graph neural network in the financial monthly settlement scheduling field, and constructs a closed-loop intelligent scheduling system integrating real-time multi-source data perception, GNN dependency relationship modeling, proximal policy optimization (PPO) decision optimization and dynamic execution control, which can perform multi-modal fusion on the environment state (time series data) and the task topology (graph data), continuously learn based on the reward signal and dynamically adjust the scheduling strategy, and realizes the paradigm shift from "static rule" to "dynamic perception and optimization".

[0022] The present application utilizes the message passing mechanism of GNN (GraphSAGE) to enable each task node to aggregate the information of its upstream and downstream tasks. This enables the system to deeply understand the propagation effect of the dependency relationship, accurately identify the key path affecting the entire monthly settlement period, and intelligently allocate the scarce resources to the tasks that can shorten the overall time the most, thereby realizing real global optimization. The bottleneck that the traditional scheduler cannot understand the task chain from a global perspective is solved.

[0023] The reward function designed by the present application considers four dimensions of key path progress, resource utilization, delay penalty and process stability, and the weights can be dynamically configured. Combined with the advantages of PPO algorithm in training stability and sample efficiency, the system can learn to make the optimal trade-off between multiple competing goals, and adapt to different business strategies, solving the problem of single goal and fixed rules that cannot adapt to different stages such as pursuing speed at the beginning of the month and pursuing stability at the end of the month.

[0024] The advantages of the additional aspects of the present application will be partially given in the following description, partially become obvious from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0025] The accompanying drawings, which form a part of the specification, are included to provide a further understanding of the application and are incorporated herein for explanation by referring to the exemplary embodiments thereof.

[0026] Figure 1 is a kind of monthly settlement task intelligent scheduling method flow chart based on reinforcement learning provided by the embodiment of the present application; Figure 2 is a kind of monthly settlement task intelligent scheduling system block diagram based on reinforcement learning provided by the embodiment of the present application. DETAILED DESCRIPTION

[0027] The present application will be further described below in conjunction with the drawings and embodiments.

[0028] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0029] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0030] Example 1 like Figure 1 As shown, this embodiment provides a method for intelligent scheduling of monthly settlement tasks based on reinforcement learning, including the following steps: Step 1: Obtain multi-dimensional status feature data of the monthly settlement task; In this embodiment, when obtaining multi-dimensional status feature data of the monthly settlement task, various data systems are connected and connected to ERP, financial reporting system, asset management system, contract management system, IT monitoring system (server / database performance), human resources system, etc. through RESTful API or message queue (such as Kafka).

[0031] Specifically, the multi-dimensional status feature data of the monthly settlement task includes task status, resource status, dependency graph and environment status; The task status includes the task queue, execution time (historical mean / variance), number of remaining tasks, priority, and task type (computing / IO / manual); Resource status includes available computing nodes, CPU / memory / disk I / O load, network bandwidth, personnel to be assigned, and skill matrix; Dependency graph includes dynamic parsing of task DAG dependencies and identifying critical paths; The environmental status includes the time until the monthly settlement deadline, historical delays, and system alarms.

[0032] Step 2: Construct a time series state vector based on the multi-dimensional state feature data of the monthly settlement task; The multi-dimensional, multi-scale, and dynamically changing system state in the process is very complex. In this embodiment, a 128-dimensional unified state vector is constructed. When creating a new image, it is not a simple splicing, but a special processing technology, including: Normalize multi-dimensional state feature data; For example, normalize the resource status CPU, memory GB, I / O MBps, etc. of different dimensions to the same scale; Embedding encoding is performed on the normalized data, for example, discrete type data such as task type and personnel skills are converted into low-dimensional dense vectors; The complex graph structure dependency relationship is encoded into a fixed-length graph embedding vector using a graph neural network. Finally, the processed multi-dimensional features are obtained into an original high-dimensional vector, and then a fully connected layer is used for dimension reduction and feature fusion. In this embodiment, a 128-dimensional time series state vector is finally output, represented as: , Among them, represents the task state feature, represents the resource state feature, represents the dependency graph feature, represents the environment state feature.

[0033] Step 3: Constructing a reinforcement learning agent, taking the time series state vector as the state space, encoding the time series state vector to obtain the time series feature, constructing the task dependency graph, extracting the global feature of the task dependency graph, and fusing the time series feature and the global feature of the task dependency graph to obtain the fusion feature, and generating the action probability distribution based on the fusion feature; The system state changes rapidly (such as a task suddenly freezing or a server unexpectedly going down), and the agent cannot obtain 100% perfect global state. The state is designed as an effective observation of the current system. The agent needs to learn to make decisions based on incomplete information. By using a recurrent network or taking the state of continuous time steps as input, the agent has certain historical information perception ability to cope with the dynamic changes of the environment.

[0034] The reinforcement learning agent is constructed, and the reinforcement learning agent adopts the Actor-Critic architecture. The Actor-Critic shares the state encoding layer. The output dimension of the Actor output layer is equal to the size of the action space, and the Softmax function is used to output the probability distribution π(a|s) of selecting each action. The output dimension of the Critic output layer is 1, and the output state value function V(s) is output, which represents the long-term expected return of the current state.

[0035] The state space of the agent is , and the action space includes: adjusting the task priority weight; dynamically allocating CPU cores and memory for computing tasks; assigning financial personnel with corresponding skills to artificial tasks; triggering the emergency mode (such as enabling the standby cluster).

[0036] Specifically, the following steps are included: Step 301: Time series state vector Encode to obtain time series features ; In this embodiment, the 128-dimensional time series state vector Encode through a fully connected network or a small Transformer encoder to obtain temporal features ; Step 302: construct a task dependency graph and extract global features of the task dependency graph; The dependencies between tasks are naturally represented by a DAG (Graph Sample and Aggregate) structure, and traditional fully connected networks cannot effectively utilize its topological information. This paper uses a graph neural network (GNN), specifically the GraphSAGE (Graph Sample and Aggregate) algorithm, to encode the task dependency graph to capture the global context and dependency constraints of the tasks.

[0037] In this embodiment, when constructing the task dependency graph, each node represents a monthly settlement task (such as "electricity bill calculation" and "rent accrual"). The initial feature vector of each node is expressed as , contains the task status information, such as: task type (encoded as one-hot), historical average execution time, priority, remaining data amount, etc. Directed edges represent the dependency relationship between tasks (such as "Task B must be started after Task A is completed"), and the direction of the edge indicates the flow of dependency; In this embodiment, a graph neural network (GNN) is used to extract global features of the task dependency graph, specifically including: GraphSAGE updates the representation of the current node by iteratively aggregating information from neighboring nodes: In the k layer: Sampling neighbors: for each node i Sampling a fixed number of neighbor nodes .

[0038] Aggregate neighbor information: Use aggregation functions (such as mean, pool, LSTM) to aggregate the features of neighbor nodes and generate neighbor messages , expressed as:

[0039] in, represents the new feature of node j at the k-1 layer, Represents the neighbor node information aggregation function, which is used to integrate the feature vectors of all neighbor nodes of the target node into an aggregate message of fixed length; Update node status: concatenate the current node features with the aggregated neighbor messages and pass a learnable weight matrix and a nonlinear activation function (like ReLU) to get the node i In the first k layer, the new feature is represented as: , where is the new feature of the i th node in the k th layer, represents the concatenation operation; After K layers of convolution, each node obtains a final embedding vector that contains K-hop neighborhood information. To obtain the global representation of the entire graph (for agent decision-making), global average pooling is performed on all node embeddings to obtain the graph global feature, represented as: , where N is the total number of nodes.

[0040] Step 303, feature fusion of time series features and task dependency graph global features to obtain environment state representation, then through a fusion network (fully connected layer) for final encoding; The environment state representation is: , where is the fusion feature, is the matrix weight, is the high-level time series feature, is the graph global feature, is the bias term; Step 304, the fusion feature is sent to the output layer of the Actor network (a fully connected layer followed by a Softmax function) to generate an action probability distribution .

[0041] Assuming there are M discrete actions (such as "raise the priority of task A" and "allocate 4 cores to task B"), the output is an M-dimensional vector, where each element represents the probability of selecting the corresponding action, and the sum of all elements is 1.

[0042] ,

[0043] where is the weight, is the bias term, representing a probability distribution, and representing the original output value generated by the policy network of the agent for the i-th, j-th action for the current state s respectively Step 4: based on the action probability distribution, introducing hard constraints and soft constraints, constructing a reward function, optimizing the reinforcement learning agent based on the proximal policy optimization algorithm, and outputting the action policy based on the optimized reinforcement learning agent; Specifically, the following steps are included: Step 401, based on the action probability distribution, introducing hard constraints and soft constraints, and constructing a reward function; To learn a better policy, the agent does not completely choose the action with the highest probability, but samples according to the probability distribution. For example, action A has a probability of 0.7 and action B has a probability of 0.3. The system has a 70% chance of choosing A and a 30% chance of choosing B, which ensures sufficient exploration.

[0044] The traditional dependency constraint is an absolute hard constraint, which cannot schedule a task whose prerequisite task is not completed; therefore, in this embodiment, after the agent outputs the action probability distribution, a action masking mechanism hard constraint is introduced, the current task dependency graph is analyzed in real time, and an action mask is dynamically generated. The mask will force the probability of all actions that violate the current dependency constraint (such as "start executing a certain prerequisite task that has not been completed") to be zero, and then the probability distribution of the remaining valid actions is re-normalized to ensure that the agent only samples from legal actions.

[0045] The innovation lies in that, instead of "hoping" the agent to learn to avoid constraints through the punishment of the reward function, the generation of illegal actions is fundamentally eliminated through the mechanism, ensuring 100% scheduling safety. This is the most efficient way to directly inject domain knowledge into the reinforcement learning loop.

[0046] The time window (such as T+5 days to complete) is a soft constraint that can be slightly violated but will cause significant business losses, and resource constraints (limited manpower, computing power) also need to be efficiently utilized.

[0047] In this embodiment, a multi-objective reward function is designed to guide the agent to meet these soft constraints. The specific multi-objective reward function includes time constraints, resource constraints, and stability constraints, and is represented as:

[0048] wherein, , , representing the adjustment coefficient, representing the time constraint, representing the resource constraint, representing the delay penalty, represents a stability reward, , , and represent weights; The term directly penalizes the delay of the critical path, and the agent will prefer to take actions that can shorten the critical path time in order to maximize the cumulative reward, thus spontaneously striving to meet the final deadline requirement.

[0049] The term rewards high resource utilization and penalizes resource idling. However, at the same time, the resource allocation action itself is constrained by the action mask (such as not being able to allocate more CPU than the total number of physical machine cores), ensuring that optimization is carried out within the upper limit of resources.

[0050] and The term penalizes frequent scheduling changes and task restarts, encouraging smooth operation.

[0051] The weights , , and are not fixed. They can be configured flexibly as hyperparameters according to the strategic focus of different companies and different monthly settlement stages (such as initial progress and final stability). This means that the same system can adapt to different optimization strategies by adjusting the weights, demonstrating strong versatility and adaptability.

[0052] Step 402, based on the proximal policy optimization algorithm, the reinforcement learning agent is optimized, and the action policy is output based on the optimized reinforcement learning agent; In this embodiment, in order to balance sample efficiency and training stability, the proximal policy optimization (PPO-Clip) algorithm is used, which is to limit the step of policy update by clipping the probability ratio to avoid training collapse.

[0053] The core of PPO is the following improved objective function for updating the Actor policy network; , where, is the probability ratio of the new and old policies, is the advantage function, which estimates the advantage of action relative to the average level, is a hyperparameter, usually set to 0.1 or 0.2, which limits the variation range of , thus stabilizing the policy update within a reliable interval, effectively preventing performance collapse due to excessive single update. minThe operation ensures that the most conservative (pessimistic) update direction is selected even in the presence of estimation errors, further enhancing stability.

[0054] To reduce variance, the advantage function Using generalized advantage estimation, denoted as: , , where γ is the discount factor that trades off immediate and future rewards, and λ is the GAE parameter that trades off bias and variance, denotes the importance sampling ratio, and denotes the future and current time-difference errors, and denotes the state and next state value functions, which are output by the Critic network. Instead of using Monte Carlo returns, the Critic network is used to estimate the value function, greatly improving sample efficiency because one trajectory can be used for multiple updates.

[0055] Value function update loss: The update of the Critic network is represented by minimizing the mean squared error loss: , where denotes the initial state value, denotes the actual cumulative return; Entropy reward loss: To encourage exploration and prevent the policy from converging to a local optimum too early, a reward for the policy entropy is added to the objective function, denoted as: , where denotes the policy function, denotes the entropy of the policy; The final total loss function is denoted as: , where and denote the weights; During training, the system samples a small batch of data from the experience replay buffer and performs multiple rounds (usually 4-10 rounds) of gradient ascent optimization on the above objectives. This repeated use of old data greatly improves sample efficiency. The optimizer is usually Adam, and its adaptive learning rate characteristics also help stabilize training.

[0056] Step 5: Convert the optimized reinforcement learning agent output action policy into specific scheduling instructions; The selected abstract action (e.g., action_id = 102) is mapped to one or more concrete, executable scheduling instructions, and the mapping relationship is predefined in an instruction library.

[0057] For example, the priority of "electricity billing" is raised to P0; Assign 4-core CPU and 8GB memory to the "electricity billing" task. Assign a financial staff with electricity billing skills.

[0058] The scheduling instructions are generated at a set time, such as every 5 minutes, and the task priority and resource configuration are dynamically modified by calling the scheduling system interface of Airflow through the REST API.

[0059] Step 6: Execute the scheduling instructions and store the decision results in the experience replay buffer. In this embodiment, the final decision result is represented as a four-tuple The experience replay buffer is used for model retraining, and a visual dashboard is provided to display the estimated completion time, key path progress, resource utilization, etc., wherein is the time series state vector, is the sampled or selected action, is the actual cumulative return is the next stage time series state vector.

[0060] Embodiment Two As shown in Figure 2 , the embodiment provides a monthly closing task intelligent scheduling system based on reinforcement learning, comprising: A data acquisition module 201 is configured to acquire multi-dimensional state feature data of the monthly closing task. A time series state vector is constructed based on the acquired multi-dimensional state feature data of the monthly closing task. An action generation module 202 is configured to construct a reinforcement learning agent, encode the time series state vector to obtain time series features, construct a task dependency graph, extract global features of the task dependency graph, fuse the time series features and the global features of the task dependency graph to obtain fused features, and generate an action probability distribution based on the fused features. An action optimization module 203 is configured to introduce hard constraints and soft constraints based on the action probability distribution, construct a multi-objective reward function, optimize the reinforcement learning agent based on a proximal policy optimization algorithm, and output an action policy based on the optimized reinforcement learning agent. A task execution module 204 is configured to convert the action policy output by the optimized reinforcement learning agent into specific monthly closing task scheduling instructions and execute the monthly closing task scheduling instructions.

[0061] It should be noted that the specific implementation of the embodiment of the application is similar to the specific implementation of the embodiment of the application, and specific please refer to the description of the method part, in order to reduce redundancy, not here.

[0062] Embodiment three The embodiment provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to realize the steps in the method for intelligent scheduling of monthly task based on reinforcement learning.

[0063] Embodiment four The embodiment provides a computer device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor realizes the steps in the method for intelligent scheduling of monthly task based on reinforcement learning when executing the program.

[0064] Those skilled in the art should understand that the embodiments of the application can be provided as a method, a system or a computer program product. Therefore, the application can be in the form of a hardware embodiment, a software embodiment or an embodiment combining software and hardware aspects. Moreover, the application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage and optical storage) containing computer usable program code.

[0065] The application is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The device for realizing the functions specified in one or more flows and / or blocks.

[0066] These computer program instructions can also be stored in a computer readable storage medium, which can guide the computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer readable storage medium produce a product including instruction devices, which realize the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The device for realizing the functions specified in one or more flows and / or blocks.

[0067] These computer program instructions can also be loaded into a computer or other programmable data processing devices, so that a series of operational steps are performed on the computer or other programmable data processing devices to generate computer-implemented processes, thus the instructions executed on the computer or other programmable data processing devices provide the function of implementing the processes specified in the flowcharts Figure 1 The processes or the functions specified in one or more flowcharts and / or one or more blocks Figure 1 The processes or the functions specified in one or more flowcharts and / or one or more blocks

[0068] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiment methods. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM), etc.

[0069] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A monthly task intelligent scheduling method based on reinforcement learning, characterized in that: The steps include: Obtain multi-dimensional status feature data of monthly settlement tasks; The time series state vector is constructed based on the multi-dimensional state feature data of the monthly settlement task; Build a reinforcement learning agent, use the time series state vector as the state space, encode the time series state vector to obtain time series features, build a task dependency graph, extract the global features of the task dependency graph, fuse the time series features with the global features of the task dependency graph to obtain fused features, and generate action probability distribution based on the fused features; Based on the action probability distribution, hard constraints and soft constraints are introduced to construct a multi-objective reward function. The reinforcement learning agent is optimized based on the proximal policy optimization algorithm, and the action strategy is output based on the optimized reinforcement learning agent. Convert the optimized reinforcement learning agent output action strategy into specific monthly task scheduling instructions and execute the monthly task scheduling instructions.

2. The method for intelligent scheduling of monthly settlement tasks based on reinforcement learning according to claim 1, characterized in that: The extraction of global features of the task dependency graph includes: For each node i Sample a fixed number of neighbor nodes; Aggregate the features of neighbor nodes to generate neighbor messages; The current node features are spliced ​​with the aggregated neighbor messages and updated through a learnable weight matrix and nonlinear activation function to obtain the node i In the k New features of layers; After K layers of convolution, the embedding vectors of all nodes are obtained, and global average pooling is performed on all node embeddings to obtain the global features of the graph.

3. The method for intelligent scheduling of monthly settlement tasks based on reinforcement learning according to claim 1, characterized in that: The hard constraint is to introduce an action masking mechanism, analyze the current task dependency graph in real time, and dynamically generate an action mask. The mask forces the probabilities of all actions that violate the current dependency constraints to zero. Subsequently, the probability distribution of the remaining valid actions is renormalized to ensure that the agent only samples from legal actions.

4. The method for intelligent scheduling of monthly settlement tasks based on reinforcement learning according to claim 1, characterized in that: The multi-objective reward function is: in, 、 、 represents the adjustment coefficient, Indicates time constraints, Represents resource constraints, represents the delay penalty, represents the stability reward, 、 , and Represents weight.

5. The method for intelligent scheduling of monthly settlement tasks based on reinforcement learning according to claim 1, characterized in that: The loss function when optimizing the reinforcement learning agent based on the proximal policy optimization algorithm is: , , , , , , in, represents the total loss, Indicates the update of Actor strategy network loss, represents the value function update loss, represents the entropy reward loss, and represents the weight, is the probability ratio of the new and old strategies, is the advantage function, is a hyperparameter, γ is the discount factor, λ is the GAE parameter, represents the importance sampling ratio, and represents the future and current timing difference error, and Represents the value function of the state and the next state, represents the initial state value, represents the actual cumulative return, represents the policy function, represents the entropy of the policy.

6. The method for intelligent scheduling of monthly settlement tasks based on reinforcement learning according to claim 1, characterized in that: When encoding the time series state vector to obtain the time series features, the time series state vector is encoded through a fully connected network or a small Transformer encoder to obtain the time series features.

7. The method for intelligent scheduling of monthly settlement tasks based on reinforcement learning according to claim 1, characterized in that: The multi-dimensional status feature data of the monthly settlement task includes task status, resource status, dependency graph and environment status; among them, the task status includes task queue, execution time, number of remaining tasks, priority and task type; the resource status includes available computing nodes, CPU / memory / disk I / O load, network bandwidth, personnel to be assigned and skill matrix; the dependency graph includes dynamically parsed task DAG dependency relationships, and the environment status includes the time to the monthly settlement deadline, historical delays, and system alarms.

8. The method for intelligent scheduling of monthly settlement tasks based on reinforcement learning according to claim 1, characterized in that: When constructing a time series state vector based on the multi-dimensional state feature data obtained for the monthly settlement task, it includes normalizing the multi-dimensional state feature data, embedding and encoding the normalized data, and using a graph neural network to encode the complex graph structure dependency relationship into a fixed-length graph embedding vector.

9. An intelligent scheduling system for monthly settlement tasks based on reinforcement learning, characterized in that: include: A data acquisition module, which is used to obtain multi-dimensional status feature data of monthly settlement tasks; The time series state vector is constructed based on the multi-dimensional state feature data of the monthly settlement task; The action generation module is used to build a reinforcement learning agent. It uses the time-series state vector as the state space, encodes the time-series state vector to obtain time-series features, constructs a task dependency graph, extracts the global features of the task dependency graph, fuses the time-series features with the global features of the task dependency graph to obtain fused features, and generates action probability distribution based on the fused features. The action optimization module is used to introduce hard and soft constraints based on the action probability distribution, construct a multi-objective reward function, optimize the reinforcement learning agent based on the proximal policy optimization algorithm, and output the action policy based on the optimized reinforcement learning agent; The task execution module is used to convert the optimized reinforcement learning agent output action strategy into specific monthly task scheduling instructions and execute the monthly task scheduling instructions.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the intelligent scheduling method for monthly settlement tasks based on reinforcement learning as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Unmanned intelligent inspection equipment cooperative scheduling method and system in photovoltaic power generation scene

    CN119358998A

  • Threat intelligence knowledge graph reasoning method based on improved GNN and reinforcement learning

    CN120218238A

  • Time-varying task scheduling method and system based on space-time constraint near-end strategy optimization

    CN120234144A

  • Computing power resource optimization method and system based on deep learning

    CN120560859A

  • Reinforcement learning and heuristic driving edge computing dependent task scheduling method

    CN120578481A

Cited By

  • Discrete production process task transfer management method and system

    CN121882640A

  • Computational graph dynamic topology reconstruction method, agent training method and related devices

    CN122021706A

  • Distributed transaction dynamic allocation method and device across heterogeneous ticket systems

    CN122395281A