Man-machine cooperative scheduling method for dynamic graph attention and multi-agent reinforcement learning

By employing a human-machine collaborative scheduling method based on dynamic graph attention and multi-agent reinforcement learning, the problem of insufficient adaptability of existing technologies in dynamic environments is solved. This enables flexible scheduling and resource optimization of the production process in intelligent manufacturing, thereby improving production efficiency and stability.

CN121477818APending Publication Date: 2026-02-06HOHAI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511683270.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing mathematical programming methods and metaheuristic algorithms are not adaptable enough to dynamic scheduling in intelligent manufacturing environments and cannot effectively cope with sudden disturbances in the production process, such as personnel changes or emergency plug-ins, which affects production efficiency and stability.

Method used

A human-machine collaborative scheduling method using dynamic graph attention and multi-agent reinforcement learning is adopted. Through the collaboration of task allocation agent, worker scheduling agent and robot management agent in the multi-agent collaborative reinforcement learning framework, the global state graph is updated in real time. Combined with graph attention network and multi-objective genetic algorithm, the task allocation scheme is optimized to achieve a balance between production efficiency, energy consumption and manpower load.

Benefits of technology

In a dynamic environment, it enables rapid adjustment of task allocation schemes, real-time response to sudden disturbances in the production process, ensuring the continuity and stability of the production process, improving production efficiency and resource utilization, and reducing the imbalance of energy consumption and worker workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121477818A_ABST
    Figure CN121477818A_ABST
Patent Text Reader

Abstract

The invention provides a man-machine cooperative scheduling method based on dynamic graph attention and multi-agent reinforcement learning, and relates to the technical field of production scheduling. The self-adaptive scheduling system is constructed by dynamically updating the global state diagram and the online learning mechanism, it is ensured that the task allocation scheme is rapidly adjusted in the dynamically changing production environment, meanwhile, the scheduling system can cope with various sudden disturbances in the production process in real time, the conditions such as equipment faults and emergency plug-ins are included, and the production efficiency is improved. The continuity and stability of the production process are ensured, and the technical problem that an existing mathematical planning method and a meta-heuristic algorithm are insufficient in adaptability in a dynamic environment can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of production scheduling technology, specifically to a human-machine collaborative scheduling method using dynamic graph attention and multi-agent reinforcement learning. Background Technology

[0002] In the field of intelligent manufacturing, human-machine collaborative dynamic scheduling is of paramount importance. By scientifically allocating human and robot resources and scheduling tasks, it can effectively improve production efficiency and reduce cost input.

[0003] Currently, in the field of production scheduling in intelligent manufacturing environments, metaheuristic algorithms improve computational efficiency, but their static optimization nature makes them unable to effectively cope with sudden disturbances in the production process, such as personnel changes or emergency plug-ins.

[0004] As can be seen from the above description, existing mathematical programming methods and metaheuristic algorithms suffer from insufficient adaptability in dynamic environments. Summary of the Invention

[0005] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a human-computer collaborative scheduling method based on dynamic graph attention and multi-agent reinforcement learning, which solves the technical problem of insufficient adaptability of existing mathematical programming methods and metaheuristic algorithms in dynamic environments.

[0006] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a human-machine collaborative scheduling method based on dynamic graph attention and multi-agent reinforcement learning. This method achieves dynamic scheduling through the collaborative efforts of a task allocation agent, a worker scheduling agent, and a robot management agent within a multi-agent collaborative reinforcement learning framework. The human-machine collaborative scheduling method includes: Real-time collection of various data from the production line, dynamic updating of the global status graph, and generation of a task feature subgraph for each task to describe the optional processing path; A graph attention network is used to perform deep encoding and feature extraction on the constructed global state graph and task feature subgraph to obtain global graph embedding features and task subgraph embedding features; Multi-agents make collaborative decisions based on global graph embedding features and task subgraph embedding features to obtain a set of candidate scheduling schemes; With the goal of balancing production efficiency, energy consumption, and manpower load, a multi-objective genetic algorithm is used to process the candidate scheduling scheme set to obtain the optimal scheduling scheme that can balance multiple objectives. The optimal scheduling scheme is issued to the production line for execution, and the next global state, the rewards of each agent and the global reward are collected and stored in the experience pool as experience data for updating the multi-agent collaborative reinforcement learning framework.

[0007] Preferably, the multi-agent cooperative reinforcement learning framework includes constructing a multi-objective scheduling model, which includes an objective function and constraints; The objective function includes: Objective 1 aims to achieve the maximum completion time. Minimize; Objective 2 aims to minimize the robot's total energy consumption. Minimize; Objective 3 aims to achieve worker load variance. Minimize; In the formula, The constraints include: Among them, constraint 5 indicates that each task must be assigned to one and only one workstation for a certain processing method; constraint 6 indicates that a single workstation can only process one independent task at a time; and constraint 7 indicates that the total number of robots assigned to all workstations does not exceed the total number of collaborative robots. Constraint 8 indicates that if the task Must be in the task If completed before, then The start time must not be earlier than The completion time; constraint 9 indicates that if the robot exist If a fault occurs, its processing tasks should not be assigned. in, The average load for all workers; Represents a set of workpieces; Indicates workpiece The set of tasks in; Represents the set of workstations, totaling... One workstation; It indicates that the workers gathered together. One worker; Represents a collection of collaborative robots, totaling... One collaborative robot; Represents a set of processing methods. , Indicates that workers are processing the materials. Indicates robot processing, This indicates human-machine collaborative processing; Indicates workpiece The set of priority relationships for the included tasks. Indicates workpiece Tasks in Need to be Completed before; Indicates workpiece Task Start time; Indicates workpiece Task Completion time; Indicates workstation Number of robots used ; Indicates workpiece Task The complexity; Indicates workpiece Task Method of adoption Processing time; Indicates the robot's fault status. 1 indicates a fault, and 0 indicates normal operation; Indicates that the robot is processing the workpiece. Task Energy consumption coefficient; Indicates workpiece Task At workstation In a way The value is 1 if the process is complete, and 0 otherwise. Indicates workpiece Task At workstation In a way The value is 1 if processing is performed, otherwise it is 0. Indicates workpiece Task At workstation In a way The value is 1 if processing is performed, otherwise it is 0. Indicates workers Total processing time.

[0008] Preferably, the global state diagram includes ; in: This represents the set of nodes in the global state diagram, including workstation nodes, task nodes, and resource nodes. Workstation nodes represent the real-time state of the physical workstation; task nodes are dynamically generated and contain task attributes; and resource nodes represent workers and robots, carrying their real-time states. This represents the set of edges in the global state graph, which includes task dependency edges and resource allocation edges. Task dependency edges represent the priority constraints between tasks, while resource allocation edges connect tasks with workstations and connect tasks with resources.

[0009] Preferably, the dynamic updating of the global state graph includes: The global state graph is updated in real time through a dynamic update mechanism. Specifically, the dynamic update mechanism is as follows: when a new task or disturbance event triggers a local graph update, only the affected nodes and related edges are modified.

[0010] Preferably, the multi-agent system makes collaborative decisions based on global graph embedding features and task subgraph embedding features to obtain a set of candidate scheduling schemes, including: The task assignment agent processes global graph embedding features and task subgraph embedding features to generate three-dimensional discrete actions. ,in: Tasks to be assigned Optional workstations For the selection of processing mode; among which, Indicates that workers are processing the materials. Indicates robot processing, This indicates human-machine collaborative processing; When the processing mode is worker processing or human-machine collaborative processing, the worker scheduling agent selects a specific worker w from the available worker pool to execute the current task; When the processing mode is robotic processing or human-robot collaborative processing, the robot scheduling agent selects a specific robot from the available robot pool. To perform the current task; In the multi-agent cooperative reinforcement learning framework, the internal policy network parameters of the three agents are randomly initialized during the initial training phase. Under the initial policy, after receiving global graph embedding features and task subgraph embedding features, the policy network of the three agents outputs a nearly uniform probability distribution. Based on the probability distribution, random sampling is performed to determine a series of combined actions jointly output by the three agents. This constitutes a set of candidate scheduling schemes.

[0011] Preferably, the worker scheduling agent is also used to update workers. Real-time fatigue level: in, Indicates rest time. Indicates workpiece Task The complexity; Indicates workpiece Task Method of adoption Processing time; Indicates workpiece Task At workstation In a way The value is 1 if the process is complete, and 0 otherwise. like If the fatigue threshold is reached, a rest period will be forcibly inserted. The robot scheduling agent is also used to monitor fault status: if Remove it from the list of available robots; This indicates a robot malfunction.

[0012] Preferably, the multi-objective genetic algorithm includes a reward function, the expression of which is as follows: in, This indicates a time-based reward. This indicates an energy consumption bonus; This indicates a load reward; , , All are weight parameters; among them, This represents the maximum completion time at time t. This represents the maximum completion time at time t-1; Indicates that the robot is processing the workpiece. Task Energy consumption coefficient; Indicates workpiece Task At workstation In a way The value is 1 if the process is complete, and 0 otherwise. This represents the variance of worker load.

[0013] Secondly, the present invention provides a human-machine collaborative scheduling system for dynamic graph attention and multi-agent reinforcement learning, wherein the dynamic job scheduling system is used to execute the human-machine collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning as described above.

[0014] Thirdly, the present invention provides a computer-readable storage medium storing a computer program for human-computer collaborative scheduling of dynamic graph attention and multi-agent reinforcement learning, wherein the computer program causes a computer to execute the human-computer collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning as described above.

[0015] Fourthly, the present invention provides an electronic device, comprising: One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including a human-machine collaborative scheduling method for performing dynamic graph attention and multi-agent reinforcement learning as described above.

[0016] (III) Beneficial Effects This invention provides a human-machine collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning. Compared with existing technologies, it has the following advantages: This invention constructs an adaptive scheduling system by dynamically updating the global state graph and using an online learning mechanism. This ensures that the task allocation scheme can be quickly adjusted in a dynamically changing production environment. At the same time, the scheduling system can respond in real time to various sudden disturbances in the production process, including equipment failures and emergency plug-ins, ensuring the continuity and stability of the production process. This invention can effectively solve the technical problem of insufficient adaptability of existing mathematical programming methods and metaheuristic algorithms in dynamic environments. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A diagram illustrating the overall framework for dynamic scheduling, where task allocation agents, worker scheduling agents, and robot management agents collaborate. Figure 2 A schematic diagram of human-machine collaborative production; Figure 3 A schematic diagram illustrating the method for calculating the weights of adjacent nodes; Figure 4 A flowchart for the improved PPO algorithm. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] This application provides a human-machine collaborative scheduling method based on dynamic graph attention and multi-agent reinforcement learning, which solves the technical problem of insufficient adaptability of existing mathematical programming methods and metaheuristic algorithms in dynamic environments, and enables rapid adjustment of task allocation schemes in dynamically changing production environments.

[0021] The technical solution in this application is to solve the above-mentioned technical problems, and the general idea is as follows: Existing scheduling methods mainly suffer from the following drawbacks: 1. Existing mathematical programming methods and metaheuristic algorithms typically employ static optimization strategies, which cannot respond in real time to dynamic disturbances in the production process, such as personnel changes or the insertion of urgent tasks. This results in scheduling schemes lacking flexibility in practical applications, making it difficult to adapt to dynamically changing production environments, thus affecting production efficiency and stability.

[0022] 2. Existing single-agent architectures employ a centralized decision-making model, requiring the simultaneous handling of complex multi-dimensional interactions. This single-decision architecture can lead to an explosion of policy network parameters, resulting in severe dimensionality curse when dealing with high-dimensional state spaces, and exhibiting poor responsiveness to dynamic perturbations.

[0023] 3. Existing metaheuristic algorithms (such as genetic algorithms and simulated annealing) have slow convergence speeds when solving large-scale problems and are prone to getting trapped in local optima. In addition, rule-oriented methods (such as priority assignment and fixed allocation strategies) widely used in actual enterprise production scheduling rely on human experience to set rules, lack adaptive optimization capabilities, have poor generalization, and are difficult to adapt to different industries or dynamically changing production scenarios.

[0024] 4. Current research on human-machine collaboration is relatively weak, lacking consideration of human factors. This leads to unreasonable allocation of human resources, where workers may experience reduced productivity due to overwork, while robot resources cannot be optimally allocated, making it difficult to achieve complementary advantages between humans and machines.

[0025] To address the aforementioned issues, this invention provides a human-machine collaborative scheduling method based on dynamic graph attention and multi-agent reinforcement learning, aiming to overcome the bottlenecks of traditional production scheduling techniques, such as poor dynamic adaptability and limitations in single-agent decision-making. This method optimizes production efficiency, energy consumption, and load balancing, significantly improving scheduling efficiency through a multi-agent collaborative decision-making mechanism. Compared to traditional static optimization algorithms, the human-machine collaborative scheduling method of this invention can perceive disturbance events in real time within a dynamic production environment and achieve adaptive optimization of objectives in a dynamic environment, thereby meeting the demands of intelligent manufacturing for high efficiency, flexibility, and stability.

[0026] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0027] This invention provides a human-machine collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning. This method is executed by a human-machine collaborative scheduling system, such as... Figure 1 As shown, dynamic scheduling is achieved through the collaborative work of the task allocation agent, worker scheduling agent, and robot management agent within a multi-agent cooperative reinforcement learning framework. The human-machine cooperative scheduling method includes: S1. Collect various data from the production line in real time, dynamically update the global status graph, and generate a task feature subgraph for each task to describe the optional processing path; S2. A graph attention network is used to perform deep encoding and feature extraction on the constructed global state graph and task feature subgraph to obtain global graph embedding features and task subgraph embedding features. S3. The multi-agent system makes collaborative decisions based on the global graph embedding features and the task subgraph embedding features to obtain a set of candidate scheduling schemes. S4. With the goal of balancing production efficiency, energy consumption and manpower load, a multi-objective genetic algorithm is used to process the candidate scheduling scheme set to obtain the optimal scheduling scheme that can balance multiple objectives. S5. The optimal scheduling scheme is issued to the production line for execution. The next global state, the rewards of each agent and the global reward are collected and stored in the experience pool as experience data for updating the multi-agent collaborative reinforcement learning framework.

[0028] This invention constructs an adaptive scheduling system by dynamically updating the global state graph and using an online learning mechanism. This ensures that the task allocation scheme can be quickly adjusted in a dynamically changing production environment. At the same time, the scheduling system can respond in real time to various sudden disturbances in the production process, including equipment failures and emergency plug-ins, ensuring the continuity and stability of the production process. This effectively solves the technical problem of insufficient adaptability of existing mathematical programming methods and metaheuristic algorithms in dynamic environments.

[0029] It should be noted that, in this embodiment of the invention, a multi-objective scheduling model is constructed within a multi-agent cooperative reinforcement learning framework. The objective function of this model is as follows: in, This represents the average load for all workers.

[0030] The constraints of the model are as follows: Objective (1) aims to achieve the maximum completion time. Minimize; Objective (2) aims to achieve the robot's total energy consumption. Minimize; Objective (3) aims to achieve worker load variance. Minimize; Formula (4) indicates that the worker load includes the time spent on the task it handles independently and the total time spent on the collaborative task it participates in; Constraint (5) indicates that each task must be assigned to one workstation in one processing mode (worker / robot / collaboration); Constraint (6) indicates that a single workstation can only handle one independent task at a time (independent processing by worker or robot); Constraint (7) indicates that the total number of robots assigned to all workstations does not exceed the total number of collaborative robots. Constraint (8) indicates that if the task Must be in the task If completed before, then The start time must not be earlier than The completion time; constraint (9) indicates that if the robot exist If a fault occurs, it is prohibited from being assigned processing tasks.

[0031] In the formula, Represents a set of workpieces; Indicates workpiece The set of tasks in; Represents the set of workstations, totaling... One workstation; It indicates that the workers gathered together. One worker; Represents a collection of collaborative robots, totaling... One collaborative robot; Represents a set of processing methods. ( Indicates that workers are processing the materials. Indicates robot processing, (This indicates human-machine collaborative processing) Indicates workpiece The set of priority relationships for the included tasks. Indicates workpiece Tasks in Need to be Completed before; Indicates workpiece Task Start time; Indicates workpiece Task Completion time; Indicates workstation Number of robots used ; Indicates workpiece Task The complexity; Indicates workpiece Task Method of adoption Processing time; Indicates the robot's fault status. 1 indicates a fault, and 0 indicates normal operation; Indicates that the robot is processing the workpiece. Task Energy consumption coefficient; Indicates workpiece Task At workstation In a way The value is 1 if the process is complete, and 0 otherwise. Indicates workpiece Task At workstation In a way The value is 1 if processing is performed, otherwise it is 0. Indicates workpiece Task At workstation In a way The value is 1 if processing is performed, otherwise it is 0. Indicates workers Total processing time.

[0032] In one embodiment, S1, various types of data on the production line are collected in real time, the global state graph is dynamically updated, and a task feature sub-graph describing the optional processing paths is generated for each task. The specific implementation process is as follows: In this embodiment of the invention, the global state diagram Used to model the overall state of a production system, including workstations, tasks, resources, and their dynamic interactions.

[0033] This represents a set of nodes in the global state graph. In this embodiment of the invention, the nodes include workstation nodes, task nodes, and resource nodes. Workstation nodes represent the real-time state of a physical workstation; task nodes are dynamically generated and contain task attributes (complexity, priority); resource nodes represent workers and robots, carrying real-time status (such as worker fatigue values ​​and robot fault flags). All nodes reflect the real-time state of the system through dynamic attribute updates.

[0034] This represents the set of edges in the global state graph. In this embodiment of the invention, the edges include task-dependent edges and resource allocation edges. Task-dependent edges represent the priority constraints between tasks; resource allocation edges connect tasks with workstations / resources and directly affect task allocation decisions.

[0035] In this embodiment of the invention, the global state graph is updated in real time through a dynamic update mechanism. Specifically, the dynamic update mechanism involves: new tasks or disturbance events (such as robot malfunctions) triggering local graph updates, modifying only the affected nodes and associated edges; non-disturbed nodes reuse historical encoding to reduce computational overhead. For example, when a new workpiece arrives, a task node and associated edges are dynamically added; after the task is completed, the corresponding node is deleted and the workstation resources are released. In other words, by collecting data such as the current workstation status, worker status, robot status, and newly added workpieces, the attributes of nodes and edges in the graph are dynamically updated.

[0036] In the task feature subgraph, each task Corresponding to a subgraph Describe its optional processing paths and resource requirements.

[0037] The nodes in the task feature subgraph include candidate processing method nodes and virtual start and end nodes. Candidate processing method nodes describe the relevant characteristics of the selectable processing methods for the task, such as processing time and energy consumption. Virtual start and end nodes serve as unified entry and exit points, without carrying specific attributes, and are primarily used to construct the complete task processing path topology. Each candidate processing method node reflects the task execution potential under different resource configurations through dynamically updated processing time parameters.

[0038] The edges in the task feature subgraph include processing mode edges. Processing mode edges connect the task with candidate processing methods, and their weights are multi-objective costs (time, energy consumption).

[0039] As described above, in this step, the system collects various data from the production line in real time, including the operating status of workstations, the current fatigue level of workers, indicators of robot malfunction, the queue of tasks to be processed and their urgency levels, and monitors for any sudden disturbances. Based on this data, a global state graph is constructed, dynamically maintaining the latest states of workstation nodes, task nodes, and resource nodes, and generating a feature subgraph describing the optional processing paths for each task. When a sudden disturbance is detected, the system only updates the affected nodes and connections, while retaining the original encoding for the rest to save computational resources.

[0040] In one embodiment, S2, a graph attention network is used to perform deep encoding and feature extraction on the constructed global state graph and task feature subgraph to obtain global graph embedding features and task subgraph embedding features. The specific implementation process is as follows: The system employs a graph attention network to deeply encode the constructed global graph and task subgraphs. Attention weights are calculated by analyzing the relationships between nodes, ultimately generating a unified state representation that integrates static topology and dynamic perturbation features. This process effectively captures complex constraints between tasks and the real-time status of production resources. Specifically, it includes: The input layer output data of the graph attention network is: node feature matrix. ,in For the number of nodes, The feature dimension is the edge feature matrix. It includes weight, type, and direction, among which Let be the total number of edges in the graph.

[0041] In graph attention networks, the input data is encoded using an incremental encoding pattern, as follows: Under normal conditions, full graph encoding is performed, and node embeddings are generated using a multi-head attention mechanism. When a disturbance event occurs, only the attention weights of the affected nodes and their neighbors are recalculated; the historical embeddings of other nodes are retained, significantly improving real-time performance. The calculation method for neighboring node weights is as follows: Figure 3 As shown.

[0042] The process of calculating attention weights in a multi-head attention mechanism is as follows: Step 1: Calculate the nodes and neighboring nodes Attention coefficient : in: Represents a node eigenvectors, Representing neighboring nodes eigenvectors, and For learnable parameters, This indicates vector concatenation.

[0043] Step 2: Normalize attention weights: Step 3: Output node representation: A 4-head attention mechanism is used, and the final node is represented as a concatenation of the outputs of each head: Step 4: Output the merged state, the specific process of which is as follows: 1. Input Source The fusion process has two main input sources: Static features - Task topology embedding: Source: This is the direct output of GAT after processing the global state graph or task feature subgraph. It is a set of node embedding vectors that capture the complex, inherent topological relationships and dependencies between nodes in the graph.

[0044] Dynamic characteristics - Disturbance event tags: Source: This is a discrete or continuous signal representing a sudden event, extracted from real-time monitoring data of the production line.

[0045] 2. Fusion process The fusion process is carried out in the following two sub-steps in sequence: Sub-step 1: Generate graph-level static embeddings The output of GAT is node-level, while decision-making requires a graph-level overall representation. Therefore, the node embeddings of the GAT output need to be aggregated first: Global graph embedding features: global state graph The embedding vectors of all nodes are aggregated to generate a single, fixed-dimensional vector. This vector summarizes the global static state of the entire production workshop.

[0046] Task subgraph embedding features: For each task, the corresponding task feature subgraph All node embeddings are aggregated in the same way to generate a vector representing all possible processing paths and their characteristics for the task.

[0047] Sub-step 2: Assembling dynamic features The static graph embedding vector generated in the previous step is concatenated with the dynamic feature vector: Global graph embedding features are concatenated with global dynamic features that affect the entire system, such as "whether any robots have malfunctioned".

[0048] The task subgraph embedding features are concatenated with local dynamic features related to the task, such as: "whether the task itself is an urgent workpiece" and "whether the pre-allocated resources for the task are available".

[0049] 3. Output Results Global graph embedding features: a comprehensive state vector that integrates global topological information and global real-time perturbations.

[0050] Task subgraph embedding features: a comprehensive state vector that integrates the characteristics of a single task and the real-time perturbations associated with that task.

[0051] In one embodiment, S3 and the multi-agent team make collaborative decisions based on global graph embedding features and task subgraph embedding features to obtain a set of candidate scheduling schemes. The specific implementation process is as follows: The system comprises three agents: a task allocation agent, a worker scheduling agent, and a robot management agent. The human-machine collaborative production model is as follows: Figure 2 As shown.

[0052] (1) The task allocation agent comprehensively considers the characteristics of the task and the status of resources to select the optimal combination of task, workstation and processing method; the worker scheduling agent monitors the fatigue value of workers in real time and automatically avoids overloaded work arrangements; the robot management agent responds to equipment failures in real time and quickly adjusts task allocation. For emergency plug-in situations, the system will automatically increase its processing priority.

[0053] The data processing procedure of the Task Assignment Agent (TAA) is as follows: Input: Global graph embedding features, task subgraph embedding features.

[0054] Action selection: Generate 3D discrete actions ,in: Tasks to be assigned Optional workstations For the selection of processing mode.

[0055] Policy function: in: Represents the policy network, For policy network parameters; state This includes global graph embedding features and task subgraph embedding features.

[0056] (2) Worker Scheduling Agent (WSA): When the proposal of the task allocation agent involves workers (independent processing by workers or human-machine collaborative processing), it is responsible for selecting a specific worker from the available worker pool. To perform this task.

[0057] Update workers Real-time fatigue level: like If a fatigue threshold is set, a rest period is forcibly inserted. In this embodiment of the invention, the fatigue threshold is set to 0.7.

[0058] (3) Robot Scheduling Agent (RSA): When the proposal of the task allocation agent involves a robot (independent robot processing or human-robot collaborative processing), the robot scheduling agent is responsible for selecting a specific robot from the available robot pool. To perform this task.

[0059] Simultaneously monitor fault status: If The robot is removed from the list of available robots, and a "robot failure" event is sent to the scheduling system. The system locates all tasks that have been "assigned to R3 but not completed", and TAA makes a rescheduling decision.

[0060] Through the above process, the internal policy network parameters of the three agents are randomly initialized during the initial training phase. Under the initial policy, after an agent receives state s, its policy network outputs a nearly uniform probability distribution. The system randomly samples based on this distribution to determine a combined action jointly output by the three agents. These random actions together constitute the candidate scheduling scheme set.

[0061] In one embodiment, S4, with the goal of balancing production efficiency, energy consumption, and manpower load, a multi-objective genetic algorithm is used to process the candidate scheduling scheme set to obtain the optimal scheduling scheme that can balance multiple objectives. The specific implementation process is as follows: The reward function for the multi-objective genetic algorithm is as follows: in, This indicates a time-based reward. This indicates an energy consumption bonus; This indicates a load reward; , , All of these are weight parameters.

[0062] A multi-objective genetic algorithm was used to combine the weights. Optimize and output the Pareto optimal solution set as the optimal scheduling scheme that can balance multiple objectives.

[0063] In one embodiment, S5, the optimal scheduling scheme is issued to the production line for execution, and the next global state, the rewards of each agent, and the global reward are collected and stored in the experience pool as experience data for updating the multi-agent collaborative reinforcement learning framework. The specific implementation process is as follows: The selected optimal scheduling plan is deployed to the production line for execution, while the system continuously monitors the plan's effectiveness. All key data from the decision-making process, including environmental conditions, actions taken, gains, and subsequent states, are categorized and stored in an experience pool, with higher priority given to experience in handling unforeseen circumstances. The system periodically extracts data from the experience pool for policy iteration, updating network parameters through an improved Proximal Policy Optimization (PPO) algorithm. The process is as follows: Figure 4 As shown, the policy update magnitude is adaptively adjusted, improving training efficiency while ensuring stability. The updated policy is deployed to the online system immediately. When the disturbance event is resolved, the system automatically switches back to the normal decision-making mode to ensure continuous optimization of production scheduling. Details are as follows: A Hybrid Priority Experience Replay (HPER) mechanism is employed for experience retrieval. HPER combines the advantages of uniform sampling and priority sampling to balance exploration and utilization, while also adapting to the efficient learning requirements of dynamic scheduling environments. Details are as follows: (1) Experience storage and priority calculation Experience storage format: Each experience tuple Includes state (Graph embedding features extracted by GAT), actions Multi-objective rewards Next state Disturbance sign (1 indicates a dynamic disturbance event, such as a fault or the arrival of an emergency workpiece.) Priority calculation: A hybrid priority system using TD error and disturbance enhancement is employed. in: It is the estimation error of the advantage function. It is a perturbation weight (preferentially learning perturbation events).

[0064] (2) The hybrid sampling strategy employs a two-stage sampling process: Uniform sampling phase: Randomly sample batches of data from the buffer to maintain exploration capabilities.

[0065] Priority sampling phase: by probability Select highly important samples, among which Control priorities.

[0066] Importance sampling correction: To avoid bias, priority sampling samples are weighted. , The correction magnitude is gradually reduced as the initial value is increased linearly to 1.0.

[0067] (3) Dynamic buffer management Segmented storage: The experience buffer is divided into a static experience area (normal scheduling data) and a dynamic disturbance area (fault / emergency workpiece data).

[0068] Automatic eviction mechanism: When the buffer is full, low-priority items in the static area are evicted first. The old sample.

[0069] In this embodiment of the invention, based on generalized advantage estimation, the task allocation agent, worker scheduling agent, and robot management agent adaptively update the parameters of their respective policy networks with the objective of maximizing the shearing objective function. Specifically: The Generalized Advantage Estimation (GAE) calculation, combined with the reward function, is as follows: in: Advantage estimation for time step t; discount factor and GAE parameters Dynamically adjusted, decaying as training progresses; For TD-error; Value network output; It serves as a multi-step forward-looking window.

[0070] The adaptive shearing objective function is as follows: in: It is the first Trainable parameters of the policy network of an agent It is the ratio of the current strategy to the old strategy. This is the shear threshold at time step t. This threshold can be dynamically adjusted, as follows: in: The initial maximum threshold, The minimum protection threshold, For the Wasserstein distance between the old and new strategies, Distance threshold The linear growth coefficient is... This is the attenuation coefficient. Therefore, when the strategy changes abruptly... (For example, when dealing with robot malfunctions), the strategy needs to be adjusted quickly, automatically widening the shearing range to accelerate adaptation; when The strategy updates are stable, gradually tightening for refined optimization.

[0071] (3) Value function update The value network outputs a comprehensive value score for the system status by comprehensively evaluating indicators such as completion time, energy consumption, and load balancing. The value network is updated by minimizing the mean squared error between its predicted values ​​and actual returns. To achieve this, the actual return value is calculated using the n-step time difference (n-step TD) method. That is, the sum of the real rewards in the next n steps plus the value estimate of the state after the nth step, and the network parameters are optimized by gradient descent algorithm to ensure that the value function can accurately evaluate the long-term expected return of the system state, providing a reliable benchmark for the advantage estimation of the policy network.

[0072] The policy network update process is an online learning process, which specifically includes: Data collection: every time completed The next scheduling decision involves collecting a batch of experience and storing it in a buffer, and marking dynamic disturbance events.

[0073] Parameter update: per Step to sample a batch from HPER and execute Round of PPO updates (per round) (a minibatch).

[0074] Policy synchronization: The updated policy is immediately deployed to the online scheduling system to achieve real-time optimization.

[0075] This invention also provides a human-machine collaborative scheduling system for dynamic graph attention and multi-agent reinforcement learning, which is used to execute the human-machine collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning as described above.

[0076] This invention also provides a computer-readable storage medium storing a computer program for human-computer collaborative scheduling of dynamic graph attention and multi-agent reinforcement learning, wherein the computer program causes a computer to execute the human-computer collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning as described above.

[0077] This invention also provides an electronic device, including: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including a human-machine collaborative scheduling method for performing dynamic graph attention and multi-agent reinforcement learning as described above.

[0078] In summary, compared with existing technologies, it has the following beneficial effects: 1. The embodiments of the present invention construct an adaptive scheduling system by dynamically updating the global state graph and using an online learning mechanism. This ensures that the task allocation scheme can be quickly adjusted in a dynamically changing production environment. At the same time, the scheduling system can respond in real time to various sudden disturbances in the production process, including equipment failures, emergency plug-ins, etc., ensuring the continuity and stability of the production process. This effectively solves the technical problem of insufficient adaptability of existing mathematical programming methods and metaheuristic algorithms in dynamic environments.

[0079] 2. Based on the improved NSGA-II algorithm and three-dimensional reward function, this invention achieves a better balance among multiple objectives such as production efficiency, energy consumption, and load balancing. The system can dynamically adjust optimization strategies according to real-time production status, significantly improving overall production efficiency while reducing energy consumption and improving the fairness of worker workload.

[0080] 3. Through a fatigue perception model and a robot energy consumption optimization mechanism, this invention achieves efficient matching of human and robot resources. The system can dynamically adjust task allocation, avoiding worker overwork while improving robot utilization and fully leveraging the advantages of human-robot collaboration.

[0081] 4. This invention, through its improved PPO algorithm and dynamic priority adjustment mechanism, enables the system to maintain stable operation even under continuous disturbances. Compared to traditional methods, this invention can restore production rhythm more quickly and reduce the impact of strategy fluctuations on the production process.

[0082] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0083] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A human-machine collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning, characterized in that, Dynamic scheduling is achieved through the collaborative work of a task allocation agent, a worker scheduling agent, and a robot management agent within a multi-agent cooperative reinforcement learning framework. The human-machine collaborative scheduling method includes: Real-time collection of various data from the production line, dynamic updating of the global status graph, and generation of a task feature subgraph for each task to describe the optional processing path; A graph attention network is used to perform deep encoding and feature extraction on the constructed global state graph and task feature subgraph to obtain global graph embedding features and task subgraph embedding features; Multi-agent teams make collaborative decisions based on global graph embedding features and task subgraph embedding features to obtain a set of candidate scheduling schemes; With the goal of balancing production efficiency, energy consumption, and manpower load, a multi-objective genetic algorithm is used to process the candidate scheduling scheme set to obtain the optimal scheduling scheme that can balance multiple objectives. The optimal scheduling scheme is issued to the production line for execution, and the next global state, the rewards of each agent and the global reward are collected and stored in the experience pool as experience data for updating the multi-agent collaborative reinforcement learning framework.

2. The human-machine collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning as described in claim 1, characterized in that, The multi-agent cooperative reinforcement learning framework includes constructing a multi-objective scheduling model, which includes an objective function and constraints. The objective function includes: Objective 1 aims to achieve the maximum completion time. Minimize; Objective 2 aims to minimize the robot's total energy consumption. Minimize; Objective 3 aims to achieve worker load variance. Minimize; In the formula, The constraints include: Among them, constraint 5 indicates that each task must be assigned to one and only one workstation for a certain processing method; constraint 6 indicates that a single workstation can only process one independent task at a time; and constraint 7 indicates that the total number of robots assigned to all workstations does not exceed the total number of collaborative robots. Constraint 8 indicates that if the task Must be in the task If completed before, then The start time must not be earlier than The completion time; constraint 9 indicates that if the robot exist If a fault occurs, its processing tasks should not be assigned. in, The average load for all workers; Represents a set of workpieces; Indicates workpiece The set of tasks in; Represents the set of workstations, totaling... One workstation; It indicates that the workers gathered together. One worker; Represents a collection of collaborative robots, totaling... One collaborative robot; Represents a set of processing methods. , Indicates that workers are processing the materials. Indicates robot processing, This indicates human-machine collaborative processing; Indicates workpiece The set of priority relationships for the included tasks. Indicates workpiece Tasks in Need to be Completed before; Indicates workpiece Task Start time; Indicates workpiece Task Completion time; Indicates workstation Number of robots used ; Indicates workpiece Task The complexity; Indicates workpiece Task Method of adoption Processing time; Indicates the robot's fault status. 1 indicates a fault, and 0 indicates normal operation; Indicates that the robot is processing the workpiece. Task Energy consumption coefficient; Indicates workpiece Task At workstation In a way The value is 1 if the process is complete, and 0 otherwise. Indicates workpiece Task At workstation In a way The value is 1 if processing is performed, otherwise it is 0. Indicates workpiece Task At workstation In a way The value is 1 if processing is performed, otherwise it is 0. Indicates workers Total processing time.

3. The human-machine collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning as described in claim 1, characterized in that, The global state diagram includes ; in: This represents the set of nodes in the global state diagram, including workstation nodes, task nodes, and resource nodes. Workstation nodes represent the real-time state of the physical workstation; task nodes are dynamically generated and contain task attributes; and resource nodes represent workers and robots, carrying their real-time states. This represents the set of edges in the global state graph, which includes task dependency edges and resource allocation edges. Task dependency edges represent the priority constraints between tasks, while resource allocation edges connect tasks with workstations and connect tasks with resources.

4. The human-machine collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning as described in claim 3, characterized in that, The dynamic update of the global state graph includes: The global state graph is updated in real time through a dynamic update mechanism. Specifically, the dynamic update mechanism is as follows: when a new task or disturbance event triggers a local graph update, only the affected nodes and related edges are modified.

5. The human-machine collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning as described in any one of claims 1 to 4, characterized in that, The multi-agent system makes collaborative decisions based on global graph embedding features and task subgraph embedding features to obtain a set of candidate scheduling schemes, including: The task assignment agent processes global graph embedding features and task subgraph embedding features to generate three-dimensional discrete actions. ,in: Tasks to be assigned Optional workstations For the selection of processing mode; among which, Indicates that workers are processing the materials. Indicates robot processing, This indicates human-machine collaborative processing; When the processing mode is worker processing or human-machine collaborative processing, the worker scheduling agent selects a specific worker w from the available worker pool to execute the current task; When the processing mode is robotic processing or human-robot collaborative processing, the robot scheduling agent selects a specific robot from the available robot pool. To perform the current task; In the multi-agent cooperative reinforcement learning framework, the internal policy network parameters of the three agents are randomly initialized during the initial training phase. Under the initial policy, after receiving global graph embedding features and task subgraph embedding features, the policy network of the three agents outputs a nearly uniform probability distribution. Based on the probability distribution, random sampling is performed to determine a series of combined actions jointly output by the three agents. This constitutes a set of candidate scheduling schemes.

6. The human-machine collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning as described in claim 5, characterized in that, The worker scheduling agent is also used to update workers. Real-time fatigue level: in, Indicates rest time. Indicates workpiece Task The complexity; Indicates workpiece Task Method of adoption Processing time; Indicates workpiece Task At workstation In a way The value is 1 if the process is complete, and 0 otherwise. like If the fatigue threshold is reached, a rest period will be forcibly inserted. The robot scheduling agent is also used to monitor fault status: if Remove it from the list of available robots; This indicates a robot malfunction.

7. The human-machine collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning as described in any one of claims 1 to 4, characterized in that, The multi-objective genetic algorithm includes a reward function, the expression of which is as follows: in, This indicates a time-based reward. This indicates an energy consumption bonus; This indicates a load reward; , , All are weight parameters; among them, This represents the maximum completion time at time t. This represents the maximum completion time at time t-1; Indicates that the robot is processing the workpiece. Task Energy consumption coefficient; Indicates workpiece Task At workstation In a way The value is 1 if the process is complete, and 0 otherwise. This represents the variance of worker load.

8. A human-machine collaborative scheduling system for dynamic graph attention and multi-agent reinforcement learning, characterized in that, The dynamic job scheduling system is used to execute the human-machine collaborative scheduling method of dynamic graph attention and multi-agent reinforcement learning as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, It stores a computer program for human-computer collaborative scheduling of dynamic graph attention and multi-agent reinforcement learning, wherein the computer program causes the computer to execute the human-computer collaborative scheduling method for dynamic graph attention and multi-agent reinforcement learning as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, include: One or more processors; Memory; as well as One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including a human-machine collaborative scheduling method for performing dynamic graph attention and multi-agent reinforcement learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-agent reinforcement learning method based on value decomposition and attention mechanism

    CN113313267A

  • Robot agent reinforcement learning training method and system in complex scene

    CN119129642A

  • Multi-agent training method based on heterogeneous dynamic graph attention mechanism

    CN119250107A

  • State representation modeling method and system based on multi-domain graph attention network

    CN120672037A

  • Flexible job shop scheduling method based on preference driven graph reinforcement learning

    CN120875285A