A multi-agent based task optimization configuration method and system
Patent Information
- Application Number
- CN202610642672.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-11
- Publication Date
- 2026-08-21
AI Technical Summary
然而,现有技术方案往往仅聚焦于单一任务类型的规划,或对任务阶段的划分不够细致,难以适应复杂动态重载协同任务下机械臂高效、精准运动规划的需求
[0034] 1. Multi-agent collaborative task optimization architecture: Construct a multi-agent system including a robotic arm agent, a task agent, and an optimal performance agent. Through interfaces between agents and learning algorithms, achieve dynamic decomposition, optimized allocation, and efficient execution of complex tasks.
Smart Images

Figure CN122606575A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial robot control and intelligent scheduling technology, specifically providing a task optimization configuration method and system based on multiple intelligences, for the intelligent decomposition, optimized scheduling and efficient execution of complex tasks in robotic arm collaborative scenarios. Background Technology
[0002] Intelligent equipment has become a core direction for the development of the nuclear industry. It must not only meet the requirements of high radiation resistance and high reliability, but also improve operational efficiency and safety through autonomous control and data-driven processes. As a typical representative of intelligent equipment, nuclear industry robots, with their irreplaceable role in scenarios such as nuclear fuel cycle and reactor decommissioning, are becoming a key carrier for technological iteration in the industry. The market size has continued to expand from 576 million yuan in 2022, confirming its strategic value. While the application prospects of nuclear industry robots are broad, technological bottlenecks in complex scenarios still need to be overcome.
[0003] Current robots generally face challenges such as high risks associated with manual operation, poor adaptability to traditional equipment, insufficient flexibility, low technological integration, and inadequate radiation resistance. These challenges are particularly pronounced in nuclear power plant maintenance and repair scenarios. The internal environment of a nuclear island is complex and subject to extreme conditions such as strong radiation, high temperature, and high pressure. For example, traditional manual entry into the containment for inspection carries the risk of high pressure, and while existing robots can replace some tasks, their positioning accuracy is only at the centimeter level, making it difficult to identify millimeter-level defects. Furthermore, robots are prone to malfunctions due to mechanical wear and sensor contamination during long-term operation, and the enclosed and hazardous nature of the nuclear environment significantly increases the difficulty of on-site maintenance, further hindering the comprehensive implementation of intelligent operation and maintenance. Therefore, developing high-performance maintenance and repair robots for the nuclear industry is of urgent practical significance in addressing these issues. On the one hand, by breaking through key technologies such as radiation-resistant design, multimodal perception and control, 5G remote collaborative control, domestic and modular design, and all-terrain obstacle crossing and operation capabilities, the robot's operational accuracy and stability in complex nuclear environments can be significantly improved, reducing the risk of human intervention and ensuring the long-term safe operation of nuclear facilities. On the other hand, intelligent robots integrate remote control and predictive maintenance functions, which can shorten overhaul periods, optimize resource allocation, help the nuclear power industry reduce costs and increase efficiency, and meet the needs of the country's large-scale nuclear power development.
[0004] Task planning, as the highest level of robotic arm motion planning, is responsible for receiving, analyzing, and breaking down tasks. It requires dividing complex task objectives into a series of action sequences that the robotic arm can directly plan and execute. In complex, dynamic, and heavy-load collaborative tasks, the robotic arm's movement includes approaching the target and manipulating it, each with different requirements for its motion patterns. Furthermore, task placement also affects the motion planning results. Therefore, the differences in competition and cooperation among different task types, the emphasis on different task phases, and task placement all place high demands on task planning. However, existing technologies often focus only on planning a single task type or lack sufficient detail in task phase division, making it difficult to meet the needs of efficient and precise motion planning for robotic arms in complex, dynamic, and heavy-load collaborative tasks. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing a multi-intelligence-based task optimization configuration method and system.
[0006] A multi-intelligence-based task optimization configuration system includes a robotic arm agent, a task agent, an optimal performance agent, a performance index agent, and a collaborative task agent.
[0007] The robotic arm agent constrains whether the task can be executed; the task agent analyzes the specific information of the task; the optimal performance agent, based on the optimal performance analysis results of the workspace, determines the placement area of the operation task in the robotic arm's workspace, and its placement position determines the upper limit of optimal performance; the performance index agent is responsible for standardizing and weighting various individual performance indicators to form composite performance indicators, providing a decision-making basis for the optimal performance agent; the collaborative task agent analyzes the task conflicts, dependencies, and resource consumption among multiple robotic arms, and dynamically adjusts the task allocation strategy according to the cooperation or competition between different tasks to achieve task collaboration and coordination.
[0008] Furthermore, the robotic arm agent receives real-time status information of the robotic arm, including joint angles, joint speeds, load conditions, and various performance indicators such as Structural Efficiency Index (SLI) and Global Condition Index (GCI). The robotic arm agent uses this data to evaluate the feasibility of each task, calculate the reachability, path feasibility, and motion constraints of the end effector, and filter out overloaded or conflicting motion sequences to generate a list of executable subtasks.
[0009] Furthermore, the task agent receives task attribute information, including task type, priority, stage division, and task dependencies. It also combines historical task execution data to intelligently break down complex tasks. The task agent divides the task into a sequence of sub-tasks, analyzes the execution order, time constraints, and dependencies of each sub-task, and then sends the preliminary task allocation plan to the robotic arm agent for verification and feasibility assessment.
[0010] Furthermore, the optimal performance agent combines global performance distribution data of the robotic arm's workspace to optimize the placement of tasks in the space. This agent analyzes the performance ceiling of each subtask in the workspace and selects the optimal operating area, feeding back the location selection and performance ceiling information to the robotic arm agent and the task agent to ensure that the task is executed in the optimal performance area.
[0011] Furthermore, a task optimization configuration method based on multi-intelligence specifically includes the following steps:
[0012] S100. Task perception and attribute extraction stage; When the system receives a collaborative transfer task, the task agent intervenes first; Based on the pre-trained semantic parsing algorithm, the task description is analyzed, key metadata is extracted, including task type, priority, time constraints, target physical attributes and temporal dependencies between subtasks, and the key metadata is encapsulated into a structured task request.
[0013] S200. Multi-dimensional state comprehensive perception stage; through the fusion of multi-source heterogeneous data, a real-time state model of the system and environment is constructed;
[0014] S300. Task intelligent decomposition stage: Based on structured task requests, the task agent combines historical case knowledge base or temporal logic model to decompose complex tasks into a sequence of sub-tasks that satisfy Markov decision process, with each sub-task corresponding to a single robotic arm action.
[0015] S400. Deep reinforcement learning coupled optimization stage: Subtask sequences, robotic arm state information, dynamic performance map and collaborative constraints are jointly input into the deep reinforcement learning model to perform multi-dimensional joint optimization decisions, including: optimal allocation between subtasks and robotic arms; generation of motion execution sequences for each robotic arm; optimization of target pose for operating subtasks; selection of the optimal spatial position that meets the needs of different stages based on the performance map, thereby achieving overall optimization of task execution efficiency and stability.
[0016] S500. Feasibility verification stage: The robotic arm agent verifies the planning results, calculates the reachability of the target pose through inverse kinematics, and combines an octree-based collision detection method to determine whether there are joint overruns, singular states, or path collisions.
[0017] S600. Task execution and closed-loop learning phase: After successful verification, control each robotic arm to execute the planned actions, and record key data during the execution process, including trajectory, force feedback and performance index changes, for subsequent analysis and learning.
[0018] Furthermore, step S200 is implemented as follows:
[0019] S210. Robotic Arm Status Perception: The robotic arm intelligent agent acquires the operating status parameters of each robotic arm in real time through a sensor network;
[0020] S220. Dynamic performance map generation: Based on the current robotic arm operating status parameters, the operational performance of each point in the workspace is calculated using a pre-established performance index model, generating the global condition index (GCI) and structural efficiency index (SLI) distributions, thereby forming a dynamic performance map that reflects the space operation capability.
[0021] S230. Confirming Cooperative Constraints and Cooperative Environmental Perception: The robotic arm agent acquires environmental information through vision or lidar, identifies the location and dynamic changes of obstacles, and detects whether there are other parallel tasks; the cooperative agent assesses the risk of shared space conflicts and forms cooperative constraints.
[0022] Furthermore, step S400 is implemented as follows:
[0023] S410. High-dimensional heterogeneous state space construction stage; This space is used to fuse task semantic features, robotic arm electromechanical state features and dynamic performance feature matrix; The task semantic features, robotic arm electromechanical state features and dynamic performance feature matrix are standardized, spliced or fused to form a composite state tensor for deep reinforcement learning decision-making;
[0024] S420. Deep Neural Network Inference Stage; The system adopts an actor-critic architecture as a deep reinforcement learning model; Taking the composite state tensor as input, it first enters the deep feature extraction network in the deep reinforcement learning model, uses a multilayer perceptron to process task and temporal features, and uses a convolutional neural network to process the voxelized performance feature matrix to extract the coupling features between task requirements and spatial physical performance; The Actor network outputs action results through pre-training, including: discrete action output, used to decide which robotic arm to assign the current subtask to; and continuous action output, used to generate the optimal position and end-effector posture parameters of the subtask in the workspace;
[0025] S430. Multi-objective reward function calculation stage: Based on the discrete and continuous actions output by the Actor network, state transitions are performed in the environment, and immediate rewards are calculated. When the subtask is heavy-load handling or fine assembly, if the selected operation pose is in the high GCI value region, a preset high reward is given; if it is in the low GCI value region, a preset low reward is given. When the subtask is no-load rapid approach, the reward is allocated according to the SLI index of the corresponding region of the path. If the selected operation pose is in the high SLI value region, a preset high reward is given; if it is in the low SLI value region, a preset low reward is given. Through the above reward mechanism, explicit optimization of global performance indicators is achieved during task execution.
[0026] S440. Policy Update Phase: The Critic network estimates the long-term cumulative reward of the current state and calculates the value error by combining the immediate reward and task completion status. The parameters of the Actor network and Critic network are updated according to the value error, so that the policy network can gradually learn to output better actions under different task and environmental states.
[0027] S450. Dynamic Trigger Monitoring Mechanism: This mechanism runs throughout the entire process of deep reinforcement learning model training and online operation. Training Mode: Repeatedly executes state construction, network inference, reward calculation, and policy update processes until the model converges. Online Inference Mode: Utilizes the trained model to generate task assignments and pose decisions in real time, while a supervision mechanism continuously monitors the difference between the current state and the training distribution. When a sudden change in task attributes, significant environmental changes, or performance degradation of the robotic arm leading to model mismatch is detected, the system immediately interrupts the original direct inference process and reactivates the dynamic optimization solution process, quickly generating new task configurations and operation plans based on the latest state information.
[0028] Furthermore, step S500 is implemented as follows:
[0029] When verification fails, the conflict information is fed back to the collaborative task agent to correct the constraints and return to the optimization stage to solve the problem again until a feasible solution is obtained.
[0030] Furthermore, step S600 is implemented as follows:
[0031] S610. Dynamic adaptive mechanism: During task execution, the system continuously monitors environmental changes and system operating status. Once an abnormal situation is detected, the current task is immediately interrupted and a replanning process is triggered to achieve dynamic adaptive adjustment.
[0032] S620. Learning and updating mechanism: Based on the task execution results, the task decomposition strategy, performance evaluation model and decision network are updated, and the system performance is continuously optimized through positive and negative sample learning to achieve adaptive evolution.
[0033] The beneficial effects of this invention are as follows:
[0034] 1. Multi-agent collaborative task optimization architecture: Construct a multi-agent system including a robotic arm agent, a task agent, and an optimal performance agent. Through interfaces between agents and learning algorithms, achieve dynamic decomposition, optimized allocation, and efficient execution of complex tasks.
[0035] 2. Task space optimization strategy based on global performance indicators: By combining the analysis results of the structural efficiency index (SLI) and global condition index (GCI) of the robotic arm, the optimal placement position of the task in the workspace of the robotic arm is determined, thereby maximizing the task execution performance and ensuring efficient and precise robotic arm collaboration.
[0036] 3. Dynamic adaptive optimization and supervision mechanism: Through learning algorithms and supervision mechanisms, when the task type or environment changes, the corresponding intelligent agent queue is automatically activated to optimize the task solution, thereby achieving real-time adaptation and high robustness of task scheduling.
[0037] In summary, this invention employs a multi-agent collaborative approach: by constructing a robotic arm agent, a task agent, and an optimal performance agent, it achieves intelligent decomposition and optimized allocation of complex tasks, improving the efficiency and accuracy of robotic arm collaboration. It also employs a task space optimization technique based on global performance indicators: utilizing the robotic arm's structural efficiency index (SLI) and global condition index (GCI), it determines the optimal placement of tasks in the workspace, maximizing task execution performance and reducing robotic arm path redundancy and operation time. Finally, it utilizes a dynamic adaptive optimization technique: learning algorithms and supervision mechanisms respond to task changes in real time, ensuring the adaptability of task scheduling and the robustness of the system.
[0038] This invention enables efficient execution of collaborative tasks by robotic arms in complex and dynamic environments, reducing task completion time and error rates while fully utilizing robotic arm resources. Its application in the intelligentization of nuclear power plant robots accelerates the progress of intelligentization. Such technological breakthroughs not only drive the intelligent transformation of the nuclear industry but also provide forward-looking technological reserves for next-generation nuclear energy systems such as high-temperature gas-cooled reactors and fusion reactors, strengthening my country's core competitiveness in the global nuclear technology field. Attached Figure Description
[0039] Figure 1 This is a schematic diagram of a task planning organization structure based on multi-agent systems.
[0040] Figure 2 This is a schematic diagram of the overall performance analysis of the robotic arm.
[0041] Figure 3 This is a flowchart of the present invention.
[0042] Figure 4 This is a flowchart of the regional optimization sub-process of the present invention. Detailed Implementation
[0043] The multi-agent task optimization configuration method provided by the present invention will be described in detail below with reference to the accompanying drawings, so as to clearly describe the data processing process of each agent, the technical means of the reinforcement learning stage, and the processing differences under different task types and stages.
[0044] like Figure 1 As shown, a multi-agent-based task optimization and configuration system comprises multiple agents, including a robotic arm agent, a task agent, an optimal performance agent, a performance index agent, and a collaborative task agent. The robotic arm agent receives real-time status information from the robotic arm, including joint angles, joint speeds, load conditions, and various performance indices such as Structural Efficiency Index (SLI) and Global Condition Index (GCI). The robotic arm agent uses this data to evaluate the executability of each task, calculates the reachability, path feasibility, and motion constraints of the end effector, filters out overloaded or conflicting motion sequences, and generates a list of executable sub-tasks, providing a basis for task allocation.
[0045] The task agent receives task attribute information, including task type, priority, stage division, and task dependencies. It also combines this information with historical task execution data to intelligently break down complex tasks. The task agent divides the task into a sequence of sub-tasks and analyzes the execution order, time constraints, and dependencies of each sub-task. Finally, it sends a preliminary task allocation plan to the robotic arm agent for verification and feasibility assessment.
[0046] Based on this, the optimal performance agent, combined with global performance distribution data of the robotic arm's workspace, optimizes the placement of tasks in the space. This agent analyzes the performance ceiling of each subtask within the workspace, selects the optimal operating area, and feeds back the location selection and performance ceiling information to both the robotic arm agent and the task agent, ensuring that tasks are executed within the optimal performance area.
[0047] The performance index agent is responsible for standardizing and weighting various individual performance indices to form composite performance indices, providing a basis for decision-making for the optimal performance agent.
[0048] The collaborative task agent analyzes the task conflicts, dependencies, and resource consumption among multiple robotic arms, and dynamically adjusts the task allocation strategy according to the cooperation or competition between different tasks to achieve task collaboration and coordination.
[0049] Furthermore, the algorithms built into the robotic arm intelligent agent described in this invention are mainly used to evaluate the executable capability and motion constraints of the robotic arm, including: a forward / inverse kinematics solution algorithm based on DH parameters, used to calculate the mapping relationship between the end effector pose and joint space; a damped least squares (DLS) inverse kinematics iterative algorithm, used to ensure the numerical stability of the solution near singular configurations; a velocity / force mapping and performance index calculation algorithm based on the Jacobian matrix, used to output operability, minimum singular value, reciprocal of condition number, etc. in real time; a collision detection algorithm based on octree space decomposition, used for rapid determination of joint self-collision and environmental collision; an improved RRT* (Rapidly-exploring Random Tree Star) path planning algorithm, used to generate collision-free feasible trajectories in joint space; and a joint limit violation and singular configuration filtering algorithm, used to screen candidate action sequences.
[0050] Furthermore, the built-in algorithm of the task agent described in this invention is geared towards task semantic parsing and decomposition, including: a pre-trained semantic parsing algorithm based on the Transformer structure (BERT / RoBERTa-like model fine-tuning), used to extract structured metadata such as task type, priority, time constraints, and target attributes from the task description; a task decomposition algorithm based on Hierarchical Task Network (HTN), used to recursively expand complex tasks into atomic subtasks; a Markov Decision Process (MDP) modeling algorithm, used to perform state-action-reward modeling on the subtask sequence; a task dependency construction and topology sorting algorithm based on Directed Acyclic Graph (DAG), used to determine the execution order of subtasks; and a historical task matching algorithm based on Case-Based Reasoning (CBR), used to retrieve and reuse decomposition templates from historical tasks.
[0051] Furthermore, the optimal performance intelligent agent's built-in algorithm described in this invention is oriented towards task space placement optimization, including: a workspace voxelization modeling algorithm for discretizing the reachable space of the robotic arm into a voxel mesh; a GCI / SLI global performance map generation algorithm for calculating the upper limit of the robotic arm's performance on a voxel-by-voxel basis; a gradient ascent-based optimal pose search algorithm for locating extreme points on the performance map; a K-Means or DBSCAN clustering algorithm for dividing the optimal operating region, suboptimal operating region, and forbidden region; and a Bayesian optimization algorithm for converging to the globally optimal placement position in a high-dimensional pose space with fewer samplings.
[0052] Furthermore, the performance index intelligent agent's built-in algorithm described in this invention is geared towards multi-index standardization and fusion, including: Min-Max normalization and Z-Score standardization algorithms to eliminate dimensional differences between individual indices; Principal Component Analysis (PCA) algorithm to reduce the dimensionality of individual indices such as minimum eigenvalue, operability, minimum condition number, and stiffness, and determine the principal component weights; a combined Analytic Hierarchy Process (AHP) and Entropy Weight Method weight determination algorithm to dynamically adjust the weights of each individual index according to the task type; a weighted linear combination algorithm to generate a composite performance index η; and a grey relational analysis algorithm to evaluate the correlation between the composite index and the task execution results.
[0053] Furthermore, the collaborative task agent built-in algorithm described in this invention is geared towards multi-robotic arm cooperative scheduling, including: a Contract Net Protocol (CNP) for task bidding, tendering, and allocation among multiple robotic arms; a Hungarian Algorithm or KM algorithm for finding the optimal matching of the task-robotic arm bipartite graph; a game theory-based Nash equilibrium solution algorithm for handling competitive task allocation; a priority and time window-based conflict resolution algorithm for resolving conflicts in the shared workspace; and multi-agent reinforcement learning (MARL) algorithms (such as MADDPG and QMIX) for learning cooperative strategies in dynamic environments.
[0054] In this system, reinforcement learning plays a crucial role in optimizing decision-making and adaptive adjustment. The sub-task execution status, candidate action list, robotic arm performance metrics, and task placement information output by each agent serve as input to the deep reinforcement learning model. Reinforcement learning employs a deep neural network as the policy network, and through iterative training, it outputs optimized schemes for robotic arm action selection, task allocation, and placement. The reward function is designed to consider task completion efficiency, execution accuracy, and resource utilization, with different weights assigned to different task types and stages. For example, in the approach stage of a handling task, the reward primarily focuses on path efficiency; while in the operation stage, the reward emphasizes grasping accuracy and success rate. Assembly and sorting tasks optimize end-effector accuracy, task order, and completion speed according to different stages. The action results output by reinforcement learning are fed back to each agent, which updates its task breakdown strategy, placement location selection, and action probability distribution based on the feedback, achieving closed-loop mutual learning between the agent and reinforcement learning.
[0055] The data processing of each agent and reinforcement learning model differs across task types and stages. In the approach phase of a handling task, the robotic arm agent focuses on path planning and collision detection, the task agent breaks down the path into sub-tasks, and the optimal performance agent selects the starting position to ensure efficient arrival. In the operation phase, the robotic arm agent precisely controls the end effector's gripping position, the optimal performance agent selects the region with the best end effector performance, and the reinforcement learning model optimizes the gripping success rate and accuracy. In the assembly task, the approach phase emphasizes the smoothness of the assembly path planning, while the operation phase emphasizes assembly accuracy and correct sequence. Sorting tasks prioritize gripping speed and accuracy, ensuring the shortest task completion time and minimal error. Finally, the system integrates the results of multi-agent analysis and reinforcement learning optimization to generate sub-task allocation schemes, robotic arm execution sequences, and task placement positions for each task. These can be directly used for robotic arm control, achieving efficient and precise multi-robotic arm collaborative task execution.
[0056] like Figure 3 As shown, a task optimization configuration method based on multi-intelligence achieves efficient scheduling of complex, dynamic, and heavy-load collaborative tasks by constructing a complete closed loop of "perception-decision-verification-execution-feedback"; specifically, it includes the following steps:
[0057] S100 Task Perception and Attribute Extraction Stage. When the system receives a collaborative transfer task triggered by the task model, the task agent first intervenes, analyzes the task description based on a pre-trained semantic parsing algorithm, and extracts key metadata, including task type, priority, time constraints (earliest start time, latest end time), target physical attributes (quality, size, grasping features), and temporal dependencies between subtasks. The above information is then encapsulated into a structured task request to provide input for subsequent planning.
[0058] S200 Multidimensional State Comprehensive Perception Stage. Through the fusion of multi-source heterogeneous data, a real-time state model of the system and its environment is constructed.
[0059] S210 Robotic Arm Status Perception: The robotic arm intelligent agent obtains the operating status parameters of each robotic arm in real time through a sensor network, including joint angle, joint speed, driving torque, end effector load and joint health status, which are used to characterize the current executable capability of the robotic arm;
[0060] S220 Dynamic Performance Map Generation: Based on the current robotic arm operating status parameters, the operation performance (such as operability and condition number) of each point in the workspace is calculated using a pre-established performance index model, and the distribution of the Global Condition Index (GCI) and Structural Efficiency Index (SLI) is further generated, thus forming a dynamic performance map that reflects the space operation capability.
[0061] Furthermore, the Global Condition Index (GCI) and Structural Efficiency Index (SLI) are used as inputs to obtain a dynamic performance map. The Global Condition Index (GCI) is calculated through composite performance indicators. Each individual indicator in the composite performance indicators corresponds to the state of each robotic arm and is synchronously associated with a Jacobian matrix.
[0062] Furthermore, the corresponding Global Condition Index (GCI) or Structural Efficiency Index (SLI) can be selected based on the task type.
[0063] S230. Confirming Cooperative Constraints and Cooperative Environmental Perception: The robotic arm agent acquires environmental information through perception means such as vision or lidar, identifies the location and dynamic changes of obstacles, and detects whether there are other parallel tasks; the cooperative agent assesses the risk of shared space conflicts and forms cooperative constraints.
[0064] S300. Task Intelligent Decomposition Stage. Based on structured task requests and combined with a historical case knowledge base or temporal logic model, the task agent decomposes complex tasks into a sequence of sub-tasks that satisfy Markov decision processes. Each sub-task corresponds to a single executable action of the robotic arm, and its performance requirements (such as high precision or high efficiency) are clearly defined, forming an executable sub-task sequence.
[0065] S400. Deep Reinforcement Learning Coupled Optimization Stage. The subtask sequence, robotic arm state information, dynamic performance map, and collaborative constraints are jointly input into the deep reinforcement learning model to perform multi-dimensional joint optimization decisions, including: optimal allocation between subtasks and robotic arms; generation of motion execution sequences for each robotic arm; optimization of target poses for operating subtasks; and selection of the optimal spatial position that meets the needs of different stages based on the performance map, thereby achieving overall optimization of task execution efficiency and stability.
[0066] S500. Feasibility Verification Phase. The robotic arm agent verifies the planning results by calculating the reachability of the target pose through inverse kinematics and using an octree-based collision detection method to determine whether there are problems such as joint overstepping, singular states, or path collisions.
[0067] S510. Conflict Feedback and Correction Phase. When verification fails, conflict information is fed back to the cooperative task agent to correct the constraints (such as setting no-entry zones or restricting the action space), and the process returns to the optimization phase to solve again until a feasible solution is obtained.
[0068] S600. Task Execution and Closed-Loop Learning Phase. After successful verification, control each robotic arm to execute the planned actions, while recording key data during execution, including trajectory, force feedback, and performance index changes, for subsequent analysis and learning.
[0069] S610. Dynamic Adaptive Mechanism. During task execution, the system continuously monitors environmental changes and system operating status. Once an anomaly is detected (such as a sudden environmental change or equipment failure), the current task is immediately interrupted and a replanning process is triggered to achieve dynamic adaptive adjustment.
[0070] S620. Learning and Update Mechanism. Based on the task execution results, the task decomposition strategy, performance evaluation model, and decision network are updated. Through positive and negative sample learning, the system performance is continuously optimized to achieve adaptive evolution.
[0071] like Figure 4 The diagram shows the sub-process for dynamic task allocation and placement optimization based on reinforcement learning, which is further refined. Figure 3 The internal mechanism of the S400 deep reinforcement learning coupling optimization stage in the core decision-making process demonstrates the micro-process of task dynamic allocation and placement area optimization based on deep reinforcement learning, clearly showing how the present invention explicitly introduces physical performance indicators into the intelligent decision-making process.
[0072] SS401. High-Dimensional Heterogeneous State Space Construction Phase. This system constructs a high-dimensional heterogeneous state space for collaborative task decision-making, which integrates task semantic features, robotic arm electromechanical state features, and dynamic performance feature matrix. This state space is not a simple superposition of the original data, but a composite state representation formed by extraction and processing by each agent. It mainly includes: Task semantic features: The task agent extracts information such as the time sequence number, task type, priority, and remaining time window of the current subtask, and encodes it into a task feature vector to represent the urgency of the task and operational requirements; Robotic arm electromechanical state features: The robotic arm agent collects real-time operating parameters such as angle, speed, temperature, and load rate of each joint to reflect the current execution capability of the robotic arm; Dynamic performance feature matrix: The optimal performance agent performs voxel modeling of the reachable workspace based on the current state of the robotic arm, obtains the GCI / SLI performance indicators and their changing trends at each voxel position, and forms a multi-dimensional dynamic performance feature matrix reflecting the space operation capability. Finally, the aforementioned task semantic features, robotic arm electromechanical state features, and dynamic performance feature matrices are standardized, spliced, or fused to form a composite state tensor for deep reinforcement learning decision-making.
[0073] SS402. Deep Neural Network Inference Stage. The system employs an Actor-Critic architecture as its deep reinforcement learning model. The composite state tensor is input into the deep feature extraction network within the reinforcement learning model. A multilayer perceptron processes task and temporal features, while a convolutional neural network processes the voxelized performance feature matrix, thereby extracting the coupling features between task requirements and spatial physical performance. Based on this, the Actor network outputs action results through pre-training, specifically including: discrete action output: used to decide which robotic arm to assign the current subtask to; and continuous action output: used to generate the optimal position and end-effector posture parameters of the subtask in the workspace. Through the combined output of these discrete and continuous actions, integrated optimization of task allocation and operational pose is achieved.
[0074] SS403. Multi-objective reward function calculation stage. Based on the discrete and continuous actions output by the Actor network, the system performs state transitions in the environment and calculates immediate rewards. The reward function adopts a multi-objective design, with weights dynamically adjusted according to different sub-task types. Performance rewards are a key design element of this invention, directly relying on a dynamic GCI / SLI performance map: when the sub-task is heavy-duty handling or fine assembly, if the selected operation pose is in a high GCI value region, a preset high reward is given; if it is in a low GCI value region, a preset low reward is given, thus guiding the model to prioritize selecting a better operation area. When the sub-task is unloaded rapid approach, rewards are allocated according to the SLI index of the corresponding area of the path. If the selected operation pose is in a high SLI value region, a preset high reward is given; if it is in a low SLI value region, a preset low reward is given, guiding the robotic arm to move towards a region with higher structural efficiency. Through the above reward mechanism, explicit optimization of global performance indicators is achieved during task execution.
[0075] The high GCI value region refers to the region that exceeds a set threshold, specifically the region that exceeds the average value; the low GCI value region refers to the region that is below the set threshold.
[0076] SS404. Policy Update Phase. The Critic network estimates the long-term cumulative reward of the current state and calculates the value error by combining the immediate reward and the task completion state (the task completion state is directly obtained by inputting discrete and continuous actions into the simulation environment E). The parameters of the Actor network and Critic network are updated based on the value error, so that the policy network can gradually learn to output better actions under different task and environment states, thereby improving the quality of task configuration and pose planning.
[0077] SS405. Dynamic Trigger Monitoring Mechanism. This sub-process runs throughout the entire training and online operation of the deep reinforcement learning model: Training mode: Repeatedly execute state construction, network inference, reward calculation, and policy update processes until the model converges; Online inference mode: Utilize the trained model to generate task assignments and pose decisions in real time, while a supervision mechanism continuously monitors the difference between the current state and the training distribution. When a sudden change in task attributes, a significant change in the environment, or a degradation in the robotic arm's performance leading to model mismatch is detected, the system immediately interrupts the original direct inference process and reactivates the dynamic optimization solution process. Based on the latest state information, it quickly generates new task configurations and operation schemes to improve the system's robustness and continuous operation capability in complex dynamic scenarios.
[0078] Example 1: Establishment of a performance index model for a robotic arm and global performance analysis
[0079] The performance of a robotic arm plays a crucial role in the overall success of a task. Typically, tasks are completed under optimal performance conditions to ensure quality execution. Therefore, this paper proposes to establish an index model using the Jacobian matrix of a six-DOF robotic arm as a medium, and analyzes the global performance of the robotic arm based on this model.
[0080] (1) Establishment of performance index model.
[0081] First, the Jacobian matrix of the robotic arm is established. Based on this, single performance indices such as minimum eigenvalue, operability, minimum condition number, and stiffness are constructed. Furthermore, a composite performance index for evaluating the performance of a heavy-duty mobile robotic arm is developed. It includes several individual performance metrics, expressed as:
[0082] (1)
[0083] In the formula, Let be the weight value of the i-th indicator, which can be determined by principal component analysis (PCA). and ; Performance values obtained by substituting specific data into different individual performance indicators.
[0084] when When calculating the Global Condition Index (GCI), the reciprocal of the condition number of the Jacobian matrix is included. Operability Minimum singular value Minimum eigenvalue and joint stiffness index Single performance indicators that are strongly correlated with the Jacobian matrix; among them, Let Jacobian matrix be the value of the robotic arm under the current joint configuration. and They are respectively Maximum and minimum singular values, Let be the joint stiffness matrix. When When used to calculate the Structural Efficiency Index (SLI), the normalized ratio of link length is included. achievable workspace volume The total length of the robotic arm is cubic The ratio (i.e.) Indicators related to the structural parameters of the robotic arm, such as the utilization rate of joint range of motion, etc.
[0085] (2) Preliminary establishment of the overall performance of the robotic arm.
[0086] The global performance index of a robotic arm is an index independent of its end-effector posture. Given the posture of the robotic arm, its global performance index has a unique fixed value. This invention uses the Structure Length Index (SLI) and Global Index (GCI) to initially establish the global performance index of the robotic arm. GCI is a global performance estimation table based on the condition number of the Jacobian matrix, independent of the robotic arm's configuration. This index is used to analyze the distribution of the condition number throughout the robotic arm's workspace, showing the distribution characteristics of the local performance index of the condition number within the entire workspace. In addition, a well-designed robotic arm should generally have a large reachable workspace. Typically, the longer the robotic arm, the larger its reachable workspace. However, even robotic arms of the same length can have significantly different reachable workspaces due to different configurations and varying link lengths. Therefore, the SLI index is used to evaluate the structural efficiency of the robotic arm, and the link dimensions are optimized based on this index to obtain the optimal robotic arm dimensions. The global optimal performance of the robotic arm varies significantly at different locations within the workspace. When performing complex tasks, regions with better global optimal performance should be selected, such as... Figure 2 As shown.
[0087] In this invention, establishing a performance index model for the robotic arm is a fundamental step in the multi-agent task optimization configuration method. First, a Jacobian matrix model is established for each robotic arm, and this model is used to calculate the end-effector motion performance under different joint configurations. Individual performance indices include minimum eigenvalue, operability, minimum condition number, and stiffness, which respectively reflect the robotic arm's motion flexibility, operability, and load-bearing capacity in specific postures. Subsequently, the individual performance indices are weighted and combined using methods such as Principal Component Analysis (PCA) to form a composite performance index, resulting in a quantitative expression of the overall motion capability of the robotic arm. The output of this model is a set of local performance characteristics of the robotic arm under different postures, such as... Figure 2As shown in (a), different curves represent the optimal and suboptimal values of each performance index, providing basic data for subsequent task placement and action selection.
[0088] Building upon this foundation, the initial establishment of the robotic arm's global performance further extends local performance to the entire workspace. Through calculations using the SLI (Structural Efficiency Index) and GCI (Global Condition Index) indices, a globally optimal performance distribution across the robotic arm's entire workspace is formed, such as... Figure 2 As shown in (b), a performance distribution map in the form of a spiral or gradient indicates the upper limit and suboptimal performance range that the robotic arm can achieve at different spatial locations. This global performance map not only reflects the reachability of the robotic arm, but also reveals which locations are more suitable for high-precision or high-load operations when performing complex collaborative tasks.
[0089] Example 2
[0090] To further verify the effectiveness of the multi-agent task optimization configuration method described in this invention, this embodiment selects a typical six-degree-of-freedom heavy-duty collaborative robotic arm as the simulation object, and its ontological parameters and simulation environment are as follows.
[0091] (1) Robotic arm model and structural parameters
[0092] The simulation object is a 6-DOF heavy-duty articulated robotic arm, mounted on an omnidirectional mobile chassis, forming a heavy-duty mobile robotic arm system. Its improved DH parameters are shown in the table below:
[0093]
[0094] The robotic arm has a rated load of 50 kg, a peak load of 80 kg, a maximum working radius of 2000 mm, and a repeatability of ±0.08 mm.
[0095] (2) Simulation environment and task settings
[0096] The simulation was built using MATLAB R2023a + Robotics Toolbox in conjunction with CoppeliaSim 4.5. The reinforcement learning model was implemented using PPO (Proximal Policy Optimization) and SAC (Soft Actor-Critic) in PyTorch 2.0, and data communication was achieved through ROS Noetic. The task scenario is a collaborative transfer scenario with two robotic arms, which includes the following three typical sub-tasks:
[0097] Task A: Heavy-duty handling task (30 kg workpiece, moving from point P1 to point P2);
[0098] Task B: Fine assembly task (end-point force-controlled assembly, tolerance ±0.1 mm);
[0099] Task C: Rapid sorting task (5 kg small items, 6-second cycle time).
[0100] The workspace was discretized into 50 mm × 50 mm × 50 mm voxels, generating approximately 26,000 reachable voxel points. GCI and SLI were calculated point by point to form a dynamic performance map.
[0101] (3) Simulation results and comparative analysis
[0102] The method of this invention was compared with a traditional single planning method (based only on RRT* path planning, excluding multi-agent cooperation and performance map guidance) on the same task set. The statistical results of each indicator are shown in the table below (each group of results is the average of 50 independent simulations):
[0103]
[0104] (4) Key process data recording
[0105] To further demonstrate the explicit optimization effect of the method of the present invention on the global performance of the robotic arm, key intermediate data during the simulation process are sampled and recorded.
[0106] ① GCI sampling data during the execution phase of Task A (heavy-load handling)
[0107] The GCI values of the voxels traversed by the trajectory at the end of the execution phase of Task A were sampled in segments (with the task start point as 0% and the target position as 100%), and the comparison results are shown in the table below:
[0108]
[0109] Data shows that the traditional method's GCI value drops significantly to 0.28 when approaching the target position, which is close to the singular configuration region and is prone to causing a sudden increase in joint velocity and end effector jitter. However, the method of this invention, guided by the optimal performance agent, keeps the end effector trajectory in the high-performance region of GCI ≥ 0.65. Even during the operation phase, it maintains GCI = 0.71, which is significantly away from the singular configuration, ensuring the stability and end effector accuracy of the heavy-load handling process.
[0110] ② Data from the reinforcement learning training process of Task B (fine assembly)
[0111] The PPO strategy was used for training on Task B. The average cumulative reward and assembly success rate were recorded every 50 rounds. The key sampling points are shown in the table below:
[0112]
[0113] Data shows that the reward function converges and stabilizes after approximately 420 training rounds, and the success rate of assembly increases from 54.0% in the initial training phase to 97.8% after convergence. This verifies the positive effect of the reward function design (explicitly binding GCI / SLI performance map) described in this invention on policy convergence speed and task success rate.
[0114] (5) Conclusion
[0115] The simulation data above show that the task optimization configuration method based on the coupling of multi-agent and deep reinforcement learning proposed in this invention is significantly better than traditional methods in key indicators such as task completion time, execution accuracy, assembly success rate, path redundancy and dual-arm conflict rate. It has obvious advantages in complex collaborative scenarios such as heavy load, fine detail and high cycle time, which verifies the effectiveness and robustness of the multi-agent architecture and global performance index guidance mechanism of this invention.
[0116] Furthermore, the output results of the performance model and the global performance distribution have the following effects on the various agents and reinforcement learning in this invention:
[0117] First, the robotic arm agent uses global performance distribution data to evaluate the feasibility and priority of candidate subtasks. During action sequence planning, the robotic arm agent references local and global performance metrics, selecting the region with the best performance to execute key actions, thereby ensuring the stability and efficiency of task execution.
[0118] Secondly, the optimal performance agent directly relies on the global performance graph for task placement optimization. Based on task type and stage, subtasks are arranged within the robotic arm's globally optimal performance region, ensuring that critical operational stages are completed within this region and avoiding tasks that impact efficiency or accuracy in suboptimal or low-performance regions. Simultaneously, the agent considers suboptimal performance regions as alternatives; when multiple tasks compete for the same optimal region, dynamic allocation ensures rational resource scheduling.
[0119] Furthermore, the performance index agent uses the output of the motion performance index model to quantitatively evaluate the execution constraints of different tasks, providing a reference for the optimal performance agent and the robotic arm agent. This agent calculates a comprehensive performance score based on the index combination results and uses it as one of the state inputs for reinforcement learning, guiding the policy network's decisions on task allocation and action selection.
[0120] During the reinforcement learning phase, the motion performance model and global performance graph are input into the deep reinforcement learning network as part of the environment state. Based on the robotic arm state, task state, and global performance metrics, the network selects action strategies (such as task allocation, sub-task order, and placement location) and executes the simulation. The reward function, combined with the global performance metrics, encourages the execution of key actions in the optimal performance region, reduces path redundancy and operational conflicts, thereby achieving closed-loop collaboration between the reinforcement learning strategy and the multi-agent system. The output of reinforcement learning is further fed back to each agent to adjust task splitting, action sequences, and task placement strategies, forming an adaptive optimization loop.
[0121] In summary, the motion performance index model and global performance distribution not only provide a quantitative basis for the execution capabilities of each agent, but also play a crucial role in state representation within reinforcement learning strategies. Their application ensures efficient task execution within the robotic arm's globally optimal performance region, improving task completion efficiency, operational accuracy, and overall system robustness, and providing reliable data support for the intelligent optimization of complex collaborative tasks.
[0122] The above embodiments fully illustrate the technical solution of the present invention, demonstrating the data processing logic, reinforcement learning optimization mechanism, and adaptive processing capability of the multi-agent system in the optimized configuration of complex tasks, enabling efficient and intelligent scheduling and execution of complex multi-robotic arm tasks.
Claims
1. A task optimization configuration system based on multi-agent systems, characterized in that, This includes robotic arm agents, task agents, optimal performance agents, performance index agents, and collaborative task agents; The robotic arm agent constrains whether the task can be executed; the task agent analyzes the specific information of the task; and the optimal performance agent, based on the optimal performance analysis results of the workspace, determines the placement area of the operation task in the robotic arm's workspace, and its placement position determines the upper limit of optimal performance. The performance index agent is responsible for standardizing and weighting various individual performance indicators to form composite performance indicators, providing a basis for decision-making for the optimal performance agent; the collaborative task agent analyzes the task conflicts, dependencies, and resource consumption among multiple robotic arms, and dynamically adjusts the task allocation strategy according to the cooperation or competition between different tasks to achieve task collaboration and coordination.
2. The multi-agent-based task optimization configuration system according to claim 1, characterized in that, The robotic arm agent receives real-time status information from the robotic arm, including joint angles, joint speeds, load conditions, and various performance indicators such as Structural Efficiency Index (SLI) and Global Condition Index (GCI). The robotic arm agent uses this data to evaluate the feasibility of each task, calculate the reachability, path feasibility, and motion constraints of the end effector, and filter out overloaded or conflicting motion sequences to generate a list of executable subtasks.
3. The multi-agent-based task optimization configuration system according to claim 1, characterized in that, The task agent receives task attribute information, including task type, priority, stage division, and task dependencies. It also combines historical task execution data to intelligently break down complex tasks. The task agent divides the task into a sequence of sub-tasks, analyzes the execution order, time constraints, and dependencies of each sub-task, and then sends the preliminary task allocation plan to the robotic arm agent for verification and feasibility assessment.
4. The multi-agent-based task optimization configuration system according to claim 1, characterized in that, The optimal performance agent combines global performance distribution data of the robotic arm's workspace to optimize the placement of tasks in the space. This agent analyzes the performance limit of each subtask in the workspace and selects the optimal operating area. It then feeds back the location selection and performance limit information to the robotic arm agent and the task agent to ensure that the task is executed in the optimal performance area.
5. A task optimization configuration method based on multi-agent systems, characterized in that, This method is used to implement any of the systems described in claims 1-4, and specifically includes the following steps: S100. Task perception and attribute extraction stage; When the system receives a collaborative transfer task, the task agent intervenes first; Based on the pre-trained semantic parsing algorithm, the task description is analyzed, key metadata is extracted, including task type, priority, time constraints, target physical attributes and temporal dependencies between subtasks, and the key metadata is encapsulated into a structured task request. S200. Multi-dimensional state comprehensive perception stage; through the fusion of multi-source heterogeneous data, a real-time state model of the system and environment is constructed; S300. Task intelligent decomposition stage: Based on structured task requests, the task agent combines historical case knowledge base or temporal logic model to decompose complex tasks into a sequence of sub-tasks that satisfy Markov decision process, with each sub-task corresponding to a single robotic arm action. S400. Deep reinforcement learning coupling optimization stage; The subtask sequence, robotic arm state information, dynamic performance map, and collaborative constraints are jointly input into the deep reinforcement learning model to perform multi-dimensional joint optimization decisions, including: optimal allocation between subtasks and robotic arms; generation of motion execution sequences for each robotic arm; optimization of target pose for operating subtasks; and selection of the optimal spatial position that meets the needs of different stages based on the performance map, thereby achieving overall optimization of task execution efficiency and stability. S500. Feasibility verification stage: The robotic arm agent verifies the planning results, calculates the reachability of the target pose through inverse kinematics, and combines an octree-based collision detection method to determine whether there are joint overruns, singular states, or path collisions. S600. Task execution and closed-loop learning phase: After successful verification, control each robotic arm to execute the planned actions, and record key data during the execution process, including trajectory, force feedback and performance index changes, for subsequent analysis and learning.
6. The task optimization configuration method based on multi-agent systems according to claim 5, characterized in that, Step S200 is implemented as follows: S210. Robotic Arm Status Perception: The robotic arm intelligent agent acquires the operating status parameters of each robotic arm in real time through a sensor network; S220. Dynamic performance map generation: Based on the current robotic arm operating status parameters, the operational performance of each point in the workspace is calculated using a pre-established performance index model, generating the global condition index (GCI) and structural efficiency index (SLI) distributions, thereby forming a dynamic performance map that reflects the space operation capability. S230. Confirming Cooperative Constraints and Cooperative Environmental Perception: The robotic arm agent acquires environmental information through vision or lidar, identifies the location and dynamic changes of obstacles, and detects whether there are other parallel tasks; the cooperative agent assesses the risk of shared space conflicts and forms cooperative constraints.
7. The task optimization configuration method based on multi-agent systems according to claim 5, characterized in that, Step S400 is implemented as follows: S410. High-dimensional heterogeneous state space construction stage; This space is used to fuse task semantic features, robotic arm electromechanical state features and dynamic performance feature matrix; The task semantic features, robotic arm electromechanical state features and dynamic performance feature matrix are standardized, spliced or fused to form a composite state tensor for deep reinforcement learning decision-making; S420. Deep Neural Network Inference Stage: The system adopts an actor-critic architecture as a deep reinforcement learning model. The composite state tensor is used as input and first enters the deep feature extraction network in the deep reinforcement learning model. The multilayer perceptron is used to process task and temporal features, and the convolutional neural network is used to process the voxelized performance feature matrix to extract the coupling features between task requirements and spatial physical performance. The Actor network outputs action results through pre-training, including: discrete action output, used to decide which robotic arm to assign the current subtask to; and continuous action output, used to generate the optimal position and end effector parameters of the subtask in the workspace. S430. Multi-objective reward function calculation stage: Based on the discrete and continuous actions output by the Actor network, state transitions are performed in the environment, and immediate rewards are calculated. When the subtask is heavy-load handling or fine assembly, if the selected operation pose is in the high GCI value region, a preset high reward is given; if it is in the low GCI value region, a preset low reward is given. When the subtask is no-load rapid approach, the reward is allocated according to the SLI index of the corresponding region of the path. If the selected operation pose is in the high SLI value region, a preset high reward is given; if it is in the low SLI value region, a preset low reward is given. Through the above reward mechanism, explicit optimization of global performance indicators is achieved during task execution. S440. Policy Update Phase: The Critic network estimates the long-term cumulative reward of the current state and calculates the value error by combining the immediate reward and task completion status. The parameters of the Actor network and Critic network are updated according to the value error, so that the policy network can gradually learn to output better actions under different task and environmental states. S450. Dynamic trigger monitoring mechanism, which runs through the entire process of deep reinforcement learning model training and online operation; Training mode: Repeat the process of state construction, network inference, reward calculation and policy update until the model converges; Online inference mode: The system generates task assignment and pose decision in real time using the trained model, while a supervision mechanism continuously monitors the difference between the current state and the training distribution. When a sudden change in task attributes, a significant change in the environment, or a degradation in the performance of the robotic arm that causes the original model to become mismatched is detected, the system immediately interrupts the original direct inference process, reactivates the dynamic optimization solution process, and quickly generates a new task configuration and operation plan based on the latest state information.
8. The task optimization configuration method based on multi-agent systems according to claim 5, characterized in that, Step S500 is implemented as follows: When verification fails, the conflict information is fed back to the collaborative task agent to correct the constraints and return to the optimization stage to solve the problem again until a feasible solution is obtained.
9. The task optimization configuration method based on multi-agent systems according to claim 5, characterized in that, Step S600 is implemented as follows: S610. Dynamic adaptive mechanism: During task execution, the system continuously monitors environmental changes and system operating status. Once an abnormal situation is detected, the current task is immediately interrupted and a replanning process is triggered to achieve dynamic adaptive adjustment. S620. Learning and updating mechanism; Based on the task execution results, the task decomposition strategy, performance evaluation model and decision network are updated, and the system performance is continuously optimized through positive and negative sample learning to achieve adaptive evolution.