Path planning system and method

By employing a multi-agent reinforcement learning method, a path planning system for a high-degree-of-freedom robotic arm is constructed. Graph neural networks are used to model the joint topology, and a strategy optimization framework combining centralized training and distributed execution is adopted. This solves the problems of discontinuous and unsmooth path planning and slow convergence speed in existing technologies, and enables efficient autonomous obstacle avoidance in complex environments.

CN121018560APending Publication Date: 2025-11-28QINGDAO UNIV OF SCI & TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202511287005.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing path planning methods for multi-degree-of-freedom robotic arms in complex environments suffer from problems such as discontinuous and uneven paths, high computational difficulty, slow convergence speed, and difficulty in achieving global optimization, especially when facing dynamic obstacles and high-dimensional state spaces.

Method used

A multi-agent reinforcement learning approach is adopted, which constructs a path planning system for a high-degree-of-freedom robotic arm through an environment perception module, a multi-agent modeling module, a collaborative reinforcement learning module, an attention mechanism fusion module, and a multi-dimensional reward function evaluation module. The topological structure between joints is modeled using a graph neural network, and a multi-dimensional reward function is constructed to achieve path optimization by combining a policy optimization framework of centralized training and distributed execution.

Benefits of technology

It improves the accuracy and efficiency of path planning, ensures stability and real-time response in complex environments, solves the problems of discontinuous and unsmooth path planning and slow convergence speed in existing technologies, and realizes the autonomous obstacle avoidance capability of high-degree-of-freedom robotic arms in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121018560A_ABST
    Figure CN121018560A_ABST
Patent Text Reader

Abstract

The invention relates to a high-degree-of-freedom mechanical arm path planning method based on multi-agent reinforcement learning. The method at least comprises six modules of environment perception, multi-agent modeling, attention mechanism fusion, collaborative reinforcement learning, and multi-dimensional reward evaluation and path execution. Mechanical arm joints are modeled through a graph neural network, and a topological structure is described; fusing a channel and a space attention mechanism to enhance feature expression; the strategy optimization efficiency is improved by adopting multiple agents of centralized training and distributed execution; a dynamically weighted multi-dimensional reward function is introduced to comprehensively consider path length, obstacle avoidance safety, joint smoothness and energy consumption. Finally, a path sequence is generated and deployed to a controller, and trajectory tracking and obstacle avoidance performance is guaranteed. The result shows that the method is superior to an existing method in the aspects of path planning precision, convergence efficiency, obstacle avoidance performance and the like, and has good engineering adaptability and popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and robot control, and particularly relates to a mechanical arm path planning system and method based on multi-agent reinforcement learning, which is particularly suitable for autonomous obstacle avoidance and efficient path generation of multi-degree-of-freedom mechanical arms in complex dynamic environments. BACKGROUND

[0002] With the continuous development of intelligent manufacturing and human-machine collaboration, mechanical arms, as the core equipment of industrial automation and intelligent manufacturing, have gradually become the key factor affecting their task completion efficiency and safety performance in autonomous path planning in complex environments. The existing multi-joint mechanical arms with six or more degrees of freedom are widely used in aviation manufacturing, medical surgery, precision assembly and other scenes with high environmental adaptability and dynamic response. However, due to the limitations in the face of dynamic obstacles, high-dimensional state space or irregular task areas, the path is often not smooth, the convergence is slow and it is easy to fall into local optimum, which is difficult to meet the requirements of real-time obstacle avoidance and global optimization. Especially in multi-degree-of-freedom mechanical arms, there is a complex coupling relationship between each joint. This not only affects the efficiency of the mechanical arm during work, but also limits the smooth path generation and overall control performance of the motion during work.

[0003] The existing path planning methods are mostly based on graph search (such as A*, Dijkstra) or sampling algorithms (such as RRT, RRT*), which have feasibility but generally have the following problems: 1. high computational difficulty, long algorithm time, low success rate and slow convergence speed, especially in high-dimensional state space; 2. discontinuous and non-smooth path; 3. insufficient modeling of mechanical arm topology and joint coupling, making it difficult to achieve global collaborative decision-making. In recent years, reinforcement learning methods have gradually attracted attention in path planning tasks due to their learning ability and environmental adaptability. However, traditional single-agent reinforcement learning methods face three challenges when dealing with multi-joint high-dimensional control problems: 1. difficulty in capturing the dynamic coupling and collaborative relationship between joints; 2. the reward function is usually designed for a single target, which cannot meet the multi-dimensional requirements of path optimization, smoothness, obstacle avoidance performance and energy consumption; 3. under distributed control, the joint strategies lack joint optimization mechanism and are prone to fall into local optimum.

[0004] Therefore, it is urgent and necessary to develop a path planning algorithm that integrates multi-agent modeling, collaborative reinforcement learning optimization mechanism, multi-dimensional dynamic reward evaluation strategy and high efficiency and stability. SUMMARY

[0005] This invention provides a multi-agent reinforcement learning robot path planning system and method, both of which address the problems of discontinuous and uneven paths, poor accuracy, poor real-time performance, and slow convergence speed in existing technologies, thereby improving the path planning performance and task execution capability of multi-degree-of-freedom robotic arms in complex environments.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0007] This invention provides a path planning system and method, characterized in that it is used to realize intelligent obstacle avoidance of a robotic arm, and includes at least an environment perception module, a multi-agent information module, a collaborative reinforcement learning module, an attention mechanism fusion module, a multi-dimensional reward function evaluation module, and a path execution module; the multi-agent information module includes at least a multi-agent modeling module; the method includes at least the following steps:

[0008] S1: The environmental perception module (10) perceives the environmental information in the workspace of the robotic arm in real time and encodes the environmental information into a high-dimensional state vector;

[0009] S2: The multi-agent modeling module (21) is a topological structure with associated information. The high-dimensional state vector is input into the multi-agent modeling module (21) to update the pose information of each agent. Based on the local state of the multi-agent and the interaction features of neighboring agents after modeling, the channel attention mechanism and spatial attention mechanism of the attention fusion module (40) are used to perform important screening of the joint physical feature group and the environmental feature group to generate joint features and provide more accurate input for decision-making.

[0010] S3: After the information is updated, the agents jointly optimize the policy network of the multi-agents through the strategy of combining centralized training and distributed execution in the collaborative reinforcement learning module (30); during the centralized training phase, each agent can access the global state information of the environment and the comprehensive reward signal R provided by the multi-dimensional reward function evaluation module (50). total By sharing experience and updating their respective policy networks synchronously, the multi-agents independently generate joint action decisions based on their own policies during the distributed execution phase, thereby constructing their respective local path segments in parallel. The execution results are evaluated by the reward function and fed back to the centralized training phase as input for policy updates, forming a closed-loop policy optimization.

[0011] S4: The multidimensional reward function module (50) constructs a weighted composite reward function based on performance indicators to dynamically adjust the reward weights; and the comprehensive reward value evaluated in real time is fed back to the collaborative reinforcement learning module (30) for policy gradient update;

[0012] S5: If the performance of the planned path does not meet the preset evaluation index threshold, repeat steps S3 to S5 to continue strategy optimization; if the set path performance target is met, execute and deploy the joint action sequence output by the collaborative reinforcement learning module (30) to the robotic arm control system to complete path tracking and obstacle avoidance action execution.

[0013] The multi-agent information module also includes attention mechanism fusion, which is used for filtering key information of the multi-agent modeling module.

[0014] The multi-agent collaborative policy optimization employs a multi-agent proximal policy optimization algorithm, introducing a shearing probability ratio and a global joint advantage function to constrain the policy gradients of each agent. This improves the overall path planning accuracy while ensuring training stability. The objective loss function is defined as follows: in The probability ratio of the strategies. The dominant function is given by the formula: β is the regularization term, β is the weight of the exploration term, and N is the number of agents.

[0015] The training phase introduces a centralized training and distributed execution architecture, where during the centralized training phase, each agent can access the global state s. global The strategy and value function are jointly optimized; during the distributed execution phase, the agent relies solely on local observations. Independently generated actions This ensures both training generalization ability and real-time execution. In this way, centralized training and distributed execution complement each other; the former provides an optimization strategy framework for the latter, while the latter ensures efficient execution in complex and dynamic environments.

[0016] The multi-agent modeling module uses a graph neural network to construct an agent topology graph G = (V, E), where nodes V = {v...} i} represents the intelligent agents of each joint of the robotic arm, and edge E = {(v i v j )} represents the physical or control coupling relationship between adjacent joints. Cross-agent feature aggregation is achieved through convolution mechanism. The node embedding update process is as follows: in, For the state embedding of the l-th layer agent, α ij For attention weights, W (l) σ represents the trainable parameters, and σ is the activation function.

[0017] The graph neural network introduces an attention weight allocation mechanism in the feature update stage, which takes the following form: Here, α and W are learning parameters used to dynamically model the influence of neighboring nodes on the choice of center point strategy.

[0018] The performance indicators of the multidimensional reward function evaluation module include, but are not limited to, the following reward items: (1) minimum path length item, which encourages the robotic arm to reach the target with the shortest distance; (2) collision penalty item, which provides a negative incentive if the end or any link comes into contact with an obstacle; (3) joint smoothness item, which ensures the continuity of the trajectory by limiting the change in joint angular velocity at adjacent moments; (4) energy consumption optimization item, which takes minimizing the power consumption of the entire execution as a secondary optimization objective.

[0019] The multidimensional reward function constructs a comprehensive reward signal R through normalization and weighted fusion. total The specific form is R total =α1·R length (t)+α2·R collision (t)+α3·R smooth (t)+α4·R energy (t), where R is a t), length (t) represents the path length reward, encouraging shorter paths; R collision (t) is the collision penalty term, which takes a negative value when a collision occurs; R smooth (t) represents the trajectory smoothness reward, measuring the change in angular velocity between adjacent joints; R energy (t) represents the energy consumption optimization term, which encourages low-power movement; α1, α2, α3, α4∈(0,1) are the weight coefficients of each reward, which satisfy the normalization condition: α1+α2+α3+α4=1.

[0020] The attention mechanism module includes a joint structure of channel attention mechanism and spatial attention mechanism, which are used to realize multi-scale feature fusion and importance modeling of information between agents.

[0021] The path execution module interpolates and dynamically warps the pose of the robotic arm end effector based on the final output path trajectory sequence, and generates a desired joint angle sequence to input into the controller for motion control. During execution, model predictive control can be combined to further correct motion errors, ensuring path tracking accuracy and obstacle avoidance robustness.

[0022] In summary, this invention proposes an intelligent system and method for a high-degree-of-freedom robotic arm to achieve autonomous obstacle avoidance path planning in complex environments. The system structure includes: an environment perception module, a multi-agent modeling module, an attention mechanism fusion module, a collaborative reinforcement learning module, a multi-dimensional reward function evaluation module, and a path execution module, forming a closed-loop architecture from state modeling to policy optimization to action execution. The system, through the construction of a structured multi-agent modeling module, introduces graph neural networks (GNNs) for the first time to accurately model the topological structure and coupling relationships between the joints of the robotic arm, achieving a refined expression of inter-joint dependencies and a structured expression of the modeling structure. Through an attention mechanism fusion module, it achieves dynamic weighting and saliency filtering of multi-source information (such as local states, adjacency information, and environmental features), improving the robustness of key state recognition and high-dimensional feature representation. By employing a collaborative reinforcement learning module, it achieves global state sharing and synchronous policy optimization during the centralized training phase, while ensuring autonomous decision-making and real-time response of the agents based on local information during the distributed execution phase, improving convergence and collaboration in multi-agent tasks. By constructing a multi-dimensional reward function evaluation module, it guides path optimization to balance safety, smoothness, and energy consumption, achieving a multi-objective collaborative trade-off between task requirements and execution efficiency. Finally, the path execution mechanism enables policy deployment and dynamic correction. This invention forms a complete closed-loop path planning system from model structure, learning mechanism, policy optimization, performance evaluation to control execution. The overall architecture improves the accuracy of path planning while significantly enhancing the stability, flexibility and generalization ability of the algorithm in real-time dynamic environments. It solves the key problems of existing technologies such as rough modeling, low training efficiency and unstable execution, and has good practical value and prospects for promotion. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Fig. 1 This diagram illustrates the workflow of a path planning system provided in an embodiment of the present invention.

[0025] Fig. 2 This diagram illustrates the structural module diagram of a path planning system provided in an embodiment of the present invention.

[0026] Fig. 3 This diagram illustrates a multidimensional reward function fusion method provided in an embodiment of the present invention.

[0027] Explanation of reference numerals in the attached figures:

[0028] 10. Environmental perception module; 20. Multi-agent modeling module; 30. Collaborative reinforcement learning module; 40. Attention mechanism fusion module; 50. Multi-dimensional reward function evaluation module; 60. Path execution module. Detailed Implementation

[0029] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive step are within the scope of protection of the present invention.

[0030] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the invention. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0031] See Figs. 1-3 This invention provides an intelligent system and method for a high-degree-of-freedom robotic arm to achieve autonomous obstacle avoidance path planning in complex environments. The system structure includes an environment perception module (10), a multi-agent information module (20), an attention mechanism fusion module (40), a collaborative reinforcement learning module (30), a multi-dimensional reward function evaluation module (50), and a path execution module (60). The main steps of the system and method are as follows:

[0032] S1: The environmental perception module (10) perceives the environmental information in the workspace of the robotic arm in real time and encodes the environmental information into a high-dimensional state vector;

[0033] S2: The multi-agent modeling module (21) is a topological structure with associated information. The high-dimensional state vector is input into the multi-agent modeling module (21) to update the pose information of each agent. Based on the local state of the multi-agent and the interaction features of neighboring agents after modeling, the channel attention mechanism and spatial attention mechanism of the attention fusion module (40) are used to perform important screening of the joint physical feature group and the environmental feature group to generate joint features and provide more accurate input for decision-making.

[0034] S3: After the information is updated, the agents jointly optimize the policy network of the multi-agents through the strategy of combining centralized training and distributed execution in the collaborative reinforcement learning module (30); during the centralized training phase, each agent can access the global state information of the environment and the comprehensive reward signal R provided by the multi-dimensional reward function evaluation module (50). total By sharing experience and updating their respective policy networks synchronously, the multi-agents independently generate joint action decisions based on their own policies during the distributed execution phase, thereby constructing their respective local path segments in parallel. The execution results are evaluated by the reward function and fed back to the centralized training phase as input for policy updates, forming a closed-loop policy optimization.

[0035] S4: The multidimensional reward function module (50) constructs a weighted composite reward function based on performance indicators to dynamically adjust the reward weights; and the comprehensive reward value evaluated in real time is fed back to the collaborative reinforcement learning module (30) for policy gradient update;

[0036] S5: If the performance of the planned path does not meet the preset evaluation index threshold, repeat steps S3 to S5 to continue strategy optimization; if the set path performance target is met, execute and deploy the joint action sequence output by the collaborative reinforcement learning module (30) to the robotic arm control system to complete path tracking and obstacle avoidance action execution.

[0037] In this invention, the following method is specifically adopted:

[0038] First, the environmental perception module (10) acquires the state information of the workspace in real time, including the position of the target point, the distribution of obstacles, and the pose state of each joint, and encodes it into a high-dimensional state vector in a unified format to represent the overall state of the current system. This high-dimensional state vector is output to the multi-agent information module (20). In the multi-agent information module (20), a topological graph of the robotic arm structure is constructed based on a graph neural network. Each joint of the robotic arm is modeled as an independent agent node, and adjacency connections are established through physical or control coupling relationships to form a structured graph network. On this basis, each agent can acquire its own state and the states of its neighboring agents. By fusing graph embedding and attention mechanisms, collaborative perception capabilities are constructed to realize joint decision input feature representation. By introducing channel attention and spatial attention through the attention mechanism fusion module (40), the dynamic filtering capability of local key joint features and environmental information is improved, enhancing the expressiveness of the strategy and the adaptability of the task.

[0039] Secondly, in the collaborative reinforcement learning module (30), a framework combining centralized training and distributed execution is adopted for policy optimization. During the training phase, each agent shares global information and jointly optimizes its own policy function and global value function through a multi-agent proximal policy optimization algorithm. Multiple agents share global experience, improving the overall coordination and collaboration efficiency of the system. During the distributed execution phase, each agent independently generates actions based on local observations, forming a coordinated path planning strategy. The system evaluates the execution effect in real time through the multi-dimensional reward function evaluation module (50) and feeds back the execution behavior and results to the centralized training to achieve closed-loop optimization of the strategy. In the multi-dimensional reward function evaluation module (50), a reward function that includes at least path length, obstacle avoidance safety, joint smoothness, energy consumption and time is designed. The sub-rewards are weighted and fused through a learnable dynamic weight mechanism to construct the final comprehensive reward signal and feed it back to the collaborative reinforcement learning module (30) to drive the gradient update of the policy network of each agent.

[0040] Finally, when the path planning performance fails to meet the set target, the system automatically executes the training optimization iteration process; when the planning result meets the requirements of path quality and obstacle avoidance performance, the path execution module (60) outputs the path trajectory and generates the desired control command, which is then deployed to the robotic arm controller for actual action execution through trajectory interpolation and dynamic time warping to ensure path tracking accuracy and safety.

[0041] It should be noted that the object involved in this invention is intelligent path planning for multi-degree-of-freedom robotic arms in complex environments, and is preferably applicable to robotic arms with six degrees of freedom or more, especially in high-precision obstacle avoidance in real-time dynamic obstacles, high-dimensional spaces and irregular task areas, as well as obstacle avoidance applications in 3D printing.

[0042] It should be noted that the encoding of environmental information into a high-dimensional state vector first determines the dimension of the environmental information. Secondly, in order to uniformly encode the environmental information into a high-dimensional vector and avoid the influence between different orders of magnitude, the data needs to be standardized. For example, the target point position is normalized so that its value range is between [0,1]. Finally, the processed environmental information is concatenated into a high-dimensional vector.

[0043] It should be noted that the high-dimensional state vector is formed by combining all features in sequence. For example, the state vector of a six-degree-of-freedom robotic arm is a 12-dimensional vector (the angle and velocity of each joint). The 3-dimensional position of the target point and the obstacle information (n-dimensional) are added together to form the final high-dimensional state vector. The combination of these features can comprehensively describe the environmental state and provide a basis for reinforcement learning algorithms to make decisions.

[0044] It should be noted that each joint of the robotic arm corresponds to a single agent, and the multiple agents are composed of the single agents. The number of agents depends on the number of joints. The number of agents is consistent with the number of joints of the robotic arm. The multiple agents are composed of all agents. For example, the six joints of a six-degree-of-freedom robotic arm correspond to six single agents, and the six single agents are combined to form a multiple agent.

[0045] It should be noted that the environmental information includes at least the target point location, obstacle distribution, and the pose state of each joint;

[0046] It should be noted that the multi-agent modeling module constructs the topology of the robotic arm based on graph neural networks, models each joint as an independent agent node, and constructs a structured graph neural network model to characterize the agent dependencies by connecting physically adjacent or functionally coupled agents through the edge connectors in the graph; each agent, based on accepting its own and neighboring states, realizes a joint decision-making strategy through graph embedding in the attention fusion mechanism (40).

[0047] It should be noted that the associated information includes, but is not limited to, spatial association (the relative position of the agent in space), topological association (the structural connection relationship between agents), and communication and information sharing (communication coupling state, task collaboration order).

[0048] Optionally, the robotic arm may include collaborative robots, biomedical robotic arms, industrial robotic arms, etc., and its structural feature is that there are complex physical and control coupling relationships between the joints, which is suitable for multi-agent modeling and collaborative strategy learning.

[0049] Preferably, the path planning system includes at least the following functional modules: an environment perception module (10), a multi-agent information module (20), a collaborative reinforcement learning module (30), an attention mechanism fusion module (40), a multi-dimensional reward function evaluation module (50), and a path execution module (60).

[0050] It should be noted that the environmental perception module (10) perceives the target point position, obstacle distribution, and pose state of each joint through devices such as visual sensors, lidar, and inertial measurement units, and encodes them uniformly into a high-dimensional state vector S. t This provides a standard input format for subsequent modules.

[0051] Preferably, the multi-agent information module (20) includes at least a multi-agent modeling module (21), which constructs a topological graph model of the robotic arm based on a graph neural network. Each joint of the robotic arm is modeled as an independent agent node, and the nodes are connected by edges to form a structured graph network. The graph neural network can perform feature propagation and joint encoding between nodes, realizing cross-agent state aggregation and policy sharing. Its state representation includes the current angle, angular velocity, force information, etc. of the joint.

[0052] Optionally, the graph neural network node embedding update process is as follows: in, For the state embedding of the l-th layer agent, α ij For attention weights, W (l) σ represents the trainable parameters, and σ is the activation function.

[0053] To further enhance state modeling capabilities, the attention mechanism fusion module (40) adopts a joint structure of channel attention (SE) and spatial attention (CBAM). Preferably, channel attention extracts global statistical features through global average pooling and max pooling, and generates channel-level attention weights using a lightweight fully connected network to highlight key feature types; the spatial attention mechanism generates a spatial attention map based on the maximum and average responses of the feature map in the spatial dimension to dynamically adjust the importance of neighboring node features. This joint mechanism effectively improves the discriminativeness and contextual relevance of multi-source feature representations, thereby enhancing the perception and modeling capabilities of multi-agent systems in dynamic environments.

[0054] It should be noted that the joint features include weighted joint features and weighted environmental features, wherein the joint features include angles, velocities, torques, etc., and the environmental features include obstacle positions, target point positions, etc.

[0055] It should be noted that the graph neural network model captures the topological structure and control coupling between each joint, enabling the system to better optimize the global path, rather than treating each joint as an independent entity and ignoring the interaction and coupling between joints in the traditional way. The attention mechanism is introduced into the graph neural network to avoid assigning the same weight to all adjacent nodes in the traditional method, even though the coupling strength between different joints may be different. The attention mechanism can dynamically adjust this influence, thereby improving the flexibility and accuracy of path planning.

[0056] It should be noted that the pose information of each agent is updated by adjusting the local state of each agent, i.e., joint pose, through the high-dimensional state vector of S1; the state can be the angle, velocity, position, and other information of each joint.

[0057] It should be noted that the reinforcement learning module (30) adopts a strategy combining centralized training and distributed execution to improve the policy convergence and generalization capabilities of the multi-agent system. During the centralized training phase, agents share global state information and comprehensive reward signals to jointly optimize the policy network, ensuring that the system achieves cooperative optimality globally. In the distributed execution phase, each agent independently selects actions based on its local state, avoiding the complex calculations caused by global information synchronization, improving computational efficiency, and reducing global optimization delays. The execution results are fed back to centralized training through a reward evaluation mechanism, achieving closed-loop policy optimization. The system can continuously learn and dynamically adapt to environmental changes in real time, thereby enhancing the policy's generalization capability and actual execution performance. Global information sharing and policy synchronization updates accelerate the learning process.

[0058] It should be noted that the dynamic adjustment of reward weights based on the comprehensive reward function enables the system to adjust its optimization priorities at different stages according to task requirements. For example, in the early stages of path planning, the system may focus more on path exploration, while in later stages, it may focus more on optimizing path length and obstacle avoidance capabilities. Dynamically adjusting reward weights ensures that the system can balance different objectives at different stages, thereby achieving global optimization.

[0059] In this invention, during the centralized training phase, each joint agent can access the global state information S. t This includes pose information for all joints, information about environmental obstacles, and the location of the target point. During training, each agent i has an independent policy network π. i (a i |h i This is used to output the motion probability distribution of the joint, while sharing a global value function. Used to evaluate the overall value of the current system state. While ensuring the stability of policy updates, it improves the accuracy and efficiency of path planning. During training, each agent's objective loss function is: in, The probability ratio of the strategies. The dominant function is given by the formula: R total It is the comprehensive reward value calculated by the multidimensional reward function module (50). Here, β is the regularization term, N is the number of agents, and ∈ is the shearing threshold for policy updates (usually 0.2). During the distributed execution phase, each agent is based on its local state. Independent selection of actions No longer relying on global state, this enables parallel and real-time execution of strategies and obstacle avoidance paths. Through a collaborative architecture of centralized training and distributed execution, the module ensures that path planning behavior in multi-degree-of-freedom mechanical structures possesses both global optimality and real-time responsiveness, significantly improving the planning system's adaptability in unknown environments.

[0060] It should be noted that the objective loss function employs a policy probability ratio and a multi-dimensional reward function to prevent excessive policy updates and to ensure long-term execution stability and efficiency, thus guaranteeing stability during training. By introducing a dominance function, the system not only evaluates the quality of individual actions but also optimizes the decision-making process through comparison with other actions, ensuring the agent takes the best action. By dynamically adjusting the weights of exploration items, the system can explore extensively in the early stages of training and focus on local optimization in later stages, ensuring the optimization effect of the system at different stages. All objective loss functions simultaneously guarantee the stability of the training process and the accuracy of path planning, ensuring the system's efficient and intelligent performance in complex environments.

[0061] It should be noted that the attention mechanism fusion module (40) adopts a combined mechanism of channel attention and spatial attention.

[0062] It should be noted that the attention weight allocation mechanism is that the attention mechanism fusion module assigns different weights to different input features (such as the angle, velocity, acceleration, etc. of each joint). By learning which features are more important to the current task, more attention is paid to these features, so that they have a greater impact on the final decision.

[0063] In this invention, a module is used to enhance the information expression capabilities among multiple agents and improve the sensitivity of local states to key environmental features and the ability to select actions. This module integrates channel attention and spatial attention mechanisms, and dynamically adjusts the fusion ratio of the two at different task stages through a gating mechanism to adapt to the varying needs of feature extraction under different planning scenarios.

[0064] It should be noted that the multidimensional reward function evaluation module (50) is used to construct a path quality evaluation system under multiple indicators, and to generate a comprehensive reward signal R by fusing the multidimensional indicators through a dynamic weighting mechanism. total This value is fed back to the collaborative reinforcement learning module (30) for gradient updates of the policy network.

[0065] The multidimensional reward function R designed in this invention total At least, but not limited to, the following four indicators: 1. Minimum path length reward in Indicates the current end-effector pose, x goal1. Target pose; 2. Collision penalty item If at any given moment the link or end effector of the robotic arm comes into contact with an obstacle, a fixed reward is awarded, where λ c , is the collision penalty coefficient; 3. Joint smoothness bonus item R smooth (t) By evaluating the degree of change in joint angular velocity at consecutive moments, continuous and abrupt trajectories are encouraged to improve path controllability and mechanical safety; 4. Energy consumption optimization items Where τ i Let α1, α2, α3, α4 ∈ (0, 1) represent the driving torque of the i-th joint, and α1, α2, α3, α4 ∈ (0, 1) be the weight coefficients of each reward, satisfying the normalization condition: α1 + α2 + α3 + α4 = 1. α1 to α4 can be adjusted according to different priorities during the training phase to achieve adaptive balance of the task objectives. The above provides accurate and adjustable numerical feedback for multi-agent policy optimization and avoids the policy bias caused by traditional single reward functions.

[0066] Preferably, the path execution module (60) is used to receive the optimal action sequence or path trajectory output by the collaborative reinforcement learning module (30), and according to the judgment conditions, when the comprehensive reward value R fed back by the multidimensional reward function evaluation module (50) is... total Once the preset performance threshold is met, the path output and control execution phase can begin.

[0067] In this invention, the judgment condition is that within each training or evaluation cycle, according to R... total ≥R threshold , where R total For comprehensive rewards, R threshold The performance evaluation threshold is typically set empirically or derived from historical data statistics. If this condition is met, it indicates that the current path strategy has reached the minimum acceptable standard in terms of smoothness, safety, efficiency, and other aspects; otherwise, the control flow will return to S3 and continue iterative learning.

[0068] It should be noted that the set threshold can be a standard such as path length, path smoothness, or execution time. For example, the path execution time can be set to be less than 5 seconds.

[0069] The beneficial effects of this invention are as follows: 1. By constructing a graph neural network agent topology, it efficiently handles local interactions and global dependencies between joints, solving the problems of weak collaborative modeling ability and unclear state dependencies in existing methods; 2. By introducing an attention mechanism fusion module, it solves the problems of insufficient fusion of local state and environmental information and limited policy expression ability, ensuring the efficiency and accuracy of path planning; 3. By adopting a centralized training and distributed execution architecture through a collaborative reinforcement learning module, it solves the problems of difficulty in synchronizing multi-agent policy learning and poor task coordination; 4. By constructing a multi-dimensional reward function and introducing a dynamic weight mechanism, it solves the problem of imbalance in the optimization of a single goal-oriented reward function; 5. By combining a path execution module with model predictive control, it solves the problems of unstable deployment between policy output and real control and low path tracking accuracy. This invention achieves full-process optimization from structural modeling, collaborative decision-making, dynamic evaluation to stable execution in multi-degree-of-freedom robotic arm path planning, effectively improving the efficiency, accuracy, and deployability of path planning, and is suitable for high-precision obstacle avoidance tasks in complex scenarios.

[0070] Example 1

[0071] To further illustrate the feasibility of this invention, the following example uses the path planning task of the six-DOF robotic arm myCobotPro630 under complex obstacles to illustrate the specific implementation process of this invention:

[0072] This embodiment constructs a three-dimensional dynamic workspace containing various types of obstacles in the simulation platform Gazebo. The environment includes both static obstacles (such as walls) and slowly moving dynamic obstacles (such as moving workbenches), with the movement speed of the dynamic obstacles controlled within 0.1 m / s. The target point is set at a random location within the workspace reachable by the robotic arm. The obstacles exhibit non-deterministic changes over time during path planning to enhance environmental complexity and verify the robustness of the system. The simulated robotic arm model is myCobotPro630, which has 6 joints, and the three-dimensional dynamic workspace is 0.5m × 0.5m × 0.5m.

[0073] During the task initialization phase, the environment perception module (10) collects data on the current working scene by accessing multi-source sensors (including RGB-D camera, lidar and IMU module), and analyzes and encodes the current end effector position of the robotic arm, the angle of each joint, the geometry of obstacles, pose information and their dynamic trajectory. The number of obstacles is set to 5 (3 static and 2 dynamic), and finally forms a high-dimensional state vector representation in a unified format. This state vector is input to the multi-agent modeling module (21), which uses a graph neural network to model the nodes of each joint of the robotic arm. Each joint is regarded as a node, and the joints are connected by physical connection and control coupling relationship to construct a topology graph, ensuring that each joint not only considers its own state, but also receives the state information of adjacent joints, realizing the expression of state dependence and collaborative perception ability between joints. The graph neural network has 3 convolutional layers, each with 128 nodes (128 input and 128 output nodes), and uses the ReLU activation function. To enhance the understanding of the global target task by the local state, an attention mechanism fusion module (40) was added, employing channel attention mechanism and spatial attention mechanism, both of which use 64 feature channels. Joint features and adjacent environment features are filtered and weighted to help the system ensure that important information (such as angle, velocity, target point position, etc.) is transmitted first, so as to achieve efficient obstacle avoidance and path planning.

[0074] This invention employs a multi-agent reinforcement learning framework, specifically based on the PPO algorithm, and constructs a centralized training and distributed execution architecture. In the centralized training phase of the collaborative reinforcement learning module (30), all agents share the global state information of the environment and the comprehensive reward signal provided by the multi-dimensional reward function evaluation module (50), and jointly optimize their respective policy networks and global value functions to enable the system to have the overall optimal policy learning capability. The training parameters are set with a learning rate of 0.0001, a batch size of 32, 500 training steps per round for a total of 1000 rounds, a discount factor of 0.99, a maximum gradient shearing value of 0.5, and the optimizer being Adam. During this process, the reward signal is dynamically generated by the multi-dimensional reward evaluation module. In this embodiment, the evaluation indicators include path length, collision penalty, trajectory smoothing, and energy consumption, which are weighted and fused to form a comprehensive reward function. The weight coefficients can be adaptively adjusted during training to balance the performance requirements of various aspects of the task and ensure global optimality during training. For example, in the reward function, when the path length is the shortest, the reward weight is 1.0, and then normalized. Each collision results in a negative reward; for example, the reward weight is -1 when a collision occurs. Similarly, if the trajectory smoothing item makes a sharp turn in the path (angle change greater than 20°), its reward weight is reduced. The weights of each item are dynamically adjusted according to the main objective of the task (e.g., if a collision occurred last time, the main objective next time is to avoid collisions). To meet the requirements of real-time performance and system decoupling, each joint agent independently generates action strategies based on its local state information, achieving distributed decision-making for path planning. For example, each joint angle (θ1, θ2, θ3, θ4, θ5, θ6) = (30°, 45°, 90°, 135°, 180°, 225°) and joint velocity (v1, v2, v3, v4, v5, v6) = (0.1 rad / s, 0.05 rad / s, 0.1 rad / s, 0.1 rad / s, 0.05 rad / s, 0.1 rad / s) in the joint state information is used to generate an action strategy based on the current joint state information through a policy network, i.e., generating the joint angle change Δθ. i For example, the target angle adjustment of joint 1, Δθ1 = 5°, determines the angle adjustment in the next time step. Subsequently, the path execution module (60) converts the action sequence output by each joint strategy network into continuous trajectory points. Through cubic spline interpolation for smoothing, combined with model predictive control, the trajectory is dynamically corrected to ensure the accuracy and smoothness of the path. For example, the interval between trajectory points is set to 0.01s, and each interval corresponds to 10 trajectory points. The total number of interpolated trajectory points is 5000 (500 steps per round, 10 trajectory points per step). The prediction time window of model predictive control is 1s, and the duration of each control optimization is 0.01s, that is, each control optimization optimizes 100 control inputs. The final output control commands are sent to the myCobotPro630 controller through the standard control interface to drive the robotic arm to complete the target path tracking and obstacle avoidance tasks. The final output is 6 joint angles, such as θ1=42.3°, θ2=50.3°, θ3=-30.1°, θ4=135.7°, θ5=55.4°, and θ6=180.0°.

[0075] After approximately 1000 training rounds, the system performance was evaluated through field tests. After the first round of training, the average path length was 2.3m and the number of collisions was 2. The results show that under different task conditions (such as changes in obstacle density, target point switching, obstacle disturbances, etc.), the system has the following performance advantages: 1. The path success rate exceeds 98%, significantly higher than traditional RRT*-based planning algorithms; 2. The average path smoothness is improved by approximately 20%, resulting in fewer abrupt movements and lower energy consumption; 3. The system response time meets the real-time requirements of robot path control.

[0076] In summary, this embodiment verifies the feasibility, effectiveness, and engineering application value of the multi-agent reinforcement learning-based path planning system and method proposed in this invention, which enables autonomous path planning of a high-degree-of-freedom robotic arm in complex environments. It has significant practical significance and industrialization potential.

[0077] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0078] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0079] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.

Claims

1. A path planning system and method, characterized in that, To achieve intelligent obstacle avoidance for a robotic arm, the system includes at least an environment perception module (10), a multi-agent information module (20), a collaborative reinforcement learning module (30), an attention mechanism fusion module (40), a multi-dimensional reward function evaluation module (50), and a path execution module (60); the multi-agent information module (20) includes at least a multi-agent modeling module (21); the method includes the following steps: S1: The environmental perception module (10) perceives the environmental information in the workspace of the robotic arm in real time and encodes the environmental information into a high-dimensional state vector; S2: The multi-agent modeling module (21) is a topological structure with associated information. The high-dimensional state vector is input into the multi-agent modeling module (21) to update the pose information of each agent. Based on the local state of the multi-agent and the interaction features of neighboring agents after modeling, the channel attention mechanism and spatial attention mechanism of the attention fusion module (40) are used to perform important screening of the joint physical feature group and the environmental feature group to generate joint features and provide more accurate input for decision-making. S3: After the information is updated, the agents use the collaborative reinforcement learning module (30) to jointly optimize the policy network of the multi-agent system using a strategy that combines centralized training and distributed execution. During the centralized training phase, each agent can access the global state information of the environment and the comprehensive reward signal R provided by the multi-dimensional reward function evaluation module (50). total By sharing experience and updating their respective policy networks synchronously, the multi-agents independently generate joint action decisions based on their own policies during the distributed execution phase, thereby constructing their respective local path segments in parallel. The execution results are evaluated by the reward function and fed back to the centralized training phase as input for policy updates, forming a closed-loop policy optimization. S4: Policy optimization uses the multi-dimensional reward function module (50) to construct a weighted composite reward function based on performance indicators to dynamically adjust the reward weights; and the comprehensive reward value evaluated in real time is fed back to the collaborative reinforcement learning module (30) for policy gradient update; S5: If the performance of the planned path does not meet the preset evaluation index threshold, return to S3 and continue to optimize the strategy; if the set path performance target is met, execute and deploy the joint action sequence output by the collaborative reinforcement learning module (30) to the robotic arm control system to complete the path tracking and obstacle avoidance action execution.

2. The path planning system and method according to claim 1, characterized in that: The multi-agent information module (20) is fused with an attention mechanism, which in turn filters key information from the multi-agent modeling module (21).

3. The path planning system and method according to claim 1, characterized in that: The multi-agent collaborative policy optimization employs a multi-agent proximal policy optimization algorithm, introducing a shearing probability ratio and a global joint advantage function to constrain the policy gradients of each agent. This improves the overall path planning accuracy while ensuring training stability. The objective loss function is defined as follows: in The probability ratio of the strategies. The dominant function is given by the formula: β is the regularization term, β is the weight of the exploration term, and N is the number of agents.

4. The path planning system and method according to claim 1, characterized in that: A centralized training and distributed execution architecture is introduced during the training phase. In the centralized training phase, each agent accesses the global state s. global The strategy and value function are jointly optimized; during the distributed execution phase, the agent relies solely on local observations. Independently generated actions This ensures both training generalization ability and real-time execution.

5. The path planning system and method according to claim 1, characterized in that: The multi-agent modeling module uses a graph neural network to construct an agent topology graph G = (V, E), where nodes V = {v...} i } represents the intelligent agents of each joint of the robotic arm, and edge E = {(v i v j )} represents the physical or control coupling relationship between adjacent joints. Cross-agent feature aggregation is achieved through convolution mechanism. The node embedding update process is as follows: in, For the state embedding of the l-th layer agent, α ij For attention weights, W (l) σ represents the trainable parameters, and σ is the activation function.

6. The path planning system and method according to claim 5, characterized in that, The graph neural network introduces an attention weight allocation mechanism in the feature update stage, which takes the following form: Where a and W are the learning parameters.

7. The path planning system and method according to claim 1, characterized in that: The performance indicators of the multidimensional reward function evaluation module include, but are not limited to, the following reward items: (1) minimum path length item, which encourages the robotic arm to reach the target with the shortest distance; (2) collision penalty item, which provides a negative incentive if the end or any link comes into contact with an obstacle; (3) joint smoothness item, which ensures the continuity of the trajectory by limiting the change in joint angular velocity at adjacent moments; (4) energy consumption optimization item, which takes minimizing the power consumption of the entire execution as a secondary optimization objective.

8. The path planning system and method according to claim 7, characterized in that: The multidimensional reward function constructs a comprehensive reward signal R through normalization and weighted fusion. total The specific form is R total =α1·R length (t)+α2·R collision (t)+α3·R smooth (t)+α4·R energy (t), where R is a t), length (t) represents the path length reward, encouraging shorter paths; R collision (t) is the collision penalty term, which takes a negative value when a collision occurs; R smooth (t) represents the trajectory smoothness reward, measuring the change in angular velocity between adjacent joints; R energy (t) represents the energy consumption optimization term, which encourages low-power movement; α1, α2, α3, α4∈(0,1) are the weight coefficients of each reward, which satisfy the normalization condition: α1+α2+α3+α4=1.

9. The path planning system and method according to claim 1: characterized in that, The attention mechanism module includes a combined structure of channel attention mechanism and spatial attention mechanism.

10. The path planning system and method according to claim 1, characterized in that: The path execution module interpolates and dynamically warps the pose of the robotic arm end effector based on the final output path trajectory sequence, and generates a desired joint angle sequence to input into the controller for motion control. During execution, model predictive control can be combined to further correct motion errors, ensuring path tracking accuracy and obstacle avoidance robustness.

Citation Information

Cited By

  • Intelligent unmanned aerial vehicle flight path planning method and system

    CN121612310A

  • Intelligent unmanned aerial vehicle flight path planning method and system

    CN121612310B

  • Aerospace equipment-oriented digital-analog fusion dual-drive digital-real fusion evaluation method

    CN121637435A

  • Single robot dynamic manipulation method based on multi-agent reinforcement learning

    CN121650025A

  • Tunnel lining trolley adaptive path planning method and system based on digital twinning

    CN121809799A