An intelligent scheduling system and application method for industrial robots with multi-robot collaboration
Through the combination of multimodal sensors and graph neural networks, combined with collaborative reinforcement learning and multi-objective trajectory optimization, the problems of insufficient environmental perception and poor adaptability of dynamic changes in collaborative motion control of multi-robots are solved, and efficient and accurate collaborative motion control and resource optimization are achieved.
Patent Information
- Application Number
- CN202510457920.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing multi-robot collaborative motion control faces problems such as insufficient environmental perception, poor adaptability to dynamic environmental changes, difficulties in taking into account multiple goals in trajectory planning, and unreasonable resource allocation, resulting in low collaboration efficiency and difficult to ensure synchronization accuracy.
Multimodal sensors are used to collect real-time motion data, feature fusion is performed through graph neural networks, and a collaborative reinforcement learning model generation collaboration strategy is combined to build a multi-objective trajectory optimization model, and trajectory planning is adopted using an improved dynamic window algorithm, and collaborative motion control is realized through a hierarchical motion control model and dynamic resource allocation module.
It realizes comprehensive perception of complex environments and adaptive responses to dynamic changes, improves the efficiency and synchronization accuracy of multi-robot collaboration, reduces energy consumption of joints, optimizes resource allocation, and improves overall production efficiency.
Smart Images

Figure CN119974019B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial automation control, and particularly to an intelligent scheduling system and application method for multi-robot collaborative industrial robots. Background Art
[0002] In modern industrial production, industrial robots are increasingly widely used. Their presence can be seen in many fields, from automobile manufacturing, electronic device assembly to logistics warehousing, etc. With the expansion of production scale and the increase in production process complexity, a single industrial robot is difficult to meet the requirements of efficient and precise production, and multi-robot collaborative operation has become an inevitable trend. However, there are currently many challenges in multi-robot collaborative motion control.
[0003] In terms of environmental perception and information fusion, traditional industrial robots often rely on a single sensor, and the environmental information obtained is limited. For example, relying only on joint encoders can only obtain the motion information of the robot's own joints, and it is impossible to sense obstacles and target objects in the surrounding environment. This is likely to lead to robot collision accidents in a complex production environment, reducing production efficiency and even damaging equipment. Even if some robots use multiple sensors, due to the lack of an effective fusion algorithm, the data between different sensors cannot form an organic whole, making it difficult to provide comprehensive and accurate environmental information, and causing the robot to lack a reliable basis when making decisions and performing motion control.
[0004] From the perspective of motion control strategies, most of the existing multi-robot cooperation strategies are based on simple preset rules and lack the ability to adapt to dynamic environmental changes. During the actual production process, obstacles in the environment may appear and move at any time, or the task requirements may change. Traditional strategies cannot adjust the robot's motion trajectory and cooperation method in a timely manner, resulting in low cooperation efficiency between robots and making it difficult to ensure synchronization accuracy. Taking an automobile assembly workshop as an example, when a temporary material transportation robot enters the working area, if the robots performing assembly operations cannot react in time, they may collide with the transportation robot, affecting the normal progress of the entire assembly process.
[0005] In terms of trajectory planning, existing trajectory planning algorithms usually only consider a single goal, such as the shortest path or the lowest energy consumption, and it is difficult to take into account multiple goals such as motion synchronization accuracy and joint energy consumption. Moreover, in the face of a dynamically changing environment, traditional algorithms cannot update the trajectory quickly, resulting in the inability to effectively guarantee the safety and stability of robot motion. For example, during the assembly of electronic devices, multiple robots need to operate on the circuit board simultaneously. It is necessary to ensure the synchronization of operations to ensure assembly accuracy, and at the same time consider the energy consumption of the robot joints to reduce operating costs. However, existing trajectory planning algorithms are difficult to meet these requirements simultaneously.
[0006] In addition, during multi-robot collaborative operations, the resource allocation problem is also very prominent. Different robots may compete for the same resources, such as computing resources, energy resources, etc. If there is a lack of a reasonable resource allocation mechanism, it is easy to cause some robots to lack resources, affecting their work efficiency, while some robots have idle resources, resulting in resource waste. For example, in a large logistics warehouse, multiple handling robots need to share computing resources for path planning and task scheduling during operation. If the resource allocation is unreasonable, some robots will wait for computing resources, reducing the operating efficiency of the entire logistics system. Summary of the Invention
[0007] The purpose of the present invention is to provide an intelligent scheduling system and application method for multi-robot collaborative industrial robots to solve the problems raised in the above background technology.
[0008] To achieve the above purpose, the present invention provides the following technical solution: An intelligent scheduling system for multi-robot collaborative industrial robots, the system includes:
[0009] Data acquisition module: Collect real-time motion data of industrial robots through multi-modal sensors, and the multi-modal sensors include joint encoders, six-axis force sensors, binocular vision sensors, and laser range sensors;
[0010] Feature fusion module: Perform spatial topological feature fusion on the real-time motion data based on a graph neural network to generate environmental dynamic feature data;
[0011] Policy generation module: Input the environmental dynamic feature data into a pre-trained cooperative reinforcement learning model. The cooperative reinforcement learning model adopts a distributed policy network structure and jointly optimizes the multi-robot cooperation strategy based on a dynamic priority reward function to generate cooperative policy parameters;
[0012] Trajectory optimization module: Construct a multi-objective trajectory optimization model according to the cooperative policy parameters. The multi-objective trajectory optimization model takes the highest motion synchronization accuracy and the lowest joint energy consumption as optimization objectives, and uses an improved dynamic window algorithm to perform real-time trajectory planning. The improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, and outputs the optimal synchronous trajectory data based on the multi-objective trajectory optimization model;
[0013] Motion control module: Establish a hierarchical motion control model according to the optimal synchronization trajectory data. The hierarchical motion control model includes a task allocation layer, a trajectory coordination layer, and an execution control layer. Among them, the task allocation layer performs global task decomposition based on the collaborative strategy parameters, the trajectory coordination layer performs local trajectory interpolation based on the optimal synchronization trajectory data, and the execution control layer realizes the end trajectory tracking of multiple robots based on the adaptive sliding mode control algorithm. Synchronization control instructions are output through the hierarchical motion control model to realize the collaborative motion control of industrial robots.
[0014] Preferably, the spatial topological feature fusion of the real-time motion data based on the graph neural network includes:
[0015] Construct a robot-environment interaction graph structure. The nodes of the graph structure include robot joint states, dynamic obstacle positions, and target point coordinates. The edge weights are calculated by the relative poses and motion trends between nodes.
[0016] Design a multi-head graph attention layer. Each attention head calculates the correlation weights between nodes through learnable parameters and performs weighted aggregation on the neighborhood node features.
[0017] Stack three groups of graph convolution modules using residual connection and layer normalization methods. Each group of modules contains two layers of graph attention layers and one layer of feature mapping layer, and the output dimensions are 512, 256, and 128 respectively.
[0018] Encode the motion trajectories of dynamic obstacles based on a spatio-temporal encoder, fuse the temporal features and graph structure features, and generate an environmental dynamic feature vector.
[0019] Preferably, the collaborative reinforcement learning model adopts a distributed policy network structure, including:
[0020] Construct a multi-agent Markov decision process. Define the state space as the joint angles of each robot, the end poses, and the distribution of environmental obstacles, and the action space as the joint speed increments and end pose adjustment amounts of each joint.
[0021] Design a dynamic priority reward function, including a synchronization error term, a collision risk term, an energy consumption penalty term, and a trajectory smoothness term. Among them, the synchronization error term is calculated by the variance of the Euclidean distance of the end poses of multiple robots, the collision risk term is calculated by the gradient norm of the obstacle distance field, and the energy consumption penalty term is calculated based on the integral of the product of joint torque and speed.
[0022] Adopt a policy parameter sharing mechanism to construct a main policy network and an auxiliary policy network. The main policy network outputs a global cooperation policy, and the auxiliary policy network generates an adaptive action correction amount based on local observations.
[0023] The collaborative experience data across robots is stored in a priority experience replay pool, and the parameters of the main policy network and the auxiliary policy network are alternately optimized using the twin-delayed deep deterministic policy gradient algorithm.
[0024] Preferably, the improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, including:
[0025] Construct a dynamic constraint relaxation model, and use a fuzzy logic controller to adjust the joint speed limit, acceleration limit, and end pose tolerance threshold of trajectory search in real time;
[0026] Design an obstacle motion prediction module to predict the future position probability distribution of dynamic obstacles based on Kalman filtering and long short-term memory network;
[0027] Introduce a probability safety corridor in the velocity space sampling, and generate a dynamic feasible velocity window according to the obstacle prediction distribution;
[0028] Adopt the Monte Carlo tree search strategy to select the optimal velocity combination within the dynamic window, and generate continuous trajectory segments through cubic spline interpolation.
[0029] Preferably, the trajectory coordination layer uses the rapidly-exploring random tree algorithm for local trajectory interpolation, and the execution steps include:
[0030] Construct a dynamic sampling area on the global trajectory reference line, and the sampling radius is positively correlated with the current velocity of the robot end; the global trajectory reference line is a reference path generated from the optimal synchronous trajectory data output by the trajectory optimization module;
[0031] Design a bidirectional expansion strategy, expand nodes simultaneously from the forward search tree and the backward search tree, and adopt dynamic weight to balance exploration and exploitation;
[0032] Introduce a kinematic feasibility detection module to verify the joint reachability of the sampled nodes through the pseudo-inverse solution of the Jacobian matrix;
[0033] Use Bezier curves to smooth the discrete path points to ensure the continuous high-order derivatives of the trajectory.
[0034] Preferably, the adaptive sliding mode control algorithm uses a non-singular terminal sliding mode surface design, including:
[0035] Construct a joint space error dynamics equation, and define the non-singular terminal sliding mode surface as a non-linear combination function of joint angle error and velocity error;
[0036] Design an adaptive reaching law to dynamically adjust the reaching speed coefficient and switching gain according to the norm of the tracking error;
[0037] The unmodeled dynamic characteristics are estimated by using a disturbance observer, and the system uncertainty is cancelled by a feedforward compensation term;
[0038] A saturation function is introduced to replace the sign function to realize the output of a continuous control quantity in the neighborhood of the sliding mode surface.
[0039] Preferably, the dynamic resource allocation module receives the cooperative policy parameters output by the policy generation module and outputs a resource allocation scheme to the task allocation layer of the motion control module, specifically including:
[0040] Construct a task urgency evaluation model, and calculate the priority weights of each subtask based on the reverse order of the task deadline and the task complexity;
[0041] Design a resource competition resolution protocol, use an improved banker's algorithm to detect the resource deadlock risk among multiple robots, and realize conflict-free scheduling through a virtual resource pre-allocation mechanism.
[0042] Preferably, the present invention further includes an application method of an intelligent scheduling system for multi-robot cooperation, and the method includes the following steps:
[0043] Step 1: Use the data acquisition module to collect the real-time motion data of the industrial robot through multi-modal sensors, and the multi-modal sensors include joint encoders, six-axis force sensors, binocular vision sensors and laser range sensors;
[0044] Step 2: With the help of the feature fusion module, perform spatial topological feature fusion on the real-time motion data based on a graph neural network to generate environmental dynamic feature data;
[0045] Step 3: Input the environmental dynamic feature data into a pre-trained cooperative reinforcement learning model in the policy generation module. The cooperative reinforcement learning model adopts a distributed policy network structure, jointly optimizes the multi-robot cooperation strategy based on a dynamic priority reward function, and generates cooperative policy parameters;
[0046] Step 4: According to the cooperative policy parameters, construct a multi-objective trajectory optimization model in the trajectory optimization module. The multi-objective trajectory optimization model takes the highest motion synchronization accuracy and the lowest joint energy consumption as optimization objectives, and uses an improved dynamic window algorithm to perform real-time trajectory planning. The improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, and outputs optimal synchronous trajectory data based on the multi-objective trajectory optimization model;
[0047] Step 5: Establish a hierarchical motion control model in the motion control module according to the optimal synchronization trajectory data. The hierarchical motion control model includes a task allocation layer, a trajectory coordination layer, and an execution control layer. The task allocation layer performs global task decomposition based on the collaborative strategy parameters. The trajectory coordination layer performs local trajectory interpolation based on the optimal synchronization trajectory data. The execution control layer realizes multi-robot end-effector trajectory tracking based on the adaptive sliding mode control algorithm. Synchronization control instructions are output through the hierarchical motion control model to achieve the collaborative motion control of industrial robots.
[0048] Preferably, the present invention further includes an electronic device, including:
[0049] A processor;
[0050] A memory for storing instructions executable by the processor;
[0051] Wherein, the processor is configured to call the instructions stored in the memory to execute the functions of the multi-robot collaborative industrial robot intelligent scheduling system described above.
[0052] Compared with the prior art, the beneficial effects of the present invention are:
[0053] The present invention collects real-time motion data of industrial robots through multi-modal sensors, including joint encoders, six-axis force sensors, binocular vision sensors, and laser range sensors, etc., and can comprehensively obtain the robot's own state and surrounding environment information. The graph neural network is used to fuse the spatial topological features of these data to generate environmental dynamic feature data. Compared with traditional single sensors or simple data fusion methods, it can more accurately and comprehensively reflect environmental changes and provide a reliable basis for subsequent decision-making. This enables the robot to quickly perceive the positions and state changes of obstacles and target objects in a complex environment, effectively avoid collisions, and improve motion safety. For example, in a complex workshop environment, the robot can adjust its motion path in a timely manner according to the fused environmental information, flexibly avoid dynamic obstacles, and ensure the smooth progress of production tasks.
[0054] The cooperative reinforcement learning model adopts a distributed policy network structure and jointly optimizes the multi-robot cooperation strategy based on a dynamic priority reward function. This method fully considers factors such as synchronization error, collision risk, energy consumption, and trajectory smoothness in the multi-robot cooperation process, and can generate more reasonable collaborative strategy parameters. Compared with traditional pre-set rule-based cooperation strategies, the strategy of the present invention enables the robot to adaptively adjust the cooperation method when facing complex and changeable tasks and environments, significantly improving the cooperation efficiency and synchronization accuracy. In the automotive parts assembly task, multiple robots can cooperate more precisely according to the optimized cooperation strategy, complete multiple assembly actions simultaneously, reduce the assembly time, and improve the assembly quality.
[0055] The trajectory optimization module constructs a multi-objective trajectory optimization model with the highest motion synchronization accuracy and the lowest joint energy consumption as the optimization objectives, and uses an improved dynamic window algorithm introducing an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism for real-time planning. This enables the robot to ensure motion synchronization among multiple robots while reducing joint energy consumption and improving energy utilization efficiency when planning the trajectory. Compared with traditional trajectory planning algorithms that only consider a single objective, the algorithm of the present invention can better adapt to complex production requirements. In an electronic device assembly production line, the robot can reduce its own energy consumption while ensuring high-precision assembly, achieving energy-saving production.
[0056] The hierarchical motion control model established by the motion control module includes a task assignment layer, a trajectory coordination layer, and an execution control layer. The task assignment layer decomposes the global task based on the collaborative strategy parameters, enabling each robot to clarify its own task; the trajectory coordination layer performs local trajectory interpolation based on the optimal synchronous trajectory data to ensure the continuity and smoothness of the robot's motion trajectory; the execution control layer realizes the end trajectory tracking of multiple robots based on the adaptive sliding mode control algorithm, with high tracking accuracy and robustness. This hierarchical structure has clear division of labor and collaborative work, and can accurately achieve the collaborative motion control of industrial robots. In a large mechanical manufacturing workshop, when multiple large robots are performing collaborative processing, high-precision processing operations can be achieved through the hierarchical motion control model, improving the product processing quality.
[0057] The dynamic resource allocation module can reasonably allocate system resources by constructing a task urgency evaluation model, designing a resource competition resolution protocol, and developing a load balancing optimizer. It evaluates the task urgency according to the task deadline and complexity, and preferentially allocates resources to tasks with high urgency; it uses an improved banker's algorithm to detect the risk of resource deadlock and avoids deadlocks through a virtual resource pre-allocation mechanism; it dynamically adjusts the calculation task allocation ratio of each control node based on the swarm intelligence algorithm to achieve load balancing. This effectively avoids resource waste and deadlock phenomena and improves the overall operation efficiency of the system. In a logistics and warehousing system, multiple handling robots can reasonably use computing resources and energy resources according to the scheduling of the dynamic resource allocation module, quickly complete the goods handling task, and improve the warehousing and logistics efficiency. Description of the Drawings
[0058] Figure 1 It is the working principle diagram of the multi-robot collaborative industrial robot intelligent scheduling system described in the present invention;
[0059] Figure 2 It is the construction flow chart of the collaborative reinforcement learning model;
[0060] Figure 3 It is the working principle diagram of the improved dynamic window algorithm;
[0061] Figure 4 It is the workflow diagram of the dynamic resource allocation module. Specific implementation manners
[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0063] Please refer to Figures 1-4 , the present invention provides a technical solution: a multi-robot collaborative industrial robot intelligent scheduling system, and the system includes:
[0064] Data acquisition module: Utilize multi-modal sensors to collect the real-time motion data of industrial robots. These sensors include joint encoders, six-axis force sensors, binocular vision sensors, and laser range sensors. The joint encoder is used to accurately measure the angle information of each joint of the robot, providing basic data for subsequent motion control; the six-axis force sensor can sense the force and torque received by the end effector of the robot in real time, enabling it to adjust according to force feedback during operation; the binocular vision sensor and the laser range sensor are responsible for obtaining the visual and distance information of the environment around the robot, which are used to detect obstacles, identify target objects, etc.
[0065] Feature fusion module: Based on graph neural networks, perform spatial topological feature fusion on the collected real-time motion data to generate environmental dynamic feature data. This process is achieved by constructing a specific graph structure, designing multi-head graph attention layers, stacking graph convolution modules, and fusing spatio-temporal features, which can effectively extract the key features in the environment and provide more comprehensive and accurate information for subsequent policy generation.
[0066] Policy generation module: Input the environmental dynamic feature data into a pre-trained cooperative reinforcement learning model. This model adopts a distributed policy network structure and jointly optimizes the multi-robot cooperation policy through a dynamic priority reward function to generate cooperative policy parameters. These parameters guide how robots cooperate in a complex environment to complete specific tasks.
[0067] Trajectory optimization module: Construct a multi-objective trajectory optimization model according to the cooperative policy parameters. This model takes the highest motion synchronization accuracy and the lowest joint energy consumption as the optimization objectives, and uses an improved dynamic window algorithm to perform real-time trajectory planning. The improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, and finally outputs the optimal synchronous trajectory data to ensure that the robot motion is both efficient and safe.
[0068] Motion control module: A hierarchical motion control model is established based on the optimal synchronization trajectory data. This model includes a task allocation layer, a trajectory coordination layer, and an execution control layer. The task allocation layer performs global task decomposition based on the collaborative strategy parameters; the trajectory coordination layer performs local trajectory interpolation according to the optimal synchronization trajectory data; the execution control layer realizes the end trajectory tracking of multiple robots based on the adaptive sliding mode control algorithm and outputs synchronization control instructions, thereby realizing the collaborative motion control of industrial robots.
[0069] The present invention will be further described below in conjunction with Embodiments 1 to 5:
[0070] Embodiment 1:
[0071] This embodiment mainly elaborates on the specific implementation method of fusing spatial topological features of real-time motion data based on a graph neural network. Its function is to effectively integrate the motion data and environmental information of the robot, generate more representative environmental dynamic feature data, and provide a more accurate basis for subsequent strategy generation.
[0072] When constructing the robot-environment interaction graph structure, the robot joint states, dynamic obstacle positions, and target point coordinates are used as the nodes of the graph structure. The robot joint states include information such as the angles and angular velocities of each joint, which directly reflect the current motion posture of the robot; the dynamic obstacle positions are obtained through binocular vision sensors and laser range sensors, and the accurate position information helps the robot avoid obstacles in a timely manner; the target point coordinates clarify the position that the robot needs to reach. For the calculation of edge weights, it is based on the relative pose and motion trend between nodes. For example, if the relative pose change between two robot joint state nodes is small and their motion trends are similar, then the edge weight between them is large, indicating a high degree of association between these two nodes.
[0073] A multi-head graph attention layer is designed, and each attention head has learnable parameters. By these parameters, the association weights between nodes are calculated to achieve weighted aggregation of the features of neighboring nodes. Suppose there are node A and its neighboring nodes B and C. The learnable parameters of the attention head will calculate the association weights between them according to the features of nodes A, B, and C , . When performing weighted aggregation on the features of neighboring nodes B and C, the new feature representation of node A is , where , are the original features of nodes B and C respectively. This way enables the model to focus on the relationships between nodes from different perspectives and extract richer feature information.
[0074] Three groups of graph convolutional modules are stacked using the residual connection and layer normalization methods. Each group of modules contains two layers of graph attention layers and one layer of feature mapping layer, with output dimensions of 512, 256, and 128 respectively. The role of the residual connection is to solve the problem of gradient disappearance during the training of deep neural networks, enabling the model to learn more effectively. Layer normalization normalizes the input of each layer to accelerate the convergence of the model. In each group of modules, the node features are first further extracted and fused through two layers of graph attention layers, and then the features are mapped to a specific dimension through the feature mapping layer. For example, after the first layer of graph attention layer weights and aggregates the input features to obtain an intermediate feature representation, the second layer of graph attention layer processes it again, and finally maps it to 512 dimensions through the feature mapping layer.
[0075] The motion trajectories of dynamic obstacles are encoded based on a spatio-temporal encoder. The spatio-temporal encoder combines time series and spatial structure information, enabling it to better capture the motion trends of dynamic obstacles. The position information of dynamic obstacles at different time points is used as the time series input, while considering their positions in space and the relationships with other nodes. Through the processing of the spatio-temporal encoder, the temporal features are fused with the graph structure features to generate an environmental dynamic feature vector. This vector integrates the robot's motion, the states of environmental obstacles, and time factors, providing comprehensive and dynamic environmental information for subsequent policy generation.
[0076] Example 2:
[0077] When constructing a multi-agent Markov decision process, the state space is defined as the joint angles of each robot, the end effector poses, and the distribution of environmental obstacles. The joint angles of the robot directly determine the shape of the robot, and different combinations of joint angles correspond to different motion postures; the end effector poses represent the position and orientation of the robot's end effector in space, which is crucial for completing specific tasks; the distribution of environmental obstacles is obtained through sensors, enabling the robot to perceive the safety status of the surrounding environment. The action space is defined as the increments of joint velocities and the adjustments of end effector poses. The increments of joint velocities control the changes in the motion speeds of the robot's joints, and the adjustments of end effector poses are used to precisely control the position and orientation of the robot's end effector to meet the task requirements.
[0078] A dynamic priority reward function is designed, including a synchronization error term, a collision risk term, an energy consumption penalty term, and a trajectory smoothing term. The synchronization error term is calculated by the variance of the Euclidean distances of the end effector poses of multiple robots, and the formula is , where is the number of robots, is the end effector pose of the th robot, is the average of the end - poses of all robots. This error term is used to measure the synchronization degree of the end - poses of multiple robots. The smaller the variance, the better the synchronization and the higher the reward. The collision - risk term is calculated by the gradient magnitude of the obstacle distance field. When the robot approaches an obstacle, the gradient magnitude of the distance field increases, the collision risk increases, and the reward decreases, prompting the robot to stay away from the obstacle. The energy - consumption penalty term is calculated based on the integral of the product of joint torque and velocity, that is where is the number of joints, is the torque of the -th joint, is the velocity of the -th joint, and are time intervals. This term encourages the robot to minimize energy consumption when completing tasks. The trajectory - smoothing term is used to ensure the smoothness of the robot's motion trajectory, avoid sudden turning or acceleration, and improve the stability and efficiency of motion.
[0079] Adopt a policy - parameter sharing mechanism to construct a main policy network and an auxiliary policy network. The main policy network outputs a global cooperation policy, which comprehensively considers the state information of all robots and environmental information, providing macro - guidance for multi - robot cooperation. The auxiliary policy network generates an adaptive action correction amount based on local observations. When the robot encounters local environmental changes or emergencies during the task execution process, the auxiliary policy network can adjust the action in a timely manner according to local observations, enabling the robot to have better adaptability.
[0080] Store the cross - robot cooperation experience data through a prioritized experience replay pool. During the multi - robot cooperation process, the actions and environmental feedback of each robot are recorded to form experience data. The prioritized experience replay pool sorts the data according to the importance of the experience data (such as the magnitude of the reward value), and preferentially replays important experience data to improve the learning efficiency of the model. Adopt the twin - delayed deep deterministic policy gradient algorithm to alternately optimize the parameters of the main policy network and the auxiliary policy network. This algorithm reduces the variance of policy updates by delaying the update of the target network and adopting a dual - network structure, improves the stability and convergence speed of the algorithm, and enables the model to learn the optimal cooperation policy faster.
[0081] Example 3:
[0082] This example details the specific implementation of introducing an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism into the improved dynamic window algorithm. Its role is to make the trajectory planning more flexible and accurate, be able to adapt to complex and changeable environments, and improve the safety and efficiency of robot motion.
[0083] Build a dynamic constraint relaxation model and use a fuzzy logic controller to adjust the joint speed limit, acceleration limit, and end - pose tolerance threshold for trajectory search in real - time. The fuzzy logic controller makes decisions based on the current motion state and environmental information of the robot. For example, when the robot approaches the target point, to improve the positioning accuracy, the fuzzy logic controller will appropriately reduce the joint speed limit and acceleration limit, and at the same time increase the end - pose tolerance threshold, allowing the robot to adjust its pose within a smaller range. When dynamic obstacles are detected around, the fuzzy logic controller will dynamically adjust the joint speed limit and acceleration limit according to information such as the motion speed and distance of the obstacles to ensure that the robot can avoid the obstacles in time.
[0084] Design an obstacle motion prediction module to predict the future position probability distribution of dynamic obstacles based on Kalman filtering and long - short - term memory networks. Kalman filtering uses a linear system state equation to optimally estimate the system state through system input - output observation data. It can effectively handle noise interference and predict the short - term motion trend of dynamic obstacles. The long - short - term memory network is good at dealing with long - term dependencies in time series and can learn the motion patterns of dynamic obstacles. Take the prediction result of Kalman filtering as the input of the long - short - term memory network, and combine historical motion data to predict the position probability distribution of dynamic obstacles at multiple future time steps. Assume the current time is , and through Kalman filtering, predict the position of the next moment as . Input and the previous historical position data into the long - short - term memory network to obtain the position probability distribution at the next time steps .
[0085] Introduce a probabilistic safety corridor in the velocity space sampling and generate a dynamic feasible velocity window according to the obstacle prediction distribution. The probabilistic safety corridor is a velocity range determined according to the predicted position of the obstacle and the robot's own safety distance. Sampling velocities within this range can ensure that the robot has a certain safety margin during motion. For example, when it is predicted that an obstacle is approaching the robot in a certain direction, the probabilistic safety corridor will expand in the direction away from the obstacle and correspondingly reduce the velocity sampling range in the direction close to the obstacle, thus generating a dynamic feasible velocity window.
[0086] The Monte Carlo tree search strategy is adopted to select the optimal speed combination within the dynamic window, and a continuous trajectory segment is generated through cubic spline interpolation. The Monte Carlo tree search strategy continuously simulates the motion trajectories of the robot under different speed combinations and evaluates the advantages and disadvantages of each speed combination. During the simulation process, according to the reward function (such as avoiding collisions and approaching the target point), the score of each trajectory is calculated, and the speed combination with the highest score is selected as the optimal speed combination. After obtaining the optimal speed combination, the cubic spline interpolation method is used to fit the discrete path points to generate a continuous and smooth trajectory segment, ensuring the smoothness of the robot's motion.
[0087] Example 4:
[0088] A dynamic sampling area is constructed on the global trajectory reference line, and the sampling radius is positively correlated with the current speed of the robot's end. When the speed of the robot's end is relatively fast, in order to ensure that the robot has enough reaction time to avoid obstacles or adjust the motion direction, the sampling radius increases accordingly; when the speed is relatively slow, the sampling radius decreases to improve the accuracy of trajectory planning.
[0089] A two-way expansion strategy is designed to expand nodes simultaneously from the forward search tree and the backward search tree, and dynamic weight balancing is adopted to explore and exploit. The forward search tree expands nodes from the current robot position towards the target point direction, and the backward search tree expands nodes from the target point towards the current robot position. During the expansion process, the weights of exploration and exploitation are dynamically adjusted according to the growth situation of the search tree. When the exploration in a certain area of the search tree is insufficient, the exploration weight is increased to encourage the search tree to expand into new areas; when the search tree has found some feasible paths, the exploitation weight is increased to optimize the path using the existing information. For example, a dynamic weight function is used to adjust the weights, where is the number of nodes in the forward search tree, is the number of nodes in the backward search tree.
[0090] A kinematic feasibility detection module is introduced to verify the joint reachability of the sampled nodes through the pseudo-inverse solution of the Jacobian matrix. The Jacobian matrix describes the kinematic relationship between the robot's joint space and the end-effector space. When generating a sampled node, the corresponding joint angles are calculated using the pseudo-inverse solution of the Jacobian matrix. If the calculated joint angles are within the joint motion range of the robot, the sampled node is kinematically feasible; otherwise, the node is discarded and resampled. Assume the Jacobian matrix is , the velocity increment of the end-effector is , and the joint velocity increment is calculated through the pseudo-inverse solution, where is the pseudo-inverse of the Jacobian matrix.
[0091] Use B-spline curves to smooth discrete path points and ensure continuous high-order derivatives of the trajectory. B-spline curves have good geometric properties and can generate smooth curves by controlling several control points. Take the discrete path points obtained by the Rapidly-exploring Random Tree (RRT) algorithm as the control points of the B-spline curve to generate a continuous and smooth trajectory. This can not only ensure the smoothness of the robot's movement but also reduce the impact and vibration during the robot's movement, improving the service life and working accuracy of the robot.
[0092] Example 5:
[0093] This example details the design of a non-singular terminal sliding mode surface for the adaptive sliding mode control algorithm, which can improve the accuracy and robustness of the robot's end-effector trajectory tracking; and the dynamic resource allocation module can optimize the allocation of system resources and improve the overall performance of the system.
[0094] For the adaptive sliding mode control algorithm, first construct the joint space error dynamics equation. Let the robot joint angle be , the desired joint angle be , the joint velocity be , the desired joint velocity be , then the joint angle error , the joint velocity error . According to the robot dynamics model, construct the joint space error dynamics equation. Define the non-singular terminal sliding mode surface as a non-linear combination function of the joint angle error and the velocity error, such as the non-linear combination function , where is a positive constant, is a real number that satisfies certain conditions. This non-singular terminal sliding mode surface design enables the system to converge to the equilibrium point within a finite time, improving the speed and accuracy of trajectory tracking.
[0095] Design an adaptive reaching law to dynamically adjust the reaching speed coefficient and the switching gain according to the norm of the tracking error. The role of the reaching law is to make the system state reach the sliding mode surface as soon as possible and maintain motion on the sliding mode surface. Let the norm of the tracking error be , and the adaptive reaching law can be expressed as , where and are coefficients adjusted according to the system performance requirements, is the sign function. When the tracking error is large, increase the reaching speed coefficient and the switching gain to enable the system to quickly approach the sliding mode surface; when the tracking error is small, reduce and to reduce the chattering of the system.
[0096] The disturbance observer is used to estimate the unmodeled dynamics, and the system uncertainties are canceled by the feedforward compensation term. In practical applications, there are various unmodeled dynamics and external disturbances in the robot system, which will affect the accuracy of trajectory tracking. The disturbance observer estimates the unmodeled dynamics and external disturbances through the analysis of the system input and output data. Then, a compensation term is added to the feedforward channel to cancel the influence of these uncertainties on the system. Let the estimated disturbance be and the feedforward compensation term be . Add it to the control input to improve the robustness of the system.
[0097] The saturation function is introduced to replace the sign function to achieve continuous control output in the neighborhood of the sliding surface. The sign function will generate high-frequency chattering near the sliding surface, which affects the performance of the system. The saturation function has a continuous derivative in the neighborhood of the sliding surface and can effectively suppress chattering. When the system state approaches the sliding surface, the saturation function is used to replace the sign function to make the control quantity change continuously and reduce the influence of chattering on the system.
[0098] For the dynamic resource allocation module, a task urgency evaluation model is constructed, and the priority weights of each subtask are calculated based on the reverse order of the deadline and the task complexity. Let the subtask have a deadline of , the task complexity be , and the task urgency be . Sort each subtask according to the task urgency, and give priority to allocating resources to tasks with high urgency to ensure that the system can complete important tasks on time.
[0099] Design a resource competition resolution protocol, use the improved banker's algorithm to detect the resource deadlock risk among multiple robots, and achieve conflict-free scheduling through the virtual resource pre-allocation mechanism. The banker's algorithm is a classic deadlock detection and avoidance algorithm. The improved banker's algorithm takes into account the characteristics of the robot system, such as the dynamic allocation and release of resources. The virtual resource pre-allocation mechanism attempts to pre-allocate virtual resources before actually allocating resources. If it is found that the pre-allocation will lead to a deadlock risk, the allocation strategy is adjusted, thus effectively avoiding resource deadlocks among multiple robots and ensuring the stable operation of the system.
[0100] Develop a load balancer optimizer to dynamically adjust the computing task allocation ratio of each control node based on swarm intelligence algorithms. Swarm intelligence algorithms simulate the behaviors of biological groups in nature, such as ant colony algorithms, particle swarm optimization algorithms, etc. Taking the particle swarm optimization algorithm as an example, regard the computing task allocation ratio of each control node as the position of a particle, and through information sharing and cooperation among particles, find the optimal task allocation scheme. Each particle adjusts its position according to its own historical optimal position and the global optimal position of the group, that is, adjusts the computing task allocation ratio, so as to balance the load of each control node and improve the overall computing efficiency of the system.
[0101] In actual industrial application scenarios, such as automotive parts assembly production lines, multiple industrial robots need to cooperate to complete multiple complex assembly tasks. Using the system of the present invention, the data acquisition module can obtain the motion data of each robot and the surrounding environment information in real time. The feature fusion module processes these data to generate environmental dynamic feature data, providing an accurate basis for the policy generation module. The policy generation module generates cooperative policy parameters through a cooperative reinforcement learning model, and the trajectory optimization module plans the optimal synchronous trajectory data based on this. The motion control module performs precise motion control based on these data to achieve the efficient cooperation of the robots. At the same time, the dynamic resource allocation module reasonably allocates system resources to ensure the stable operation of the entire production line, greatly improving the assembly efficiency and quality.
[0102] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0103] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-machine collaborative industrial robot intelligent scheduling system, characterized in that: include: Data acquisition module: collects real-time motion data of industrial robots through multimodal sensors, which include joint encoders, six-dimensional force sensors, binocular vision sensors and laser ranging sensors; Feature fusion module: performs spatial topological feature fusion on the real-time motion data based on graph neural network to generate environmental dynamic feature data; Strategy generation module: input the environmental dynamic feature data into a pre-trained collaborative reinforcement learning model, the collaborative reinforcement learning model adopts a distributed strategy network structure, jointly optimizes the multi-robot collaborative strategy based on a dynamic priority reward function, and generates collaborative strategy parameters; Trajectory optimization module: construct a multi-objective trajectory optimization model according to the collaborative strategy parameters, wherein the multi-objective trajectory optimization model takes the highest motion synchronization accuracy and the lowest joint energy consumption as optimization goals, and adopts an improved dynamic window algorithm to plan the trajectory in real time, wherein the improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, and outputs the optimal synchronization trajectory data based on the multi-objective trajectory optimization model; Motion control module: A hierarchical motion control model is established according to the optimal synchronous trajectory data. The hierarchical motion control model includes a task allocation layer, a trajectory coordination layer and an execution control layer. The task allocation layer performs global task decomposition based on the collaborative strategy parameters, the trajectory coordination layer performs local trajectory interpolation based on the optimal synchronous trajectory data, and the execution control layer realizes multi-robot terminal trajectory tracking based on an adaptive sliding mode control algorithm. The hierarchical motion control model outputs synchronous control instructions to realize collaborative motion control of industrial robots.
2. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: The performing spatial topological feature fusion on the real-time motion data based on the graph neural network includes: Construct a robot environment interaction graph structure, wherein the nodes of the graph structure include robot joint states, dynamic obstacle positions and target point coordinates, and edge weights are calculated by relative postures and motion trends between nodes; Design a multi-head graph attention layer, where each attention head calculates the association weights between nodes through learnable parameters and performs weighted aggregation of neighborhood node features; Three groups of graph convolutional modules are stacked using residual connection and layer normalization methods. Each group of modules contains two graph attention layers and one feature mapping layer, with output dimensions of 512, 256, and 128 respectively. The motion trajectory of dynamic obstacles is encoded based on the spatiotemporal encoder, and the temporal features are fused with the graph structure features to generate the environment dynamic feature vector.
3. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: The collaborative reinforcement learning model adopts a distributed strategy network structure, including: Construct a multi-agent Markov decision process, define the state space as the joint angles of each robot, the terminal posture and the distribution of environmental obstacles, and define the action space as the velocity increment of each joint and the terminal posture adjustment amount; Design a dynamic priority reward function, including a synchronization error term, a collision risk term, an energy penalty term, and a trajectory smoothing term, wherein the synchronization error term is calculated by the Euclidean distance variance of the multi-robot terminal postures, the collision risk term is calculated by the gradient modulus of the obstacle distance field, and the energy penalty term is calculated based on the product integral of the joint torque and velocity; A strategy parameter sharing mechanism is adopted to construct a main strategy network and an auxiliary strategy network. The main strategy network outputs a global collaborative strategy, and the auxiliary strategy network generates an adaptive action correction based on local observations. The collaborative experience data across robots is stored in a priority experience replay pool, and the double-delayed deep deterministic policy gradient algorithm is used to alternately optimize the parameters of the main policy network and the auxiliary policy network.
4. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: The improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, including: A dynamic constraint relaxation model is constructed, and a fuzzy logic controller is used to adjust the joint velocity limit, acceleration limit and terminal position tolerance threshold of trajectory search in real time. Design an obstacle motion prediction module to predict the future position probability distribution of dynamic obstacles based on Kalman filtering and long short-term memory network; A probabilistic safety corridor is introduced into speed space sampling to generate a dynamic feasible speed window according to the predicted obstacle distribution; The Monte Carlo tree search strategy is used to select the optimal velocity combination within the dynamic window, and continuous trajectory segments are generated by cubic spline interpolation.
5. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: The trajectory coordination layer uses a fast exploration random tree algorithm to perform local trajectory interpolation, and the execution steps include: A dynamic sampling area is constructed on the global trajectory reference line, and the sampling radius is positively correlated with the current speed of the robot end; the global trajectory reference line is a reference path generated by the optimal synchronous trajectory data output by the trajectory optimization module; A bidirectional expansion strategy is designed to expand nodes from both the forward search tree and the backward search tree, using dynamic weights to balance exploration and development.
6. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: The adaptive sliding mode control algorithm adopts the design of non-singular terminal sliding mode surface. The specific methods include: Construct the joint space error dynamic equation and define the non-singular terminal sliding surface as a nonlinear combination function of joint angle error and velocity error; An adaptive reaching law is designed to dynamically adjust the reaching speed coefficient and switching gain according to the norm of the tracking error.
7. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: It also includes a dynamic resource allocation module, which receives the coordination strategy parameters output by the strategy generation module and outputs a resource allocation plan to the task allocation layer of the motion control module, specifically including: Construct a task urgency assessment model to calculate the priority weight of each subtask based on the reverse order of task deadlines and task complexity; A resource contention resolution protocol is designed, an improved banker's algorithm is used to detect the risk of resource deadlock among multiple robots, and conflict-free scheduling is achieved through a virtual resource pre-allocation mechanism.
8. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 7 is characterized in that: The dynamic resource allocation module also includes: Build a load balancing optimizer to dynamically adjust the computing task allocation ratio of each control node based on the swarm intelligence algorithm.
9. An application method of a multi-machine collaborative industrial robot intelligent scheduling system, characterized in that: The following steps are involved: Step 1: Using a data acquisition module, collect real-time motion data of the industrial robot through a multimodal sensor, wherein the multimodal sensor includes a joint encoder, a six-dimensional force sensor, a binocular vision sensor, and a laser ranging sensor; Step 2: With the help of a feature fusion module, the real-time motion data is fused with spatial topological features based on a graph neural network to generate environmental dynamic feature data; Step 3: Input the environmental dynamic feature data into the pre-trained collaborative reinforcement learning model in the strategy generation module, wherein the collaborative reinforcement learning model adopts a distributed strategy network structure, jointly optimizes the multi-robot collaborative strategy based on a dynamic priority reward function, and generates collaborative strategy parameters; Step 4: Based on the collaborative strategy parameters, a multi-objective trajectory optimization model is constructed in the trajectory optimization module. The multi-objective trajectory optimization model takes the highest motion synchronization accuracy and the lowest joint energy consumption as optimization goals, and uses an improved dynamic window algorithm to plan the trajectory in real time. The improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, and outputs the optimal synchronization trajectory data based on the multi-objective trajectory optimization model. Step 5: Establish a hierarchical motion control model in the motion control module according to the optimal synchronous trajectory data, wherein the hierarchical motion control model includes a task allocation layer, a trajectory coordination layer and an execution control layer, wherein the task allocation layer performs global task decomposition based on the collaborative strategy parameters, the trajectory coordination layer performs local trajectory interpolation based on the optimal synchronous trajectory data, and the execution control layer realizes multi-robot terminal trajectory tracking based on an adaptive sliding mode control algorithm, and outputs synchronous control instructions through the hierarchical motion control model to realize collaborative motion control of industrial robots.
10. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to call the instructions stored in the memory to execute the functions of the multi-machine collaborative industrial robot intelligent scheduling system as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Distributed mobile mechanical arm task layered optimization control method based on generalized coordinates
CN109079780A
Industrial robot trajectory tracking control algorithm
CN111673742A