Multi-machine collaborative industrial robot intelligent scheduling system and application method

Through the combination of multimodal sensors and graph neural networks, combined with collaborative reinforcement learning and improved dynamic window algorithms, the problems of insufficient environmental perception and unreasonable resource allocation in multi-robot collaborative motion control are solved, and efficient and safe collaborative motion and resource optimization are achieved.

CN119974019AActive Publication Date: 2025-05-13SHANGHAI WANTULIN ROBOT TECH CO LTD

Patent Information

Application Number
CN202510457920.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Existing multi-robot collaborative motion control faces problems such as insufficient environmental perception, lack of adaptability in motion control strategies, difficulty in taking into account synchronization accuracy and energy consumption, and unreasonable resource allocation.

Method used

Multimodal sensors are used to collect real-time motion data, feature fusion is performed through graph neural networks, and collaborative reinforcement learning model is used to generate collaborative strategy parameters, multi-objective trajectory optimization model is built, and trajectory planning is adopted using improved dynamic window algorithms, and collaborative motion control and resource optimization are realized through hierarchical motion control model and dynamic resource allocation module.

Benefits of technology

The robot quickly perceives obstacles and goals in complex environments, improves motion safety and collaboration efficiency, ensures synchronization accuracy and energy consumption optimization, avoids resource waste and deadlocks, and improves overall production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119974019A_ABST
    Figure CN119974019A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial automation control, and discloses a multi-machine collaborative industrial robot intelligent scheduling system and an application method.The system comprises a data acquisition module for acquiring real-time motion data through a multi-modal sensor; the feature fusion module is used for generating environment dynamic feature data based on graph neural network fusion data; the strategy generation module is used for generating collaborative strategy parameters by utilizing a collaborative reinforcement learning model; the trajectory optimization module is used for constructing a multi-target trajectory optimization model to plan an optimal synchronous trajectory; the motion control module is used for establishing a layered motion control model to realize cooperative motion control; a dynamic resource allocation module can also be included to optimize resource allocation. The application method sequentially executes corresponding steps according to the modules of the system. According to the method, the environment can be accurately sensed, the cooperation strategy can be optimized, the efficient track can be planned, accurate motion control and reasonable resource allocation can be realized, and the cooperative motion control capability and the production efficiency of the industrial robot in the complex environment can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial automation control technology, and in particular to a multi-machine collaborative industrial robot intelligent scheduling system and an application method. Background Art

[0002] In modern industrial production, industrial robots are increasingly used, from automobile manufacturing, electronic equipment assembly to logistics warehousing and many other fields. With the expansion of production scale and the increase in the complexity of production processes, a single industrial robot can no longer meet the needs of efficient and precise production, and multi-robot collaborative operation has become an inevitable trend. However, the current multi-robot collaborative motion control faces many challenges.

[0003] In terms of environmental perception and information fusion, traditional industrial robots often rely on a single sensor and obtain limited environmental information. For example, relying solely on joint encoders can only obtain the motion information of the robot's own joints, and cannot perceive obstacles and target objects in the surrounding environment. This can easily lead to robot collision accidents in complex production environments, reduce production efficiency, and even damage equipment. Even if some robots use multiple sensors, due to the lack of effective fusion algorithms, the data from different sensors cannot form an organic whole, and it is difficult to provide comprehensive and accurate environmental information, which makes the robot lack a reliable basis for decision-making and motion control.

[0004] From the perspective of motion control strategies, most of the existing multi-robot collaboration strategies are based on simple preset rules and lack the ability to adapt to dynamic environmental changes. In the actual production process, obstacles in the environment may appear and move at any time, or task requirements may change. Traditional strategies cannot adjust the robot's motion trajectory and collaboration mode in time, resulting in low efficiency of collaboration between robots and difficulty in ensuring synchronization accuracy. Taking the automobile assembly workshop as an example, when a temporary material transport robot enters the work area, if the robot performing assembly work cannot respond in time, it may collide with the transport robot, affecting the normal progress of the entire assembly process.

[0005] In terms of trajectory planning, existing trajectory planning algorithms usually only consider a single goal, such as the shortest path or the lowest energy consumption, and it is difficult to take into account multiple goals such as motion synchronization accuracy and joint energy consumption. Moreover, in the face of dynamically changing environments, traditional algorithms cannot quickly update trajectories, resulting in the inability to effectively guarantee the safety and stability of robot motion. For example, in the process of assembling electronic equipment, multiple robots need to operate circuit boards at the same time. It is necessary to ensure the synchronization of operations to ensure assembly accuracy, and to consider the energy consumption of robot joints to reduce operating costs. However, it is difficult for existing trajectory planning algorithms to meet these requirements at the same time.

[0006] In addition, when multiple robots work together, the resource allocation problem is also very prominent. Different robots may compete for the same resources, such as computing resources, energy resources, etc. If there is a lack of a reasonable resource allocation mechanism, it is easy to cause some robots to have insufficient resources, affecting their work efficiency, while some robots are idle, resulting in resource waste. For example, in a large logistics warehouse, multiple handling robots need to share computing resources to perform path planning and task scheduling during operation. If the resource allocation is unreasonable, some robots will wait for computing resources, reducing the operating efficiency of the entire logistics system. Summary of the invention

[0007] The purpose of the present invention is to provide a multi-machine collaborative industrial robot intelligent scheduling system and application method to solve the problems raised in the above background technology.

[0008] To achieve the above object, the present invention provides the following technical solution: a multi-machine collaborative industrial robot intelligent scheduling system, the system comprising: Data acquisition module: collects real-time motion data of industrial robots through multimodal sensors, which include joint encoders, six-dimensional force sensors, binocular vision sensors and laser ranging sensors; Feature fusion module: performs spatial topological feature fusion on the real-time motion data based on graph neural network to generate environmental dynamic feature data; Strategy generation module: input the environmental dynamic feature data into a pre-trained collaborative reinforcement learning model, the collaborative reinforcement learning model adopts a distributed strategy network structure, jointly optimizes the multi-robot collaborative strategy based on a dynamic priority reward function, and generates collaborative strategy parameters; Trajectory optimization module: construct a multi-objective trajectory optimization model according to the collaborative strategy parameters, wherein the multi-objective trajectory optimization model takes the highest motion synchronization accuracy and the lowest joint energy consumption as optimization goals, and adopts an improved dynamic window algorithm to plan the trajectory in real time, wherein the improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, and outputs the optimal synchronization trajectory data based on the multi-objective trajectory optimization model; Motion control module: A hierarchical motion control model is established according to the optimal synchronous trajectory data. The hierarchical motion control model includes a task allocation layer, a trajectory coordination layer and an execution control layer. The task allocation layer performs global task decomposition based on the collaborative strategy parameters, the trajectory coordination layer performs local trajectory interpolation based on the optimal synchronous trajectory data, and the execution control layer realizes multi-robot terminal trajectory tracking based on an adaptive sliding mode control algorithm. The hierarchical motion control model outputs synchronous control instructions to realize collaborative motion control of industrial robots.

[0009] Preferably, the spatial topological feature fusion of the real-time motion data based on the graph neural network includes: Construct a robot environment interaction graph structure, wherein the nodes of the graph structure include robot joint states, dynamic obstacle positions and target point coordinates, and edge weights are calculated by relative postures and motion trends between nodes; Design a multi-head graph attention layer, where each attention head calculates the association weights between nodes through learnable parameters and performs weighted aggregation of neighborhood node features; Three groups of graph convolutional modules are stacked using residual connection and layer normalization methods. Each group of modules contains two graph attention layers and one feature mapping layer, with output dimensions of 512, 256, and 128 respectively. The motion trajectory of dynamic obstacles is encoded based on the spatiotemporal encoder, and the temporal features are fused with the graph structure features to generate the environment dynamic feature vector.

[0010] Preferably, the collaborative reinforcement learning model adopts a distributed strategy network structure, including: Construct a multi-agent Markov decision process, define the state space as the joint angles of each robot, the terminal posture and the distribution of environmental obstacles, and define the action space as the velocity increment of each joint and the terminal posture adjustment amount; Design a dynamic priority reward function, including a synchronization error term, a collision risk term, an energy penalty term, and a trajectory smoothing term, wherein the synchronization error term is calculated by the Euclidean distance variance of the multi-robot terminal postures, the collision risk term is calculated by the gradient modulus of the obstacle distance field, and the energy penalty term is calculated based on the product integral of the joint torque and velocity; A strategy parameter sharing mechanism is adopted to construct a main strategy network and an auxiliary strategy network. The main strategy network outputs a global collaborative strategy, and the auxiliary strategy network generates an adaptive action correction based on local observations. The collaborative experience data across robots is stored in a priority experience replay pool, and the double-delayed deep deterministic policy gradient algorithm is used to alternately optimize the parameters of the main policy network and the auxiliary policy network.

[0011] Preferably, the improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, including: A dynamic constraint relaxation model is constructed, and a fuzzy logic controller is used to adjust the joint velocity limit, acceleration limit and terminal position tolerance threshold of trajectory search in real time. Design an obstacle motion prediction module to predict the future position probability distribution of dynamic obstacles based on Kalman filtering and long short-term memory network; A probabilistic safety corridor is introduced into speed space sampling to generate a dynamic feasible speed window according to the predicted obstacle distribution; A Monte Carlo tree search strategy is used to select the optimal velocity combination within a dynamic window, and continuous trajectory segments are generated by cubic spline interpolation.

[0012] Preferably, the trajectory coordination layer uses a fast exploration random tree algorithm to perform local trajectory interpolation, and the execution steps include: A dynamic sampling area is constructed on the global trajectory reference line, and the sampling radius is positively correlated with the current speed of the robot end; the global trajectory reference line is a reference path generated by the optimal synchronous trajectory data output by the trajectory optimization module; Design a bidirectional expansion strategy to expand nodes from both the forward search tree and the backward search tree, using dynamic weights to balance exploration and development; A kinematic feasibility detection module is introduced to verify the joint reachability of the sampling nodes through the pseudo-inverse solution of the Jacobian matrix; Bezier curves are used to smooth the discrete path points to ensure continuous high-order derivatives of the trajectory.

[0013] Preferably, the adaptive sliding mode control algorithm adopts a non-singular terminal sliding mode surface design, including: Construct the joint space error dynamic equation and define the non-singular terminal sliding surface as a nonlinear combination function of joint angle error and velocity error; Design an adaptive approaching law to dynamically adjust the approaching speed coefficient and switching gain according to the norm of the tracking error; A disturbance observer is used to estimate the unmodeled dynamic characteristics, and the system uncertainty is offset by a feedforward compensation term. The saturation function is introduced to replace the sign function to achieve continuous control quantity output in the neighborhood of the sliding surface.

[0014] Preferably, the dynamic resource allocation module receives the coordination strategy parameters output by the strategy generation module, and outputs a resource allocation scheme to the task allocation layer of the motion control module, specifically including: Construct a task urgency assessment model to calculate the priority weight of each subtask based on the reverse order of task deadlines and task complexity; A resource contention resolution protocol is designed, an improved banker's algorithm is used to detect the risk of resource deadlock among multiple robots, and conflict-free scheduling is achieved through a virtual resource pre-allocation mechanism.

[0015] Preferably, the present invention also includes an application method of a multi-machine collaborative industrial robot intelligent scheduling system, the method comprising the following steps: Step 1: Using a data acquisition module, collect real-time motion data of the industrial robot through a multimodal sensor, wherein the multimodal sensor includes a joint encoder, a six-dimensional force sensor, a binocular vision sensor, and a laser ranging sensor; Step 2: With the help of a feature fusion module, the real-time motion data is fused with spatial topological features based on a graph neural network to generate environmental dynamic feature data; Step 3: Input the environmental dynamic feature data into the pre-trained collaborative reinforcement learning model in the strategy generation module, wherein the collaborative reinforcement learning model adopts a distributed strategy network structure, jointly optimizes the multi-robot collaborative strategy based on a dynamic priority reward function, and generates collaborative strategy parameters; Step 4: Based on the collaborative strategy parameters, a multi-objective trajectory optimization model is constructed in the trajectory optimization module. The multi-objective trajectory optimization model takes the highest motion synchronization accuracy and the lowest joint energy consumption as optimization goals, and uses an improved dynamic window algorithm to plan the trajectory in real time. The improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, and outputs the optimal synchronization trajectory data based on the multi-objective trajectory optimization model. Step 5: Establish a hierarchical motion control model in the motion control module according to the optimal synchronous trajectory data, wherein the hierarchical motion control model includes a task allocation layer, a trajectory coordination layer and an execution control layer, wherein the task allocation layer performs global task decomposition based on the collaborative strategy parameters, the trajectory coordination layer performs local trajectory interpolation based on the optimal synchronous trajectory data, and the execution control layer realizes multi-robot terminal trajectory tracking based on an adaptive sliding mode control algorithm, and outputs synchronous control instructions through the hierarchical motion control model to realize collaborative motion control of industrial robots.

[0016] Preferably, the present invention further includes an electronic device, comprising: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to call the instructions stored in the memory to execute the functions of the above-mentioned multi-machine collaborative industrial robot intelligent scheduling system.

[0017] Compared with the prior art, the present invention has the following beneficial effects: The present invention collects real-time motion data of industrial robots through multimodal sensors, including joint encoders, six-dimensional force sensors, binocular vision sensors, and laser ranging sensors, etc., which can fully obtain the robot's own state and surrounding environment information. The spatial topological features of these data are fused using graph neural networks to generate environmental dynamic feature data. Compared with traditional single sensors or simple data fusion methods, it can more accurately and comprehensively reflect environmental changes and provide a reliable basis for subsequent decision-making. This enables the robot to quickly perceive the position and state changes of obstacles and target objects in complex environments, effectively avoid collisions, and improve movement safety. For example, in a complex workshop environment, the robot can adjust the motion path in time according to the fused environmental information, flexibly avoid dynamic obstacles, and ensure the smooth progress of production tasks.

[0018] The collaborative reinforcement learning model adopts a distributed strategy network structure and jointly optimizes the multi-robot collaboration strategy based on a dynamic priority reward function. This approach fully considers factors such as synchronization error, collision risk, energy consumption, and trajectory smoothness in the multi-robot collaboration process, and can generate more reasonable collaboration strategy parameters. Compared with the traditional collaboration strategy with preset rules, the strategy of the present invention enables robots to adaptively adjust the collaboration mode when facing complex and changeable tasks and environments, significantly improving collaboration efficiency and synchronization accuracy. In the task of assembling automotive parts, multiple robots can cooperate more accurately according to the optimized collaboration strategy, complete multiple assembly actions at the same time, reduce assembly time, and improve assembly quality.

[0019] The trajectory optimization module constructs a multi-objective trajectory optimization model with the highest motion synchronization accuracy and the lowest joint energy consumption as the optimization goals, and uses an improved dynamic window algorithm that introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism for real-time planning. This allows the robot to ensure motion synchronization between multiple robots while reducing joint energy consumption and improving energy efficiency when planning trajectories. Compared with traditional trajectory planning algorithms that only consider a single goal, the algorithm of the present invention can better adapt to complex production needs. In electronic equipment assembly production lines, robots can reduce their own energy consumption while ensuring high-precision assembly, thereby achieving energy-saving production.

[0020] The hierarchical motion control model established by the motion control module includes a task allocation layer, a trajectory coordination layer, and an execution control layer. The task allocation layer performs global task decomposition based on the collaborative strategy parameters, so that each robot can clearly understand its own tasks; the trajectory coordination layer performs local trajectory interpolation based on the optimal synchronous trajectory data to ensure the continuity and smoothness of the robot's motion trajectory; the execution control layer realizes multi-robot terminal trajectory tracking based on the adaptive sliding mode control algorithm, which has high tracking accuracy and robustness. This hierarchical structure has a clear division of labor and works in collaboration, and can accurately realize the collaborative motion control of industrial robots. In large-scale machinery manufacturing workshops, when multiple large robots are performing collaborative processing, high-precision processing operations can be achieved through the hierarchical motion control model, thereby improving product processing quality.

[0021] The dynamic resource allocation module can reasonably allocate system resources by building a task urgency assessment model, designing a resource competition resolution protocol, and developing a load balancing optimizer. The task urgency is assessed according to the deadline and complexity of the task, and resources are allocated to tasks with high urgency first; the improved banker's algorithm is used to detect resource deadlock risks, and deadlock is avoided through a virtual resource pre-allocation mechanism; based on the swarm intelligence algorithm, the computing task allocation ratio of each control node is dynamically adjusted to achieve load balancing. This effectively avoids resource waste and deadlock, and improves the overall operating efficiency of the system. In the logistics warehousing system, multiple handling robots can reasonably use computing resources and energy resources according to the scheduling of the dynamic resource allocation module, quickly complete cargo handling tasks, and improve warehousing and logistics efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a working principle diagram of the multi-machine collaborative industrial robot intelligent scheduling system of the present invention; Figure 2 A flowchart for building a collaborative reinforcement learning model; Figure 3 To improve the working principle diagram of the dynamic window algorithm; Figure 4 Workflow diagram for the dynamic resource allocation module. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] See also Figure 1-4 The present invention provides a technical solution: a multi-machine collaborative industrial robot intelligent scheduling system, the system comprising: Data acquisition module: Use multimodal sensors to collect real-time motion data of industrial robots. These sensors include joint encoders, six-dimensional force sensors, binocular vision sensors, and laser ranging sensors. Joint encoders are used to accurately measure the angle information of each joint of the robot, providing basic data for subsequent motion control; six-dimensional force sensors can sense the force and torque of the robot's end effector in real time, so that it can adjust according to force feedback during operation; binocular vision sensors and laser ranging sensors are responsible for obtaining visual and distance information of the robot's surrounding environment, which is used to detect obstacles and identify target objects.

[0025] Feature fusion module: Based on the graph neural network, the spatial topological features of the collected real-time motion data are fused to generate dynamic feature data of the environment. This process is achieved by building a specific graph structure, designing a multi-head graph attention layer, stacking graph convolution modules, and fusing spatiotemporal features. It can effectively extract key features in the environment and provide more comprehensive and accurate information for subsequent strategy generation.

[0026] Strategy generation module: Input the dynamic feature data of the environment into the pre-trained collaborative reinforcement learning model. This model adopts a distributed policy network structure, jointly optimizes the multi-robot collaborative strategy through a dynamic priority reward function, and generates collaborative strategy parameters. These parameters guide how robots collaborate in complex environments to complete specific tasks.

[0027] Trajectory optimization module: A multi-objective trajectory optimization model is constructed based on the collaborative strategy parameters. The model takes the highest motion synchronization accuracy and the lowest joint energy consumption as the optimization goals, and uses an improved dynamic window algorithm to plan the trajectory in real time. The improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, and ultimately outputs the optimal synchronization trajectory data to ensure that the robot movement is both efficient and safe.

[0028] Motion control module: A hierarchical motion control model is established based on the optimal synchronous trajectory data. The model includes a task allocation layer, a trajectory coordination layer, and an execution control layer. The task allocation layer performs global task decomposition based on the collaborative strategy parameters; the trajectory coordination layer performs local trajectory interpolation based on the optimal synchronous trajectory data; the execution control layer implements multi-robot terminal trajectory tracking based on the adaptive sliding mode control algorithm and outputs synchronous control instructions, thereby realizing the collaborative motion control of industrial robots.

[0029] The present invention will be further described below in conjunction with Examples 1 to 5: Embodiment 1: This embodiment mainly describes the specific implementation method of spatial topological feature fusion of real-time motion data based on graph neural network. Its role is to effectively integrate the robot's motion data and environmental information, generate more representative environmental dynamic feature data, and provide a more accurate basis for subsequent strategy generation.

[0030] When constructing the robot environment interaction graph structure, the robot joint state, dynamic obstacle position and target point coordinates are used as nodes of the graph structure. The robot joint state contains information such as the angle and angular velocity of each joint, which directly reflects the current motion posture of the robot; the dynamic obstacle position is obtained through binocular vision sensors and laser ranging sensors, and accurate position information helps the robot avoid obstacles in time; the target point coordinates clarify the position that the robot needs to reach. The calculation of edge weights is based on the relative posture and motion trend between nodes. For example, if the relative posture between two robot joint state nodes changes little and the motion trend is similar, then the edge weight between them is large, which means that the two nodes are highly associated.

[0031] Design a multi-head graph attention layer, each attention head has learnable parameters. These parameters are used to calculate the association weights between nodes, and weighted aggregation of neighboring node features is achieved. Assuming there is a node A and its neighboring nodes B and C, the learnable parameters of the attention head will calculate the association weights between them based on the features of nodes A, B, and C. , When weighted aggregation is performed on the features of neighboring nodes B and C, the new feature of node A is expressed as ,in , are the original features of nodes B and C respectively. This approach allows the model to focus on the relationship between nodes from different angles and extract richer feature information.

[0032] Three groups of graph convolution modules are stacked using residual connection and layer normalization methods. Each group of modules contains two graph attention layers and one feature mapping layer, with output dimensions of 512, 256 and 128 respectively. The role of residual connection is to solve the gradient vanishing problem in the deep neural network training process, so that the model can learn more effectively. Layer normalization normalizes the input of each layer to accelerate the convergence of the model. In each group of modules, the node features are further extracted and fused through two layers of graph attention layers, and then the features are mapped to specific dimensions through the feature mapping layer. For example, the first layer of graph attention layer performs weighted aggregation on the input features to obtain the intermediate feature representation, and the second layer of graph attention layer processes it again on this basis, and finally maps it to 512 dimensions through the feature mapping layer.

[0033] The motion trajectory of dynamic obstacles is encoded based on the spatiotemporal encoder. The spatiotemporal encoder combines time series and spatial structure information to better capture the motion trend of dynamic obstacles. The position information of dynamic obstacles at different time points is used as the time series input, while considering the relationship between its position in space and other nodes. Through the processing of the spatiotemporal encoder, the time series features are fused with the graph structure features to generate the environment dynamic feature vector. This vector integrates the robot motion, the state of environmental obstacles and the time factor, providing comprehensive and dynamic environmental information for subsequent strategy generation.

[0034] Embodiment 2:

[0035] When constructing a multi-agent Markov decision process, the state space is defined as the joint angles, terminal postures, and environmental obstacle distribution of each robot. The robot joint angles directly determine the robot's shape, and different joint angle combinations correspond to different motion postures; the terminal posture represents the position and posture of the robot's end effector in space, which is crucial for completing specific tasks; the distribution of environmental obstacles is obtained through sensors, allowing the robot to perceive the safety of the surrounding environment. The action space is defined as the velocity increment of each joint and the end posture adjustment. The joint velocity increment controls the change in the movement speed of the robot's joints, and the end posture adjustment is used to accurately control the position and posture of the robot's end effector to meet the task requirements.

[0036] Design a dynamic priority reward function, including synchronization error term, collision risk term, energy penalty term and trajectory smoothing term. The synchronization error term is calculated by the Euclidean distance variance of the multi-robot terminal postures, and the formula is: ,in is the number of robots, For the The end position of the robot, is the average value of all robot terminal postures. This error term is used to measure the degree of synchronization of multiple robot terminal postures. The smaller the variance, the better the synchronization and the higher the reward. The collision risk term is calculated by the gradient modulus of the obstacle distance field. When the robot approaches the obstacle, the gradient modulus of the distance field increases, the collision risk increases, and the reward decreases, prompting the robot to stay away from the obstacle. The energy consumption penalty term is calculated based on the product integral of the joint torque and velocity, that is, ,in is the number of joints, For the The torque of the joint, For the The speed of each joint, and This item encourages the robot to minimize energy consumption when completing tasks. The trajectory smoothing item is used to ensure the smoothness of the robot's motion trajectory, avoid sudden turns or acceleration, and improve the stability and efficiency of the motion.

[0037] The main strategy network and the auxiliary strategy network are constructed by adopting the strategy parameter sharing mechanism. The main strategy network outputs the global collaboration strategy, which comprehensively considers the state information and environmental information of all robots and provides macro guidance for multi-robot collaboration. The auxiliary strategy network generates adaptive action corrections based on local observations. When the robot encounters local environmental changes or emergencies during the execution of the task, the auxiliary strategy network can adjust the action in time according to the local observations, making the robot more adaptable.

[0038] The collaborative experience data across robots is stored through the priority experience replay pool. During the multi-robot collaboration process, the actions and environmental feedback of each robot are recorded to form experience data. The priority experience replay pool sorts the data according to the importance of the experience data (such as the size of the reward value), and replays important experience data first to improve the efficiency of model learning. The dual-delayed deep deterministic policy gradient algorithm is used to alternately optimize the parameters of the main policy network and the auxiliary policy network. By delaying the update of the target network and adopting a dual network structure, the algorithm reduces the variance of the policy update, improves the stability and convergence speed of the algorithm, and enables the model to learn the optimal collaborative strategy faster.

[0039] Embodiment 3:

[0040] This embodiment introduces in detail the specific implementation of improving the dynamic window algorithm by introducing an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, which makes trajectory planning more flexible and accurate, can adapt to complex and changing environments, and improves the safety and efficiency of robot movement.

[0041] A dynamic constraint relaxation model is constructed, and a fuzzy logic controller is used to adjust the joint velocity limit, acceleration limit, and terminal posture tolerance threshold of trajectory search in real time. The fuzzy logic controller makes decisions based on the robot's current motion state and environmental information. For example, when the robot approaches the target point, in order to improve the positioning accuracy, the fuzzy logic controller will appropriately reduce the joint velocity limit and acceleration limit, and increase the terminal posture tolerance threshold, allowing the robot to adjust the posture within a smaller range. When dynamic obstacles are detected around, the fuzzy logic controller will dynamically adjust the joint velocity limit and acceleration limit according to information such as the movement speed and distance of the obstacle to ensure that the robot can avoid the obstacle in time.

[0042] Design an obstacle motion prediction module to predict the future position probability distribution of dynamic obstacles based on Kalman filtering and long short-term memory networks. Kalman filtering uses the linear system state equation to optimally estimate the system state through system input and output observation data. It can effectively handle noise interference and predict the short-term motion trend of dynamic obstacles. Long short-term memory networks are good at handling long-term dependencies in time series and can learn the motion patterns of dynamic obstacles. The prediction results of Kalman filtering are used as the input of the long short-term memory network, combined with historical motion data, to predict the position probability distribution of dynamic obstacles in multiple time steps in the future. Assume that the current moment is , predict the next moment through Kalman filtering The location is ,Will And the previous historical position data is input into the long short-term memory network to obtain the future The probability distribution of the position at each time step .

[0043] A probabilistic safety corridor is introduced in the speed space sampling to generate a dynamic feasible speed window according to the predicted distribution of obstacles. The probabilistic safety corridor is a speed range determined according to the predicted position of the obstacle and the robot's own safety distance. Sampling the speed within this range can ensure that the robot has a certain safety margin during movement. For example, when it is predicted that an obstacle is approaching the robot in a certain direction, the probabilistic safety corridor will expand in the direction away from the obstacle, and the speed sampling range in the direction close to the obstacle will be correspondingly reduced, thereby generating a dynamic feasible speed window.

[0044] The Monte Carlo tree search strategy is used to select the optimal speed combination in the dynamic window, and a continuous trajectory segment is generated by cubic spline interpolation. The Monte Carlo tree search strategy evaluates the pros and cons of each speed combination by continuously simulating the robot's motion trajectory under different speed combinations. During the simulation process, the score of each trajectory is calculated according to the reward function (such as avoiding collisions, approaching the target point, etc.), and the speed combination with the highest score is selected as the optimal speed combination. After obtaining the optimal speed combination, the discrete path points are fitted using the cubic spline interpolation method to generate continuous and smooth trajectory segments to ensure the smoothness of the robot's motion.

[0045] Embodiment 4:

[0046] A dynamic sampling area is constructed on the global trajectory reference line, and the sampling radius is positively correlated with the current speed of the robot end. When the robot end speed is fast, in order to ensure that the robot has enough reaction time to avoid obstacles or adjust the direction of movement, the sampling radius increases accordingly; when the speed is slow, the sampling radius decreases to improve the accuracy of trajectory planning.

[0047] Design a two-way expansion strategy, expand nodes from both the forward search tree and the backward search tree, and use dynamic weights to balance exploration and development. The forward search tree expands nodes from the current robot position toward the target point, and the backward search tree expands nodes from the target point toward the current robot position. During the expansion process, the weights of exploration and development are dynamically adjusted according to the growth of the search tree. When the search tree does not explore a certain area sufficiently, increase the exploration weight to encourage the search tree to expand to new areas; when the search tree has found some feasible paths, increase the development weight to optimize the path using the existing information. For example, through a dynamic weight function To adjust the weight, is the number of nodes in the forward search tree, is the number of nodes in the backward search tree.

[0048] The kinematic feasibility detection module is introduced to verify the joint accessibility of the sampling node through the pseudo-inverse solution of the Jacobian matrix. The Jacobian matrix describes the kinematic relationship between the robot joint space and the end effector space. When a sampling node is generated, the corresponding joint angle is calculated using the pseudo-inverse solution of the Jacobian matrix. If the calculated joint angle is within the joint motion range of the robot, the sampling node is kinematically feasible; otherwise, the node is discarded and resampled. Assume that the Jacobian matrix is , the velocity increment of the end effector is , calculate the joint velocity increment through pseudo-inverse solution ,in is the pseudo-inverse of the Jacobian matrix.

[0049] The Bezier curve is used to smooth the discrete path points to ensure the continuous high-order derivative of the trajectory. The Bezier curve has good geometric properties and can generate a smooth curve by controlling several control points. The discrete path points obtained by the rapid exploration random tree algorithm are used as the control points of the Bezier curve to generate a continuous and smooth trajectory. This not only ensures the stability of the robot's movement, but also reduces the impact and vibration of the robot during movement, and improves the robot's service life and working accuracy.

[0050] Embodiment 5:

[0051] This embodiment introduces in detail that the adaptive sliding mode control algorithm adopts a non-singular terminal sliding mode surface design, which can improve the accuracy and robustness of the robot terminal trajectory tracking; and the dynamic resource allocation module can optimize the allocation of system resources and improve the overall performance of the system.

[0052] For the adaptive sliding mode control algorithm, the joint space error dynamics equation is first constructed. Assume that the robot joint angle is , the expected joint angle is , the joint velocity is , the expected joint velocity is , then the joint angle error , joint velocity error According to the robot dynamics model, the joint space error dynamics equation is constructed. The non-singular terminal sliding surface is defined as a nonlinear combination function of the joint angle error and the velocity error, such as the nonlinear combination function ,in is a positive constant, is a real number that satisfies certain conditions. This non-singular terminal sliding surface design can make the system converge to the equilibrium point in a finite time and improve the speed and accuracy of trajectory tracking.

[0053] Design an adaptive reaching law to dynamically adjust the reaching speed coefficient and switching gain according to the norm of the tracking error. The function of the reaching law is to make the system state reach the sliding surface as quickly as possible and keep moving on the sliding surface. Suppose the norm of the tracking error is , the adaptive reaching law can be expressed as ,in and is a coefficient adjusted according to system performance requirements. is a sign function. When the tracking error is large, increase the approach speed coefficient and switching gain , so that the system can quickly approach the sliding surface; when the tracking error is small, reduce and , in order to reduce the system chattering.

[0054] The disturbance observer is used to estimate the unmodeled dynamic characteristics, and the system uncertainty is offset by the feedforward compensation term. In practical applications, the robot system has various unmodeled dynamic characteristics and external interferences, which will affect the accuracy of trajectory tracking. The disturbance observer estimates the unmodeled dynamic characteristics and external interference by analyzing the input and output data of the system. Then, the compensation term is added to the feedforward channel to offset the impact of these uncertainties on the system. Let the estimated disturbance be , the feedforward compensation term is , adding it to the control input can improve the robustness of the system.

[0055] The saturation function is introduced to replace the sign function to achieve continuous control output in the neighborhood of the sliding surface. The sign function will produce high-frequency chattering near the sliding surface, affecting the performance of the system. The saturation function has a continuous derivative in the neighborhood of the sliding surface and can effectively suppress chattering. When the system state is close to the sliding surface, the saturation function is used to replace the sign function to make the control variable change continuously and reduce the impact of chattering on the system.

[0056] For the dynamic resource allocation module, a task urgency evaluation model is constructed to calculate the priority weight of each subtask based on the reverse order of deadline and task complexity. The deadline is , the task complexity is , task urgency . Sort each subtask according to its urgency, and prioritize resources to tasks with high urgency to ensure that the system can complete important tasks on time.

[0057] A resource contention resolution protocol is designed, and an improved banker algorithm is used to detect the risk of resource deadlock among multiple robots. Conflict-free scheduling is achieved through a virtual resource pre-allocation mechanism. The banker algorithm is a classic deadlock detection and avoidance algorithm. The improved banker algorithm takes into account the characteristics of the robot system, such as the dynamic allocation and release of resources. The virtual resource pre-allocation mechanism attempts to pre-allocate virtual resources before actually allocating resources. If it is found that pre-allocation will lead to a deadlock risk, the allocation strategy is adjusted to effectively avoid resource deadlock among multiple robots and ensure stable operation of the system.

[0058] Develop a load balancing optimizer to dynamically adjust the computing task allocation ratio of each control node based on the swarm intelligence algorithm. Swarm intelligence algorithms simulate the behavior of biological groups in nature, such as ant colony algorithms and particle swarm optimization algorithms. Taking the particle swarm optimization algorithm as an example, the computing task allocation ratio of each control node is regarded as the position of particles, and the optimal task allocation scheme is found through information sharing and collaboration between particles. Each particle adjusts its position according to its own historical optimal position and the global optimal position of the group, that is, adjusts the computing task allocation ratio, so that the load of each control node is balanced and the overall computing efficiency of the system is improved.

[0059] In actual industrial application scenarios, such as automobile parts assembly production lines, multiple industrial robots need to collaborate to complete multiple complex assembly tasks. Using the system of the present invention, the data acquisition module obtains the motion data of each robot and the surrounding environment information in real time. The feature fusion module processes these data to generate environmental dynamic feature data, providing an accurate basis for the strategy generation module. The strategy generation module generates collaborative strategy parameters through a collaborative reinforcement learning model, and the trajectory optimization module plans the optimal synchronization trajectory data based on this. The motion control module performs precise motion control based on these data to achieve efficient collaboration of robots. At the same time, the dynamic resource allocation module reasonably allocates system resources to ensure the stable operation of the entire production line, greatly improving assembly efficiency and quality.

[0060] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0061] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-machine collaborative industrial robot intelligent scheduling system, characterized in that: include: Data acquisition module: collects real-time motion data of industrial robots through multimodal sensors, which include joint encoders, six-dimensional force sensors, binocular vision sensors and laser ranging sensors; Feature fusion module: performs spatial topological feature fusion on the real-time motion data based on graph neural network to generate environmental dynamic feature data; Strategy generation module: input the environmental dynamic feature data into a pre-trained collaborative reinforcement learning model, the collaborative reinforcement learning model adopts a distributed strategy network structure, jointly optimizes the multi-robot collaborative strategy based on a dynamic priority reward function, and generates collaborative strategy parameters; Trajectory optimization module: construct a multi-objective trajectory optimization model according to the collaborative strategy parameters, wherein the multi-objective trajectory optimization model takes the highest motion synchronization accuracy and the lowest joint energy consumption as optimization goals, and adopts an improved dynamic window algorithm to plan the trajectory in real time, wherein the improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, and outputs the optimal synchronization trajectory data based on the multi-objective trajectory optimization model; Motion control module: A hierarchical motion control model is established according to the optimal synchronous trajectory data. The hierarchical motion control model includes a task allocation layer, a trajectory coordination layer and an execution control layer. The task allocation layer performs global task decomposition based on the collaborative strategy parameters, the trajectory coordination layer performs local trajectory interpolation based on the optimal synchronous trajectory data, and the execution control layer realizes multi-robot terminal trajectory tracking based on an adaptive sliding mode control algorithm. The hierarchical motion control model outputs synchronous control instructions to realize collaborative motion control of industrial robots.

2. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: The performing spatial topological feature fusion on the real-time motion data based on the graph neural network includes: Construct a robot environment interaction graph structure, wherein the nodes of the graph structure include robot joint states, dynamic obstacle positions and target point coordinates, and edge weights are calculated by relative postures and motion trends between nodes; Design a multi-head graph attention layer, where each attention head calculates the association weights between nodes through learnable parameters and performs weighted aggregation of neighborhood node features; Three groups of graph convolutional modules are stacked using residual connection and layer normalization methods. Each group of modules contains two graph attention layers and one feature mapping layer, with output dimensions of 512, 256, and 128 respectively. The motion trajectory of dynamic obstacles is encoded based on the spatiotemporal encoder, and the temporal features are fused with the graph structure features to generate the environment dynamic feature vector.

3. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: The collaborative reinforcement learning model adopts a distributed strategy network structure, including: Construct a multi-agent Markov decision process, define the state space as the joint angles of each robot, the terminal posture and the distribution of environmental obstacles, and define the action space as the velocity increment of each joint and the terminal posture adjustment amount; Design a dynamic priority reward function, including a synchronization error term, a collision risk term, an energy penalty term, and a trajectory smoothing term, wherein the synchronization error term is calculated by the Euclidean distance variance of the multi-robot terminal postures, the collision risk term is calculated by the gradient modulus of the obstacle distance field, and the energy penalty term is calculated based on the product integral of the joint torque and velocity; A strategy parameter sharing mechanism is adopted to construct a main strategy network and an auxiliary strategy network. The main strategy network outputs a global collaborative strategy, and the auxiliary strategy network generates an adaptive action correction based on local observations. The collaborative experience data across robots is stored in a priority experience replay pool, and the double-delayed deep deterministic policy gradient algorithm is used to alternately optimize the parameters of the main policy network and the auxiliary policy network.

4. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: The improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, including: A dynamic constraint relaxation model is constructed, and a fuzzy logic controller is used to adjust the joint velocity limit, acceleration limit and terminal position tolerance threshold of trajectory search in real time. Design an obstacle motion prediction module to predict the future position probability distribution of dynamic obstacles based on Kalman filtering and long short-term memory network; A probabilistic safety corridor is introduced into speed space sampling to generate a dynamic feasible speed window according to the predicted obstacle distribution; A Monte Carlo tree search strategy is used to select the optimal velocity combination within a dynamic window, and continuous trajectory segments are generated by cubic spline interpolation.

5. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: The trajectory coordination layer uses a fast exploration random tree algorithm to perform local trajectory interpolation, and the execution steps include: A dynamic sampling area is constructed on the global trajectory reference line, and the sampling radius is positively correlated with the current speed of the robot end; the global trajectory reference line is a reference path generated by the optimal synchronous trajectory data output by the trajectory optimization module; A bidirectional expansion strategy is designed to expand nodes from both the forward search tree and the backward search tree, using dynamic weights to balance exploration and development.

6. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: The adaptive sliding mode control algorithm adopts the design of non-singular terminal sliding mode surface. The specific methods include: Construct the joint space error dynamic equation and define the non-singular terminal sliding surface as a nonlinear combination function of joint angle error and velocity error; An adaptive reaching law is designed to dynamically adjust the reaching speed coefficient and switching gain according to the norm of the tracking error.

7. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 1 is characterized in that: It also includes a dynamic resource allocation module, which receives the coordination strategy parameters output by the strategy generation module and outputs a resource allocation plan to the task allocation layer of the motion control module, specifically including: Construct a task urgency assessment model to calculate the priority weight of each subtask based on the reverse order of task deadlines and task complexity; A resource contention resolution protocol is designed, an improved banker's algorithm is used to detect the risk of resource deadlock among multiple robots, and conflict-free scheduling is achieved through a virtual resource pre-allocation mechanism.

8. The multi-machine collaborative industrial robot intelligent scheduling system according to claim 7 is characterized in that: The dynamic resource allocation module also includes: Build a load balancing optimizer to dynamically adjust the computing task allocation ratio of each control node based on the swarm intelligence algorithm.

9. An application method of a multi-machine collaborative industrial robot intelligent scheduling system, characterized in that: The following steps are involved: Step 1: Using a data acquisition module, collect real-time motion data of the industrial robot through a multimodal sensor, wherein the multimodal sensor includes a joint encoder, a six-dimensional force sensor, a binocular vision sensor, and a laser ranging sensor; Step 2: With the help of a feature fusion module, the real-time motion data is fused with spatial topological features based on a graph neural network to generate environmental dynamic feature data; Step 3: Input the environmental dynamic feature data into the pre-trained collaborative reinforcement learning model in the strategy generation module, wherein the collaborative reinforcement learning model adopts a distributed strategy network structure, jointly optimizes the multi-robot collaborative strategy based on a dynamic priority reward function, and generates collaborative strategy parameters; Step 4: Based on the collaborative strategy parameters, a multi-objective trajectory optimization model is constructed in the trajectory optimization module. The multi-objective trajectory optimization model takes the highest motion synchronization accuracy and the lowest joint energy consumption as optimization goals, and uses an improved dynamic window algorithm to plan the trajectory in real time. The improved dynamic window algorithm introduces an adaptive constraint relaxation factor and a dynamic obstacle prediction mechanism, and outputs the optimal synchronization trajectory data based on the multi-objective trajectory optimization model. Step 5: Establish a hierarchical motion control model in the motion control module according to the optimal synchronous trajectory data, wherein the hierarchical motion control model includes a task allocation layer, a trajectory coordination layer and an execution control layer, wherein the task allocation layer performs global task decomposition based on the collaborative strategy parameters, the trajectory coordination layer performs local trajectory interpolation based on the optimal synchronous trajectory data, and the execution control layer realizes multi-robot terminal trajectory tracking based on an adaptive sliding mode control algorithm, and outputs synchronous control instructions through the hierarchical motion control model to realize collaborative motion control of industrial robots.

10. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to call the instructions stored in the memory to execute the functions of the multi-machine collaborative industrial robot intelligent scheduling system as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Distributed mobile mechanical arm task layered optimization control method based on generalized coordinates

    CN109079780A

  • Industrial robot trajectory tracking control algorithm

    CN111673742A

  • Multi-robot unknown environment exploration method and system based on asymmetric topological representation

    CN118372260A

  • Autonomous robot decision-making system based on multi-modal perception fusion and method thereof

    CN119295883A

  • Desilting robot intelligent control method and system based on deep learning

    CN119392782A

Cited By

  • Adaptive robot trajectory planning method and system based on deep reinforcement learning

    CN120095834A

  • Adaptive Robot Trajectory Planning Method and System Based on Deep Reinforcement Learning

    CN120095834B

  • Heterogeneous robot control system with intelligent and multi-mode perceptual driving functions

    CN120206538A

  • AGENT equipment health assessment system and method based on knowledge graph

    CN120296527A

  • AGENT equipment health assessment system and method based on knowledge graph

    CN120296527B