Robot motion planning method, device and equipment and storage medium

Through the deep reinforcement learning network combined with multimodal information to plan the robot motion path, the problem of inaccurate and low efficiency in the existing technology is solved, and more efficient and accurate assembly task execution is achieved.

CN119952730AActive Publication Date: 2025-05-09SINOTRANS +1
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510438712.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the prior art, robot motion planning implemented by means of dynamic models or pre-set robot joint sequences, etc., has problems of insufficient accuracy and low efficiency.

Method used

The deep reinforcement learning network is used to combine multimodal information (such as RGB images, depth images, infrared images, force sensor information and haptic sensor information) for robot motion path planning, and train it in multiple parallel environments through distributed proximity strategy optimization algorithms to improve the efficiency and accuracy of path planning.

Benefits of technology

It realizes more accurate motion trajectory planning, improves the efficiency and accuracy of assembly tasks, is more adaptable, and does not require accurate robot dynamics models, reducing system maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119952730A_ABST
    Figure CN119952730A_ABST
Patent Text Reader

Abstract

The invention provides a robot motion planning method, device and equipment and a storage medium, and is applied to the technical field of robot assembling.The method comprises the steps that multi-modal information collected by a robot at the current moment, the pose corresponding to the current moment of a to-be-assembled object and environment information corresponding to the to-be-assembled object are obtained; the environment information comprises whether obstacles exist around the to-be-assembled object or not; determining the state of the robot at the current moment according to the multi-modal information at the current moment; according to the state of the robot at the current moment, the pose of the to-be-assembled object and the environment information, a deep reinforcement learning network is adopted to carry out motion path planning on the robot, and the motion track of the robot moving from the pose of the robot at the current moment to the to-be-assembled object is determined; wherein in the deep reinforcement learning network, a plurality of parallel environments are adopted to jointly plan a motion path. By adopting the technical scheme of the invention, the path planning accuracy and planning efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot assembly technology, and in particular to a robot motion planning method, device, equipment and storage medium. Background Art

[0002] Robotic assembly refers to the process of grabbing, inserting, and tightening industrial parts by robots to assemble parts and complete assembly tasks. Usually, motion planning is required before the robot performs assembly tasks, and after motion planning, the robot is controlled to perform assembly tasks according to the motion planning.

[0003] In related technologies, when planning the motion of a robot to perform assembly tasks, some technologies use dynamic models to accurately grasp the robot's kinematic and dynamic parameters, and calculate and control the robot's motion trajectory through forward and inverse kinematics to complete the assembly task. Other technologies use pre-set information such as the robot's joint motion sequence and visual feedback to achieve motion planning and thus complete the assembly task.

[0004] However, the above technology has the problem that the motion planning obtained is not accurate enough and has low efficiency. Summary of the invention

[0005] The present invention provides a robot motion planning method, device, equipment and storage medium, which are used to solve the defects in the prior art that the motion planning obtained by implementing robot motion planning based on a dynamic model or a pre-set robot joint sequence is not accurate and the efficiency is low. The method determines the state of the robot through the multimodal information collected by the robot, and combines the position and posture of the object to be assembled and the environmental information, and uses a deep reinforcement learning network to plan the motion path of the robot to obtain a more accurate planned path. The deep reinforcement learning network uses multiple parallel environments to jointly perform motion path planning, which can significantly improve the efficiency of path planning.

[0006] The present invention provides a robot motion planning method, comprising: Obtaining the multimodal information collected by the robot at the current moment, the position and posture of the object to be assembled at the current moment, and the environmental information corresponding to the object to be assembled; the above environmental information includes whether there are obstacles around the object to be assembled; Determine the state of the robot at the current moment according to the multimodal information at the current moment; the state of the robot at the current moment includes the position and posture of the robot at the current moment; According to the current state of the robot, the position and posture of the object to be assembled, and the environmental information, a deep reinforcement learning network is used to plan the robot's motion path and determine the robot's motion trajectory from its current position to the object to be assembled. Among them, the deep reinforcement learning network is trained using a distributed proximal strategy optimization algorithm, in which multiple parallel environments are used to jointly plan the motion path.

[0007] According to a robot motion planning method provided by the present invention, the motion path planning of the robot is performed using a deep reinforcement learning network according to the current state of the robot, the position and posture of the object to be assembled, and the environmental information, and the motion trajectory of the robot moving from its current position and posture to the object to be assembled is determined, including: According to the current state of the robot, the position and posture of the object to be assembled, and the environmental information, a deep reinforcement learning network is used to plan the motion path of the robot to determine the state of the robot at the next moment; the state at the next moment includes the position and posture of the robot at the next moment; According to the various assembly factors corresponding to the object to be assembled during the assembly process, determining the reward functions corresponding to the various assembly factors; According to the reward function, the state of the robot at the next moment, the posture of the object to be assembled and the environmental information, the deep reinforcement learning network is used to iteratively plan the robot's motion path and determine the motion trajectory.

[0008] According to a robot motion planning method provided by the present invention, the above-mentioned reward functions corresponding to the various assembly factors corresponding to the object to be assembled in the assembly process are determined, including: Obtain assembly requirements and various assembly factors corresponding to the object to be assembled; Determine the weight corresponding to each assembly factor according to assembly requirements; The reward function is determined according to each assembly requirement and the weight corresponding to each assembly requirement.

[0009] According to a robot motion planning method provided by the present invention, the method further includes: When the robot moves to the vicinity of the object to be assembled according to the motion trajectory, a visual image of the object to be assembled is acquired after the camera captures the image, and a target position and posture corresponding to the object to be assembled is determined according to the visual image; Get the first position corresponding to the current end effector of the robot; According to the first posture and the target posture, the first posture of the robot is adjusted to determine the second posture; the accuracy of the second posture is higher than the accuracy of the first posture; Send a control instruction to the robot; the control instruction includes a second posture, which is used to control the robot to adjust the posture of the end effector according to the second posture.

[0010] According to a robot motion planning method provided by the present invention, the method further includes: The assembly task corresponding to the object to be assembled is divided into multiple subtasks, and multiple policy networks are configured in the deep reinforcement learning network; each subtask corresponds to a policy network, and the multiple subtasks include a first subtask of moving the robot from its current position to the object to be assembled; When it is determined that the robot has completed the first subtask, the current position and posture of the robot is fed back to the deep reinforcement learning network, so that the deep reinforcement learning network optimizes the policy network corresponding to the second subtask based on the current position and posture of the robot; the second subtask is the next subtask to be performed after the first subtask is completed among the multiple subtasks; The second subtask is performed according to the policy network corresponding to the second subtask optimized by the deep reinforcement learning network.

[0011] According to a robot motion planning method provided by the present invention, the updating method of the strategy network of each subtask includes: Determine the degree of difference between the first strategy and the second strategy according to the first strategy output by the strategy network at the current time and the second strategy output at the previous time; Determine the target action currently selected according to the first strategy, and evaluate the advantage of the target action according to the advantage estimation function; The policy network is updated based on the degree of difference, the advantage of the target action, the number of parallel environments currently used, and the number of time steps to determine an updated policy network.

[0012] According to a robot motion planning method provided by the present invention, the multimodal information includes a red, green, and blue (RGB) image of an object to be assembled, a depth image of the object to be assembled, an infrared image of the object to be assembled, mechanical feature information collected by a force sensor of the robot, and tactile feature information collected by a tactile sensor of the robot. The state of the robot at the current moment is determined according to the multimodal information at the current moment, including: A convolutional neural network is used to extract features from the RGB image and the infrared image at the current moment, respectively, to determine a first feature corresponding to the RGB image and a second feature corresponding to the infrared image; the first feature is used to characterize the surface texture and color information of the object to be assembled, and the second feature is used to characterize the contour information of the object to be assembled; A point cloud processing network is used to perform feature extraction processing on the depth image to determine a third feature corresponding to the depth image; the third feature is used to characterize the distance between the robot and the object to be assembled and the three-dimensional shape of the object to be assembled; A feature fusion operation is performed on the first feature, the second feature, the third feature, the mechanical feature information, and the tactile feature information to determine the state of the robot at the current moment.

[0013] The present invention also provides a robot motion planning device, comprising the following modules: An acquisition module is used to acquire the multimodal information collected by the robot at the current moment, the position and posture of the object to be assembled at the current moment, and the environmental information corresponding to the object to be assembled; the above environmental information includes whether there are obstacles around the object to be assembled; A state determination module, used to determine the state of the robot at the current moment according to the multimodal information at the current moment; the state of the robot at the current moment includes the position and posture of the robot at the current moment; The motion planning module is used to plan the motion path of the robot using a deep reinforcement learning network according to the current state of the robot, the position and posture of the object to be assembled, and the environmental information, and to determine the motion trajectory of the robot moving from its current position to the object to be assembled; Among them, the deep reinforcement learning network is trained using a distributed proximal strategy optimization algorithm, in which multiple parallel environments are used to jointly plan the motion path.

[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the robot motion planning method described above is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the robot motion planning method as described in any one of the above is implemented.

[0016] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the robot motion planning method described above is implemented.

[0017] The robot motion planning method, device, equipment and storage medium provided by the present invention obtain multimodal information collected by the robot at the current moment, the position and posture corresponding to the object to be assembled at the current moment, and the environmental information corresponding to the object to be assembled, determine the state of the robot at the current moment according to the multimodal information at the current moment, and use a deep reinforcement learning network to plan the motion path of the robot according to the state of the robot at the current moment, the position and posture of the object to be assembled and the environmental information, so as to determine the motion trajectory of the robot moving from the position and posture at the current moment to the object to be assembled; wherein the state of the robot at the current moment includes its position and posture at the current moment, and the deep reinforcement learning network is trained by using a distributed proximal strategy optimization algorithm, and the distributed proximal strategy optimization algorithm uses multiple parallel environments to jointly plan the motion path. In this method, since the state of the robot can be determined by the multimodal information collected by the robot, and combined with the position and environmental information of the object to be assembled, a deep reinforcement learning network is used to plan the motion path of the robot. In this way, through multimodal information and combined with other more information, the robot can plan a high-precision motion trajectory in the motion planning process, improve the accuracy of motion planning and subsequent object assembly, and at the same time, this method does not require an accurate robot dynamics model, can adapt to the changes in the physical characteristics and environment of different robots, and has stronger adaptability to assembly parts of various shapes and sizes and different assembly scenes, so it can reduce the system maintenance cost; at the same time, when the position of the assembly object is offset, obstacles appear, or the lighting conditions change, the path planning and assembly tasks can still be effectively completed. In addition, since the deep reinforcement learning network in this method uses multiple parallel environments to jointly plan the motion path, the efficiency of path planning can be significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 It is one of the flow charts of the robot motion planning method provided by the present invention.

[0020] Figure 2 This is the second flow chart of the robot motion planning method provided by the present invention.

[0021] Figure 3 It is a structural schematic diagram of the robot motion planning device provided by the present invention.

[0022] Figure 4It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] At present, the motion planning of robot assembly is usually based on accurate robot dynamics models or manual feature-based methods. For methods based on dynamics models, it is necessary to accurately master the robot's kinematic and dynamic parameters, such as joint speed, acceleration, force and torque, and calculate and control the robot's motion trajectory through forward and inverse kinematics to complete the assembly task. In terms of visual servoing, some existing technologies use visual sensor feedback to adjust the robot's motion according to the deviation between the target and the current position through classical control theory (such as proportional-integral-differential (PID) control). In terms of reinforcement learning, some work has begun to explore the use of deep reinforcement learning (DRL) for robot task planning, but when it is combined with visual servoing for motion planning of assembly tasks, it is still at a relatively early stage, and most of them do not fully consider the characteristics of assembly tasks and the advantages of visual servoing.

[0025] In some traditional robot assembly systems, the robot's motion planning and control mainly rely on pre-programmed trajectories and simple perception of the environment, and its adaptability and flexibility to the environment are limited. In terms of visual servoing, traditional methods based on manual features require manually designed features, and feature extraction may not be accurate and robust enough for complex assembly parts and scenes. For example, in some automated assembly production lines, assembly is completed through pre-set robot joint motion sequences and simple visual feedback (such as the two-dimensional position deviation of the target), which makes it difficult to handle changes in part shape and posture and complex environmental interference. In the exploration based on deep reinforcement learning, some studies have only applied it to simple robot operation tasks such as grasping and placing, and have not deeply involved complex assembly tasks. In addition, it is not perfect in state representation, action space design, reward function design, and integration with visual servoing. For example, the use of simple state representation may not fully reflect the complexity of the assembly task, and the reward function may not fully consider multiple factors such as assembly accuracy, time efficiency, and collision risk, resulting in poor learning results.

[0026] The above-mentioned method based on precise dynamic model requires precise modeling of the robot. Once the physical characteristics of the robot change (such as wear, replacement of parts) or the environment changes, the model needs to be readjusted. The system maintenance cost is high and the adaptability is poor. The traditional visual servo method is based on manual features. For complex assembly scenes, the extraction and design of manual features are difficult and easily affected by factors such as illumination and occlusion, resulting in insufficient motion planning accuracy and robustness. Simple deep reinforcement learning is applied in robot assembly. The state and action space design may not be reasonable enough, and high-dimensional visual information cannot be effectively processed. The reward function is not comprehensive, the learning efficiency is low, and it is difficult to achieve accurate and efficient assembly motion planning, especially in multi-component assembly, high-precision assembly requirements and complex environments. Performance is limited. In short, the motion planning methods in the above-mentioned technologies face many challenges when facing complex assembly parts, diverse environmental conditions and high-precision assembly requirements. For example, in the case of multi-component assembly, variable shapes and sizes of parts, and the presence of environmental interference and obstacles, the above-mentioned motion planning methods based on precise dynamic models or manual features are difficult to guarantee accuracy and adaptability, and the motion planning efficiency is also low. Based on this, an embodiment of the present invention provides a robot motion planning method, device, equipment and storage medium, which can solve the above technical problems.

[0027] Combine the following Figure 1-Figure 2 The robot motion planning method of the present invention is described.

[0028] It should be noted that the execution subject of the embodiment of the present invention may be a robot motion planning device, or an electronic device including the robot motion planning, or other devices or equipment. The electronic device here may be a terminal or a server. In the case of a terminal, the electronic device may be a robot or an electronic device in a robot. The following embodiments are described by taking the electronic device as the execution subject as an example.

[0029] Figure 1 is one of the flow charts of the robot motion planning method provided by the present invention, such as Figure 1 As shown, the method comprises the following steps: Step 102, obtaining the multimodal information collected by the robot at the current moment, the position and posture of the object to be assembled at the current moment, and the environmental information corresponding to the object to be assembled; the above environmental information includes whether there are obstacles around the object to be assembled.

[0030] The multimodal information collected by the robot at the current moment may be different types of information, including image information collected by a camera or video camera on the robot, or sensor information collected by mechanical or tactile sensors on the robot. Optionally, the multimodal information may include a red, green, and blue (RGB) image of the object to be assembled, a depth image of the object to be assembled, an infrared image of the object to be assembled, mechanical feature information collected by the robot's force sensor, and tactile feature information collected by the robot's tactile sensor.

[0031] The object to be assembled may be a component to be assembled, such as a screw, a bolt, etc., which may be determined according to the actual assembly scene, and the shape and size of the object to be assembled may be arbitrary. The object to be assembled may be imaged at the current moment by a camera or a video camera or other acquisition device on the robot, and the pose of the object to be assembled at the current moment may be obtained by calculating the acquired image. The pose of the object to be assembled at the current moment may include the position and posture at the current moment, and the pose of the object to be assembled at the current moment may be the pose perceived by the robot.

[0032] The environmental information corresponding to the object to be assembled can be obtained by collecting images of the object to be assembled and the environment around the robot through a camera or video camera on the robot, and identifying the environmental information of the robot and the object to be assembled through the collected environmental images. The environmental information may include whether there are obstacles around the robot and the object to be assembled, the distance to the obstacles when there are obstacles, the number or shape of the obstacles, and other information.

[0033] Step 104, determining the state of the robot at the current moment according to the multimodal information at the current moment; the state of the robot at the current moment includes the position and posture of the robot at the current moment.

[0034] In this step, the robot can also obtain its own position at the current moment through the positioning module, and can also obtain the robot's posture at the current moment through other sensors (such as speed sensors, angle sensors, etc.). The robot's current position and posture can be used to obtain the robot's current posture.

[0035] After obtaining the multimodal information collected by the robot at the current moment, all the information in the multimodal information can be combined with the position and posture of the robot at the current moment to directly serve as the state of the robot at the current moment, or a part of the multimodal information can be selected to obtain the state of the robot at the current moment in combination with the position and posture of the robot at the current moment, or part of the information in the modal information can be further processed and then combined with the position and posture of the robot at the current moment to obtain the state of the robot at the current moment, etc. In short, the state of the robot at the current moment can be obtained through the multimodal information collected by the robot at the current moment.

[0036] Step 106, based on the current state of the robot, the position and posture of the object to be assembled, and environmental information, a deep reinforcement learning network is used to plan the motion path of the robot to determine the motion trajectory of the robot moving from its current position and posture to the object to be assembled.

[0037] Among them, the specific architecture and type of the deep reinforcement learning network are not specifically limited here, for example, it can be a deep Q network. In this step, the deep reinforcement learning network can be pre-trained. During the specific training, the deep reinforcement learning network is trained using the distributed proximal policy optimization (DPPO) algorithm, and the distributed proximal policy optimization algorithm uses multiple parallel environments to jointly plan the motion path. That is, in each training cycle of the deep reinforcement learning network, multiple environments are used for parallel training to accelerate the gradient update of the deep reinforcement learning network to improve the training efficiency, thereby improving the efficiency of subsequent motion path planning and object assembly through the trained deep reinforcement learning network.

[0038] The input of the above-mentioned deep reinforcement learning network can be the state at a certain moment and other information, and the output can be the strategy distribution and / or the executed actions and / or the value estimation, etc.

[0039] Specifically, after the deep reinforcement learning network is trained, the robot's current state, the position and posture of the object to be assembled at the current moment, and the environmental information obtained above can be input into the deep reinforcement learning network for motion path planning, and the action that the robot needs to perform at the next moment can be output, and then the robot's state at the next moment is determined based on this, and then input into the deep reinforcement learning network for motion path planning, and the actions performed by the robot at each moment are obtained through repeated iterations, and the robot's motion trajectory is obtained through the position combination corresponding to these actions. The motion trajectory here can be the motion trajectory of the robot moving from the current posture to the posture of the object to be assembled. In this way, the robot's motion path planning can be achieved.

[0040] As can be seen from the above description, the state data of the above robot combines information of multiple different modes, so that the information used by the robot in path planning, motion planning or object assembly is more diverse and richer, thereby improving the accuracy of subsequent path planning, motion planning or object assembly. At the same time, such richer information does not require an accurate robot dynamics model, so that the motion planning process in this embodiment can adapt to changes in the physical properties of different robots and environmental changes, and has stronger adaptability to assembly parts of various shapes and sizes and different assembly scenarios.

[0041] In this embodiment, by acquiring the multimodal information collected by the robot at the current moment, the position and posture corresponding to the object to be assembled at the current moment, and the environmental information corresponding to the object to be assembled, the state of the robot at the current moment is determined according to the multimodal information at the current moment, and according to the state of the robot at the current moment, the position and posture of the object to be assembled and the environmental information, a deep reinforcement learning network is used to plan the motion path of the robot to determine the motion trajectory of the robot moving from the position and posture at the current moment to the object to be assembled; wherein, the state of the robot at the current moment includes its position and posture at the current moment, and the deep reinforcement learning network is trained by using a distributed proximal strategy optimization algorithm, and the distributed proximal strategy optimization algorithm uses multiple parallel environments to jointly plan the motion path. In this method, since the state of the robot can be determined by the multimodal information collected by the robot, and combined with the position and environmental information of the object to be assembled, a deep reinforcement learning network is used to plan the motion path of the robot. In this way, through multimodal information and combined with other more information, the robot can plan a high-precision motion trajectory in the motion planning process, improve the accuracy of motion planning and subsequent object assembly, and at the same time, this method does not require an accurate robot dynamics model, can adapt to the changes in the physical characteristics and environment of different robots, and has stronger adaptability to assembly parts of various shapes and sizes and different assembly scenes, so it can reduce the system maintenance cost; at the same time, when the position of the assembly object is offset, obstacles appear, or the lighting conditions change, the path planning and assembly tasks can still be effectively completed. In addition, since the deep reinforcement learning network in this method uses multiple parallel environments to jointly plan the motion path, the efficiency of path planning can be significantly improved.

[0042] The following embodiment describes a possible implementation method for motion path planning.

[0043] In some embodiments, the above step 106 uses a deep reinforcement learning network to plan the motion path of the robot according to the current state of the robot, the position and posture of the object to be assembled, and the environmental information, and determines the motion trajectory of the robot moving from its current position and posture to the object to be assembled, which may include the following steps: Step A1, according to the current state of the robot, the posture of the object to be assembled and the environmental information, a deep reinforcement learning network is used to plan the motion path of the robot to determine the state of the robot at the next moment; the state at the next moment includes the posture of the robot at the next moment.

[0044] In this step, the current state of the robot, the position and posture of the object to be assembled, and the environmental information can be input into the deep reinforcement learning network for motion path planning, and the strategy distribution and / or the action to be performed and / or the value estimation of the action to be performed by the robot at the next moment can be output. Then, the position and posture of the robot are adjusted according to the strategy distribution and / or the action to be performed and / or the value estimation of the action to be performed by the robot at the next moment, and the multimodal information collected by the robot at the next moment is re-collected, and then the state of the robot at the next moment is determined according to the multimodal information collected at the next moment, and the position and posture of the robot at the next moment are obtained.

[0045] Step A2, determining reward functions corresponding to the various assembly factors according to the various assembly factors corresponding to the object to be assembled during the assembly process.

[0046] In this step, before assembling the object, an assembly reward function can be set. The reward function here is related to multiple assembly factors. As an optional embodiment, the assembly requirements and each assembly factor corresponding to the object to be assembled can be obtained; the weight corresponding to each assembly factor can be determined according to the assembly requirements; and the reward function can be determined according to each assembly requirement and the weight corresponding to each assembly requirement.

[0047] Specifically, the assembly requirements may represent requirements for various assembly factors, where the various assembly factors may include assembly accuracy, assembly time, assembly energy consumption, collision risk during assembly, operational stability during assembly, etc. The assembly requirements may be determined in advance according to the assembly task, for example, the assembly requirements may be the requirement for higher assembly accuracy or the requirement for faster assembly time, etc.

[0048] Each assembly factor can be pre-set with a default weight. After the assembly requirements are determined, the default weights of each assembly factor can be adjusted according to the assembly requirements to obtain the final weight. For example, in high-precision electronic component assembly tasks, the assembly requirements indicate extremely high requirements for assembly accuracy. The weight of assembly accuracy can be increased to make the robot pay more attention to assembly accuracy; in large-scale production scenarios where assembly requirements have high production efficiency requirements, in order to improve overall production efficiency, the weight of assembly time can be appropriately increased. In this way, the final weight of the assembly factor can be determined. After that, each assembly factor can be multiplied by its corresponding weight, and the products can be added to obtain the reward function. For example, the reward function can be expressed as follows: .

[0049] in, R represents the reward function, P Indicates assembly accuracy, which is a key indicator for measuring assembly quality, ensuring that parts can be accurately installed in designated locations; TIt indicates the assembly time, which can reflect the assembly efficiency. In large-scale production, shortening the assembly time can significantly improve production efficiency. E Indicates assembly energy consumption. In the context of increasingly tight energy supply, reducing energy consumption is crucial for the sustainable development of enterprises. C Indicates the collision risk during assembly, avoiding collisions between the robot and the surrounding environment or parts, which can reduce the risk of equipment damage and production interruption; S Indicates operational stability during assembly. Stable operation helps improve assembly quality and consistency. , , , , They represent the weights of assembly accuracy, assembly time, assembly energy consumption, collision risk during assembly, and operational stability during assembly, respectively. They can be flexibly adjusted according to the needs of specific assembly tasks. Usually .

[0050] The design of the above-mentioned reward function can motivate the robot to pursue higher assembly accuracy, shorter assembly time, lower energy consumption, smaller collision risk and higher operational stability while ensuring assembly success, thereby achieving more optimized motion planning and control.

[0051] Step A3, based on the reward function, the state of the robot at the next moment, the position and posture of the object to be assembled, and the environmental information, the deep reinforcement learning network is used to iteratively plan the motion path of the robot and determine the motion trajectory.

[0052] In this step, a reward function can be designed for the deep reinforcement learning network according to the above-mentioned reward function. When the current state of the robot, the position and posture of the object to be assembled, and the environmental information are input into the deep reinforcement learning network for motion path planning, the deep reinforcement learning network can obtain the reward obtained after performing an action based on the current state based on the above-mentioned reward function. Then, the strategy in the deep reinforcement learning network can be adjusted based on the reward to make the current strategy optimal, and the strategy distribution and / or the action to be performed and / or the value estimate of the action to be performed by the robot at the next moment can be output based on the optimal strategy. Then, in this way, the strategy can be updated and the motion path can be planned by continuously iterating the reward function and the state, and finally the motion trajectory of the robot moving from the current position to the object to be assembled can be obtained.

[0053] In this embodiment, the robot's current posture, the posture of the object to be assembled, and the environmental information are combined with a deep reinforcement learning network to obtain the robot's state at the next moment, and the reward function constructed by combining multiple assembly factors is used to plan the robot's motion path to obtain the robot's motion trajectory. In this way, the robot's motion path is planned by using the reward function of multiple assembly factors, so that the final planned path can be more in line with the actual assembly requirements, with better accuracy and higher efficiency. In addition, the weight of each assembly factor is determined by the assembly requirements, and a reward function including multiple assembly factors is constructed accordingly. In this way, the robot can ensure the success of the assembly while meeting the assembly requirements corresponding to various assembly factors as much as possible, and achieve more optimized motion path planning and control.

[0054] The above embodiment describes the process of the robot moving from its current position to the position of the object to be assembled. This process implements coarse-grained path planning. In order to achieve more refined path planning, the following embodiment describes the process of performing refined motion path planning based on deep reinforcement learning combined with visual servoing.

[0055] Figure 2 This is the second flow chart of the robot motion planning method provided by the present invention, such as Figure 2 As shown, the above method may further include the following steps: Step 202 , when the robot moves to the vicinity of the object to be assembled according to the motion trajectory, a visual image of the object to be assembled acquired by a camera is obtained, and a target posture corresponding to the object to be assembled is determined according to the visual image.

[0056] In this step, when the robot moves from the current position to the vicinity of the object to be assembled, the camera installed on the robot can be used to capture images of the object to be assembled to obtain a captured visual image, which can be an RGB image, a depth image, etc. Then the electronic device can obtain the visual image captured by the camera, and process the visual image, such as by identifying the position and posture of the object to be assembled in the image, obtaining the posture in the image, and then converting the posture in the image into the physical space to obtain the target posture of the object to be assembled in the physical space.

[0057] Step 204, obtaining the first pose currently corresponding to the end effector of the robot.

[0058] In this step, the position data, speed data, and angle data of the robot can be collected by the position sensor, speed sensor, angle sensor, and other devices at the end effector of the robot, and the current position and posture of the end effector of the robot can be obtained through the collected position data, speed data, angle data, etc., and recorded as the first position. For example, the currently collected position, speed, and angle of the robot are used as its first position.

[0059] Step 206, adjusting the first posture of the robot according to the first posture and the target posture to determine a second posture; the accuracy of the second posture is higher than the accuracy of the first posture.

[0060] In this step, after obtaining the first pose of the robot end effector and the target pose of the object to be assembled, the deviation between the first pose and the target pose can be calculated, and then the first pose can be accurately adjusted according to the deviation and the preset gain to obtain a second pose with higher accuracy.

[0061] For example, the formula Adjust the posture (including position and attitude) of the robot's end effector, where u is the adjusted second posture, and K is the preset gain, which is the gain factor in the robot's visual servo system. It mainly determines the range of motion of the robot when making fine adjustments and can be predetermined by the robot's visual servo system. is the deviation between the first pose and the target pose.

[0062] Step 208, sending a control instruction to the robot; the control instruction includes a second posture, which is used to control the robot to adjust the posture of the end effector according to the second posture.

[0063] In this step, after adjusting a more precise second posture for the robot, the second posture can be encapsulated in a control instruction and sent to the robot. After obtaining the control instruction, the robot can control its end effector to adjust the current first posture to a more precise second posture, so as to better assemble the assembly object in the future, such as grasping, putting down, tightening, etc.

[0064] After the robot has accurately adjusted its posture according to the control instructions of the visual servo system, the adjusted posture can be fed back to the deep reinforcement learning network, so that the deep reinforcement learning network can analyze the execution effect of the current output strategy based on the adjusted posture, and continuously optimize the strategy network according to the execution effect, thereby realizing a closed loop of the entire assembly process.

[0065] For example, in the case of automobile parts assembly, the deep reinforcement learning network first plans a rough grasping and placement path based on environmental information to guide the robot to approach the target parts. In the process of approaching the target, the robot's visual servo system obtains the target position and posture information of the parts in real time through the camera, and calculates the deviation between its end effector and the target position. Then, according to the control law , visual servo sends precise adjustment instructions to the robot, fine-tuning the position and posture of the robot's end effector to ensure accurate grasping and placement of parts. Afterwards, deep reinforcement learning analyzes the execution effect of the strategy based on the feedback results of visual servo, summarizes the experience and lessons, and optimizes the subsequent strategies. Through this closed-loop optimization mechanism, the robot can continuously improve the accuracy and efficiency of assembly.

[0066] In this embodiment, when the robot moves roughly around the object to be assembled, the camera on the robot captures the visual image of the object to obtain its target posture and the first posture of the robot's end effector, and the posture of the robot's end effector is adjusted in combination with these two postures so that it can approach the object to be assembled more accurately. In this way, precise assembly of the object to be assembled can be achieved through the close collaboration of the visual servo system and the deep reinforcement learning network.

[0067] The following embodiment describes a process of setting multiple strategy networks based on an assembly task to execute an assembly task.

[0068] In some embodiments, the above method may further include the following steps: The assembly task corresponding to the object to be assembled is divided into multiple subtasks, and multiple policy networks are configured in the deep reinforcement learning network; each subtask corresponds to a policy network, and the multiple subtasks include a first subtask of moving the robot from its current position to the object to be assembled; When it is determined that the robot has completed the first subtask, the current position and posture of the robot is fed back to the deep reinforcement learning network, so that the deep reinforcement learning network optimizes the policy network corresponding to the second subtask based on the current position and posture of the robot; the second subtask is the next subtask to be performed after the first subtask is completed among the multiple subtasks; The second subtask is performed according to the policy network corresponding to the second subtask optimized by the deep reinforcement learning network.

[0069] After determining the assembly task of the object to be assembled, the assembly task can be divided into multiple subtask levels, such as coarse positioning subtask, precise positioning subtask, insertion subtask and tightening subtask, etc. The coarse positioning subtask is the task of the robot moving from its current position to the vicinity of the object to be assembled according to the planned motion trajectory, which can be recorded as the first subtask.

[0070] In this embodiment, a corresponding strategy network and value network are pre-configured for each subtask in the deep learning network according to multiple subtasks, so as to execute the corresponding motion planning subtask.

[0071] After the robot completes the first subtask, it can continue to perform the second subtask, which is a subtask that is performed continuously after the first subtask. For example, the first subtask is a coarse positioning task, the second subtask is a precise positioning subtask, and the third subtask can be a grasping and inserting subtask, etc. After completing the first subtask, the electronic device can obtain the current posture of the robot and feed it back to the deep reinforcement learning network, so that the deep reinforcement learning network adjusts and optimizes the policy network of the second subtask according to the input current posture, so as to obtain the optimal policy network of the second subtask, and after determining the optimal policy network of the second subtask, the second subtask can be performed according to the optimal policy network of the second subtask.

[0072] It is understandable that the policy network of each subtask can be continuously optimized and updated. For each subtask policy network, its update method may include the following steps: According to the first strategy output by the policy network at the current time and the second strategy outputted at the previous time, the degree of difference between the first strategy and the second strategy is determined; the target action currently selected is determined according to the first strategy, and the advantage of the target action is evaluated according to the advantage estimation function; according to the degree of difference, the advantage of the target action, the number of parallel environments currently used, and the number of time steps, the policy network is updated to determine the updated policy network.

[0073] Among them, the policy network of each subtask can be updated according to the following formula: .

[0074] in, represents the policy loss function, represents the policy loss function Parameters The gradient of N Indicates the number of parallel environments used to calculate the loss function in each training cycle, by N Run in parallel in multiple environments to speed up gradient updates and improve training efficiency; T is the number of time steps, which indicates the number of times the agent interacts with the environment in each training cycle; n Indicates the number of parallel environments, ranging from 1 to N , t Indicates time / moment, ranging from 1 to T ; It represents the probability ratio, which is used to measure the difference between the first strategy output by the strategy network in the current time and the second strategy output in the previous time. It can be calculated using the following formula: .

[0075] in,s t Indicates that the robot is t The state of the moment, Indicates in status s t Next, select the action; Indicates the first strategy of the current output; Represents the second strategy of the previous output; each strategy output includes the current state and the action selected under the current state.

[0076] represents the advantage estimate, i.e., the estimated advantage, used to evaluate the advantage of the current action; clip () function means limiting the probability ratio within a certain range to prevent its value from exceeding the preset range, thereby avoiding excessive strategy updates that lead to training failure; Represents the pruning parameter, which is used to limit the magnitude of the strategy update and ensure the stability of the training.

[0077] Specifically, the above can obtain the estimated advantage through the strategy currently output by the strategy network, which includes the current state of the robot and the target action selected in the current state, and then estimate the advantage of the current target action; at the same time, calculate the probability ratio between the current output strategy and the previous output strategy, and update the strategy network in combination with the current number of time steps and the number of parallel environments, clipping parameters, etc., to obtain an updated strategy network, and the performance of the updated strategy network is better.

[0078] In the actual assembly process of complex mechanical products, such as assembling an engine, the policy network at the coarse positioning level will quickly guide the robot to the approximate location of the target part based on environmental information, laying the foundation for subsequent precise operations. Subsequently, the policy network at the precise positioning level will use more detailed environmental perception information to further adjust the robot's position so that it approaches the part more accurately. At key operation levels such as insertion and tightening, the policy network at the corresponding level will generate precise motion control instructions based on the shape, size and assembly requirements of the part.

[0079] In addition, the policy networks and value networks corresponding to these multiple levels or subtasks can be pre-trained, and the distributed proximal policy optimization DPPO algorithm can be used for training during the training process. The DPPO algorithm combines the advantages of distributed training and can make full use of multiple parallel environments for training at the same time, greatly accelerating the learning speed, so that the robot can learn effective strategies for different assembly stages (coarse positioning, precise positioning, insertion, tightening, etc.) more quickly, so as to better cope with complex high-dimensional states and action spaces.

[0080] In this embodiment, by dividing the assembly task into multiple subtasks and configuring a corresponding strategy network and value network for each subtask, it is convenient to provide a better strategy for each subtask to perform the corresponding task, thereby improving the accuracy of subtask execution and the efficiency of the entire assembly task execution.

[0081] The following embodiment illustrates a process of obtaining the state of the robot at the current moment based on the multimodal information collected by the robot at the current moment.

[0082] In some embodiments, determining the state of the robot at the current moment according to the multimodal information at the current moment in the above step 104 may include the following steps: A convolutional neural network is used to extract features from the RGB image and the infrared image at the current moment, respectively, to determine a first feature corresponding to the RGB image and a second feature corresponding to the infrared image; the first feature is used to characterize the surface texture and color information of the object to be assembled, and the second feature is used to characterize the contour information of the object to be assembled; A point cloud processing network is used to perform feature extraction processing on the depth image to determine a third feature corresponding to the depth image; the third feature is used to characterize the distance between the robot and the object to be assembled and the three-dimensional shape of the object to be assembled; A feature fusion operation is performed on the first feature, the second feature, the third feature, the mechanical feature information, and the tactile feature information to determine the state of the robot at the current moment.

[0083] Among them, the above-mentioned convolutional neural network can be, for example, a CNN network, and the point cloud processing network can be, for example, a PointNet network, etc.

[0084] Specifically, after acquiring the RGB image at the current moment, the key features such as the surface texture and color information of the object to be assembled in the RGB image can be extracted through a convolutional neural network to obtain the extracted features, which are recorded as the first features. At the same time, after acquiring the infrared image, the same convolutional neural network as above or a different convolutional neural network can be used to extract the contour feature information of the object to be assembled in the infrared image to obtain the extracted features, which are recorded as the second features.

[0085] The robot can also collect the depth image of the object to be assembled at the current moment through devices such as lidar or depth camera, and then use the point cloud processing network to extract the spatial structure features of the depth image to obtain the distance between the robot and the object to be assembled, the three-dimensional shape of the object to be assembled, the position of the object to be assembled and other feature information, and then use these feature information as the third feature.

[0086] At the same time, sensor devices such as force sensors and tactile sensors can also be installed on the robot's end effector (such as a manipulator). When the robot grasps an object to be assembled, the mechanical sensor can be used to collect the force information at the current moment when the robot grasps the object to be assembled, and the tactile sensor can be used to collect the tactile information at the current moment when the robot grasps the object to be assembled. The force information can be recorded as mechanical characteristic information, and the tactile information can be recorded as tactile characteristic information.

[0087] After obtaining the first feature, the second feature, the third feature, the mechanical feature information, and the tactile feature information at the current moment, a feature fusion operation can be performed on these features. For example, the following formula can be used to perform the feature fusion operation to obtain the fused high-dimensional feature, and the high-dimensional feature is used as the current state of the robot: .

[0088] in, Indicates the state of the robot at the current moment; represents feature fusion operation; Represents an RGB image; represents a depth image; Indicates infrared image; Represents mechanical characteristic information and tactile characteristic information.

[0089] Taking the complex assembly workshop environment as an example, when the robot needs to grasp a part with complex surface texture and changing surrounding light, RGB images can provide surface texture and color information of the part, depth images can accurately measure the distance between the part and the robot and the three-dimensional shape of the part, infrared images can assist in identifying the contour of the part when the light is poor, and force sensors and tactile sensors can provide feedback on the force conditions when the robot grasps the part.

[0090] In this embodiment, the state of the robot at the current moment is obtained by extracting and fusing the multimodal information collected by the robot at the current moment. By fusing the multimodal feature information, the robot can perceive the assembly scene more comprehensively and accurately, greatly improving the accuracy and robustness of the state representation, and providing a solid and reliable data basis for the robot's motion planning.

[0091] It can be seen from the description of the above embodiments that in the embodiments of the present invention, the robot collects information about the environment and assembly parts through multimodal sensors, including RGB images, depth images, infrared images and other sensor information, and extracts and fuses features through these information to obtain the state of the robot, which is a high-dimensional state representation. After that, a deep reinforcement learning network / architecture is used to generate a preliminary motion plan based on the state information, and the assembly task is decomposed into multiple levels, each of which uses a different strategy network and value network for learning. In this process, a reward function related to each assembly factor of the assembly requirement is added to give feedback based on the performance of the assembly, guiding the strategy optimization of deep reinforcement learning. At the same time, the motion trajectory of the deep reinforcement learning plan is finely adjusted in real time through visual servoing, and the feedback of the visual servoing reacts to the strategy update of the deep reinforcement learning, forming a closed-loop optimization, and finally achieving precise assembly operations.

[0092] The embodiments of the present invention have the following technical effects: 1. Improve assembly efficiency and precision.

[0093] Through multimodal information fusion and improved deep reinforcement learning architecture, the robot can learn effective assembly strategies more quickly and maintain high assembly accuracy in complex environments, reducing assembly time and errors to meet the requirements of high-precision assembly tasks.

[0094] 2. Enhance adaptability and robustness.

[0095] Without the need for an accurate robot dynamics model, the system can adapt to changes in the physical characteristics of different robots and environmental changes, and has stronger adaptability to assembly parts of various shapes and sizes and different assembly scenarios, reducing system maintenance costs. Even when the position of the parts is offset, obstacles appear, or the lighting conditions change, the assembly task can still be completed effectively.

[0096] 3. Optimize resource utilization.

[0097] The reward function that comprehensively considers various assembly factors can enable the robot to optimize resource utilization, reduce energy consumption and costs while completing the assembly task. For large-scale automated assembly production lines, it can significantly save energy and operating costs.

[0098] The robot motion planning device provided by the present invention is described below. The robot motion planning device described below and the robot motion planning method described above can be referenced to each other.

[0099] Figure 3 is a schematic diagram of the structure of the robot motion planning device provided by the present invention, see Figure 3 As shown, the device may include: The acquisition module 310 is used to acquire the multimodal information collected by the robot at the current moment, the position and posture of the object to be assembled at the current moment, and the environmental information corresponding to the object to be assembled; the above environmental information includes whether there are obstacles around the object to be assembled; A state determination module 320, used to determine the state of the robot at the current moment according to the multimodal information at the current moment; the state of the robot at the current moment includes the position and posture of the robot at the current moment; The motion planning module 330 is used to plan the motion path of the robot using a deep reinforcement learning network according to the current state of the robot, the position and posture of the object to be assembled, and the environmental information, and determine the motion trajectory of the robot moving from its current position and posture to the object to be assembled; Among them, the deep reinforcement learning network is trained using a distributed proximal strategy optimization algorithm, in which multiple parallel environments are used to jointly plan the motion path.

[0100] In some embodiments, the motion planning module 330 is specifically used to plan the motion path of the robot using a deep reinforcement learning network according to the current state of the robot, the position and posture of the object to be assembled, and the environmental information, and determine the state of the robot at the next moment; the state at the next moment includes the position and posture of the robot at the next moment; According to the various assembly factors corresponding to the object to be assembled during the assembly process, determining the reward functions corresponding to the various assembly factors; According to the reward function, the state of the robot at the next moment, the posture of the object to be assembled and the environmental information, the deep reinforcement learning network is used to iteratively plan the robot's motion path and determine the motion trajectory.

[0101] Optionally, the motion planning module 330 is specifically used to obtain assembly requirements and assembly factors corresponding to the object to be assembled; determine the weights corresponding to the assembly factors according to the assembly requirements; and determine the reward function according to the assembly requirements and the weights corresponding to the assembly requirements.

[0102] In some embodiments, the above apparatus further comprises: A target posture determination module is used to obtain a visual image of the object to be assembled after the camera captures the image when the robot moves around the object to be assembled according to the motion trajectory, and determine the target posture corresponding to the object to be assembled according to the visual image; The robot posture acquisition module is used to obtain the first posture corresponding to the robot's end effector; A posture adjustment module is used to adjust the first posture of the robot according to the first posture and the target posture to determine the second posture; the accuracy of the second posture is higher than the accuracy of the first posture; The control module is used to send control instructions to the robot; the above control instructions include a second posture, which is used to control the robot to adjust the posture of the end effector according to the second posture.

[0103] In some embodiments, the above apparatus further comprises: A division module is used to divide the assembly task corresponding to the object to be assembled into multiple subtasks, and configure multiple policy networks in the deep reinforcement learning network; each subtask corresponds to a policy network, and the multiple subtasks include a first subtask of moving the robot from its current position to the object to be assembled; A feedback module is used to feed back the current posture of the robot to the deep reinforcement learning network when it is determined that the robot has completed the first subtask, so that the deep reinforcement learning network optimizes the policy network corresponding to the second subtask based on the current posture of the robot; the second subtask is the next subtask to be performed after the first subtask is completed among the multiple subtasks; An execution module is used to execute the second subtask according to the policy network corresponding to the second subtask optimized by the deep reinforcement learning network.

[0104] In some embodiments, the apparatus further comprises an updating module, the updating module being used to update the policy network of each subtask; The update module is specifically used to determine the degree of difference between the first strategy and the second strategy output by the policy network at the current time and the second strategy outputted at the previous time; determine the target action currently selected according to the first strategy, and evaluate the advantage of the target action according to the advantage estimation function; update the policy network according to the degree of difference, the advantage of the target action, the number of parallel environments currently used, and the number of time steps, and determine the updated policy network.

[0105] In some embodiments, the multimodal information includes an RGB image of the object to be assembled, a depth image of the object to be assembled, an infrared image of the object to be assembled, mechanical feature information collected by the force sensor of the robot, and tactile feature information collected by the tactile sensor of the robot. The state determination module 320 is specifically used to use a convolutional neural network to perform feature extraction on the RGB image and the infrared image at the current moment, respectively, to determine the first feature corresponding to the RGB image and the second feature corresponding to the infrared image; the first feature is used to characterize the surface texture and color information of the object to be assembled, and the second feature is used to characterize the contour information of the object to be assembled; the point cloud processing network is used to perform feature extraction processing on the depth image to determine the third feature corresponding to the depth image; the third feature is used to characterize the distance between the robot and the object to be assembled and the three-dimensional shape of the object to be assembled; the first feature, the second feature, the third feature, the mechanical feature information and the tactile feature information are subjected to feature fusion operation to determine the state of the robot at the current moment.

[0106] It should be noted here that the above-mentioned device provided in the embodiment of the present invention can implement all the method steps implemented in the above-mentioned method embodiment, and can achieve the same technical effect. The parts and beneficial effects that are the same as the method embodiment in this embodiment will not be described in detail here.

[0107] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor (processor) 410 , a communication interface (Communications Interface) 420 , a memory (memory) 430 and a communication bus 440 , wherein the processor 410 , the communication interface 420 , and the memory 430 communicate with each other through the communication bus 440 . The processor 410 can call the logic instructions in the memory 430 to execute the robot motion planning method, which includes: obtaining the multimodal information collected by the robot at the current moment, the posture corresponding to the object to be assembled at the current moment, and the environmental information corresponding to the object to be assembled; the above-mentioned environmental information includes whether there are obstacles around the object to be assembled; determining the state of the robot at the current moment according to the multimodal information at the current moment; the above-mentioned state of the robot at the current moment includes the posture of the robot at the current moment; according to the state of the robot at the current moment, the posture of the object to be assembled and the environmental information, using a deep reinforcement learning network to plan the motion path of the robot, and determine the motion trajectory of the robot moving from its current posture to the object to be assembled; wherein the deep reinforcement learning network is trained using a distributed proximal strategy optimization algorithm, and the distributed proximal strategy optimization algorithm uses multiple parallel environments to jointly plan the motion path.

[0108] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0109] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the robot motion planning method provided by the above methods, which includes: obtaining multimodal information collected by the robot at the current moment, the posture corresponding to the object to be assembled at the current moment, and the environmental information corresponding to the object to be assembled; the above environmental information includes whether there are obstacles around the object to be assembled; determining the state of the robot at the current moment according to the multimodal information at the current moment; the state of the robot at the current moment includes the posture of the robot at the current moment; according to the state of the robot at the current moment, the posture of the object to be assembled and the environmental information, using a deep reinforcement learning network to plan the motion path of the robot, and determining the motion trajectory of the robot moving from its current posture to the object to be assembled; wherein the deep reinforcement learning network is trained using a distributed proximal strategy optimization algorithm, and multiple parallel environments are used in the distributed proximal strategy optimization algorithm to jointly plan the motion path.

[0110] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the robot motion planning method provided by the above-mentioned methods, the method comprising: obtaining multimodal information collected by the robot at the current moment, the posture corresponding to the object to be assembled at the current moment, and the environmental information corresponding to the object to be assembled; the above-mentioned environmental information includes whether there are obstacles around the object to be assembled; determining the state of the robot at the current moment according to the multimodal information at the current moment; the state of the robot at the current moment includes the posture of the robot at the current moment; according to the state of the robot at the current moment, the posture of the object to be assembled and the environmental information, using a deep reinforcement learning network to plan the motion path of the robot, and determining the motion trajectory of the robot moving from its current posture to the object to be assembled; wherein the deep reinforcement learning network is trained using a distributed proximal strategy optimization algorithm, and the distributed proximal strategy optimization algorithm uses multiple parallel environments to jointly plan the motion path.

[0111] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0112] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A robot motion planning method, characterized in that: include: Acquire the multimodal information collected by the robot at the current moment, the posture corresponding to the object to be assembled at the current moment, and the environmental information corresponding to the object to be assembled; The environmental information includes whether there are obstacles around the object to be assembled; Determining the state of the robot at the current moment according to the multimodal information at the current moment; the state of the robot at the current moment includes the position and posture of the robot at the current moment; According to the current state of the robot, the position and posture of the object to be assembled, and the environmental information, a deep reinforcement learning network is used to plan the motion path of the robot to determine the motion trajectory of the robot moving from its current position and posture to the object to be assembled; Among them, the deep reinforcement learning network is trained by adopting a distributed proximal strategy optimization algorithm, and the distributed proximal strategy optimization algorithm adopts multiple parallel environments to jointly plan the motion path.

2. The robot motion planning method according to claim 1, characterized in that: The method of using a deep reinforcement learning network to plan a motion path for the robot according to the current state of the robot, the position and posture of the object to be assembled, and the environmental information, and determining a motion trajectory of the robot moving from its current position and posture to the object to be assembled, comprises: According to the current state of the robot, the position and posture of the object to be assembled and the environmental information, a deep reinforcement learning network is used to plan the motion path of the robot to determine the state of the robot at the next moment; the state at the next moment includes the position and posture of the robot at the next moment; According to the multiple assembly factors corresponding to the object to be assembled during the assembly process, determining the reward functions corresponding to the multiple assembly factors; According to the reward function, the state of the robot at the next moment, the position and posture of the object to be assembled and the environmental information, a deep reinforcement learning network is used to iteratively plan the motion path of the robot to determine the motion trajectory.

3. The robot motion planning method according to claim 2, characterized in that: The step of determining the reward functions corresponding to the various assembly factors in the assembly process of the object to be assembled comprises: Obtaining assembly requirements and assembly factors corresponding to the object to be assembled; Determine the weight corresponding to each of the assembly factors according to the assembly requirements; The reward function is determined according to each of the assembly requirements and the weight corresponding to each of the assembly requirements.

4. The robot motion planning method according to claim 1, characterized in that: The method further comprises: When the robot moves to the vicinity of the object to be assembled according to the motion trajectory, a visual image of the object to be assembled is acquired after being captured by a camera, and a target posture corresponding to the object to be assembled is determined according to the visual image; Obtaining the first position currently corresponding to the end effector of the robot; According to the first posture and the target posture, the first posture of the robot is adjusted to determine a second posture; the accuracy of the second posture is higher than the accuracy of the first posture; A control instruction is sent to the robot; the control instruction includes the second posture, which is used to control the robot to adjust the posture of the end effector according to the second posture.

5. The robot motion planning method according to claim 1, characterized in that: The method further comprises: The assembly task corresponding to the object to be assembled is divided into a plurality of subtasks, and a plurality of policy networks are configured in the deep reinforcement learning network; each of the subtasks corresponds to a policy network, and the plurality of subtasks include a first subtask of moving the robot from its current position to the object to be assembled; When it is determined that the robot has completed the first subtask, the current posture of the robot is fed back to the deep reinforcement learning network, so that the deep reinforcement learning network optimizes the policy network corresponding to the second subtask based on the current posture of the robot; the second subtask is the next subtask to be performed after the first subtask is completed among the multiple subtasks; The second subtask is performed according to the policy network corresponding to the second subtask optimized by the deep reinforcement learning network.

6. The robot motion planning method according to claim 5, characterized in that: The updating method of the policy network for each subtask includes: Determine the degree of difference between the first strategy and the second strategy according to the first strategy output by the strategy network at the current time and the second strategy output by the strategy network at the previous time; Determine a target action currently selected according to the first strategy, and evaluate the advantage of the target action according to an advantage estimation function; The policy network is updated according to the degree of difference, the advantage of the target action, the number of parallel environments currently used, and the number of time steps to determine an updated policy network.

7. The robot motion planning method according to claim 1, characterized in that: The multimodal information includes a red, green, and blue (RGB) image of the object to be assembled, a depth image of the object to be assembled, an infrared image of the object to be assembled, mechanical feature information collected by a force sensor of the robot, and tactile feature information collected by a tactile sensor of the robot. Determining the state of the robot at the current moment according to the multimodal information at the current moment includes: A convolutional neural network is used to extract features from the RGB image and the infrared image at the current moment, respectively, to determine a first feature corresponding to the RGB image and a second feature corresponding to the infrared image; the first feature is used to characterize the surface texture and color information of the object to be assembled, and the second feature is used to characterize the contour information of the object to be assembled; Performing feature extraction processing on the depth image using a point cloud processing network to determine a third feature corresponding to the depth image; the third feature is used to characterize the distance between the robot and the object to be assembled and the three-dimensional shape of the object to be assembled; A feature fusion operation is performed on the first feature, the second feature, the third feature, the mechanical feature information, and the tactile feature information to determine the state of the robot at a current moment.

8. A robot motion planning device, characterized in that: include: An acquisition module is used to acquire the multimodal information collected by the robot at the current moment, the posture corresponding to the object to be assembled at the current moment, and the environmental information corresponding to the object to be assembled; The environmental information includes whether there are obstacles around the object to be assembled; A state determination module, used to determine the state of the robot at the current moment according to the multimodal information at the current moment; the state of the robot at the current moment includes the position and posture of the robot at the current moment; A motion planning module, for planning the motion path of the robot using a deep reinforcement learning network according to the current state of the robot, the position and posture of the object to be assembled, and the environmental information, and determining the motion trajectory of the robot moving from its current position and posture to the object to be assembled; Among them, the deep reinforcement learning network is trained by adopting a distributed proximal strategy optimization algorithm, and the distributed proximal strategy optimization algorithm adopts multiple parallel environments to jointly plan the motion path.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the robot motion planning method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the robot motion planning method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Artificial potential field and reinforced learning based man-machine co-fusion assembly line implementation method

    CN111515932A

  • Reinforcement learning awarding method suitable for movable mechanical arm

    CN111515961A

  • Distributed near-end strategy optimization method based on cognitive behavior knowledge and application thereof

    CN112906233A

  • Navigation decision-making method combining curiosity mechanism and self-imitation learning

    CN116892932A

  • Autonomous mobile robot path planning method based on deep reinforcement learning

    CN118259669A