Motion planning method and device of mechanical arm, storage medium and computer program product
Through real-time perception and dynamic planning, the problem that the robotic arm cannot be accurately inserted into the charging interface in a dynamic environment is solved, and a safe and fast charging task is achieved.
Patent Information
- Application Number
- CN202510671635.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing robotic arm motion planning algorithm cannot perceive environmental changes in real time in a dynamic environment, resulting in the charging robot being unable to accurately plug into the vehicle charging interface, posing a collision risk, and it is impossible to complete the charging task safely and quickly.
By setting sensors on the robotic arm to obtain environment and motion information in real time, building a local grid map, using the action strategy network to generate action instructions, and updating it through the evaluation network to achieve dynamic planning.
It realizes that the robot can reach the target position safely and quickly in a dynamic environment, accurately complete the charging task, and provides efficient and convenient charging services.
Smart Images

Figure CN120363199A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for motion planning and control of a robot in a complex environment, and particularly to a motion planning method and device for a robotic arm of a charging robot, a storage medium, and a computer program product. Background Art
[0002] Motion planning of a robot is an important issue in autonomous unmanned systems, which involves how to plan a safe and fast task execution strategy through sensors in an unknown environment. For example, in the scenario of automatic charging, the robotic arm of a charging robot needs to manipulate a charging plug and accurately insert the charging plug into the charging interface of a vehicle, and reach a specified pose only relying on vision in a cluttered environment to complete the charging task.
[0003] If there is no precise environmental perception and sufficient obstacle information cannot be extracted, the robotic arm may collide during the task execution process, damaging the equipment; if a reasonable motion strategy cannot be planned for the robotic arm, the charging task cannot be completed safely and quickly either.
[0004] Existing motion planning algorithms perform path planning in a known environment, ignoring the uncertainty of the environment in the real scenario. For example, in a dynamic environment, the position of the charging interface may change due to different vehicle models, different parking postures or different parking positions on the same parking space, different environments of the parking space, etc. Some of the existing technologies preset fixed trajectories and cannot perform motion planning in real time, which may cause cumulative errors. Therefore, the accuracy of the existing robotic arm motion planning methods in a dynamic environment is insufficient.
[0005] As mentioned above, existing motion planning algorithms all plan a global path in a known environment and perform path tracking control. Such a method can effectively avoid obstacles in the known environment, but in the face of a complex application environment, such as the change of the parking pose caused by different parking habits of different users, and the charging interface positions of different vehicle models are also different. At this time, if a preset fixed path or a path planned in a known environment is used, the charging robot cannot accurately insert the charging plug into the charging interface of the vehicle. Therefore, how to plan a path in real time according to the dynamic environment and accurately execute tasks in an uncertain complex environment is an urgent problem to be solved currently. Summary of the Invention
[0006] In view of the above problems in the prior art, the present invention provides a motion planning method for dynamically planning the motion trajectory of a robotic arm according to the real-time perceived environmental information and the motion information of the robotic arm itself in a dynamic environment.
[0007] To achieve the above object, according to the first aspect of the present invention, there is provided a motion planning method for a robotic arm, and the motion planning method includes the following steps: An information acquisition step of acquiring environmental information around the robotic arm and motion information of the robotic arm through one or more sensors provided on the robotic arm; A positioning step of obtaining the current pose information of the robotic arm according to the acquired environmental information and motion information; A local map construction step of constructing a local grid map including obstacle information around the robotic arm and the current pose of the robotic arm according to the environmental information and the current pose information of the robotic arm; An action instruction generation step of inputting the current pose information, the local grid map, and the body perception information of the robotic arm into a pre-constructed action policy network, and generating an action instruction through the action policy network, where the action instruction is used to guide the action of the robotic arm; An update step of generating evaluation information on the action of the robotic arm through an evaluation network, and updating the action policy network and the evaluation network according to the evaluation information.
[0008] As a preferred solution, in the information acquisition step, the one or more sensors include an image sensor and an inertial measurement unit; the information acquisition step includes: Collecting image information around the robotic arm as the environmental information through the image sensor, and collecting inertial data of the robotic arm as the motion information through the inertial measurement unit.
[0009] As a preferred solution, the positioning step includes the following: A preprocessing step of extracting features from the image information to obtain feature point tracking data, and pre-integrating the motion information to obtain the pose, velocity, and rotation angle at the current moment, and simultaneously calculating the pre-integration increment, the covariance matrix of the pre-integration, and the Jacobian matrix between adjacent frames as pre-integration data; An initialization step of calculating the relative pose between adjacent frames according to the feature point tracking data, and aligning the relative pose with the pre-integration data to obtain the initial pose in the world coordinate; A local optimization step of locally optimizing the initial pose by using a sliding window optimization method to obtain a local pose; A loop detection step of matching the current frame with historical key frames, and screening out the frames matching the historical key frames according to the matching relationship as loop constraints; A global optimization step of globally optimizing the local pose according to the loop constraints to obtain a global pose.
[0010] As a preferred solution, the local map construction step includes: The depth image of the current frame obtained by the image acquisition device; According to the depth information of each two-dimensional pixel point in the depth image, project the two-dimensional pixel point into the three-dimensional space to obtain the local grid map; where the projection formula is defined as follows: , , where is the camera internal parameter; u is the pixel abscissa; v is the pixel ordinate, Z is the depth of each pixel point obtained by the depth image; X and Y respectively represent the abscissa and ordinate after projection into the three-dimensional space.
[0011] As a preferred solution, in the updating step, the evaluation network rewards the actions of the robotic arm based on the reward function as the evaluation information, and the reward function is defined as follows: , where represents the position reward function, which is used to reward the end effector of the robotic arm for approaching the specified position in the Cartesian space; represents the attitude reward function, which is used to reward the end effector of the robotic arm for reaching the specified attitude; represents the collision reward function, which is used to punish the collision of the robotic arm with the environment; is the weight coefficient of the position reward; is the weight coefficient of the attitude reward; is the collision penalty coefficient; The position reward function , the attitude reward function and the collision reward function are defined as follows: , , , where is the target attitude, that is, the specified attitude that the robotic arm needs to reach; is the attitude of the end effector of the robotic arm at time t, that is, the current attitude information; represents the two-norm, represents the difference between two quaternions.
[0012] As a preferred solution, in the action instruction generation step, the action policy network generates the action instruction by calculating a policy function that maximizes the cumulative reward, and the policy function is defined as follows: , where, is the reward decay factor; is the policy function to be optimized; is the reward obtained by the action policy network at time t; is the expectation of all possible trajectories under the policy .
[0013] As a preferred solution, the update step includes an action policy network update step, and the action policy network update step includes: A policy gradient calculation step, which uses the policy gradient algorithm to calculate the gradient of the optimization metric with respect to the network parameters of the action policy network for the action policy network. The policy gradient calculation formula is defined as follows: , where, is the state; J is the optimization metric; is the gradient operator; is the reward decay factor; is the total number of samples; are the network parameters of the action policy network; are the network parameters of the evaluation network; is the policy gradient; represents the state value estimation; is the reward obtained at time t; An action policy network parameter update step, which uses the method of gradient descent to update the network parameters of the action policy network. The gradient calculation formula of the network parameters of the action policy network is as follows: , where, are the network parameters of the action policy network; are the network parameters of the action policy network .
[0014] As a preferred solution, the update step further includes an evaluation network update step, and the evaluation network update step includes: A loss function construction step, which defines the loss function of the value function of the evaluation network based on the temporal difference learning method: , where, is the state; is the loss function; The reward obtained at time t; The reward decay factor; Denote the state value estimation; The network parameters of the evaluation network; Value gradient calculation step, calculate the gradient of the loss function with respect to the network parameters of the evaluation network The gradient is defined as follows: , where, Is the gradient operator; Is the state; Denote the state value estimation; The reward obtained at time t; Evaluation network parameter update step, update the network parameters of the evaluation network according to the gradient, and the gradient calculation formula of the network parameters of the evaluation network is as follows: , where, Is the network parameter of the evaluation network The update step size.
[0015] According to the second aspect of the present invention, there is provided a motion planning device for a robotic arm, and the motion planning device includes the following steps: An information acquisition unit, configured to acquire the environmental information around the robotic arm and the motion information of the robotic arm through one or more sensors arranged on the robotic arm; A positioning unit, obtaining the current pose information of the robotic arm according to the acquired environmental information and motion information; A local map construction unit, configured to construct a local grid map including the obstacle information around the robotic arm and the current pose of the robotic arm according to the environmental information and the current pose information of the robotic arm; An action instruction generation unit, configured to input the current pose information, the local grid map, and the body perception information of the robotic arm into a pre-constructed action policy network, and generate an action instruction through the action policy network, and the action instruction is used to guide the action of the robotic arm; An update unit, configured to generate evaluation information for the action of the robotic arm through an evaluation network, and update the action policy network and the evaluation network according to the evaluation information.
[0016] According to the third aspect of the present invention, there is provided a non-transitory storage medium storing a computer program, and when the computer program is executed by a processor, it can implement the motion planning method for a robotic arm according to the first aspect of the present invention.
[0017] According to a fourth aspect of the present invention, there is provided a computer program product comprising computer instructions which, when executed by a processor, are capable of implementing the motion planning method of the robotic arm according to the first aspect of the present invention.
[0018] The beneficial effects of the present invention are as follows: By enabling the robot to interact with the surrounding environment in a simulation environment, the motion planning method of the present invention can real-time sense the environmental information around the robot and the motion information of the robot itself in a dynamic environment. Through the real-time collected sensing information, reinforcement learning is performed to train and update the action policy network, and the action policy network is deployed to run in the automatic charging system. According to the sensing information of the dynamic environment, the motion path of the robot is dynamically planned, so that the robot can safely and quickly reach the target pose and more accurately complete the charging task, providing a more efficient, convenient and safe charging service for electric vehicle users. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic diagram illustrating an exemplary structure of an automatic charging system for implementing the motion planning method of a robotic arm according to the present invention.
[0020] Figure 2 A flowchart illustrating the motion planning method of the robotic arm according to the present invention.
[0021] Figure 3 A flowchart illustrating the positioning step in the motion planning method of the robotic arm according to the present invention.
[0022] Figure 4 A schematic diagram illustrating an exemplary structure of the motion planning device of the robotic arm according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangements of components, numerical representations and numerical values described in these embodiments do not limit the scope of the present invention.
[0024] In the present invention, the term "unit" may refer to a software environment, a hardware environment, or a combination of a software and a hardware environment. In a software environment, the term "unit" refers to a functionality, an application, a software module, a function, a routine, a set of instructions, or a program that can be executed by a programmable processor, such as a microprocessor, a central processing unit (CPU), or a specially designed programmable device, or a controller. The memory contains instructions or programs that, when executed by the CPU, cause the CPU to perform operations corresponding to the unit or function. In a hardware environment, the term "unit" refers to a hardware element, a circuit, a component, a physical structure, a system, a module, or a subsystem. According to a specific embodiment, the term "unit" may include mechanical, optical, or electrical components, or any combination thereof. The term "unit" may include active (e.g., transistors) or passive (e.g., capacitors) components. The term "unit" may include semiconductor devices having a substrate and other material layers with various conductive concentrations. It may include a CPU or a programmable processor that can execute a program stored in the memory to perform a specified function. The term "unit" may include logic elements (e.g., AND, OR) implemented by transistor circuits or any other switching circuits. In a combination of a software and a hardware environment, the term "unit" or "circuit" refers to any combination of the software and hardware environments as described above. In addition, the terms "element", "component", "part", or "device" may also refer to a "circuit" integrated or not integrated with a packaging material.
[0025] Taking the motion planning and control of the robotic arm of a charging robot as an example, the motion planning method of the present invention is applied to control the hanging rail type robotic arm of an automatic charging system. With reference to the accompanying drawings, the system architecture of the automatic charging system of the present invention will be first described.
[0026] [Architecture of the Automatic Charging System of the Present Invention] The automatic charging system of the present invention includes a track network composed of a main track and branch tracks that covers each parking space in a parking lot and an agent component that can move on the track network, so as to ensure that the agent component can move to any designated parking space efficiently and safely. Among them, the main track can be in a closed state to run through the entire parking lot, and the branch tracks extending from the main track can lead to each parking space to allow the agent component to stay beside the parking space without affecting the passage of the main track.
[0027] The agent component in the automatic charging system of the present invention includes a charging pile and a robotic arm, both of which have the ability to move autonomously on the main track and the branch tracks. The charging pile can move to a specific parking space according to a scheduling instruction to provide power supply for an electric vehicle; the robotic arm is responsible for the plugging and unplugging operations of inserting and pulling out the charging gun from the charging pile and also has the ability to move to any parking space to ensure the automation and seamless docking of the charging process.
[0028] For an exemplary structure of the automatic charging system in the embodiments of the present invention, reference may be made to Figure 1 as shown. Figure 1 It shows the shape of the track of the automatic charging system and its positional relationship with the parking spaces in the parking lot. The automatic charging system of the present invention includes a main track 1 and a branch track 2, as well as a charging pile and a robotic arm that can move autonomously on these two tracks. As Figure 1 shown, the main track 1 presents a closed loop and is arranged to cover the entire parking lot. Branch tracks 2 extend from the main track 1, and charging spaces 3 for charging are distributed on both sides of each branch track. The branch tracks 2 allow the charging pile and the robotic arm to stay and operate at the parking spaces without affecting the smoothness of the main track 1.
[0029] Although Figure 1 the main track 1 shown in
[0030] presents a closed loop, the present invention is not limited thereto, and the shape of the main track 1 can also be other closed shapes such as circular or oval.
[0031] In addition, although in the following description, the main track and the branch track in the automatic charging system of the present invention adopt a preferred suspended type to facilitate passage and operation while saving space, the present invention is not limited thereto. According to the height and layout of the parking lot, the automatic charging system of the present invention can also be applied to a ground type track laid on the ground.
[0031] The charging operation process of the automatic charging system of the present invention will be described in detail below.
[0032] [Charging Operation Process of Automatic Charging System] After the electric vehicle docks at any one of the parking spaces in the automatic charging system, the user can send a charging request to the automatic charging system. The automatic charging system selects a suitable suspended track type charging pile and a suspended track type robotic arm. Specifically, after receiving the charging request from the electric vehicle, the automatic charging system selects the most suitable suspended track type charging pile and suspended track type robotic arm to undertake the charging task for this parking space. For example, it preferentially selects the suspended track type charging pile and suspended track type robotic arm that are the closest to this parking space and in an idle state. The present invention does not limit the selection method of the suspended track type charging pile and suspended track type robotic arm.
[0033] Furthermore, the automatic charging system moves the selected suspended track type charging pile and suspended track type robotic arm to the designated parking space respectively. Here, it should be understood that the present invention does not make a special limitation on the order in which the suspended track charging pile and the suspended track type robotic arm reach the designated parking space, and the two can reach successively or simultaneously.
[0034] After the hanging-rail charging pile and the hanging-rail robotic arm both move along the rail to the corresponding branch rail of the designated parking space, an action instruction for the current environment of the robotic arm is generated according to the motion planning method of the robotic arm described later, so that the hanging-rail robotic arm automatically grabs the charging gun on the hanging-rail charging pile according to the action instruction generated by the action strategy network, inserts it into the electric vehicle, and the hanging-rail charging pile starts to charge the electric vehicle. Next, until the charging of the charging pile is completed, the hanging-rail robotic arm moves along the rail to the current parking space, pulls out the charging gun and puts it back on the charging pile. Thus, the charging operation is completed, and the charging pile and the robotic arm leave the current parking space and perform other charging tasks.
[0035] In the above automatic charging system, for the steps of controlling the robotic arm to automatically grab the charging gun and insert it into the electric vehicle, and controlling the robotic arm to put the charging gun back into the charging pile again, it involves the motion planning and control of the robotic arm. The motion planning method adopted for the robotic arm control will be described in detail below.
[0036] [Motion Planning Method of Robotic Arm] [Design of Motion Planning Algorithm] The motion planning method of the robotic arm of the present invention can be applied to an automatic charging system, and is implemented by a processor in the automatic charging system (such as a charging pile) or the robotic arm executing a computer program stored on a memory in the automatic charging system or the robotic arm. As an alternative solution, it can also be implemented by the automatic charging system or the robotic arm communicating with a server, and the processor in the server executes a computer program stored in the server or the cloud and feeds back the program execution result to the automatic charging system in real time.
[0037] In order to meet the needs of the hanging-rail robotic arm to adapt to an uncertain dynamic environment and complete the charging task, the present invention discloses a motion planning method of the robotic arm based on real-time perception information, which can real-time perceive the environmental information around the robotic arm and the motion information of the robotic arm itself, generate dynamic action instructions according to the perceived environmental information and motion information, control the hanging-rail robotic arm to make obstacle avoidance actions, and achieve safe and fast arrival at the specified pose.
[0038] Reinforcement learning generally refers to an intelligent agent interacting with the environment, executing actions and receiving feedback (usually rewards or punishments), so as to learn the optimal strategy in a Markov process to maximize the cumulative reward, and the concepts involved include the state, action and reward of the intelligent agent.
[0039] For the convenience of understanding, some technical terms used in the present invention are explained as follows.
[0040] Agent: It refers to an entity that can perceive the environment and take actions to achieve specific goals. It can be software, hardware, or a system, with autonomy, adaptability, and interaction capabilities. An agent perceives changes in the environment (such as through sensors or data input), makes judgments and decisions based on the knowledge and algorithms it has learned, and then executes actions to affect the environment or reach a predetermined goal.
[0041] State of the agent: When the robotic arm moves, its states include the current position of the end effector , pose ; the position of the target pose , pose ; the angles of each joint of the current robotic arm ; the current local grid map . The position of the end effector is represented by coordinates in the world coordinate system, and the pose is represented by quaternions in the world coordinate system. The local grid map is the grid state quantity around the current position , represented by a three-dimensional tensor. It decomposes the environment into a series of discrete grids, each grid has a value, and the grid contains two basic types of information: coordinates and whether it is an obstacle. The environmental information is represented by the probability value of each grid being occupied, where the value of each element is one of -1, 0, 1. -1 represents an unknown state, 0 represents unoccupied, and 1 represents occupied.
[0042] Action of the agent: The decision made by the agent that affects its interaction with the environment. For example, the increments of each joint of the robotic arm in the embodiments of the present invention .
[0043] Reward of the agent: The reward obtained by the robotic arm through interaction with the environment, which reflects the desirability of the actions performed in a specific state.
[0044] Value function of the agent: A function that estimates the expected cumulative reward that the agent can obtain, starting from a given state and following a specific policy.
[0045] The agent involved in the present invention is a suspended rail-type robotic arm, which independently interacts with the environment in the environment of the automatic charging system created thereon, improves its own policy using the rewards feedback from the environment to obtain a higher cumulative reward. And through a pre-set reward function, continuously improves the agent's policy using the rewards feedback from the environment.
[0046] In the present invention, the motion planning method of the robotic arm mainly includes an information acquisition step, a positioning step, a local map construction step, an action instruction generation step, and an update step. The following refers toFigure 2 The motion planning method of the robotic arm of the present invention will be described.
[0047] As Figure 2 shown, first, in the information acquisition step S100, environmental information around the robotic arm and motion information of the robotic arm are acquired through one or more sensors provided on the robotic arm.
[0048] The sensors in the embodiments of the present invention may include visual sensors. For example, a target image around the robotic arm is acquired using a camera or other image acquisition devices with image acquisition functions as environmental information. The target image contains obstacle information in the surrounding environment, the position of the charging gun, and the position of the charging interface of the electric vehicle. The sensors may also include an Inertial Measurement Unit (IMU) for measuring parameters such as the acceleration, angular velocity, and tilt angle of the robotic arm as motion information of the robotic arm. In addition to using visual sensors to perceive environmental information, other sensing devices such as lidar may also be used to acquire environmental information.
[0049] Before performing the motion planning method of the robotic arm, first, a simulation environment of the charging robot in a complex environment is built, and a physical simulation model of the charging robot (six-degree-of-freedom robotic arm) is created. A visual sensor for perceiving environmental information is installed on the end effector of the robotic arm, and the internal parameters of the camera are set according to the actual device. An IMU is installed on the camera, and the external parameters between the camera and the IMU are set according to the actual device. In the simulation environment, obstacles of different shapes such as rectangles and circles are randomly set to simulate obstacles in the actual complex environment. After building the simulation environment, the information acquisition step S100 is performed to acquire the environmental information around the robotic arm and the motion information of the robotic arm.
[0050] Furthermore, the positioning step S200 is performed to obtain the current pose information of the robotic arm based on the acquired environmental information and motion information.
[0051] Step S200 is a visual positioning process of the end effector of the robotic arm: Based on the image information collected by the visual sensor and the motion information collected by the IMU, a Visual Inertial Odometry is built. The Visual Inertial Odometry is used to estimate the pose of the end effector in the world coordinate system to facilitate the creation of a local map in combination with depth images in the later stage.
[0052] Simultaneous localization and mapping (SLAM) technology is often used for the autonomous navigation problem of robots. The positioning and mapping data provided by the SLAM module help the robot obtain its own state estimation and obstacle information of the environment. Visual-inertial odometry (VIO) is one of the SLAM technologies, which estimates the state of the robot by fusing visual information with an inertial measurement unit. After the pose of the camera is known, the local map can be constructed through depth information.
[0053] Local map construction step S300: Construct a local grid map containing the obstacle information around the robotic arm and the current pose of the robotic arm according to the environmental information and the current pose information of the robotic arm.
[0054] Step S300 is used to construct a local grid map. In this embodiment, the visual sensor is a binocular camera. The binocular camera obtains a depth image in the current frame. According to the depth information on the two-dimensional pixel points, the points on an entire ray are projected into three-dimensional space to update the local grid map. The local grid map can be used for the motion trajectory planning in the next step.
[0055] Action instruction generation step S400: Input the current pose information, the local grid map, and the proprioceptive information of the robotic arm into a pre-constructed action policy network, and generate an action instruction through the action policy network. The action instruction is used to guide the action of the robotic arm.
[0056] Step S400 is a motion planning algorithm based on reinforcement learning. According to the current pose information of the robotic arm (i.e., the end effector pose) obtained in step S200, the local grid map obtained in step S300, and the proprioceptive information of the robotic arm itself (such as the angles of each joint of the robotic arm), reinforcement learning outputs the current action policy through a neural network model. In reinforcement learning, the neural network model of the present invention adopts an Actor-Critic structure. The Actor is the action policy network, which is used to receive environmental information and output the corresponding action variables, such as outputting the increments of each joint of the robotic arm.
[0057] Update step S500: Generate evaluation information for the action of the robotic arm through an evaluation network, and update the action policy network and the evaluation network according to the evaluation information.
[0058] Specifically, the evaluation network is the Critic network, which is used to output an action instruction in step S400. After guiding the robotic arm to complete the corresponding action, it determines the quality of the action executed by the robotic arm according to the current state of the robotic arm, that is, generates evaluation information, and guides the Actor-Critic structure to update the network parameters according to the evaluation information, especially helps the Actor network to update the strategy, so that the Actor network continuously optimizes the action strategy.
[0059] In summary, the motion planning method of the present invention enables the robot to interact with the surrounding environment in a simulation environment, real-time sense the environmental information around the robot and the motion information of the robot itself in a dynamic environment, and through the real-time collected perception information for reinforcement learning, trains and updates the action strategy network, and deploys the action strategy network to run in the automatic charging system, dynamically plans the motion path of the robot according to the perception information of the dynamic environment, enables the robot to reach the target pose safely and quickly, and more accurately completes the charging task, providing a more efficient, convenient and safe charging service for electric vehicle users.
[0060] The following will describe each step in the motion planning method of the present invention in detail.
[0061] As an optional implementation manner, in the information acquisition step S100, one or more sensors include an image sensor and an inertial measurement unit; the information acquisition step includes: Collect the image information around the robotic arm as environmental information through the image sensor, and collect the inertial data of the robotic arm as motion information through the inertial measurement unit.
[0062] As described above, the present invention can collect image information through a vision sensor and collect motion information through an IMU. Compared with traditional positioning methods, the advantage of visual positioning lies in high precision and high stability. Visual positioning can real-time sense environmental changes and quickly and accurately locate the target. This characteristic enables the robot to autonomously navigate and avoid obstacles in a complex environment, greatly improving the safety and reliability of the operation.
[0063] In the present invention, the positioning step S200 further includes a preprocessing step, an initialization step, a local optimization step, a loop detection step, and a global optimization step. The following will refer to Figure 3 Describe the positioning step S200 of the present invention.
[0064] As Figure 3 shown, in the preprocessing step S201, feature extraction is performed on the image information to obtain feature point tracking data, and pre-integration is performed on the motion information to obtain the pose, velocity, and rotation angle at the current moment. At the same time, the pre-integration increment, the covariance matrix of the pre-integration, and the Jacobian matrix between adjacent frames are calculated as pre-integration data.
[0065] Step S201 preprocesses the image data and IMU data: For the image, feature points are extracted and optical flow tracking is performed using the KLT pyramid to prepare for solving the camera pose. For the IMU data, the IMU data is pre-integrated to obtain the pose, velocity, and rotation angle at the current moment, and at the same time, the pre-integration increment between adjacent frames, as well as the covariance matrix and Jacobian matrix of the pre-integration, are calculated.
[0066] In the initialization step S202, the relative pose between adjacent frames is calculated based on the feature point tracking data, and the relative pose is aligned with the pre-integration data to obtain the initial pose in the world coordinate system.
[0067] In step S202, first, only vision-based initialization is performed to solve the relative pose of the camera, and then it is aligned with the IMU pre-integration to solve the initialization parameters.
[0068] In the local optimization step S203, the initial pose is locally optimized using the sliding window optimization method to obtain the local pose.
[0069] Specifically, local non-linear optimization optimizes the visual constraints and IMU constraints in a common objective function. Here, local optimization only optimizes the variables in the window of the current frame and the previous n frames. The local non-linear optimization outputs a more accurate local pose.
[0070] In the loop detection step S204, the current frame is matched with the historical key frames, and the frames that match the historical key frames are selected as loop constraints according to the matching relationship.
[0071] Loop detection saves the previously detected image key frames. When returning to the same place passed before, it is judged whether this frame has been visited through the matching relationship of feature points. The key frames mentioned above are the selected camera frames that can be recorded while avoiding redundancy. Among them, the selection criterion for key frames is that the displacement between the current frame and the previous frame exceeds a certain threshold or the number of matching feature points is less than a certain threshold.
[0072] In the global optimization step S205, the local pose is globally optimized according to the loop constraints to obtain the global pose.
[0073] Global optimization is to perform non-linear optimization using visual constraints, IMU constraints, and the loop constraints obtained from loop detection when loop detection occurs, and output a more accurate global pose based on local optimization.
[0074] Furthermore, the local map construction step S300 includes: The depth image of the current frame obtained by the image acquisition device. Among them, the image sensor can adopt a binocular camera. The binocular camera is similar to the human eyes and relies on two captured pictures (color RGB or grayscale images) to calculate the depth. Specifically, first, it is necessary to calibrate the binocular camera to obtain the internal and external parameters of the two cameras and the homography matrix; then, correct the original images according to the calibration results. The two corrected images are on the same plane and parallel to each other; perform pixel point matching on the corrected two images; calculate the depth of each pixel according to the matching results to obtain the depth image.
[0075] According to the depth information of each two-dimensional pixel point in the depth image, project the two-dimensional pixel points into the three-dimensional space to obtain a local grid map; among them, the projection formula is defined as follows: , (1) , (2) Among them, is the internal parameter of the binocular camera; u is the abscissa of the pixel; v is the ordinate of the pixel, Z is the depth of each pixel point obtained by the depth image; X and Y respectively represent the abscissa and ordinate after projection into the three-dimensional space.
[0076] When the robot enters a new environment, it does not know the indoor obstacle information. This requires the robot to be able to traverse the entire environment, detect the position of the obstacle, find the corresponding serial number value in the grid map according to the obstacle position, and modify the corresponding grid value. The free grid assigns a value of 0 to the grid that does not contain an obstacle, and the obstacle grid assigns a value of 1 to the grid that contains an obstacle. Thus, update the obstacle information in the local grid map.
[0077] The embodiment of the present invention utilizes the visual information real-time sensed by the image acquisition device, fully extracts effective obstacle information, and can also achieve autonomous obstacle avoidance planning by constructing a map in an unknown environment. Through the random initialization of the environment in the simulation, the policy network has generalization in the actual scenario.
[0078] [Reward function construction] In order to enable the robotic arm to quickly, flexibly and safely reach the specified pose to complete the charging task, it is necessary to set a reasonable reward function for reinforcement learning. In the update step S500, the evaluation network rewards the actions of the robotic arm based on the reward function as evaluation information. The reward function is defined as follows: , (3) Among them, represents the position reward function, which is used to reward that the end effector of the robotic arm approaches the specified position in the Cartesian space; Represents the pose reward function, which is used to reward the end effector of the robotic arm for reaching the specified pose; Represents the collision reward function, which is used to penalize the collision between the robotic arm and the environment; Is the weight coefficient of the position reward; Is the weight coefficient of the pose reward; Is the collision penalty coefficient; Position reward function , pose reward function And the collision reward function Are defined as follows: , (4) , (5) , (6) Among them, Is the target pose, that is, the specified pose that the robotic arm needs to reach; Is the pose of the end effector of the robotic arm at time t, that is, the current pose information; Represents the two-norm, Represents the difference between two quaternions.
[0079] The present invention constructs a reward function based on multiple dimensions of position, pose, and collision, evaluates the actions of the robotic arm in the current environment from different dimensions, comprehensively evaluates the current actions, and gives corresponding rewards, so as to continuously optimize the action strategy of the policy network. The reward function (Reward Function) is an immediate feedback signal given by the environment to the agent after performing a certain action in the current state. It is a scalar value that directly quantifies the "good or bad" of the current action and guides the agent to learn and make decisions towards the expected goal. By accumulating these immediate rewards, the agent can optimize its strategy.
[0080] [Policy function construction] In the action instruction generation step S400, the action policy network generates action instructions by calculating the policy function that maximizes the cumulative reward. The policy function is defined as follows: , (7) Among them, Is the reward decay factor; Is the policy function to be optimized; Is the reward obtained by the action policy network at time t; Is the expectation of all possible trajectories under the policy Below.
[0081] Specifically, the policy function that maximizes the cumulative reward can be regarded as maximizing the expected value of the value function. The policy It is the strategy that maximizes this expected value. By adjusting the strategy parameters and maximizing the cumulative reward, the optimal strategy is found. , making the agent more inclined to choose actions that can bring higher cumulative rewards.
[0082] In this embodiment, the neural network model of reinforcement learning includes two parts: the Actor network and the Critic network. In reinforcement learning, the Actor network is , indicating that it outputs action a in state s, is the network parameters of. The Actor network needs to interact with the environment and learn a better strategy using policy gradients under the guidance of the value function of the Critic network. The Critic network is expressed as , indicating the value function in state s, representing the expected cumulative reward starting from state s and following policy π. What the Critic network does is to learn a value function through the data collected by the Actor network interacting with the environment. This value function will be used to judge what actions are good and what actions are not good in the current state, thereby helping the Actor network to update the strategy. The value function of the Critic network is the expected value of the agent's long-term cumulative reward, helping the agent to plan long-term and weigh immediate rewards against future benefits.
[0083] As mentioned above, a reward function is introduced in the reinforcement learning of the present invention. The reward function is the immediate feedback of the environment on the quality of actions. The policy function in the Actor network determines the action policy based on the rewards calculated by the above reward function, so as to output the optimal action policy. The reward function is the driving force for the policy function to optimize the policy.
[0084] [Update of Action Policy Network] The update step S500 includes the action policy network update step, that is, updating the network parameters of the Actor network. The action policy network update step includes: Policy gradient calculation step: Using the policy gradient algorithm to calculate the gradient of the optimization metric with respect to the network parameters of the action policy network for the action policy network. The policy gradient calculation formula is defined as follows: , (8) where, is the state; J is the optimization metric; is the gradient operator; is the reward decay factor; is the total number of samples; are the network parameters of the action policy network; are the network parameters of the evaluation network; is the policy gradient; Represents the state value estimation; Is the reward obtained at time t.
[0085] The update step of the action policy network parameters, using the gradient descent method to update the network parameters of the action policy network. The gradient calculation formula of the network parameters of the action policy network is as follows: , (9) Among them, Are the network parameters of the action policy network; Are the network parameters of the action policy network Is the update step size of
[0086] After calculating through the policy gradient calculation formula, according to Update the network parameters of the action policy network . In the present invention, the update of the Actor network adopts the principle of Policy Gradient. The policy gradient algorithm introduces the temporal difference error to guide the update of the Actor network gradient, which can play a role in reducing variance and improving robustness.
[0087] [Evaluation network update] The update step S500 further includes an evaluation network update step, that is, updating the network parameters of the Critic network. The evaluation network update step includes: Loss function construction step, defining the loss function of the value function of the evaluation network based on the temporal difference learning method: , (10) Among them, s is the state; Is the loss function; Is the reward obtained at time t; Is the reward decay factor; Represents the state value estimation; Are the network parameters of the evaluation network; Value gradient calculation step, calculating the gradient of the loss function with respect to the network parameters of the evaluation network. The value gradient calculation formula is defined as follows: , (11) Among them, Is the gradient operator; s Is the state; Represents the state value estimation; Is the reward obtained at time t; Evaluation network parameter update step, updating the network parameters of the evaluation network according to the gradient. The gradient calculation formula of the network parameters of the evaluation network is as follows: , (12) Among them, is the network parameter of the evaluation network and is the update step size.
[0088] In terms of the structure of the neural network model, the present invention is divided into three types: a data preprocessing network, an evaluation network, and a policy network. Among them, the evaluation network belongs to the Critic network part, the policy network belongs to the Actor network part, and the data preprocessing network is shared by the Actor network and the Critic network. The data preprocessing network is composed of a convolutional neural network (Convolutional Neural Networks, abbreviated as CNN) and a multilayer perceptron network (Multilayer Perceptron, MLP). Among them, the convolutional neural network processes the local grid map, extracts the features of the downsampled grid map in the form of a three-dimensional tensor through the CNN, and the multilayer perceptron network processes other states. The outputs of the two are fused into a vector and respectively input into the Critic network and the Actor network. The Critic network outputs the value function, and the Actor network outputs the action policy. During training, in order to prevent the parameter update from being unstable due to the backpropagation of the two parts, the parameters of the preprocessing network are only updated with the gradient of the Critic network, and the gradient of the Actor network only updates the policy network.
[0089] When collecting data, reinforcement learning will use the current Actor network to interact with the environment, and then train after obtaining N trajectories of the agent. The present invention sets a turn-based system to collect motion trajectories, and there are three conditions for the end of a turn: the interaction time of the agent reaches the maximum value; the agent collides with the environment; the agent reaches the specified pose.
[0090] The motion planning method of the present invention obtains real-time perceived environmental information and motion information, and the robot in the simulation environment conducts a large number of interactions with the environment. Reinforcement learning can train a better policy network, deploy the policy network to the actual system for operation, set the specified pose, and the robot will reach safely and quickly to complete the charging task.
[0091] The robotic arm of the present invention performs real-time optimization control through a neural network model, and changes its own control strategy in real time according to the change of visual information, and can achieve adaptive adjustment to the environment. Even if the environment changes, the robotic arm can dynamically adjust its own action strategy. Compared with the existing static path planning, it has stronger practicability and flexibility. In addition, the neural network model only needs to perform forward calculation, and the real-time performance of the algorithm is better.
[0092] It should be noted that applying this motion planning method to the motion planning of the robotic arm of a charging robot is just one of the implementation manners, and it can also be applied to other types of robots, such as cleaning robots, restaurant service robots, hotel service robots, medical robots, etc. The motion planning method of the present invention does not limit the application field of the robot.
[0093] In addition, the present invention also provides a motion planning device for a robotic arm.
[0094] [Motion Planning Device for Robotic Arm of the Present Invention] As Figure 4 shown, the motion planning device 400 of the present invention includes an information acquisition unit 401, a positioning unit 402, a local map construction unit 403, an action instruction generation unit 404, and an update unit 405. Each unit will be described in detail below.
[0095] The information acquisition unit 401 is configured to acquire the environmental information around the robotic arm and the motion information of the robotic arm through one or more sensors disposed on the robotic arm.
[0096] The sensors in the embodiments of the present invention may include visual sensors. For example, a target image around the robotic arm is acquired using a camera or other image acquisition devices with image acquisition functions as environmental information. The target image includes obstacle information in the surrounding environment, the position of the charging gun, and the position of the charging interface of the electric vehicle. The sensor may also include an IMU for measuring parameters such as the acceleration, angular velocity, and tilt angle of the robotic arm as the motion information of the robotic arm. In addition to using visual sensors to perceive environmental information, other sensing devices such as lidar may also be used to acquire environmental information.
[0097] Before executing the motion planning method of the robotic arm, first build a simulation environment of the charging robot in a complex environment and create a physical simulation model of the charging robot (a six-degree-of-freedom robotic arm). A visual sensor for perceiving environmental information is installed on the end effector of the robotic arm, and the internal parameters of the camera are set according to the actual device. An IMU is installed on the camera, and the external parameters between the camera and the IMU are set according to the actual device. In the simulation environment, obstacles of different shapes such as rectangles and circles are randomly set to simulate the obstacles in the actual complex environment. After building the simulation environment, the information acquisition unit 401 acquires the environmental information around the robotic arm and the motion information of the robotic arm.
[0098] The positioning unit 402 obtains the current pose information of the robotic arm according to the acquired environmental information and motion information.
[0099] In this embodiment, the positioning unit 402 uses a visual positioning method to position the end effector of the robotic arm: Based on the image information collected by the visual sensor and the motion information collected by the IMU, a visual inertial odometer is built. The visual inertial odometer is used to estimate the pose of the end effector in the world coordinate system, so as to create a local map in combination with the depth image in the later stage. The visual inertial odometer (VIO) is a type of SLAM technology, which estimates the state of the robot by fusing visual information with the inertial measurement unit. After knowing the pose of the camera, the local map can be constructed through the depth information.
[0100] The local map construction unit 403 is configured to construct a local grid map including the obstacle information around the robotic arm and the current pose information of the robotic arm according to the environmental information and the current pose information of the robotic arm.
[0101] Specifically, the local map construction unit 403 is used to construct a local grid map. The visual sensor in this embodiment is a binocular camera. The binocular camera obtains a depth image in the current frame. According to the depth information on the two-dimensional pixel points, the points on an entire ray are projected into the three-dimensional space to update the local grid map, and the local grid map can be used for the motion trajectory planning in the next step.
[0102] The action instruction generation unit 404 is configured to input the current pose information, the local grid map, and the self-perception information of the robotic arm into a pre-constructed action policy network, and generate an action instruction through the action policy network. The action instruction is used to guide the action of the robotic arm.
[0103] Based on the motion planning algorithm of reinforcement learning, the action instruction generation unit 404, according to the current pose information of the robotic arm (i.e., the pose of the end effector) calculated by the positioning unit 402, the local grid map constructed by the local map construction unit 403, and the self-perception information of the robotic arm itself (such as the angles of each joint of the robotic arm), reinforcement learning outputs the current action policy through a neural network model. In reinforcement learning, the neural network model of the present invention adopts an Actor-Critic structure. The Actor is the action policy network, which is used to receive environmental information and output corresponding action variables, such as outputting the increments of each joint of the robotic arm.
[0104] The update unit 405 is configured to generate evaluation information for the action of the robotic arm through the evaluation network, and update the action policy network and the evaluation network according to the evaluation information.
[0105] Specifically, the evaluation network is the Critic network, which is used to generate an evaluation message according to the current state of the robotic arm after the action instruction generation unit 404 outputs an action instruction to guide the robotic arm to complete the corresponding action, that is, to generate evaluation information, and guide the Actor-Critic structure to update network parameters according to the evaluation information. In particular, it helps the Actor network to update the policy, so that the Actor network continuously optimizes the action policy.
[0106] In summary, the motion planning device of the present invention enables the robot to interact with the surrounding environment in a simulation environment, real-time sense the environmental information around the robot and the motion information of the robot itself in a dynamic environment, train and update the action policy network through reinforcement learning based on these real-time acquired perception information, and deploy the action policy network to run in the automatic charging system, and dynamically plan the motion path of the robot according to the perception information of the dynamic environment, so that the robot can safely and quickly reach the target pose and more accurately complete the charging task, providing a more efficient, convenient and safe charging service for electric vehicle users.
[0107] [Other Embodiments] Embodiments of the present invention can also be implemented by a computer of a system or device that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium (which can also be more completely referred to as a "non-transitory computer-readable storage medium") to perform one or more of the functions of the above embodiments and / or includes one or more circuits (e.g., an application specific integrated circuit (ASIC)) for performing one or more of the functions of the above embodiments. Moreover, embodiments of the present invention can be implemented by a method that, by the computer of the system or device, for example, reads and executes the computer-executable instructions from the storage medium to perform one or more of the functions of the above embodiments and / or controls the one or more circuits to perform one or more of the functions of the above embodiments. The computer can include one or more processors (e.g., a central processing unit (CPU), a microprocessing unit (MPU)), and can include a network of separate computers or separate processors to read and execute the computer-executable instructions. The computer-executable instructions can be provided to the computer, for example, from a network or the storage medium. The storage medium can include, for example, one or more of a hard disk, a random access memory (RAM), a read-only memory (ROM), a memory of a distributed computing system, an optical disc (such as a compact disc (CD), a digital versatile disc (DVD) or a Blu-ray Disc (BD)™), a flash device, and a memory card, etc.
[0108] Although the present invention has been described above with reference to exemplary embodiments, the above embodiments are only for explaining the technical concept and features of the present invention and should not be used to limit the protection scope of the present invention. Any equivalent variation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A motion planning method for a robotic arm, characterized in that, The described motion planning method includes the following steps: An information acquisition step, where environmental information around the robotic arm and motion information of the robotic arm are acquired through one or more sensors set on the robotic arm; A positioning step, where the current pose information of the robotic arm is obtained based on the acquired environmental information and motion information; A local map construction step, where a local grid map containing obstacle information around the robotic arm and the current pose of the robotic arm is constructed based on the environmental information and the current pose information of the robotic arm; An action instruction generation step, where the current pose information, the local grid map, and the proprioceptive information of the robotic arm are input into a pre-constructed action policy network, and an action instruction is generated through the action policy network, and the action instruction is used to guide the action of the robotic arm; An update step, where evaluation information for the action of the robotic arm is generated by an evaluation network, and the action policy network and the evaluation network are updated according to the evaluation information.
2. The motion planning method according to claim 1, wherein In the information acquisition step, the one or more sensors include an image sensor and an inertial measurement unit; the information acquisition step includes: Collecting image information around the robotic arm as the environmental information through the image sensor, and collecting inertial data of the robotic arm as the motion information through the inertial measurement unit.
3. The motion planning method according to claim 2, wherein The positioning step includes the following: A preprocessing step, where feature extraction is performed on the image information to obtain feature point tracking data, and pre-integration is performed on the motion information to obtain the pose, velocity, and rotation angle at the current moment, and at the same time, the pre-integration increment, the covariance matrix of the pre-integration, and the Jacobian matrix between adjacent frames are calculated as pre-integration data; An initialization step, where the relative pose between adjacent frames is calculated based on the feature point tracking data, and the relative pose is aligned with the pre-integration data to obtain the initial pose in the world coordinate; A local optimization step, where the initial pose is locally optimized using a sliding window optimization method to obtain the local pose; A loop detection step, where the current frame is matched with historical key frames, and frames that match the historical key frames are selected according to the matching relationship as loop constraints; A global optimization step, where the local pose is globally optimized according to the loop constraints to obtain the global pose.
4. The motion planning method according to claim 2, characterized in that The local map construction step includes: The depth image of the current frame obtained by the image acquisition device; According to the depth information of each two-dimensional pixel point in the depth image, the two-dimensional pixel points are projected into three-dimensional space to obtain a local grid map; where the projection formula is defined as follows: , , Among them, is the camera internal parameter; u is the horizontal pixel coordinate; v is the vertical pixel coordinate, Z is the depth of each pixel obtained from the depth image; X and Y respectively represent the horizontal and vertical coordinates after projection into the three-dimensional space.
5. The motion planning method according to claim 1, characterized in that, In the update step, the evaluation network rewards the action of the robotic arm based on a reward function as the evaluation information, and the reward function is defined as follows: , Among them, represents the position reward function, which is used to reward the end effector of the robotic arm for approaching a specified position in Cartesian space; represents the attitude reward function, which is used to reward the end effector of the robotic arm for reaching a specified attitude; represents the collision reward function, which is used to penalize the collision between the robotic arm and the environment; is the weight coefficient of the position reward; is the weight coefficient of the attitude reward; is the collision penalty coefficient; The position reward function 、the attitude reward function and the collision reward function are defined as follows: , , , wherein, is the target pose, that is, the specified pose that the robotic arm needs to reach; is the pose of the end effector of the robotic arm at time t, that is, the current pose information; represents the two-norm, represents the difference between two quaternions.
6. The motion planning method according to claim 1, wherein In the action instruction generation step, the action policy network generates the action instruction by calculating a policy function that maximizes the cumulative reward, and the policy function is defined as follows: , Among them, is the reward decay factor; is the policy function to be optimized; is the reward obtained by the action policy network at time t; is the expectation of all possible trajectories under the policy below.
7. The motion planning method according to claim 1, wherein The update step includes an action policy network update step, and the action policy network update step includes: Steps for calculating the policy gradient: Using the policy gradient algorithm, calculate the gradient of the optimization metric with respect to the network parameters of the action policy network. The formula for the policy gradient is defined as follows: , Among them, s is the state; J is the optimization index; is the gradient operator; is the reward decay factor; is the total number of samples; are the network parameters of the action policy network; are the network parameters of the evaluation network; is the policy gradient; represents the state value estimation; is the reward obtained at time t; Steps for updating the network parameters of the action policy network: Use the gradient descent method to update the network parameters of the action policy network. The formula for calculating the gradient of the network parameters of the action policy network is as follows: , Among them, are the network parameters of the action policy network; are the network parameters of the action policy network is the update step size.
8. The motion planning method according to claim 1, characterized in that The update step further includes an evaluation network update step, which includes: Steps for constructing the loss function: Based on the temporal difference learning method, define the loss function of the value function of the evaluation network. , wherein, s is the state; is the loss function; is the reward obtained at time t; is the reward decay factor; represents the state value estimation; are the network parameters of the evaluation network; Value gradient calculation step, calculating the gradient of the loss function with respect to the network parameters of the evaluation network The value gradient calculation formula is defined as follows: , Among them, is the gradient operator; s is the state; represents the state value estimation; is the reward obtained at time t; Steps for updating the network parameters of the evaluation network: Update the network parameters of the evaluation network according to the gradient. The formula for calculating the gradient of the network parameters of the evaluation network is as follows: , Among them, are the network parameters of the evaluation network is the update step size.
9. A motion planning device for a robotic arm, characterized in that, The motion planning device includes the following steps: An information acquisition unit, configured to acquire the environmental information around the robotic arm and the motion information of the robotic arm through one or more sensors provided on the robotic arm. A positioning unit, configured to obtain the current pose information of the robotic arm according to the acquired environmental information and motion information. A local map construction unit, configured to construct a local grid map including the obstacle information around the robotic arm and the current pose of the robotic arm according to the environmental information and the current pose information of the robotic arm. An action instruction generation unit, configured to input the current pose information, the local grid map, and the body perception information of the robotic arm into a pre-constructed action policy network, and generate an action instruction through the action policy network, where the action instruction is used to guide the action of the robotic arm. An update unit, configured to generate evaluation information for the action of the robotic arm through an evaluation network, and update the action policy network and the evaluation network according to the evaluation information.
10. A non-transitory storage medium storing a computer program, which when executed by a processor, is capable of implementing the motion planning method of the robotic arm according to any one of claims 1-8.
11. A computer program product including computer instructions, which when executed by a processor, is capable of implementing the motion planning method of the robotic arm according to any one of claims 1-8.
Citation Information
Patent Citations
Seven-axis mechanical arm autonomous obstacle avoidance motion control method based on improved SAC algorithm
CN117697755A
Path planning method and device for charging robot, storage medium and program product
CN119806166A
Optimization method for whole-body motion planning of mobile mechanical arm
CN119871459A