Robot hierarchical reinforcement learning control system based on target guidance
Through a layered reinforcement learning control system, visual pose prediction and depth model are used for real-time detection and control, the problem of insufficient flexibility and adaptability of robots in complex environments is solved, and efficient robot operation is achieved.
Patent Information
- Application Number
- CN202510404626.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
When facing complex and dynamic environments, existing robot systems lack flexibility and adaptability, making it difficult to effectively deal with multi-step operational tasks in unstructured environments. The existing reinforcement learning methods are cost-effective in high-dimensional spaces, slow convergence speed, and lack inter-level coordination, resulting in limited system performance.
A hierarchical reinforcement learning control system based on goal guidance, including visual pose prediction module, upper-level decision-making subsystem, middle-level planning subsystem and lower-level control subsystem are adopted. Real-time detection and tracking are carried out through the improved YOLOv8 algorithm, combined with deep kinematics and dynamic models, and precise motion control is achieved using impedance control.
It significantly reduces the computational complexity, improves the overall performance of the robot in complex tasks, enhances the scalability and adaptability of the system, and realizes the efficient operation of the robot in an unstructured environment.
Smart Images

Figure CN120326599A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot control, and particularly to a hierarchical reinforcement learning control system for robots based on target guidance. Background Art
[0002] With the rapid development of artificial intelligence and robot technology, robots are increasingly widely used in industrial, logistics, service and other fields. However, existing robot systems still face many challenges when dealing with complex and dynamic environments. Most traditional robot control methods rely on pre-programmed paths and fixed operation processes, lacking flexibility and adaptability, and are difficult to handle multi-step operation tasks in unstructured environments. This results in low efficiency of robots in practical applications, poor operation accuracy and stability, and it is difficult to meet the rapidly changing task requirements.
[0003] As a machine learning method that can self-optimize, reinforcement learning shows great potential in robot control. However, when directly applied to complex tasks in high-dimensional spaces, it has high computational costs, slow convergence speed, and lacks practical operability. In addition, existing reinforcement learning methods often have difficulty achieving effective coordination between different levels when dealing with multi-level task decomposition and coordination, resulting in limited overall system performance.
[0004] In view of this, the present invention is specifically proposed. Summary of the Invention
[0005] The purpose of the present invention is to provide a hierarchical reinforcement learning control system for robots based on target guidance, thereby solving the above technical problems existing in the prior art.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] A hierarchical reinforcement learning control system for robots based on target guidance, comprising:
[0008] A visual pose prediction module, an upper-level decision-making subsystem, a middle-level planning subsystem, and a lower-level control subsystem; wherein,
[0009] The visual pose prediction module can detect and track the pose of target objects in the environment in real time through an improved YOLOv8 algorithm to obtain real-time environmental data;
[0010] The upper-level decision-making subsystem is communicatively connected to the visual pose prediction module, and can evaluate the task progress according to the upper-level decision-making target and the current state of the obtained real-time environmental data through the upper-level decision-making target guidance strategy, and select the corresponding middle-level operation target according to the task progress;
[0011] The middle - layer planning subsystem is communicatively connected to the upper - layer decision - making subsystem. It can generate specific robot planning actions according to the current state and the middle - layer operation objectives through the planning - objective guiding strategy, and transmit the expected motion trajectory of the planning action to the lower - layer control subsystem.
[0012] The lower - layer control subsystem is communicatively connected to the middle - layer planning subsystem. It can combine the corresponding depth kinematic model and depth dynamic model of the robot. After performing trajectory interpolation on the robot motion trajectory output by the middle - layer planning subsystem, it combines impedance control to control the actual motion trajectory of the robot to accurately execute the operation actions.
[0013] Compared with the prior art, the robot hierarchical reinforcement learning control system based on goal - guiding provided by the present invention has the following beneficial effects:
[0014] By setting up the upper - layer decision - making subsystem, the middle - layer planning subsystem and the lower - layer control subsystem, and controlling in a hierarchical manner, the computational complexity of each module is significantly reduced, avoiding the difficulties of directly modeling and solving in a high - dimensional space. The interface definition between each level is clear, which is conducive to the design and optimization of each module, enabling each level to focus on dealing with specific problems, thereby improving the overall performance of the robot when performing complex tasks. In addition, this framework has good scalability. The vision module, decision - making module and control module can be replaced with different algorithms as long as they comply with the unified interface. Brief Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a schematic diagram of the composition of the robot hierarchical reinforcement learning control system based on goal - guiding provided by the embodiments of the present invention.
[0017] Figure 2 It is a schematic diagram of the composition of the upper - layer decision - making subsystem of the robot hierarchical reinforcement learning control system based on goal - guiding provided by the embodiments of the present invention.
[0018] Figure 3 It is a schematic diagram of the cable - driven flexible robot of the robot hierarchical reinforcement learning control system based on goal - guiding provided by the embodiments of the present invention. Detailed Embodiments
[0019] Next, in combination with the specific content of the present invention, the technical solutions in the embodiments of the present invention will be described clearly and completely; obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments, which does not constitute a limitation to the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the protection scope of the present invention.
[0020] First, the following explanations will be made for the terms that may be used in this article:
[0021] The term "and / or" means that either or both of the two can be achieved. For example, X and / or Y means that it includes both the case of "X" or "Y" and the three cases of "X and Y".
[0022] The description of terms such as "comprising", "including", "containing", "having" or other similar semantics should be interpreted as non-exclusive inclusion. For example: including a certain technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction condition, processing condition, parameter, algorithm, signal, data, product or article, etc.) should be interpreted as not only including the clearly listed certain technical feature element, but also including other technical feature elements well-known in the art that are not clearly listed.
[0023] The term "consisting of" means excluding any technical feature element that is not clearly listed. If this term is used in a claim, this term will make the claim a closed type, making it not include technical feature elements other than the clearly listed technical feature elements, except for related conventional impurities. If this term only appears in a certain clause of a claim, then it only limits the elements clearly listed in that clause, and the elements recorded in other clauses are not excluded from the overall claim.
[0024] Unless otherwise clearly stipulated or limited, terms such as "install", "connect", "join", "fix" and other terms should be understood in a broad sense. For example: it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this article can be understood according to specific situations.
[0025] The orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of description and simplification, and does not expressly or implicitly imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this text.
[0026] The following provides a detailed description of the solution provided by the present invention. The content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those skilled in the art. In the embodiments of the present invention, those not specified in specific conditions are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. For the reagents or instruments not specified in the embodiments of the present invention for the manufacturer, they are all conventional products that can be obtained through commercial purchase.
[0027] As Figure 1 shown, the embodiment of the present invention provides a robot hierarchical reinforcement learning control system based on target guidance, including:
[0028] a visual pose prediction module, an upper-layer decision-making subsystem, a middle-layer planning subsystem, and a lower-layer control subsystem; among them,
[0029] the visual pose prediction module can detect and track the pose of the target object in the environment in real time through an improved YOLOv8 algorithm to obtain real-time environmental data;
[0030] the upper-layer decision-making subsystem is communicatively connected to the visual pose prediction module, and can evaluate the task progress according to the upper-layer decision-making target and the current state of the obtained real-time environmental data through the upper-layer decision-making target guidance strategy, and select the corresponding middle-layer operation target according to the task progress;
[0031] the middle-layer planning subsystem is communicatively connected to the upper-layer decision-making subsystem, and can generate specific robot planning actions according to the current state and the middle-layer operation target through the planning target guidance strategy, and transmit the expected motion trajectory of the planning action to the lower-layer control subsystem;
[0032] the lower-layer control subsystem is communicatively connected to the middle-layer planning subsystem, and can combine the corresponding depth kinematic model and depth dynamic model of the robot, and after interpolating the robot motion trajectory output by the middle-layer planning subsystem, combine impedance control to control the actual motion trajectory of the robot to precisely execute the operation action.
[0033] Referring to Figure 3 , preferably, in the above system, the robot is a cable-driven flexible robot.
[0034] Preferably, in the above system, the visual pose prediction module includes:
[0035] A data acquisition sub-module, a preprocessing sub-module, feature extraction, an object detection sub-module, a depth data processing sub-module, and a corner-based pose prediction sub-module; wherein,
[0036] The data acquisition sub-module is communicatively connected to a camera installed on the robot and can collect RGB-D images in the environment in real time;
[0037] The preprocessing sub-module is communicatively connected to the data acquisition sub-module and can perform preprocessing of denoising, enhancement, and correction on the RGB-D images collected by the data acquisition sub-module to obtain preprocessed images;
[0038] The feature extraction and object detection sub-module is communicatively connected to the preprocessing sub-module and can perform feature extraction on the preprocessed images, detect and identify target objects in the environment and the preliminary localization regions of the target objects;
[0039] The depth data processing sub-module is communicatively connected to the preprocessing sub-module and can use the depth information obtained by a depth sensor installed on the robot, combine it with the preprocessed images to generate point cloud data of the target object, and estimate the initial pose of the target object in the world coordinate system through geometric calculations;
[0040] The corner-based pose prediction sub-module is communicatively connected to the feature extraction and object detection sub-module, the depth data processing sub-module, and the upper-level decision-making subsystem respectively, and can detect the precise pose of the target object within the preliminary localization region of the target object given by the feature extraction and object detection sub-module through an improved YOLOv8 algorithm, and output the obtained precise pose of the target object to the upper-level decision-making subsystem.
[0041] Preferably, in the above system, the corner-based pose prediction sub-module detects the precise pose of the target object within the preliminary localization region of the target object given by the feature extraction and object detection sub-module through an improved YOLOv8 algorithm in the following manner, including:
[0042] Step 11, boundary expansion: Expand the boundary of the preliminary localization region of the target object given by the feature extraction and object detection sub-module to obtain an expanded localization region;
[0043] Step 12, corner detection: Use the ORB algorithm to detect the corners of the target object in the expanded localization region;
[0044] Step 13, Color Clustering: Apply the K-means clustering algorithm to the pixels within the extended positioning area obtained in Step 11 for color clustering processing, group pixels with similar colors, and separate the target object from the complex background;
[0045] Step 14, Corner Point Matching: Match according to the relative positions and spatial distances of the corner points of each target object obtained in Step 12 to identify all the corner points of the same target object;
[0046] Step 15, Pose Calculation: Calculate the center point coordinates of the set of corner points of each target object after matching, convert the image coordinates into the corresponding coordinates in the world coordinate system, and calculate the preliminary pose of the target object according to the geometric relationship;
[0047] Step 16, Kalman Filtering: Perform Kalman filtering processing on the preliminary pose to remove the prediction errors caused by occlusion and illumination changes to obtain the accurate pose of the target object.
[0048] See Figure 2 , Preferably, in the above system, the upper-level decision-making subsystem includes:
[0049] A task decomposition and mapping module and an upper-level reinforcement learning network; where
[0050] The task decomposition and mapping module can decompose the operation tasks to be executed by the robot into a step sequence N of N ordered steps, and obtain a mapping from the state space to the step sequence N according to the step sequence N. Each step in the mapping corresponds to an expected middle-level operation target;
[0051] The upper-level reinforcement learning network is communicatively connected to the task decomposition and mapping module. In the upper-level decision-making target guiding policy of the upper-level reinforcement learning network, when an action needs to be selected, it will judge the progress of the robot operation task according to the current state of the system, and select a suitable step from the mapping according to the task progress, and output it to the middle-level planning subsystem as the middle-level operation target.
[0052] Preferably, in the above system, the upper-level reward function of the upper-level reinforcement learning network of the upper-level decision-making subsystem is:
[0053]
[0054] where s t represents the current state; a H is the action output by the upper-level decision-making target guiding policy, that is, the number g of the middle-level operation target M ; Φ(g H ) represents the state corresponding to the upper-level decision-making target g H ; ε His a preset threshold value;
[0055] The L2 norm is adopted in the reward function to measure the current state s t and the upper-level decision-making goal g H corresponding to the state Φ(g H ). If the distance between the two is less than the preset threshold value ε H , it is considered that the upper-level decision-making goal of the operation task has been achieved, and at this time the reward function returns 1; otherwise, it returns 0.
[0056] Preferably, in the above system, the upper-level reinforcement learning network includes: a current value network and a target value network. The target-guided double deep Q learning algorithm is adopted, and experience replay, mini-batch, and soft update are combined to update the target value network. The definition of the experience sample ξ of the upper-level decision-making goal guidance strategy of the upper-level reinforcement learning network is:
[0057]
[0058] where s t represents the current state; g H represents the expected upper-level decision-making goal; a H is the action output by the upper-level decision-making goal guidance strategy, that is, the number g M of the middle-level operation goal; is the upper-level reward function; is an indicator function used to evaluate the progress of the selected middle-level operation goal, s t+1 represents the next state, represents the distance error between the position of the end effector of the robotic arm and the target position, ε x represents the distance error threshold, represents the speed of the end effector of the robotic arm, ε v represents the speed threshold;
[0059] In the discrete action space, the update rule of the target value function Q(s t , g H ) is:
[0060]
[0061] where v represents the learning rate used to control the update step size, and the superscript H represents the parameters belonging to the upper-level decision-making subsystem; represents the completed target value function, and its definition is:
[0062]
[0063] where are the parameters of the target value network; a H ′ represents the after-action of the upper-level decision-making subsystem, Represents the action space of the upper-level decision-making subsystem;
[0064] Introduce two independent value networks, each network is equipped with a target value network, and the minimum value of the two target value networks is used as the estimated value, and the gradient is calculated through the following loss function and the network parameter θ is updated:
[0065]
[0066] Wherein, represents the experience replay buffer, ξ represents the experience sample of the upper-level decision-making target guiding policy of the upper-level reinforcement learning network; when updating the target value network parameters through soft update, the parameter τ is used to update the target value network parameters:
[0067] θ target ←τθ+(1 - τ)θ target ;
[0068] Select an action according to the maximum value output by the current value network, the action is the target number, and select the corresponding middle-level operation target in the target library according to the target number and transfer it to the middle-level planning subsystem.
[0069] Preferably, in the above system, the middle-level planning subsystem adopts a flexible actor-critic algorithm to generate the planning action of the robot according to the current state and the middle-level operation target through the planning target guiding policy, and transfer the expected motion trajectory of the planning action to the lower-level control subsystem;
[0070] The objective function J(π) of this flexible actor-critic algorithm is:
[0071]
[0072] Wherein, t represents the current moment; T represents the total duration of task execution; represents the expectation; ρ π represents the distribution of the state-action pair generated by the policy π; r(s t ,a t ) represents the reward function; is an adaptive temperature parameter, used to control the weight of entropy in the objective function to control the randomness of the policy π, represents the set of real numbers; represents the entropy of the policy π in the state s t under, and this entropy is used to measure the uncertainty of the policy π;
[0073] The flexible actor-critic algorithm includes: a flexible value function, a flexible Q function and an adaptive temperature parameter, wherein,
[0074] The flexible value function V(s) represents in the state st Under the current policy, considering the expected return obtained by executing an action and the entropy of the action, the flexible value function V(s) is expressed as:
[0075]
[0076] The flexible Q-function evaluation represents the expected return of taking action a t in state s t and following the current policy. The flexible Q-function is expressed as:
[0077]
[0078] The flexible actor-critic algorithm adjusts the adaptive temperature parameter α to make the policy entropy close to a desired target entropy The representation of the adaptive temperature parameter α is:
[0079]
[0080] The flexible actor-critic algorithm uses the gradient ascent method to maximize the objective function to optimize the policy π. At the same time, it introduces flexible policy updates and gradually updates the policy parameters by minimizing the KL divergence. The update formula for the policy parameters π new is:
[0081]
[0082] where π′ is the sampling in the control policy distribution; Π is the distribution of the control policy represented by the actor network; D KL is the KL divergence of the minimized policy; s t is the state of the robotic arm trajectory tracking task at time t; α represents the adaptive temperature parameter; is the state-action value function before update; is the partition function used to normalize the distribution.
[0083] See Figure 2 , preferably, in the above system, the middle-level planning subsystem also adopts the Hindsight Experience Replay algorithm, i.e., the HER algorithm. Assume that for any state s t , there exists a corresponding completed goal g. This assumption forms the basis for the application of the HER algorithm.
[0084] The core concept of the HER algorithm is to use failed experiences for knowledge acquisition, reinterpret past failure experiences to solve current problems, and regard them as successful achievements of different goals, thereby improving learning efficiency. Specifically, in each sampling process, the HER algorithm not only considers the original goal g M , but also considers a series of hindsight goals gM ′. These post - hoc goals are generated based on the current state or trajectory and are used to calculate an additional reward signal.
[0085] For each experience sample sampled from the environment a series of corresponding post - hoc experience samples can be generated where represents the reward signal calculated according to the post - hoc goal g M ′. Subsequently, these post - hoc experience samples are stored in the replay buffer together with the original experience samples, and a batch of experience samples are randomly selected for training during the update process of the SAC algorithm. The policy trained through this integration strategy can output appropriate actions according to the current state and the middle - layer operation goal g M and the appropriate action is then passed to the lower - layer control subsystem to achieve precise control. The action is then passed to the lower - layer control subsystem to achieve precise control.
[0086] Preferably, in the above - mentioned system, the lower - layer control subsystem includes:
[0087] an interpolation processing module and an impedance controller in the joint space; where
[0088] the interpolation processing module is communicatively connected to the middle - layer planning subsystem, can take the action output by the middle - layer planning subsystem as input, and can perform high - precision interpolation calculation between adjacent data points of the action data through the position trajectory function S p (k) to obtain the interpolated action trajectory data. The position trajectory function S p (k) within the data gap H is:
[0089]
[0090] where T represents the total duration of the task, that is, the time interval; k represents the interpolation time point, with a variation range from t to t + T, and t is the current moment; by applying the trajectory function S p (k), the end - effector of the robotic arm generates a smooth pose p k , velocity and acceleration trajectory
[0091] The impedance controller in the joint space is communicatively connected to the interpolation processing module, can take the trajectory output by the interpolation processing module as input, and can adjust the motion of the robotic arm according to the trajectory points in the trajectory to perform smooth and precise trajectory tracking and force control on the robotic arm.
[0092] Preferably, in the above - mentioned system, the joint control torque corresponding to the impedance controller in the joint space is:
[0093]
[0094] where τ M is the joint control torque vector; x d represents the desired gripper pose. The gripper position is adjusted in real time during the actual movement according to the actions output by the middle - layer planning subsystem, and the pose remains constant; is the predicted value of the Jacobian matrix. The superscript T represents the transpose of the matrix, and the superscript - 1 represents the inverse of the matrix; M d is the desired inertia matrix; B d is the desired damping matrix, K d is the desired stiffness matrix; is the desired acceleration; e are the velocity error and position error respectively; is the predicted value of the Coriolis force and centrifugal force terms; θ, and are the joint angle, angular velocity and angular acceleration respectively; is the inertia matrix.
[0095] In summary, the system of the embodiment of the present invention, by setting the upper - layer decision - making subsystem, the upper - layer decision - making subsystem and the lower - layer control subsystem, controls in a hierarchical manner, significantly reducing the computational complexity of each module and avoiding the difficulties of direct modeling and solving in a high - dimensional space. The interface definition between each layer is clear, which is beneficial to the design and optimization of each module, enabling each layer to focus on dealing with specific problems, thereby improving the overall performance of the robot when performing complex tasks. In addition, the framework has good scalability, and the vision module, decision - making module and control module can be replaced with different algorithms as long as the unified interface is followed.
[0096] In order to more clearly show the technical solutions provided by the present invention and the technical effects produced, the following takes specific embodiments to describe in detail the solutions provided by the embodiments of the present invention.
[0097] Embodiment 1
[0098] The embodiment of the present invention provides a hierarchical reinforcement learning control system for a robot based on target guidance, which significantly reduces the computational complexity of each module in a hierarchical manner and avoids the difficulties of direct modeling and solving in a high - dimensional space. The interface definition between each layer is clear, which is beneficial to the design and optimization of each module, enabling each layer to focus on dealing with specific problems, thereby improving the overall performance of the robot when performing complex tasks. In addition, the framework has good scalability, and the vision module, decision - making module and control module can be replaced with different algorithms as long as the unified interface is followed. The control system of the present invention includes the following main modules:
[0099] (1) Visual pose prediction module: An improved YOLOv8 algorithm is adopted to detect and track the pose of target objects in the environment in real time, providing key environmental data support for upper-layer decision-making.
[0100] (2) Upper-layer decision-making subsystem: Based on a preset target library and real-time environmental data, it selects appropriate middle-layer operation targets and continuously plans the next action. The upper-layer decision-making subsystem uses a goal-based reinforcement learning algorithm to improve the accuracy and adaptability of decision-making by continuously optimizing the strategy.
[0101] (3) Middle-layer planning subsystem: It converts the middle-layer operation targets provided by the upper-layer decision-making subsystem into specific robotic arm motion paths. The middle-layer planning subsystem performs path planning through a flexible actor-critic reinforcement learning algorithm to ensure the smoothness and accuracy of the robotic arm motion.
[0102] (4) Lower-layer control subsystem: It precisely controls the motion trajectory of the robotic arm through interpolation and impedance control to ensure the accurate execution of operation actions. The lower-layer control subsystem combines the corresponding depth kinematic model and depth dynamic model of the robot to make real-time adjustments to the actual motion of the robotic arm, improving the stability and response speed of control.
[0103] The specific composition of each module will be further described below.
[0104] (1) Visual pose prediction module:
[0105] The described visual pose prediction module is mainly used to detect and track objects in the robot operation environment, obtain their pose information, and thus provide accurate data support for the upper-layer decision-making subsystem. The specific implementation method is as follows:
[0106] Data acquisition sub-module: Through a camera installed on the robot (such as RealSense D435i), it acquires RGB-D image data in the environment in real time. The camera can simultaneously obtain color images and depth information, thus providing comprehensive visual data.
[0107] Preprocessing sub-module: It preprocesses the acquired RGB-D images, including operations such as image denoising, enhancement, and correction, to improve the accuracy of subsequent processing. The preprocessed image data is used to generate aligned point cloud data, providing a basis for pose prediction.
[0108] Feature Extraction and Object Detection Sub-module: An improved YOLOv8 (You Only Look Once version 8) algorithm is used to extract features from the pre-processed image, detect and identify target objects in the environment. The YOLOv8 algorithm can efficiently perform real-time object detection and output the position information of the objects.
[0109] Depth Data Processing Sub-module: Using the depth information obtained by the depth sensor and combining with the RGB image, the point cloud data of the target object is generated. Through geometric calculations, the initial pose of the object in the world coordinate system is estimated.
[0110] Corner-based Pose Prediction Sub-module: In the preliminary positioning area provided by the YOLOv8 algorithm, the following steps for precise pose detection are further carried out:
[0111] Step 11, Boundary Expansion: Expand the boundary of the preliminary positioning area to ensure that no important corners are missed;
[0112] Step 12, Corner Detection: Use the ORB (Oriented FAST and Rotated BRIEF) algorithm for corner detection to enhance the algorithm's adaptability to scale and rotation changes;
[0113] Step 13, Color Clustering: Use the K-means clustering algorithm for color clustering processing, group pixels with similar colors, separate building blocks from the complex background, and reduce noise and irrelevant interference;
[0114] Step 14, Corner Matching: Match according to the relative position and spatial distance of the corners to identify all corners of the same building block;
[0115] Step 15, Pose Calculation: Calculate the center point coordinates of the set of corner points of each matched building block, convert the image coordinates to the corresponding coordinates in the world coordinate system, and calculate the pose of the building block according to the geometric relationship;
[0116] Step 16, Kalman Filtering: Perform Kalman filtering on the preliminarily estimated pose to remove prediction errors caused by factors such as occlusion and lighting changes, and improve the stability and accuracy of pose prediction;
[0117] Output the optimized pose information of the target object to the upper-layer decision-making subsystem for path planning and motion control when the robot performs complex operation tasks.
[0118] Through the above steps, the visual pose prediction module realizes the real-time detection and precise tracking of target objects in the environment, provides key pose information support for the robot system, and significantly improves the operation performance of the robot in complex environments.
[0119] (2) Hierarchical Reinforcement Learning Framework:
[0120] 21) Problem Modeling:
[0121] For complex operation tasks, the concept of the goal space is introduced. The complex operation task is modeled as a goal-guided discrete-time Markov decision process (Goal Guided Markov Decision Process, G-MDP), that is:
[0122]
[0123] Among them, the state space represents the set of all possible states, where each state s t describes the situation of the robot itself and its environment at time step t. The action space represents the set of all possible actions that can be executed. At time step t, the robot selects an action a t to execute according to the current state s t . The goal space represents the set of all possible goals, and each goal g describes the state that the robot needs to reach or the task to be completed. The state transition probability p(s t+1 |s t , a t ) describes the probability that the robot transfers to the next state s t after executing the action a t . t+1
[0124] The goal-guided reward function r g (s t , a t ) calculates the reward according to the current state s t , the executed action a t and the goal g to guide the robot to learn how to achieve the goal most effectively. The discount factor γ is used to discount future rewards when calculating the cumulative reward.
[0125] The goal of reinforcement learning is to find a goal-guided policy, π(a t |s t , g) to maximize the expected value of the cumulative discounted reward in the operation task, that is:
[0126]
[0127] Among them, represents the task execution process, T represents the total execution duration, and t represents the current moment.
[0128] For the above-mentioned target-guided discrete-time Markov decision process problem, the present invention proposes a target-guided hierarchical reinforcement learning framework. This framework divides the entire system into three main levels: the upper-level decision-making subsystem, the middle-level planning subsystem, and the lower-level control subsystem. Through the close cooperation between levels, efficient and intelligent control is achieved. Among them, the core of the upper-level decision-making subsystem is the upper-level decision-making target guiding strategy Among them represents the upper-level decision-making target. This upper-level decision-making target guiding strategy evaluates the task progress according to the current state s t and the upper-level decision-making target g H , and accordingly selects an appropriate middle-level operation target g M , and transmits it to the middle-level planning subsystem to achieve the overall decision-making of the task. The core of the middle-level planning subsystem is the middle-level operation target guiding strategy According to the current state s t and the middle-level operation target g M , generate specific planning actions and transmit its expected motion trajectory to the lower-level control subsystem to guide the manipulator to execute the corresponding actions. The lower-level control subsystem ensures that the motion trajectory of the manipulator is both smooth and stable through trajectory interpolation smoothing and impedance control, and combines with the depth kinematics and dynamics models to accurately execute the required motion. The trajectory interpolation smoothing technology makes the motion of the manipulator smoother and more natural, reducing unnecessary jitters and mutations; while the impedance control ensures that the manipulator maintains a certain compliance when interacting with the environment.
[0129] 22) Decomposition and mapping of tasks:
[0130] Once the visual pose prediction module provides the necessary information, the upper-level decision-making subsystem begins to play its core role. This upper-level decision-making subsystem is responsible for refining and decomposing the abstract high-level decision-making target g H (such as completing a certain assembly task or reaching a certain state). This process involves an in-depth understanding of the task requirements, a comprehensive assessment of the current state of the system, and the generation and screening of potential operation sequences. Specifically, first, assume that a complex operation task target g H can be systematically decomposed into N ordered steps, which means that there is a mapping ψ N from the state space to the step sequence, that is:
[0131]
[0132] Among them, each step corresponds to an expected middle-level operation target g MThis decomposition is reasonable and conforms to the usual way of task execution in the real world, that is, step by step and in an orderly manner.
[0133] In the upper-level decision objective guidance strategy of the upper-level decision-making subsystem, when an action needs to be selected, it will, according to the current system state s t judge the task progress, and select a suitable step from the mapping ψ N and output the corresponding middle-level operation objective g M to the middle-level operation objective guidance strategy. This process is carried out in accordance with the logical order of the task. The upper-level decision objective guidance strategy gradually guides the middle-level operation objective guidance strategy to achieve different task steps by sequentially outputting actions, so as to achieve the overall task decision-making. After receiving the middle-level operation objective g M output by the upper-level decision objective guidance strategy, the middle-level operation objective guidance strategy will perform detailed path planning and action generation according to this middle-level operation objective and the current state s t It does not need to know the global information of the overall task, nor does it need to know the position or role of the current step in the entire task. It only needs to complete the task of this step according to the received objective.
[0134] Table 1.1 Mapping ψ of the two-block stacking task N Example
[0135]
[0136] Table 1.1 gives a mapping ψ of the two-block stacking task N Example. In this example, the first row represents the current state s t , which is mapped to middle-level operation objectives related to four different steps. The following rows respectively correspond to these steps: the second row describes the action of "reaching block 1", the third row describes the process of "grabbing block 1", the fourth row describes the step of "putting block 1 on block 2", and finally, the fifth row represents the "end" step of the task. In Table 1.1, h represents the height of the block, represents the real-time position of the gripper. Similarly, and represent the real-time positions of block 1 and block 2 respectively. represents the starting position of block 1. In addition, represents the state of the gripper, including open and closed two possible states.
[0137] In this system, the strategy π of the upper-level decision-making subsystem gH is responsible for selecting actions related to the four middle - level steps according to the current state. Meanwhile, the middle - level operation target guiding strategy needs to learn how to plan the movement of the end - effector of the robotic arm to achieve these steps. Through this hierarchical decision - making and planning structure, the entire block - stacking task can be completed efficiently and accurately.
[0138] In the task of stacking two blocks, an additional step "grab block 1" is introduced. Block 1 is grabbed and lifted to a height of 2 times, i.e., 2×h. At the same time, a height offset ε is introduced in the step of "place block 1 on block 2". The purpose of these two designs is to effectively reduce the possible collisions between blocks during the stacking process, thereby significantly increasing the success rate of the multi - block stacking task. Through such a design, it can be ensured that when placing block 1 on block 2, there is enough space between the two for safe and accurate stacking, thus reducing the potential for collisions and failures.
[0139] Similarly, adding the "end" step is also to emphasize the clarity and standardization of task completion, ensuring the stability and controllability of the entire stacking process. This step - splitting that incorporates human experience not only helps to guide the successful execution of the current task but also provides a basis and reference for the expansion of the multi - block stacking task.
[0140] 23) State - space design:
[0141] In the system of the present invention, the upper - level decision - making subsystem and the middle - level planning subsystem share the same task - state representation, although their pursued goals are significantly different. This consistent design enables the system to work collaboratively at different levels while allowing each level to focus on its specific goal, thereby improving the overall learning efficiency and adaptability.
[0142] Specifically, for the stacking task of M blocks, the definition of the state space can be expressed as:
[0143]
[0144] where, represents the position of the end - effector of the robotic arm, represents the speed of the end - effector of the robotic arm, represents the opening - and - closing state of the end - effector of the robotic arm, represents the contact force received by the gripper; and represents the position of the block relative to the gripper.
[0145] To accelerate the learning of the upper-level decision-making target guidance strategy of the upper-level decision-making subsystem and the middle-level operation target guidance strategy of the middle-level planning subsystem, the relative positions of the building blocks rather than the absolute positions are used to represent the task states. The relative position information helps to reduce the complexity of the state space, enabling the strategy to understand and adapt to different task scenarios more quickly. In addition, the relative position information provides an intuitive and concise state description, making it easier for the policy network to learn the effective state-to-action mapping relationship, thus improving the learning efficiency. To achieve transfer between different manipulators and improve the generalization of the strategy, the joint space state is not used in the state information of the manipulator itself. Instead, the position and velocity of the end effector operation space of the manipulator and its opening / closing state are used. This choice makes the strategy more independent of the specific manipulator structure, facilitating transfer between different manipulators. At the same time, due to the particularity of the building block stacking task, the pose information is omitted to further reduce the dimensions of the state space and the action space, improving the convergence speed of the strategy. In addition, to achieve more compliant control and enhance the system's perception ability, the data of the end effector torque sensor of the manipulator is introduced as tactile perception information. These data can reflect the interaction force between the manipulator and the environment in real time, providing important feedback for the middle-level operation target guidance strategy. Combining this tactile perception information, the middle-level operation target guidance strategy can more finely adjust the action output to adapt to different environmental changes and task requirements. The combination of this perception and control enables the system to complete tasks more flexibly and accurately, thus improving the overall control performance.
[0146] 24) Action Space and Target Space Design:
[0147] As shown in Table 1.1, for the stacking task of two building blocks, a fine-grained task decomposition is carried out, dividing it into four logically coherent and relatively independent steps. These steps together construct a target library, providing a clear and well-defined action space for the upper-level decision-making subsystem. The upper-level decision-making target guidance strategy dynamically selects the most appropriate middle-level operation target from the target library by comprehensively evaluating the current environmental state and task requirements, and then conveys it to the middle-level operation target guidance strategy for execution.
[0148] It is worth emphasizing that the adopted task decomposition method exhibits excellent scalability. Specifically, when the number of building blocks involved in the stacking task increases to M, the size of the target library will be expanded accordingly in proportion to 4M. This scalability ensures that the proposed method can flexibly handle stacking tasks of different scales and complexities.
[0149] Due to the artificial task division, the action space of the upper-level decision-making subsystem is abstracted as a series of target numbers. There is a clear mapping relationship between these numbers and the specific targets listed in Table 1.1, defined as Φ.
[0150] The upper-level decision-making target guiding strategy only needs to output the corresponding target number, and then quickly converts it into a specific target instruction through the mapping relationship Φ for the middle-level operation target guiding strategy to execute. This abstraction of the action space reduces the complexity of the strategy output. In addition, this design method of the action space also provides convenience for the teaching process. By giving the order of numbers instead of the specific execution path, it can guide the upper-level decision-making target guiding strategy to learn in a more intuitive way. At the same time, this number teaching method also provides clear guidance for the middle-level operation target guiding strategy, enabling it to complete the target in a reasonable order and with gradually increasing difficulty.
[0151] In hierarchical reinforcement learning, the non-stationarity problem is a core challenge. Non-stationarity refers to the changes in the state transition probability or reward function faced by the strategy during the learning process due to the dynamics of the environment or policy updates, resulting in unexpected changes. Such changes make the previously collected experience may no longer be applicable, thus triggering fluctuations and instability in the learning process.
[0152] However, this design of the discrete action space and target space, as well as the convenient teaching design, helps to alleviate the trouble of the non-stationarity problem. This design can improve the convergence speed of the strategy by reducing the complexity of the strategy output, thereby enhancing the execution efficiency of the overall task. This simple and effective teaching method can reduce the exploration time and cost in the learning process, helping each layer of the strategy to understand the core logic and execution key points of the task more quickly, so as to achieve more efficient and accurate decision-making and planning. Specifically, a binary coding scheme is used to define the action space, and this definition also applies to the representation of the upper-level decision-making target g H In this coding scheme, each binary bit represents a specific action or target. For example, the coding pair [1,0,0,0] corresponds to the target g1, that is, "reach block 1"; while the coding [0,0,1,0] represents the target g3 of "put block 1 on block 2". The action space of the middle-level operation target guiding strategy is defined as the position change amount of the end effector and the opening and closing actions of the gripper, which can be expressed as
[0153] 25) Reward function design:
[0154] To evaluate the performance of the upper-level decision-making target guiding strategy and guide the learning process, the reward function is designed as follows:
[0155]
[0156] Among them, where s t represents the current state; a H is the action output by the upper-level decision-making target guiding strategy, that is, the number g of the middle-level operation targetM ; Φ(g H ) represents the corresponding state of the upper - layer decision - making goal g H ; ε H is a preset threshold; the L2 norm is used in this reward function to measure the proximity between the current state s t and the corresponding state Φ(g H ) of the upper - layer decision - making goal. If the distance between the two is less than the preset threshold ε H , it is considered that the upper - layer decision - making goal of the task has been achieved, and the reward function returns 1 at this time; otherwise, it returns 0.
[0157] For the middle - layer planning subsystem, a detailed reward function is defined to guide the movement of the robotic arm. The specific expression of this reward function is as follows:
[0158]
[0159] where, represents the position error at the current time t, which is composed of the position error of the end - effector and the position error of the object being operated; W x represents the corresponding weight parameter; is used to encourage the gripper to reach the desired opening and closing state; is used to punish excessive contact force during the operation process. When the contact force exceeds the preset threshold ε ext , it will be punished; represents the reward for task completion. When the position error is less than ε x and the speed of the gripper is less than the threshold ε v , a reward will be obtained.
[0160] (3) Upper - layer decision - making subsystem:
[0161] To train the upper - layer decision - making goal - guiding strategy, a Goal - Guided Double Deep Q - Learning (GDDQN) algorithm is proposed, and combined with the Experience Replay mechanism, Mini - Batch, and Soft Update methods to update the target - value network. Specifically, the definition of the experience sample of the upper - layer decision - making goal - guiding strategy is:
[0162]
[0163] where, g H represents the desired upper - layer decision - making goal; is the action output by the upper - layer decision - making goal - guiding strategy and is also the number g M of the middle - layer operation goal; is the upper - layer reward function; is an indicator function used to evaluate the completion of the selected middle - layer operation target. Through this definition, empirical samples can be effectively utilized to train the upper - layer decision - making strategy, and the completion of the middle - layer operation target can be considered during the training process.
[0164] In the discrete action space, the update rule of the target value function Q(s t , g H ) is:
[0165]
[0166] where V represents the learning rate, which is used to control the step size of the update. The definition of the completed target value function is:
[0167]
[0168] where are the parameters of the target value network.
[0169] To enhance the stability of learning and the accuracy of neural network approximation, two independent value networks are introduced, and each value network is equipped with a target value network to reduce the risk of over - estimating the value. The minimum value of these two target networks is used as the estimated value, and the gradient is calculated through the following loss function and the network parameters θ are updated:
[0170]
[0171] where represents the experience replay buffer. When updating the target network parameters, a soft - update method is adopted, and the parameters are updated through the parameter τ:
[0172] θ target ←τθ+(1 - τ)θ target ;
[0173] This method helps to smooth the update process of the network parameters, reduce the mutation during parameter update, and thus improve the stability of learning.
[0174] Finally, the action (i.e., the target number) is selected according to the maximum value output by the current value network, and the corresponding middle - layer operation target is selected in the target library and passed to the middle - layer planning subsystem to implement the upper - layer decision - making. Figure 1 shows the overall process of the GDDQN algorithm.
[0175] (4) Middle - layer planning subsystem:
[0176] To train the middle-level planning subsystem, the Soft Actor-Critic (SAC) algorithm is used for online optimization. The SAC algorithm is a model-free deep reinforcement learning algorithm, especially suitable for tasks with continuous action spaces. Its core idea is to introduce entropy regularization in the traditional actor-critic framework to achieve more effective exploration. Specifically, the goal of SAC is to learn a policy that maximizes the cumulative reward while also maximizing its entropy at each time step. The objective function of the SAC algorithm consists of two parts: cumulative reward and policy entropy. The cumulative reward encourages the agent to take actions that can obtain higher returns, while the policy entropy encourages the agent to maintain a certain degree of exploration and avoid falling into a deterministic policy prematurely. The form of the objective function is as follows:
[0177]
[0178] where \(t\) represents the current time step; \(T\) represents the total duration of task execution; denotes the expectation; \(\rho\) π denotes the distribution of state-action pairs generated by the policy \(\pi\); \(r(s t , a t ) represents the reward function; is an adaptive temperature parameter used to control the weight of entropy in the objective function to control the randomness of the policy \(\pi\), denotes the set of real numbers; denotes the entropy of the policy \(\pi\) in the state \(s t , which is used to measure the uncertainty of the policy \(\pi\);
[0179] The core components of the SAC algorithm include a soft value function, a soft Q function, and an adaptive temperature parameter. The soft value function \(V(s)\) represents the expected return that can be obtained by taking actions according to the current policy in the state \(s t , taking into account the entropy of the actions. Its definition is as follows:
[0180]
[0181] The soft Q function is used to evaluate the expected return of taking the action \(a t in the state \(s t and following the current policy:
[0182]
[0183] The choice of the temperature parameter \(\alpha\) has a significant impact on the algorithm performance. The SAC algorithm adaptively adjusts \(\alpha\) to balance exploration and exploitation. The goal is to make the policy entropy close to a desired target entropy The representation of the adaptive temperature parameter \(\alpha\) is:
[0184]
[0185] Finally, the SAC algorithm maximizes the objective function and optimizes the policy π by using the gradient ascent method. At the same time, the algorithm introduces a flexible policy update method to gradually update the policy parameters by minimizing the KL divergence. This method ensures the stability and continuity of the learning process.
[0186]
[0187] where π′ is the sampling in the control policy distribution; Π is the distribution of the control policy represented by the actor network; D KL is to minimize the KL divergence of the policy; s t is the state of the robotic arm trajectory tracking task at time t as a robot; α represents the adaptive temperature parameter; is the state-action value function before update; is the partition function for normalizing the distribution.
[0188] By introducing entropy as an additional objective, the SAC algorithm not only enhances the exploration of the policy but also ensures the diversity of the policy during the learning process, thus effectively improving the learning efficiency and the adaptability of the policy. In addition, by adjusting the expected value of the temperature coefficient, SAC can flexibly balance between exploration and exploitation to adapt to different task requirements.
[0189] To further improve the sample efficiency and learning speed, the Hindsight Experience Replay (HER) algorithm is introduced. Here, a hypothesis is proposed: for any state s t , there exists a corresponding achieved goal g. This hypothesis forms the basis for the application of the HER algorithm.
[0190] The core concept of the HER algorithm is to use failed experiences for knowledge acquisition, to reinterpret past failure experiences to solve current problems, and to regard them as successful realizations of different goals, thereby improving the learning efficiency. Specifically, in each sampling process, the HER algorithm not only considers the original goal g M , but also considers a series of hindsight goals g M ′. These hindsight goals are generated based on the current state or trajectory and are used to calculate additional reward signals.
[0191] For each experience sample sampled from the environment a series of corresponding hindsight experience samples can be generated where represents according to the hindsight goal g MThe calculated reward signal. Subsequently, these post hoc experience samples are stored in the replay buffer together with the original experience samples, and a batch of experience samples are randomly selected for training during the update process of the SAC algorithm. The policy trained through this integration strategy can output appropriate actions according to the current state and the middle-level operation target g M This action is then passed to the lower-level control subsystem to achieve precise control.
[0192] (5) Lower-level control subsystem:
[0193] First, a robotic arm (i.e., a cable-driven flexible robot) is modeled using a deep kinematic model and a deep dynamic model as follows:
[0194] The forward kinematic model of the robotic arm is determined as: x = f(q); where, x is the pose of the end effector of the robotic arm in the operation space, m is the degree of freedom of the operation space of the robotic arm; q is the joint angle of the robotic arm, n is the degree of freedom of the joint space of the robotic arm;
[0195] The velocity relationship between the joints of the robotic arm and the end effector of the robotic arm is expressed as the derivative of the kinematic equation f with respect to time where, \(\dot{x}\) is the derivative of the pose of the end effector of the robotic arm in the operation space with respect to time; J is the actual Jacobian matrix of the robotic arm, m is the degree of freedom of the operation space of the robotic arm, n is the degree of freedom of the joint space of the robotic arm; \(\dot{q}\) is the angular velocity of the joints of the robotic arm;
[0196] The pose of the end effector of the robotic arm in the operation space is determined as x = [p, α], where, p is the position of the end effector of the robotic arm; α is the quaternion representing the attitude of the end effector of the robotic arm, α = [η, ∈] = [η, ∈1, ∈2, ∈3], where ∈ represents the rotation axis of the quaternion, ∈1, ∈2, and ∈3 are the three components of the rotation axis respectively, η represents the angle of rotation around the rotation axis, and α satisfies the constraint The relationship between the derivative of the quaternion α with respect to time and the angular velocity ω of the end effector of the robotic arm in the operation space is i.e., this relationship is:
[0197]
[0198] where, ω x 、ω y and ω z They are the components of the angular velocity of the end effector of the robotic arm on the x-axis, y-axis, and z-axis respectively;
[0199] A fully connected deep neural network is established to learn the forward kinematic model of the robotic arm, and the corresponding deep kinematic model of the robotic arm is obtained as: where θ are the network parameters of the deep kinematic model; f θ is the kinematic equation obtained by deep learning; q is the joint angle of the robotic arm;
[0200] Apply the chain rule to the derivative of each layer in the deep kinematic model, and iteratively calculate the derivative of the output of the deep neural network with respect to the input to recover and learn the Jacobian matrix. The learned Jacobian matrix J θ (q) is expressed as:
[0201]
[0202] where θ are the network parameters of the deep kinematic model; is the partial derivative of the kinematic equation obtained by deep learning with respect to the joint angle; h i is the i-th network layer of the deep kinematic model, h i =σ i (W i h i-1 +b i ), i = 1,..., N; W i is the weight applied to the input of the previous network layer of the deep neural network in the deep kinematic model; b i is the bias; σ i and σ′ i are the non-linear activation function of the deep neural network and its derivative respectively.
[0203] The training dataset of the deep kinematic model where q is the joint angle of the robotic arm; is the joint angular velocity of the robotic arm, x is the pose of the end effector of the robotic arm in the operating space; is the derivative of the pose of the end effector of the robotic arm in the operating space with respect to time;
[0204] The loss function during the training of the deep kinematic model is:
[0205]
[0206] where θ are the network parameters of the deep kinematic model; is the training dataset of the deep kinematic model; is the size of the training dataset; is the generalized inverse matrix of the Jacobian matrix; the first loss term ||x - f θ (q)|| 2 is the mean square error between the actual pose and the predicted pose of the manipulator's operational space; the second loss term is the mean square error between the actual velocity and the predicted velocity of the manipulator's operational space; the last loss term is the Jacobian matrix that makes the deep neural network more accurate by fitting the joint angular velocity of the manipulator.
[0207] Establish a Lagrangian-based deep dynamics model corresponding to the manipulator through deep learning in the following way, including:
[0208] Determine that the forward dynamics model and the inverse dynamics model of the manipulator are respectively:
[0209]
[0210] where, are respectively the joint angle, joint angular velocity and joint angular acceleration of the manipulator; τ is the torque acting on the manipulator joint;
[0211] Select the generalized coordinates of the manipulator as the joint angles of the manipulator, and define the Lagrangian function as where, is the kinetic energy of the manipulator; V(q) is the potential energy of the manipulator; M(q) is the mass matrix of the manipulator; the superscript T represents the transpose matrix;
[0212] To ensure the positive definite symmetry of the mass matrix M(q) of the manipulator, perform a Cholesky decomposition on the mass matrix M(q) to obtain where, is a lower triangular matrix with non-negative diagonal elements, the superscript T represents the transpose matrix, and use a deep neural network to learn the lower triangular matrix with non-negative diagonal elements and the potential energy V(q) of the manipulator to fit the Lagrangian function Combine the Euler-Lagrange equation where τ is the torque acting on the manipulator joint, and obtain that the deep forward dynamics model and the deep inverse dynamics model that make up the corresponding deep dynamics model of the manipulator are respectively:
[0213]
[0214] where, q, and are respectively the joint angle, joint angular velocity and joint angular acceleration of the manipulator; is the mass matrix of the manipulator; is the derivative of the mass matrix of the manipulator with respect to time; are the Coriolis force and the centripetal force; is the conservative force including gravity and spring elastic force; τ M is the joint output torque of the robotic arm; is the joint friction force of the robotic arm, and the joint friction force of the robotic arm is obtained through the introduced prior joint friction force model of the robotic arm composed of Coulomb friction, viscous friction and Stribeck friction force. The prior joint friction force model of the robotic arm is:
[0215]
[0216] where, f c is the Coulomb friction force; f υ is the viscous friction coefficient; f s is the maximum static friction force; υ and δ are the correlation coefficients of the Stribeck friction force model respectively; is the angular velocity of the i-th joint of the robotic arm.
[0217] The training data set of the deep dynamics model where, q is the joint angle of the robotic arm; is the joint angular velocity of the robotic arm; is the joint angular acceleration of the robotic arm; τ is the joint torque of the robotic arm;
[0218] The loss function of the deep dynamics model is:
[0219]
[0220] where, is the training data set of the deep dynamics model; is the size of the training data set of the deep dynamics model; is the L2 regression loss of the robotic arm joint angular acceleration, are respectively the predicted value and the actual value of the robotic arm joint angular acceleration at time t; is the L2 regression loss of the joint torque, τ t are respectively the predicted value and the actual value of the robotic arm joint torque at time t; is the multi-step prediction loss obtained through numerical integration. H is the total number of prediction steps, q t+i 、 are respectively the actual values of the robotic arm joint angle and joint angular velocity at time t + 1, and are respectively the predicted values of the robotic arm joint angle and joint angular velocity at time t + 1.
[0221] For the cable-driven flexible robot designed and developed for the laboratory, the deep kinematic model and the deep dynamic model were correspondingly modified. Figure 3 The cable-driven flexible robot used in this embodiment is shown. The robotic arm of this robot has 10 revolute joints, among which 7 are active revolute joints and 3 pairs of coupled joints. Therefore, this robotic arm is actually a 7-degree-of-freedom, 10-link serial robotic arm driven by 7 motors and simultaneously subject to 3 sets of independent kinematic constraints. Define the mapping relationship from the motor angle to the joint angle as θ = f s (q). According to the design of the robotic arm, two joints θ4 and θ5 of the elbow are driven by one motor q4, while four joints θ6 to θ9 of the wrist are respectively driven by two motors q5 and q6. Therefore, the following joint constraint relationships exist:
[0222] θ4 = θ5, θ6 = θ9, θ7 = θ8;
[0223] A shared network was added to the deep kinematic model and the deep dynamic model to learn the mapping from the joint space to the motor space. For the constraint relationships in the joint space, a loss term for the joint angle coupling constraint was added to the loss function to ensure that the model can better comply with the actual physical constraints during the training process. The expression of the loss function is:
[0224]
[0225] where λ1, λ2, and λ3 represent the weights used to adjust the constraint loss term.
[0226] In the deep dynamic model, the mapping relationship between the joint space torque and the motor space torque is:
[0227]
[0228] where τ M represents the joint torque, and τ q represents the motor torque; while J s represents the Jacobian matrix of the speed conversion relationship between the motor space and the joint space:
[0229]
[0230] where and represent the joint angular velocity and the motor angular velocity respectively. This Jacobian matrix can be calculated through a shared network in the deep kinematic and dynamic models.
[0231] During the actual data acquisition process, the motor angle q, the angular velocity and the torque τ qThe information is recorded by the host computer. Meanwhile, the end pose of the robot is captured by the camera system Optitrack.
[0232] An impedance controller in the joint space is designed, and its corresponding joint control torque is:
[0233]
[0234] where τ M is the joint control torque vector; x d represents the desired gripper pose. The gripper position is adjusted in real time during the actual movement according to the actions output by the middle-level planning subsystem, and the pose remains constant; is the predicted value of the Jacobian matrix. The superscript T represents the transpose of the matrix, and the superscript -1 represents the inverse of the matrix; M d is the desired inertia matrix; B d is the desired damping matrix, K d is the desired stiffness matrix; is the desired acceleration; e are the velocity error and position error respectively; is the predicted value of the Coriolis force and centrifugal force terms; θ, and are the joint angle, angular velocity and angular acceleration respectively; is the inertia matrix.
[0235] For the lower-level control subsystem, in order to maintain the continuity and smoothness of the actions, first, the actions generated by the middle-level operation target guiding strategy are interpolated. By performing high-precision interpolation calculations between adjacent data points, the data gaps caused by frequency differences are effectively filled. This processing method ensures the continuity and smoothness of the pose trajectory in time and space, thus avoiding control errors and jitters caused by frequency mismatches.
[0236] For the change in position, the cubic spline interpolation algorithm is adopted. Given the known position change Δ p , the current position p t and the velocity , the desired end-effector position p des = p t +Δp can be calculated. Assuming the desired velocity is zero, the position trajectory function S p (k) within the time interval T can be derived, and its expression is:
[0237]
[0238] Among them, T represents the total duration of task execution, that is, the time interval; k represents the interpolation time point, with a variation range from t to t + T, where t is the current moment; by applying the trajectory function S p (k), a smooth pose p can be generated for the end effector of the robotic arm within the data gap H k , velocity and acceleration trajectories These trajectories will then be used as inputs and provided to the subsequent impedance controller. The impedance controller will adjust the movement of the robotic arm according to these trajectory points to achieve smooth and precise trajectory tracking and force control.
[0239] The control system of the present invention decomposes multiple key levels such as visual perception, intelligent decision-making, path planning, and precise control, strengthens the information interaction and collaborative effects between levels, and realizes the control of the robot. It can significantly improve the performance of the robot in completing multi-step operation tasks in complex and unstructured environments, and solve the limitations that the current robot control system often shows when facing dynamic changing environments and diverse task requirements. Through the application of the hierarchical reinforcement learning method, each level can effectively cooperate to improve the decision-making efficiency and execution effect of the robot in complex tasks. At the same time, the system has good scalability, can adapt to different scenarios and task requirements, and has broad application prospects.
[0240] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.
[0241] As described above, only the specific and preferred embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of suggestion that this information constitutes prior art known to those skilled in the art.
Claims
1. A target-guided hierarchical reinforcement learning control system for a robot, characterized in that, Including: A visual pose prediction module, an upper-level decision-making subsystem, a middle-level planning subsystem, and a lower-level control subsystem; among them, The visual pose prediction module can detect and track the pose of target objects in the environment in real time through an improved YOLOv8 algorithm to obtain real-time environmental data; The upper-level decision-making subsystem is communicatively connected to the visual pose prediction module and can evaluate the task progress according to the upper-level decision-making target and the current state of the obtained real-time environmental data through the upper-level decision-making target guidance strategy, and select the corresponding middle-level operation target according to the task progress; The middle-level planning subsystem is communicatively connected to the upper-level decision-making subsystem and can generate specific robot planning actions according to the current state and the middle-level operation target through the planning target guidance strategy, and transmit the expected motion trajectory of the planning action to the lower-level control subsystem; The lower-level control subsystem is communicatively connected to the middle-level planning subsystem and can combine the corresponding depth kinematic model and depth dynamic model of the robot. After interpolating the robot motion trajectory output by the middle-level planning subsystem, it combines impedance control to control the actual motion trajectory of the robot to accurately execute the operation action.
2. The target-guided robot hierarchical reinforcement learning control system according to claim 1, wherein The robot is a cable-driven flexible robot; The visual pose prediction module includes: A data acquisition sub-module, a preprocessing sub-module, feature extraction, a target detection sub-module, a depth data processing sub-module, and a corner-based pose prediction sub-module; among them, The data acquisition sub-module is communicatively connected to a camera installed on the robot and can collect RGB-D images in the environment in real time; The preprocessing sub-module is communicatively connected to the data acquisition sub-module and can perform preprocessing of denoising, enhancement, and correction on the RGB-D images collected by the data acquisition sub-module to obtain preprocessed images; The feature extraction and target detection sub-module is communicatively connected to the preprocessing sub-module and can extract features from the preprocessed images, detect and identify target objects in the environment and the preliminary positioning regions of the target objects; The depth data processing sub-module is communicatively connected to the preprocessing sub-module and can use the depth information obtained by a depth sensor installed on the robot, combine the preprocessed images to generate point cloud data of the target object, and estimate the initial pose of the target object in the world coordinate system through geometric calculations; The corner-based pose prediction sub-module is communicatively connected to the feature extraction and target detection sub-module, the depth data processing sub-module, and the upper-level decision-making subsystem respectively, and can detect the accurate pose of the target object within the preliminary positioning region of the target object given by the feature extraction and target detection sub-module through an improved YOLOv8 algorithm, and output the obtained accurate pose of the target object to the upper-level decision-making subsystem.
3. The robot hierarchical reinforcement learning control system based on target guidance according to claim 2, characterized in that, The corner-based pose prediction sub-module detects the accurate pose of the target object within the preliminary positioning region of the target object given by the feature extraction and target detection sub-module through the improved YOLOv8 algorithm in the following manner, including: Step 11, boundary expansion: Expand the boundary of the preliminary positioning area of the target object given by the feature extraction and target detection sub-module to obtain an expanded positioning area; Step 12, corner detection: Use the ORB algorithm to detect the corners of the target object within the expanded positioning area, and obtain the corners of the target object within the expanded positioning area; Step 13, color clustering: Perform color clustering processing on the pixels within the expanded positioning area obtained in Step 11 using the K-means clustering algorithm, group the pixels with similar colors, and separate the target object from the complex background; Step 14, corner matching: Match according to the relative positions and spatial distances of the corners of each target object obtained in Step 12, and identify all the corners of the same target object; Step 15, pose calculation: Calculate the center point coordinates of the corner set of each matched target object, convert the image coordinates to the corresponding coordinates in the world coordinate system, and calculate the preliminary pose of the target object according to the geometric relationship; Step 16, Kalman filtering: Perform Kalman filtering processing on the preliminary pose to remove the prediction errors caused by occlusion and light changes to obtain the accurate pose of the target object.
4. The target-guided robot hierarchical reinforcement learning control system according to any one of claims 1-3, characterized in that The upper-layer decision-making subsystem includes: A task decomposition and mapping module and an upper-layer reinforcement learning network; where, The task decomposition and mapping module can decompose the operation task to be executed by the robot into a step sequence N of N ordered steps, and obtain a mapping from the state space to the step sequence N according to the step sequence N. Each step in the mapping corresponds to an expected middle-layer operation target; The upper-layer reinforcement learning network is communicatively connected to the task decomposition and mapping module. In the upper-layer decision-making target-guided policy of this upper-layer reinforcement learning network, when an action needs to be selected, it will judge the progress of the robot operation task according to the current state of the system, and select a suitable step from the mapping according to the task progress, and output it to the middle-layer planning subsystem as the middle-layer operation target.
5. The target-guided robot hierarchical reinforcement learning control system according to claim 4, wherein The upper reward function of the upper reinforcement learning network of the upper decision-making subsystem is as follows: Among them, s t represents the current state; a H is the action output by the upper-layer decision-making target guiding strategy, that is, the number g of the middle-layer operation target M ; Φ(g H ) represents the state corresponding to the upper-layer decision-making target g H ; ε H is a preset threshold value; The L2 norm is used in this reward function to measure the current state s t and the upper-level decision-making goal g H the proximity to the corresponding state Φ(g H ). If the distance between the two is less than the preset threshold ε H , it is considered that the upper-level decision-making goal of the operation task has been achieved, and at this time the reward function returns 1; otherwise, it returns 0.
6. The robot hierarchical reinforcement learning control system based on target guidance according to claim 5, wherein The upper-layer reinforcement learning network includes: a current value network and a target value network, adopts a target-guided double deep Q-learning algorithm, and combines experience replay, mini-batch and soft update to update the target value network. The definition of the experience sample ξ of the upper-layer decision-making target-guided policy of this upper-layer reinforcement learning network is: Among them, s t represents the current state; g H represents the desired upper-level decision-making goal; a H is the action output by the upper-level decision-making goal guiding strategy, that is, the number of the middle-level operation goal g M ; is the upper-level reward function; is an indicator function used to evaluate the progress of the selected middle-level operation goal, s t+1 represents the next state, represents the distance error between the position of the end effector of the robotic arm and the target position, ε x represents the distance error threshold, represents the speed of the end effector of the robotic arm, ε v represents the speed threshold; In the discrete action space, the update rule for the target value function Q(s t , g H ) is as follows: where ν represents the learning rate used to control the update step size; represents the completion of the objective value function, which is defined as: Among them, are the parameters of the target value network; Introduce two independent value networks, each equipped with a target value network. The minimum value of the two target value networks is used as the estimated value, and the gradient is calculated through the following loss function and the network parameters are updated Among them, represents the experience replay buffer, and ξ represents the experience sample of the upper decision target guiding policy of the upper reinforcement learning network. When updating the target value network parameters through soft update, the parameter τ is used to update the target value network parameters, which is: θ target ← τθ+(1 - τ)θ target ; Select an action according to the maximum value output by the current value network. The action is the middle-layer operation target number, and select the corresponding middle-layer operation target in the target library according to the middle-layer operation target number and transfer it to the middle-layer planning subsystem.
7. The target-guided robot hierarchical reinforcement learning control system according to claim 1 or 2, characterized in that The middle-layer planning subsystem uses a flexible actor-critic algorithm to generate the planning action of the robot according to the current state and the middle-layer operation target through the planning target-guided policy, and transfer the expected motion trajectory of the planning action to the lower-layer control subsystem; The objective function J(π) of this flexible actor-critic algorithm is: where, \(t\) represents the current time; \(T\) represents the total duration of task execution; denotes the expectation; \(\rho\) π represents the distribution of state - action pairs generated by the policy \(\pi\); \(r(s\) t , \(a\) t ) represents the reward function; is an adaptive temperature parameter used to control the weight of entropy in the objective function to control the randomness of the policy \(\pi\), denotes the set of real numbers; represents the entropy of the policy \(\pi\) in the state \(s\) t which is used to measure the uncertainty of the policy \(\pi\); The flexible actor-critic algorithm includes: a flexible value function, a flexible Q function and an adaptive temperature parameter, where, The flexible value function V(s) represents the expected return that can be obtained by executing an action according to the current policy in state s t while considering the entropy of the action. The flexible value function V(s) is expressed as: The flexible Q - function evaluation represents the expected return when taking action a t in state s t and following the current policy. The flexible Q - function is represented as: The flexible actor-critic algorithm adjusts the adaptive temperature parameter α to make the policy entropy close to a desired target entropy The representation of the adaptive temperature parameter α is as follows: The flexible actor-critic algorithm uses the gradient ascent method to optimize the policy π by maximizing the objective function. At the same time, flexible policy updates are introduced, and the policy parameters are gradually updated by minimizing the KL divergence. The update formula for the policy parameters π new is as follows: Among them, π′ is the sampling in the control policy distribution; Π is the distribution of the control policy represented by the actor network; D KL is to minimize the KL divergence of the policy; s t is the state of the robotic arm trajectory tracking task at time t as a robot; α represents the adaptive temperature parameter; is the state-action value function before update; is the partition function for normalizing the distribution; The reward function of the middle-layer planning subsystem is: Among them, represents the position error at the current time t, which is composed of the position error of the end effector and the position error of the object to be manipulated; W x represents the corresponding weight parameter; is used to motivate the gripper to reach the desired opening and closing state; is used to penalize excessive contact forces that occur during the operation. When the contact force exceeds the preset threshold ε ext , it will be penalized; represents the reward for task completion. When the position error is less than ε x and the speed of the gripper is less than the threshold ε v , a reward will be obtained.
8. The target-guided robot hierarchical reinforcement learning control system according to claim 7, wherein The middle-level planning subsystem also adopts the off-policy experience replay algorithm. For any state, there exists a corresponding completed goal. During each sampling process, not only the original goal is considered, but also a series of off-policy goals are considered. These off-policy goals are generated based on the current state or trajectory and are used to calculate additional reward signals; For each experience sample sampled from the environment, a series of corresponding off-policy experience samples are generated. Subsequently, these off-policy experience samples and the original experience samples are stored in the replay buffer together, and a batch of experience samples are randomly selected for training during the update process of the flexible actor-critic algorithm.
9. The target-guided robot hierarchical reinforcement learning control system according to claim 1 or 2, characterized in that The lower-level control subsystem includes: an interpolation processing module and an impedance controller in the joint space; where The interpolation processing module is communicatively connected to the middle-level planning subsystem, and can use the actions output by the middle-level planning subsystem as inputs, and can perform high-precision interpolation calculations between adjacent data points of the action data through the position trajectory function S p (k) to obtain the interpolated action trajectory data, and the position trajectory function S p (k) within the time interval T is as follows: Among them, T represents the total duration of task execution, that is, the time interval; k represents the interpolation time point, with a variation range from t to t + T, where t is the current moment; by applying the trajectory function S p (k), the end effector of the robotic arm generates a smooth pose p k , velocity and acceleration trajectories the impedance controller in the joint space is communicatively connected to the interpolation processing module, can take the trajectory output by the interpolation processing module as input, and adjust the movement of the robotic arm according to the trajectory points in the trajectory, performing smooth and precise trajectory tracking and force control on the robotic arm.
10. The target-guided robot hierarchical reinforcement learning control system according to claim 9, characterized in that, The joint control torque corresponding to the impedance controller in the joint space is: where τ M is the joint control torque vector; x d represents the desired gripper pose. The gripper position is adjusted in real time during actual motion according to the actions output by the middle-level planning subsystem, and the pose remains constant; is the predicted value of the Jacobian matrix, the superscript T represents the transpose of the matrix, and the superscript -1 represents the inverse of the matrix; M d is the desired inertia matrix; B d is the desired damping matrix, K d is the desired stiffness matrix; is the desired acceleration; e are the velocity error and position error respectively; is the predicted value of the Coriolis force and centrifugal force terms; θ, and are the joint angle, angular velocity and angular acceleration respectively; is the inertia matrix.
Citation Information
Cited By
Long-time-domain robot autonomous operation method based on hierarchical reinforcement learning
CN121696986A
Multi-modal data acquisition and intelligent control system for unmanned aerial vehicle
CN122064000A