Demonstration learning-based redundant robot task priority hierarchical structure learning method
Through the redundant robot task priority hierarchical learning method based on demonstration learning, the challenge of task priority scheduling and coordination of redundant robots in complex environments is solved, automatic learning and dynamic adjustment of priority hierarchy are realized, and the efficiency and flexibility of task processing are improved.
Patent Information
- Application Number
- CN202510177435.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-13
AI Technical Summary
Existing redundant robots face task priority scheduling and coordination challenges when performing complex tasks, especially in dynamic environments, and prior art is difficult to effectively deal with task priority changes and conflicts.
The redundant robot task priority hierarchy learning method is adopted based on demonstration learning, and automatic learning and dynamic adjustment of priority hierarchy of priority hierarchy is achieved by establishing a task priority parameterized matrix, constructing a generalized zero-space projection matrix, calculating a continuous shaping priority Jacobacter matrix, and using nonlinear optimization and Gaussian hybrid model for learning.
This method can effectively solve the task priority scheduling and coordination problems of redundant robots in complex environments, improve the efficiency and flexibility of multi-task scheduling, and reduce the difficulty of demonstration personnel's teaching.
Smart Images

Figure CN119989923A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of redundant robot demonstration learning, and in particular to a redundant robot task priority hierarchy structure learning method based on demonstration learning. Background Art
[0002] Nowadays, redundant robots are widely used to perform complex operations involving multiple tasks. Although the redundant structure enables robots to handle multiple tasks simultaneously, conflicts between tasks may occur when the task objectives cannot be met simultaneously. To solve this problem, it is usually necessary to assign priorities to tasks. Therefore, complex robot controllers must be able to handle task prioritization and follow various constraints. Currently, a variety of hierarchical control frameworks have been proposed in the field of robotics to manage task objectives. Some frameworks use a strict hierarchy based on null space projection to ensure that critical tasks are completed first, while low-priority tasks are only executed in the null space of high-priority tasks. This method is often used in priority inverse kinematics, acceleration control, and joint torque control. The strict hierarchy ensures that the task priorities are clear, but it also restricts the constraints of the tasks and reduces the number of tasks that can be executed simultaneously when the control degrees of freedom are limited. In addition, the switching of the hierarchy may cause control discontinuity. Another type of framework uses a non-strict task hierarchy to resolve task conflicts through weight tradeoffs. In this structure, low-priority tasks are not restricted to the null space of high-priority tasks, but may affect the execution effect of high-priority tasks.
[0003] In a broader context, robots need to handle both strict and non-strict hierarchies, especially in dynamically changing environments, where the priorities between tasks may change from non-strict priorities to strict priorities. Existing strict hierarchies usually parameterize task priorities through lexicographic order, dividing tasks into strict high priority and low priority. However, this approach has limitations. First, strict priority is only an extreme form of task priority. In practical applications, there is often no clear priority between tasks, especially in complex scenarios, where it is difficult to strictly define priorities. Second, strict priority may be too conservative, restricting the execution of low-priority tasks and affecting system flexibility. In contrast, continuous priority parameterization can more accurately describe the priority relationship between tasks, overcome the limitations of strict priority, and improve the robot's task processing capabilities in dynamic environments. Therefore, how to flexibly handle changes in task priorities has become an important topic of current research.
[0004] As the capabilities of redundant robots increase and the number and complexity of tasks increase, task priority management becomes a bottleneck. Traditional methods of manually pre-defining priorities have many limitations. For example, although safety-critical tasks are easy to distinguish, it is more difficult to prioritize non-safety tasks and cannot effectively cope with dynamic environments. Overly strict priority settings may lead to task conflicts and limit the flexibility of robots. Therefore, it is particularly important to use learning methods to allow robots to autonomously decide on the priority hierarchy structure. Through learning, robots can dynamically adjust priorities according to actual conditions and environmental changes, improving the efficiency and flexibility of multi-task scheduling, especially in complex and changing environments.
[0005] Existing priority hierarchy learning methods can be divided into learning methods based on strict hierarchy and methods related to task soft weighting. The former assumes that low-priority tasks will be projected into the null space of high-priority tasks. Some methods teach task priorities by encoding the relationship between the end position and the desired configuration through neural networks, while others optimize the null space task strategy at runtime through task switching control. However, these methods assume strict task priorities and are difficult to adapt to changes in dynamic environments. Some priority identification methods identify tasks by observing joint trajectories and action space behaviors, but rely on prior knowledge of tasks. Among soft weighting methods, some studies adjust task weights through derivative-free random optimization. Although priorities can be adjusted dynamically, fitness functions and basic tasks need to be predefined. Other methods deal with incompatibility between tasks through random optimization, but fail to fully consider parallel tasks and priority coordination issues. There are also studies that adjust priorities based on Gaussian kernel variance calculations, but rely on predefined kernel centers. Existing priority learning methods rely on prior knowledge or predefined parameters, and fail to effectively extract task variability from demonstrations, resulting in the inability to adaptively adjust priorities during execution. Some methods have also failed to generate a strict task hierarchy, affecting the precise management of task priorities and execution effects. Therefore, it is of great significance to propose a hybrid priority parameter demonstration learning method that can extract task variability, which not only enhances the robot's task processing capability in complex environments, but also provides a more intelligent and reliable solution for multi-task scheduling and priority management. Summary of the invention
[0006] In order to address the deficiencies in the above-mentioned prior art, the purpose of the present invention is to provide a redundant robot task priority hierarchy learning method based on demonstration learning, which can effectively solve the challenges of task priority scheduling, task coordination, etc. faced by redundant robot systems when performing complex tasks, while realizing automatic learning and dynamic adjustment of the priority hierarchy, reducing the difficulty of demonstrator teaching, thereby improving the performance and applicability of the learning method in complex multi-task environments.
[0007] The technical solution adopted by the present invention to solve the technical problem is:
[0008] A method for learning a hierarchy of redundant robot task priorities based on demonstration learning is provided, the method comprising the following steps:
[0009] S1: Establish a parameterized matrix of task priorities;
[0010] S2: Construct a generalized null space projection matrix based on the matrix of S1, and use the generalized projection matrix to project the task Jacobian matrix into a continuous shaping priority Jacobian matrix;
[0011] S3: According to the constraint optimization theory, the continuous shaping priority Jacobian matrix of S2 is used to solve the optimal analytical expression satisfied by each subtask under the hierarchical structure as the hierarchical consistency error;
[0012] S4: Combining the continuous shaping priority Jacobian matrix of S2 and the hierarchical consistency error criterion of S3, a task priority hierarchy learning scheme based on demonstration learning is designed for the redundant robotic arm system. The optimal continuous priority parameters that can minimize the hierarchical consistency error of the demonstration data are solved using a nonlinear optimization method. The obtained continuous priority parameters are modeled and learned through a Gaussian mixture model, and generalized using Gaussian mixture regression. Finally, a probabilistic model mapping from environmental states to priority parameters is obtained, thereby realizing priority hierarchy demonstration learning.
[0013] Furthermore, in S1, a new priority parameterization method is constructed to realize the insertion and removal of tasks at any priority level, as well as the smooth online reorganization of priorities. The complex priority network reflecting the absolute importance of each task in the entire hierarchy can be encoded in matrix form:
[0014]
[0015] Among them, α i,j ∈[0,1], which indicates the absolute priority of task j in the task stack, and satisfies that the elements in each column are arranged in non-descending order, that is, for i1≤i2,
[0016] Furthermore, for strict priority, if the priority of task j is lower than that of task i, that is, the level index h j >i, then α i,j =0; if the priority of task j is not lower than i, that is, the level index h j ≤i, then α i,j =1;
[0017] In the soft priority during transition, 0<α i,j <1, task j is allowed to move insufficiently on level i, α i,jThe larger the value, the more degrees of freedom task j occupies at the i-th level; consider two special cases: (i) if tasks j1 and j2 have the same priority level, then the j1-th column and j2-th column of Ψ have the same elements; (ii) if task j is removed from the stack, then all elements of the j-th column of Ψ are zero.
[0018] Furthermore, in S2, the continuous shaping priority Jacobian matrix is implemented as follows:
[0019] S21: Select the task to be processed at level i: When the algorithm reaches level i, only tasks that satisfy α i,j ≠α i-1,j (for i>1) or α i,j ≠0 (for i=1) will participate in the task and N; the selected task index and priority parameters are expressed as and J ti represents the augmented matrix obtained by concatenating the Jacobian matrices of the selected tasks:
[0020] S22: Calculate the selected task in the i-th layer For each selected task ti(k) (k=1,…,n i ), the continuous shaping priority Jacobian is computed as follows:
[0021]
[0022] For level i, only the continuous shaping priority Jacobian matrix corresponding to the selected task is updated according to the above formula, while other continuous shaping priority Jacobian matrices remain unchanged. This iterative process continues until the lowest priority level is reached, and the initial value of this iteration is Where h = 1,…,r;
[0023] S23: Sort the selected α in descending order i,j : vector Ξ i The elements in are arranged in descending order, and the sorted vectors and their corresponding task indexes are represented as and Then, It will be further expanded according to the dimension of the task:
[0024]
[0025] Among them, 1 m represents an m-dimensional vector consisting of 1s, and m ini(k) represents the dimension of the operation space of task ind(k). In this way, The dimension of will match the dimension of the extended task space;
[0026] S24: The Jacobian matrix of the task corresponding to the sorting: The augmented Jacobian matrix is rearranged according to the vector ind:
[0027]
[0028] In this way, the task Jacobian matrix will be arranged in descending order of task importance: the most important tasks appear in the first few rows, while the least important tasks appear in the last few rows;
[0029] S25: Obtain a truncated orthogonal matrix by QR decomposition: For a positive definite weight matrix W = L T L performs Cholesky decomposition to obtain L, and then, Perform QR decomposition to obtain the truncated orthogonal matrix Q, permutation vector p and rank t, where N i-1 is the continuous shaping projection of the i-1th layer; the t orthogonal column vectors of Q are related to the t linearly independent task space directions (i.e. By doing this, linearly dependent rows are removed from the calculation of the continuous shaping projection because they are not important. The row space orthogonal basis has no effect;
[0030] S26: Obtain a truncated diagonal matrix: by The elements in are placed on the diagonal in order to construct a diagonal matrix:
[0031]
[0032] Among them, I m represents the m×m identity matrix, and then, by selecting The diagonal elements of The diagonal elements of are driven by Ψ during the transition and modulate the activation of the corresponding column vector in Q;
[0033] S27: Obtain the truncated diagonal matrix: The continuous shaping projections of the first i layers are calculated as follows:
[0034]
[0035] Furthermore, in S3, the process of using the constrained optimization theory to solve the optimal analytical expression satisfied by each subtask under the hierarchical structure includes: applying the constrained optimization lemma to the optimization problem. According to the lemma, on the hierarchical consistent trajectory, the following formula holds:
[0036]
[0037] By Nature Finally you can get This is the optimal analytical expression that each subtask satisfies under the hierarchical structure.
[0038] Furthermore, in S4, the method of implementing the priority hierarchy demonstration learning is as follows:
[0039] S41: Using nonlinear optimization method to solve the optimal continuous priority parameter that can minimize the hierarchical consistency error of the demonstration data: The optimization problem of solving the optimal priority hierarchy structure of a certain state is expressed as:
[0040]
[0041] Among them, any priority parameterization matrix Ψ is about the set The center of gravity coordinates are defined as a set of real coefficients w = [w1,…,w l ] T , is the hierarchical consistency error obtained in S3;
[0042] S42: The obtained continuous priority parameters are modeled and learned through a Gaussian mixture model, and generalized using Gaussian mixture regression, and finally a probabilistic model mapping from environmental states to priority parameters is obtained to achieve demonstration learning of priority hierarchy.
[0043] Furthermore, in S42, the method for implementing the demonstration learning of the priority hierarchy is: calculating the priority parameter distribution, that is, the conditional distribution of the priority parameter after a given state:
[0044]
[0045] in represents the weight of the kth sub-model under state s, which is calculated by the following formula:
[0046]
[0047] The calculation formula of the mean and covariance matrix is expressed as:
[0048]
[0049] From this, we can get the expected value of the priority parameter distribution p(w|s) under state s:
[0050]
[0051] From the obtained optimal center of gravity coordinates w, we can calculate Ψ and the corresponding continuous shaping priority Jacobian matrix
[0052] Compared with the prior art, the present invention has the following beneficial effects:
[0053] 1. The redundant robot task priority hierarchy learning method based on demonstration learning in the example of the present invention constructs a task priority hierarchy learning scheme based on statistical machine learning and constrained optimization theory, which includes the establishment of a task priority parameterization matrix, the calculation of a continuous shaping priority Jacobian matrix, the solution of the optimal priority parameters, and the statistical learning of the priority parameters; first, by establishing a task priority parameterization matrix and constructing a generalized null space projection matrix, the continuous shaping priority Jacobian matrix is calculated; then, the hierarchical consistency error of the demonstration data is minimized by a nonlinear optimization method, and the continuous priority parameters are optimized; finally, the priority mapping is generalized and learned through a Gaussian mixture regression model, thereby realizing automatic learning and dynamic adjustment of task priorities;
[0054] 2. The redundant robot task priority hierarchy learning method based on demonstration learning exemplified in the present invention, the entire task priority hierarchy learning method based on demonstration learning consists of two steps: extracting the optimal priority hierarchy from the demonstration and performing generalization learning based on the extracted hierarchy. This method can realize task priority scheduling and task coordination of redundant robots facing complex dynamic environments and unknown interference situations in actual human-machine collaboration. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0056] Figure 1 This is an example of a planar robot conflict task, used to illustrate the meaning of hierarchical consistency error;
[0057] Figure 2 A planar 6-DOF robotic arm for experimental verification;
[0058] Figure 3 is the optimal priority parameter at each moment extracted from the demonstration;
[0059] Figure 4 The optimal hierarchical consistency error corresponding to the optimization problem for extracting the optimal priority parameters;
[0060] Figure 5 is the generalized learning result of the priority parameter and the current state. DETAILED DESCRIPTION
[0061] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant inventions, rather than to limit the inventions. It should also be noted that, for ease of description, only the parts related to the invention are shown in the accompanying drawings.
[0062] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0063] Embodiment 1:
[0064] This embodiment provides a redundant robot task priority hierarchy learning method based on demonstration learning, comprising the following steps:
[0065] S1: Establish a parameterized matrix of task priorities;
[0066] S2: Construct a generalized null space projection matrix based on the matrix of S1, and use the generalized projection matrix to project the task Jacobian matrix into a continuous shaping priority Jacobian matrix;
[0067] S3: According to the constraint optimization theory, the continuous shaping priority Jacobian matrix of S2 is used to solve the optimal analytical expression satisfied by each subtask under the hierarchical structure as the hierarchical consistency error;
[0068] S4: Combining the continuous shaping priority Jacobian matrix of S2 and the hierarchical consistency error criterion of S3, a task priority hierarchy learning scheme based on demonstration learning is designed for the redundant robotic arm system. The optimal continuous priority parameters that can minimize the hierarchical consistency error of the demonstration data are solved using a nonlinear optimization method. The obtained continuous priority parameters are modeled and learned through a Gaussian mixture model, and generalized using Gaussian mixture regression. Finally, a probabilistic model mapping from environmental states to priority parameters is obtained, thereby realizing priority hierarchy demonstration learning.
[0069] In this embodiment, the specific process of establishing the task priority parameterization matrix in S1 is:
[0070] Traditional hierarchical tracking control methods based on strict hierarchy use standard lexicographical hierarchy to parameterize the importance level of each task. Assume that given r tasks assigned by the user, each task has a corresponding hierarchy h i As the level index increases, the priorities of tasks are arranged in descending order (e.g., h i =1 indicates that task i has the highest priority). In addition, different tasks may have the same priority, which means that their level indexes may be equal. These tasks are represented by r task coordinate vectors x i where each task is represented by a differentiable function fi To define:
[0071]
[0072] Index the level h i Sort in ascending order, that is, sort the tasks from high to low priority, and get the corresponding priority task index set {n1,…,n r In order to inertially decouple tasks with different priorities, the original Jacobian matrix is projected to obtain the priority Jacobian matrix corresponding to i=1,…,r:
[0073]
[0074] Among them, N i (q) is the dynamically consistent null space projection matrix. N i (q) can be calculated by the following recursive algorithm:
[0075]
[0076] where I is the identity matrix of appropriate size; represents the weighted generalized inverse The dynamically consistent pseudo-inverse of , the weighted matrix is the inertia matrix M(q).
[0077] However, the lexicographic order hierarchy cannot describe the transition state from one strict hierarchy to another, so a new priority parameterization method needs to be designed to realize the insertion and removal of tasks at any priority hierarchy and the smooth online reorganization of priorities.
[0078] Embodiment 2:
[0079] In this embodiment, the specific process of establishing the task priority parameterization matrix in S1 is:
[0080] During the task priority reorganization process, the continuous change of the priority parameter drives the task stack from one strict priority to a non-strict priority and then to another different strict priority. The complex priority network reflecting the absolute importance of each task in the entire hierarchy can be encoded in the matrix form:
[0081]
[0082] where l is not necessarily equal to r, because there may be multiple tasks at the same priority level. In addition, α i,j ∈[0,1] represents the absolute priority of task j in the task stack, and satisfies that the elements in each column are arranged in non-descending order, that is, for i1≤i2,
[0083] For strict priority, if task j has a lower priority than task i, that is, level index h j >i, then α i,j =0; if the priority of task j is not lower than i, that is, the level index h j ≤i, then α i,j = 1. In the soft priority during the transition period, 0 < α i,j <1, task j is allowed to move insufficiently on level i, α i,j The larger the value, the more degrees of freedom task j occupies at the i-th level; consider two special cases: (i) if tasks j1 and j2 have the same priority level, then the j1-th column and j2-th column of Ψ have the same elements; (ii) if task j is removed from the stack, then all elements of the j-th column of Ψ are zero.
[0084] Using this priority parameterization method, three typical strict lexicographic order levels T1>T2>T3>T4, T2>[T1,T4], and T2>T3>[T1,T4] can be expressed as:
[0085]
[0086] In this implementation example S2, a generalized null space projection matrix is constructed according to the priority parameterization matrix in S1, and the generalized projection matrix is used to calculate the continuous shaping priority Jacobian matrix. The specific process is:
[0087] This step recursively calculates the continuous shaping priority Jacobian matrix based on the matrix Ψ in order of priority And the projection matrix N. In each recursion, the first i layers and N are obtained through the following 7 steps:
[0088] S21: Select the task to be processed at level i: When the algorithm reaches level i, only tasks that satisfy α i,j ≠α i-1,j (for i>1) or α i,j ≠0 (for i=1) will participate in the task and N; the selected task index and priority parameters are expressed as and J ti represents the augmented matrix obtained by concatenating the Jacobian matrices of the selected tasks:
[0089] S22: Calculate the selected task in the i-th layer For each selected task ti(k) (k=1,…,n i ), the continuous shaping priority Jacobian is computed as follows:
[0090]
[0091] For level i, only the continuous shaping priority Jacobian matrix corresponding to the selected task is updated according to the above formula, while other continuous shaping priority Jacobian matrices remain unchanged. This iterative process continues until the lowest priority level is reached, and the initial value of this iteration is Where h = 1,…,r;
[0092] S23: Sort the selected α in descending order i,j : vector Ξ i The elements in are arranged in descending order, and the sorted vectors and their corresponding task indexes are represented as and Then, It will be further expanded according to the dimension of the task:
[0093]
[0094] Among them, 1 m represents an m-dimensional vector consisting of 1s, and m ini(k) represents the dimension of the operation space of task ind(k). In this way, The dimension of will match the dimension of the extended task space;
[0095] S24: The Jacobian matrix of the task corresponding to the sorting: The augmented Jacobian matrix is rearranged according to the vector ind:
[0096]
[0097] S25: Obtain a truncated orthogonal matrix by QR decomposition: For a positive definite weight matrix W = L T L performs Cholesky decomposition to obtain L, and then, Perform QR decomposition to obtain the truncated orthogonal matrix Q, permutation vector p and rank t, where N i-1 is the continuous shaping projection of the i-1th layer; the t orthogonal column vectors of Q are related to the t linearly independent task space directions (i.e. By doing this, linearly dependent rows are removed from the calculation of continuous shaping projections because they are not related to The row space orthogonal basis has no effect;
[0098] S26: Obtain a truncated diagonal matrix: by The elements in are placed on the diagonal in order to construct a diagonal matrix:
[0099]
[0100] Among them, Im represents the m×m identity matrix, and then, by selecting The diagonal elements of The diagonal elements of are driven by Ψ during the transition and modulate the activation of the corresponding column vector in Q;
[0101] S27: Obtain the truncated diagonal matrix: The continuous shaping projections of the first i layers are calculated as follows:
[0102]
[0103] In S3, the process of using constrained optimization theory to solve the optimal analytical expression for each subtask in the hierarchical structure includes:
[0104] If all control tasks are fully hierarchically feasible, then we can eventually perfect every task in the hierarchy. This can be demonstrated by the tracking error in task space. and the corresponding augmented tracking error The control objective can be written as when t→∞, However, this is not true for tasks that are feasible at some levels. This goal can be described by the concept of optimization, that is, the objective function Under the constraints Minimize, where represents the optimal value of task 1. The final robot configuration is the solution to the constrained optimization problem:
[0105]
[0106]
[0107] If the task dimension is smaller than the degree of freedom of the robot system, the solution of the above equation is not unique. We can further extend the above equation to each level, that is:
[0108]
[0109] For j=1,…,i-1, we have and Assume that all degrees of freedom are used to perform the tasks and that each individual task is feasible due to the absence of kinematic singularities. The feasibility of the equality constraints at level i depends on the solutions to the first i-1 optimization problems, so the feasibility of the entire constrained optimization problem can be obtained recursively. By recursively solving the optimization problem for i=1,…,r, the solution is the only certainty. Under the hierarchical constraints, all The constrained local minimum, and we call the corresponding a hierarchically consistent trajectory. If x i is not completely hierarchically feasible, then is different from x i,ids To avoid solving the non - linear optimization problem in each optimization cycle to calculate we introduce the following lemma: Let q = q0 be a joint configuration, at which V i (q) is minimized / maximized under the constraint x = f(q). Then, the inner product is zero, where Z(q) satisfies
[0110] Applying the constrained optimization lemma to our optimization problem can nullify and the matrix with the same dimension as can be chosen as because when i < j According to the lemma, on the hierarchically consistent trajectory, the following formula holds:
[0111]
[0112] By property Finally, we can obtain This is the optimal analytical expression satisfied by each subtask under the hierarchical structure.
[0113] In S4, our method consists of two components: 1) Hierarchical structure identification: Through a set of candidate task priorities and the reference trajectories of each subtask, we identify the priority hierarchy corresponding to the demonstration based on the demonstrated trajectories. 2) Motion synthesis: Using the designed learning algorithm, reproduce the taught priority behavior in a new scenario.
[0114] The hierarchically consistent error proposed in S3 describes the optimal performance that each subtask can achieve under the priority hierarchy. When the hierarchy is a strict priority, each subtask can achieve the optimal performance without affecting the high - priority tasks (i.e., in the null space of the high - priority tasks), and at this time That is to say, when the priority parameterization matrix Ψ of S1 correctly describes the current strict priority hierarchy, the continuous shaping priority Jacobian matrix calculated by S2 can make the hierarchically consistent error zero. For the sake of illustration, we take a planar robot as an example, such as Figure 1. Task 1 has a higher priority and requires controlling the position of the end effector; Task 2 has a lower priority and requires controlling the position of the last joint. Obviously, there is a conflict between Task 1 and Task 2, because the end effector and the last joint are connected by a rigid link, and their absolute distance is constant, which can be regarded as a constraint. In hierarchical control, Task 2 will not interfere with the execution of Task 1. In this case, Task 1 will asymptotically track the target trajectory (blue circle), while we want the tracking error of Task 2 to be as small as possible. As a result, the actual position of the last joint is on the straight line between the target point of Task 2 and the actual position of the end effector and satisfies .
[0115]
[0116] Based on this point of view, we hope to use nonlinear optimization to identify the priority parameterization matrix. By optimizing the priority parameters, the hierarchical consistency error is minimized. The optimal priority parameters at this time are the priority hierarchy structure corresponding to the demonstration trajectory. The priority parameterization matrix needs to satisfy various nonlinear constraints in S1 at all times during optimization, which is not conducive to solving nonlinear optimization problems. For this reason, we introduce barycentric coordinates to make the constraints linear with respect to the optimization variables. Assume that the basic task set is T = (T1,…,T r ), the candidate strict hierarchy set is [T1, T2] represents the strict priority of T1 over T2. The set calculated by S1 The corresponding strict priority parameterized matrix set is At this time, any priority parameterization matrix Ψ is about the set The center of gravity coordinates are defined as a set of real coefficients w = [w1,…,w l ] T , the center coordinates satisfy the following properties: 1) Non-negativity: w i ≥0, 2) Linear: and At this time, the optimization variables of the new optimization problem are converted into the barycentric coordinates w i Therefore, the optimization problem of solving the optimal priority hierarchy for a certain state is expressed as:
[0117]
[0118]
[0119] Where r is the total number of tasks. For each time t = 1, ..., T, there is a corresponding robot and environment state s t And the priority parameterization matrix Ψ obtained from the above optimization problem tFor the priority learning task, assume that the demonstration dataset contains N trajectories, represented as The i-th trajectory X i ∈X consists of the states {w t ,s t}. If different demonstration trajectories have different lengths T, the data is preprocessed using the generalized time warping method. For the demonstration data set, first train the corresponding Gaussian mixture model The state quantity is y = [w T ,s T ] T , β k is a weight coefficient and satisfies The mean vector is The covariance matrix is Model unknown parameters Estimation is performed using the expectation maximization algorithm. After the probability model training is completed, the probability model of the mapping from state to priority parameters can be obtained through Gaussian mixture regression. This model can calculate the priority parameter distribution, that is, the conditional distribution of the priority parameter after a given state. in represents the weight of the kth sub-model under state s, which is calculated by the following formula:
[0120]
[0121] The calculation formula of the mean and covariance matrix is expressed as:
[0122]
[0123] From this, we can get the expected value of the priority parameter distribution p(w|s) under state s:
[0124]
[0125] From the obtained optimal center of gravity coordinates w, we can calculate Ψ and the corresponding continuous shaping priority Jacobian matrix
[0126] Verification analysis:
[0127] In order to verify the effectiveness of the proposed priority hierarchy learning method, we conducted a simulation experiment on a 6-DOF planar robot, as shown in the schematic diagram. Figure 2As shown. To simplify the model, we assume that the mass of each robot link is concentrated at its center point and has a mass of 1 kg. The length of the first link is 0.25 m, and the lengths of the remaining links are all 0.5 m. In this simulation, the initial configuration is q(0) = [0m, 40°, -40°, -50°, -40°, -50°] T , the gravitational acceleration is simulated as In this simulation, the following five tasks are arranged: Task 1: x-coordinate of TCP, y-coordinate of TCP; Task 2: x-coordinate of J6, y-coordinate of J6; Task 3: Angle of J2 about z-axis; Task 4: Angle of J3 about z-axis, angle of J2 / J3 about z-axis, angle of TCP about z-axis; Task 5: x-coordinate of J1, x-coordinate of J1 / TCP. TCP represents the tool center point, and J1 to J6 represent the 1st to 6th joints respectively. J2 / J3 represents the coupling joint angle q(2)+q(3). J1 / TCP represents the position difference between TCP and joint J1 in the x-axis direction, as shown in Figure 2 The origin of the task space is defined as the task coordinate of the zero configuration, and its zero configuration is q0 = [0m, 45°, -45°, -45°, -45°] T (Right now Figure 2 The tasks are assigned to different levels and their priorities are adjusted in the following order: [T1, T2, T3, T4, T5] → [T3, T2, T1, T5, T4] → [T3, T2, T5, T4] → [T3, T2, T1, T5, T4]. The time-based smooth transition of the priority parameter matrix is driven by the parameter β, which is expressed as Ψ = (1-β)Ψ1 + βΨ2, where Ψ1 and Ψ2 represent the priority parameter matrices before and after the transition, respectively. The parameter β is a time-modulated continuous transition parameter, expressed as β = 6Δt 5 -15Δt 4 +10Δt 3 , where Δt represents the transition duration, and the transition ends at 1 second. All joint angles and task coordinates of the demonstration trajectory are recorded for training the learning model. We performed four demonstrations in total, and input the obtained demonstration data into the optimization model to obtain the optimal priority parameter w at each moment. oot =[w1,w2,w3] T , the results are as follows Figure 3 The corresponding objective function value is shown as Figure 4 shown. Figure 3 This is consistent with our settings. It can be seen that our framework can accurately find the optimal priority parameters for each moment of the demonstration trajectory. The priority parameters and time components of the four demonstration trajectories are obtained. A Gaussian mixture model was trained as training data with Gaussian components K = 25, and then Gaussian mixture regression was used to obtain the reproducible priority parameters, such as Figure 5 The experimental results show that the reproduced priority parameters reproduce the priority change characteristics in the demonstration, and realize the extraction of the hierarchical structure and generalization learning from the demonstration.
[0128] Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A redundant robot task priority hierarchy learning method based on demonstration learning, characterized in that: The method comprises the following steps: S1: Establish a parameterized matrix of task priorities; S2: Construct a generalized null space projection matrix based on the matrix of S1, and use the generalized projection matrix to project the task Jacobian matrix into a continuous shaping priority Jacobian matrix; S3: According to the constraint optimization theory, the continuous shaping priority Jacobian matrix of S2 is used to solve the optimal analytical expression satisfied by each subtask under the hierarchical structure as the hierarchical consistency error; S4: Combining the continuous shaping priority Jacobian matrix of S2 and the hierarchical consistency error criterion of S3, a task priority hierarchy learning scheme based on demonstration learning is designed for the redundant robotic arm system. The optimal continuous priority parameters that can minimize the hierarchical consistency error of the demonstration data are solved using a nonlinear optimization method. The obtained continuous priority parameters are modeled and learned through a Gaussian mixture model, and generalized using Gaussian mixture regression. Finally, a probabilistic model mapping from environmental states to priority parameters is obtained, thereby realizing priority hierarchy demonstration learning.
2. The redundant robot task priority hierarchy learning method based on demonstration learning according to claim 1 is characterized in that: In S1, a new priority parameterization method is constructed to realize the insertion and removal of tasks at any priority level, as well as the smooth online reorganization of priorities. The complex priority network reflecting the absolute importance of each task in the entire hierarchy can be encoded in matrix form: Among them, α i,j ∈[0,1], which indicates the absolute priority of task j in the task stack, and satisfies that the elements in each column are arranged in non-descending order, that is, for i1≤i2, 3. The redundant robot task priority hierarchy learning method based on demonstration learning according to claim 2 is characterized in that: For strict priority, if task j has a lower priority than task i, that is, level index h j >i, then α i,j =0; if the priority of task j is not lower than i, that is, the level index h j ≤i, then α i,j =1; In the soft priority during transition, 0<α i,j <1, task j is allowed to move insufficiently on level i, α i,j The larger the value, the more degrees of freedom task j occupies at the i-th level; consider two special cases: (i) if tasks j1 and j2 have the same priority level, then the j1-th column and j2-th column of Ψ have the same elements; (ii) if task j is removed from the stack, then all elements of the j-th column of Ψ are zero.
4. The redundant robot task priority hierarchy learning method based on demonstration learning according to claim 2 is characterized in that: In S2, the continuous shaping priority Jacobian matrix is implemented as follows: S21: Select the task to be processed at level i: When the algorithm reaches level i, only tasks that satisfy α i,j ≠α i-1,j (for i>1) or α i,j ≠0 (for i=1) will participate in the task and N; the selected task index and priority parameters are expressed as and J ti represents the augmented matrix obtained by concatenating the Jacobian matrices of the selected tasks: S22: Calculate the selected task in the i-th layer For each selected task ti(k) (k=1,…,n i ), the continuous shaping priority Jacobian is computed as follows: For level i, only the continuous shaping priority Jacobian matrix corresponding to the selected task is updated according to the above formula, while other continuous shaping priority Jacobian matrices remain unchanged. This iterative process continues until the lowest priority level is reached, and the initial value of this iteration is Where h = 1,…,r; S23: Sort the selected α in descending order i,j : vector Ξ i The elements in are arranged in descending order, and the sorted vectors and their corresponding task indexes are represented as and Then, It will be further expanded according to the dimension of the task: Among them, 1 m represents an m-dimensional vector consisting of 1s, and m ind(k) represents the dimension of the operation space of task ind(k). In this way, The dimension of will match the dimension of the extended task space; S24: The Jacobian matrix of the task corresponding to the sorting: The augmented Jacobian matrix is rearranged according to the vector ind: S25: Obtain a truncated orthogonal matrix by QR decomposition: For a positive definite weight matrix W = L T L performs Cholesky decomposition to obtain L, and then, Perform QR decomposition to obtain the truncated orthogonal matrix Q, permutation vector p and rank t, where N i-1 is the continuous shaping projection of the i-1th layer; S26: Obtain a truncated diagonal matrix: by The elements in are placed on the diagonal in order to construct a diagonal matrix: Among them, I m represents the m×m identity matrix, and then, by selecting The diagonal elements of The diagonal elements of are driven by Ψ during the transition and modulate the activation of the corresponding column vector in Q; S27: Obtain the truncated diagonal matrix: The continuous shaping projections of the first i layers are calculated as follows:
5. The redundant robot task priority hierarchy learning method based on demonstration learning according to claim 4 is characterized in that: In S3, the process of using constrained optimization theory to solve the optimal analytical expression satisfied by each subtask under the hierarchical structure includes: applying the constrained optimization lemma to the optimization problem. According to the lemma, on the hierarchical consistent trajectory, the following formula holds: By Nature Finally you can get This is the optimal analytical expression that each subtask satisfies under the hierarchical structure.
6. The redundant robot task priority hierarchy learning method based on demonstration learning according to claim 5 is characterized in that: In S4, the way to implement priority hierarchy demonstration learning is: S41: Using nonlinear optimization method to solve the optimal continuous priority parameter that can minimize the hierarchical consistency error of the demonstration data: The optimization problem of solving the optimal priority hierarchy structure of a certain state is expressed as: Among them, any priority parameterization matrix Ψ is about the set The center of gravity coordinates are defined as a set of real coefficients w = [w1,…,w l ] T , is the hierarchical consistency error obtained in S3; S42: The obtained continuous priority parameters are modeled and learned through a Gaussian mixture model, and generalized using Gaussian mixture regression, and finally a probabilistic model mapping from environmental states to priority parameters is obtained to achieve demonstration learning of priority hierarchy.
7. The redundant robot task priority hierarchy learning method based on demonstration learning according to claim 6 is characterized in that: In S42, the method for implementing the demonstration learning of the priority hierarchy is: calculating the priority parameter distribution, that is, the conditional distribution of the priority parameter after a given state: in represents the weight of the kth sub-model under state s, which is calculated by the following formula: The calculation formula of the mean and covariance matrix is expressed as: From this, we can get the expected value of the priority parameter distribution p(w|s) under state s: From the obtained optimal center of gravity coordinates w, we can calculate Ψ and the corresponding continuous shaping priority Jacobian matrix