Robot teaching obstacle avoidance method based on neural network and dynamic system model

By introducing neural networks and dynamic system models into robot teaching technology, the decision function, modulation matrix and energy function are designed, and the problem of robots being difficult to avoid obstacles independently in unstructured environments in the existing technology is solved, and higher anti-interference ability and task completion accuracy are achieved.

CN120095814AActive Publication Date: 2025-06-06SOUTH CHINA UNIV OF TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510295853.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-06
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

Existing robot teaching technology is difficult to avoid obstacles independently in an unstructured environment, and it is highly time-dependent, making it prone to interruption and damage to tasks due to sudden interference.

Method used

The robot teaching obstacle avoidance method based on neural network and dynamic system models is adopted. The human demonstration trajectory is learned through a nonlinear dynamic system model, and the decision function and modulation matrix are designed to extract task characteristics and obstacle avoidance characteristics, and the energy function is used to ensure that the trajectory converges to the unique target point.

Benefits of technology

The ability to continue to complete tasks after being interrupted or interfered is realized, the system's anti-interference ability is enhanced, and it can respond to changes in locations of dynamic obstacles in real time, ensuring that the robot can complete tasks accurately and stably.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120095814A_ABST
    Figure CN120095814A_ABST
Patent Text Reader

Abstract

The invention discloses a robot teaching obstacle avoidance method based on a neural network and a dynamic system model. The method comprises the following steps: acquiring a demonstration data set; coding the current robot state into a high-dimensional feature vector; nonlinear feature components are extracted from the encoded features, weighted summation is carried out on the features, quadratic terms are added, and an energy function is constructed; constructing a decision function, and multiplying the state and environment information of the robot by corresponding weights to obtain the output of the decision function; inputting the output of the decision function into a multi-task hybrid neural network, generating a parameter vector, constructing a lower triangular matrix based on the parameter vector, and constructing a modulation matrix through cholosky decomposition; constructing a loss function, and obtaining an optimal model parameter through a nonlinear programming method; and taking the output of the neural dynamic system model as a robot control strategy to complete a specified task. According to the invention, by introducing the dynamic system model and the neural network, the bottleneck problems of time dependence and dynamic obstacle avoidance are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot control technology, and in particular to a robot teaching obstacle avoidance method based on a neural network and a dynamic system model. Background Art

[0002] As the production environment of robots becomes more and more complex and interference increases, it is difficult to meet the needs of robots mastering human skills through offline programming in an unstructured environment. Therefore, it is necessary to use teaching programming to enable robots to master human skills, such as drag teaching.

[0003] Robots can master human skills more quickly through teaching, but many current robot teaching technologies are time-dependent and cannot autonomously avoid obstacles or return to the task status after avoiding obstacles. If interference or obstacles occur in the robot assembly space due to sudden factors or accidents in actual production scenarios, and conventional teaching methods do not have the ability to avoid obstacles, the robot will be damaged due to collisions, and even interfere with the operation of the production line, causing huge economic losses.

[0004] For example, the patent publication number CN119077728A is a robot autonomous learning and adaptive trajectory planning method for human-machine collaboration. This method uses dynamic primitives to learn human teaching samples to obtain generalized trajectories, uses a particle swarm algorithm to quickly search for local obstacle avoidance trajectories based on generalized trajectories, and combines generalized trajectories and obstacle avoidance trajectories of particle swarm algorithms to achieve task completion and obstacle avoidance integration. However, this method is based on a dynamic primitive teaching algorithm with high time dependence, and if it encounters interference and causes interruption of operation, it will cause task failure;

[0005] For example, patent publication number CN109702744A is a robot imitation learning method based on a dynamic system model. This method uses a dynamic system model to learn human teaching samples and generate generalized trajectories. It uses control theory to construct a Gaussian mixture model as a nonlinear dynamic system model. By adding stability constraints, it ensures that the model converges stably to the task target point and ensures that the task can be completed. However, although this method uses a dynamic system teaching algorithm with low time dependence, it does not detect and avoid obstacles in the environment.

[0006] It can be seen that the limitations of mainstream robot teaching schemes such as the time-dependent dynamic primitive (DMP) method are particularly obvious in unstructured environments. For example, when the robot is interrupted by a sudden interference (such as temporary obstacle insertion or sensor noise), the DMP method needs to re-plan the time parameters to resume the task, which not only increases the computational complexity, but may also cause the robot to lose control and cause serious losses due to time synchronization problems. In addition, although the traditional dynamic system method ensures the convergence of the target through stability constraints, its obstacle avoidance function usually relies on the static environment assumption and cannot respond to the position changes of dynamic obstacles in real time. Summary of the invention

[0007] In order to overcome the defects and shortcomings of the prior art, the present invention provides a robot teaching obstacle avoidance method based on a neural network and a dynamic system model. The present invention learns human demonstration trajectories based on a nonlinear dynamic system model, designs a modulation matrix combined with a decision function in the dynamic system model to simultaneously learn the task characteristics and obstacle avoidance characteristics in the demonstration, designs an energy function containing a neural network to ensure that the trajectory has a unique task target point, designs constraints that ensure that the model can converge stably to ensure that all generalized trajectories can reach the task target point, and obtains the optimal model parameters by converting the parameter learning problem in the model into an optimization problem containing constraints, inputs the task starting point and obstacle information into the dynamic system model to obtain a generalized trajectory to guide the robot system to complete the task. The present invention solves the bottleneck problems of time dependence and dynamic obstacle avoidance by introducing a dynamic system model and a neural network.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] The present invention provides a robot teaching obstacle avoidance system based on a neural network and a dynamic system model, comprising: a demonstration data acquisition module, a decision function building module, a modulation matrix building module and an energy function building module, and a control strategy output module;

[0010] The demonstration data acquisition module is used to collect human demonstration data and obstacle information to generate a demonstration data set;

[0011] The decision function building module is used to fuse state and environmental information to build a decision function, dynamically allocate the weights of task tracking and obstacle avoidance behaviors using the decision function, and generate an adaptive control strategy based on the robot state and environmental information;

[0012] The modulation matrix building module is used to build a modulation matrix through a multi-task hybrid neural network. The modulation matrix dynamically adjusts the direction and amplitude of the robot's movement and learns to reproduce complex trajectory characteristics and obstacle avoidance behaviors.

[0013] The energy function construction module is used to extract nonlinear feature components from the features encoded by the robot state through a neural network, weight the features and add quadratic terms to construct an energy function, and make the trajectory converge to a unique target point based on the energy function;

[0014] The control strategy output module is used to construct a loss function, obtain the optimal parameters of the neural dynamic system model through a nonlinear programming method, and use the output of the neural dynamic system model as a robot control strategy to complete a specified task.

[0015] The present invention also provides a robot teaching obstacle avoidance method based on a neural network and a dynamic system model, provided with the robot teaching obstacle avoidance system based on a neural network and a dynamic system model, comprising the following steps:

[0016] Collect human demonstration data and obstacle information to generate a demonstration data set;

[0017] Encode the current robot state into a high-dimensional feature vector;

[0018] Extract nonlinear feature components from the encoded features through a neural network, sum the features by weights and add quadratic terms to construct an energy function;

[0019] The decision function is constructed by integrating the state and environmental information. The robot's state and environmental information are multiplied by the corresponding weights to obtain the output of the decision function.

[0020] The output of the decision function is input into the multi-task hybrid neural network to generate a parameter vector, a lower triangular matrix is ​​constructed based on the parameter vector, and a modulation matrix is ​​constructed through cholosky decomposition;

[0021] Construct a loss function, transform the parameter learning problem into a constrained optimization problem, and obtain the optimal model parameters through nonlinear programming methods;

[0022] The output of the neural dynamic system model is used as the robot control strategy to complete the specified task.

[0023] As a preferred technical solution, human demonstration data and obstacle information are collected to generate a demonstration data set, which is specifically expressed as follows:

[0024]

[0025] Among them, x t,n is the position state of the robot at time t in the nth demonstration, is the speed state, x obj Obstacle information in the environment.

[0026] As a preferred technical solution, a nonlinear feature component is extracted from the encoded features through a neural network, the features are weighted and summed, and a quadratic term is added to construct an energy function, which specifically includes:

[0027] The nonlinear feature components are extracted from the encoded features and expressed as:

[0028]

[0029] Among them, a k and b k is a learnable feature parameter, σ is an activation function, c(x) represents the encoded feature, g k (x) represents the nonlinear characteristic component;

[0030] All nonlinear characteristic components g k (x) Through the weight vector ω, we can get the neural network part P 1 (x):

[0031]

[0032] Construct the quadratic term P 2 (x):

[0033] P2(x)=κxTx(κ>0)

[0034] Among them, κ is a positive number;

[0035] The energy function is expressed as:

[0036] V(x)=P 1 (x)-P 1 (0)+P 2 (x).

[0037] As a preferred technical solution, the encoded feature c(x) is specifically expressed as:

[0038]

[0039] Among them, ∈>0 is the adjustment parameter, and x represents the robot state.

[0040] As a preferred technical solution, the parameter constraints of the energy function are:

[0041]

[0042] Among them, ω represents the weight vector.

[0043] As a preferred technical solution, the decision function is constructed by integrating the state and environmental information. The state and environmental information of the robot are multiplied by the corresponding weights to obtain the output of the decision function, which specifically includes:

[0044] Calculate the distance between the robot and the obstacle, and dynamically allocate the parameter α for controlling task tracking through the Sigmoid function 1 and obstacle avoidance parameter α 2 The weight of

[0045] Multiply the robot's state x and environment information z by the corresponding weight α 1 and α 2 , and get the output of the decision function, which is specifically expressed as:

[0046] D(x,z)=[α 1 x,α 2 z].

[0047] As a preferred technical solution, the environmental information z is expressed as:

[0048]

[0049] Among them, x represents the current state of the robot, x obj represents the center position of the obstacle, represents the speed of the robot, Represents the gradient of the energy function.

[0050] As a preferred technical solution, a lower triangular matrix is ​​constructed based on the parameter vector and a modulation matrix is ​​constructed by cholosky decomposition, which is specifically expressed as:

[0051]

[0052] Where D(x,z) represents the output of the decision function, ∈>0 is the adjustment parameter, and I is the unit matrix.

[0053] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0054] (1) The present invention realizes the robot teaching and obstacle avoidance functions by introducing a dynamic system model. Different from other teaching methods with strong time dependence, the present invention can continue to complete the task after being interrupted or interfered, thereby enhancing the anti-interference ability of the system.

[0055] (2) The present invention extracts task tracking features and obstacle avoidance features from human demonstrations through a decision function module, so that when the robot autonomously generalizes, it can respectively implement task tracking and obstacle avoidance functions according to whether it encounters obstacles.

[0056] (3) The present invention constructs an energy function and a modulation matrix through a neural network. The energy function ensures that the dynamic system can have a unique minimum point and that the task can be completed at the target position. The modulation matrix ensures the stability of the system, so that the robot can accurately and stably converge to the minimum point, which is the task target point. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a schematic diagram of the overall implementation framework of the robot teaching obstacle avoidance method based on neural network and dynamic system model of the present invention;

[0058] Figure 2 A schematic diagram of a process for constructing a decision function for the present invention;

[0059] Figure 3 A schematic diagram of a process for constructing a modulation matrix according to the present invention;

[0060] Figure 4 A schematic diagram of a process for constructing an energy function for the present invention;

[0061] Figure 5 Schematic diagram comparing the obstacle avoidance performance of the present invention with other methods. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0063] Example 1

[0064] This embodiment provides a robot teaching obstacle avoidance system based on a neural network and a dynamic system model, including: a demonstration data acquisition module, a decision function construction module, a modulation matrix construction module and an energy function construction module, and a control strategy output module;

[0065] Among them, the demonstration data acquisition module collects human demonstration data and obstacle information to generate a demonstration data set;

[0066] The decision function building module is used to distinguish the characteristics of task trajectory tracking and obstacle avoidance trajectories in the human demonstration dataset, and dynamically adjust the system's behavior between trajectory tracking and obstacle avoidance tasks by combining the robot's current state and environmental information. The decision function consists of two parts: one is the robot's current state information, and the other is environmental information, including the relative position and speed of the robot and the obstacle, as well as the gradient of the energy function. By introducing a smooth switching mechanism, the decision function can dynamically adjust the weight according to the distance between the robot and the obstacle, ensuring that the system prioritizes obstacle avoidance when approaching an obstacle, and focuses on trajectory tracking tasks when far away from obstacles.

[0067] The role of the decision function is mainly reflected in two aspects: first, it enhances the obstacle avoidance capability of the system, enabling the robot to autonomously avoid obstacles in complex environments; second, it improves the generalization performance of the system, enabling the robot to adapt to different obstacle layouts and initial conditions, thereby completing tasks efficiently in a variety of scenarios.

[0068] The modulation matrix building module is used to construct a modulation matrix through a multi-task hybrid neural network. Its input combines the current state of the robot, environmental obstacle information, and the gradient of the energy function, thereby comprehensively perceiving the task requirements and environmental constraints. The modulation matrix dynamically adjusts the direction and amplitude of the robot's motion, enabling it to accurately learn and reproduce complex trajectory characteristics and obstacle avoidance behaviors. Specifically, the modulation matrix extracts trajectory shape features from demonstration data based on the nonlinear fitting ability of the neural network, and integrates obstacle distance information to generate an adaptive motion correction strategy. The advantage of the modulation matrix lies in its flexibility and robustness. On the one hand, it can automatically learn the local details of the trajectory through the training of the neural network, while capturing the global impact of obstacles on the trajectory; on the other hand, its symmetrical positive definite structure design ensures the stability of the dynamic system and avoids trajectory deviations caused by local disturbances.

[0069] The energy function building module provides a stable motion guidance mechanism for the system by constructing an energy function NEUM that is continuously differentiable, radially unbounded, and has a unique global minimum point. The energy function consists of two parts: one is a nonlinear feature term extracted based on a neural network, which is used to capture local details of complex trajectories; the other is a quadratic term, which is used to ensure the radial unboundedness of the energy function, thereby ensuring that the system can still converge when it is far away from the target. The advantage of the energy function lies in its flexibility and robustness. Through the nonlinear fitting ability of the neural network, the energy function can accurately model the shape characteristics of complex trajectories, and at the same time combine the quadratic term to ensure the global stability of the system. In addition, the unique minimum point design of the energy function avoids the problem of local minima, allowing the system to always converge to the target state.

[0070] The control strategy output module uses the output of the neural dynamic system model as the robot control strategy to complete the specified task.

[0071] Example 2

[0072] like Figure 1 As shown, this embodiment provides a robot teaching obstacle avoidance method based on a neural network and a dynamic system model, comprising the following steps:

[0073] Step 1: Collect human demonstration data and obstacle information. Human experts complete robot tasks through teaching and generate demonstration data sets where x t,n is the position state of the robot at time t in the nth demonstration, is the speed state, x obj Obstacle information in the environment;

[0074] Step 2: State encoding and feature expansion: encode the current robot state x into a high-dimensional feature vector c(x) to enhance nonlinear expression capabilities;

[0075] Step 3: Construct the energy function. First, extract the nonlinear feature components from the encoded feature c(x) through a neural network. where a k ,b k is the feature parameter that needs to be trained and learned, σ is the activation function, and then the features are weighted and summed and quadratic terms are added to construct the energy function V(x);

[0076] like Figure 4 As shown, we first receive a human-demonstrated trajectory dataset It contains the robot state x and its corresponding dynamic changes

[0077] The input state x passes through the state encoding module c(x) and performs nonlinear feature expansion:

[0078]

[0079] Among them, ∈>0 is a tuning parameter that enhances the nonlinear expression ability of the state.

[0080] Feature extraction and weighted combination: The encoded state c(x) is input into the feature extraction module to generate a set of nonlinear feature components g k (x):

[0081]

[0082] Among them, a k and b k is a learnable feature parameter, σ is an activation function (such as a hyperbolic tangent function), and all feature components g k (x) is weighted summed through the weight vector ω to form the neural network part P 1 (x):

[0083]

[0084] Energy function structure integration: The energy function V(x) is composed of the neural network part P that captures the nonlinear characteristics of complex trajectories. 1 (x) and the quadratic term P that ensures radial unboundedness of the energy function 2 (x) consists of two parts, P 2 (x) is defined as:

[0085] P2(x)=κxTx(κ>0)

[0086] κ is a small positive number, and the final energy function form is:

[0087] V(x)=P 1 (x)-P 1 (0)+P 2 (x)

[0088] Among them, P 1 (0) is used for normalization to ensure that the energy is minimum at the target point x = 0;

[0089] Energy function parameter constraints: In the process of learning and optimizing the energy function parameters, the parameters in the energy function must meet the following conditions to ensure that the system can converge to the target position:

[0090]

[0091] Parameter optimization: This embodiment transforms the parameter learning problem of the model into an optimization problem with additional constraints, where the optimal parameters of the model can be solved by optimization methods and nonlinear programming methods, where the objective function is set to:

[0092]

[0093] Step 4: Fusion state and environment information to build a decision function and calculate the distance between the robot and the obstacle:

[0094] d=∥xx obj ∥

[0095] The parameter α that controls task tracking is dynamically allocated through the Sigmoid function 1 and obstacle avoidance parameter α 2 The weight of

[0096] like Figure 2 As shown, the input of the decision function includes the current state of the robot x, the center position of the obstacle x obj , the speed of the robot And the gradient of the energy function These information together constitute the environmental information z, which is used in the subsequent decision-making process;

[0097] Distance calculation: Calculate the distance between the robot's current position and the center of the obstacle d = ∥ xx obj ∥, this distance is used to determine whether the robot is close to an obstacle, and thus decide whether it needs to switch to obstacle avoidance mode;

[0098] Weight calculation: Calculate two weight coefficients α through the Sigmoid function 1 and α 2Among them, α 1 represents the weight of task tracking, α 2 Represents the weight of obstacle avoidance. The calculation of these two weights is based on the distance d and a safety threshold β, ensuring that the obstacle avoidance weight gradually increases when approaching obstacles, while the task tracking weight dominates when moving away from obstacles;

[0099] Construction of decision function: multiply the robot's state x and environment information z by the corresponding weight α 1 and α 2 , and get the output of the decision function. Specifically, the output of the decision function is D(x,z)=[α 1 x,α 2 z], where

[0100] Step 5: Generate the modulation matrix. The decision function output D(x,z) is input into the multi-task hybrid neural network Net(·) to generate the parameter vector n. The lower triangular matrix is ​​constructed through n and the modulation matrix M is constructed through cholosky decomposition.

[0101] like Figure 3 As shown, the input of the modulation matrix includes the output of the decision function It contains the robot's current state x and environment information z;

[0102] The core of the modulation matrix is ​​a multi-task hybrid neural network (Multi-task Hybrid Neural Network), denoted as The task of the neural network is to learn the characteristics of the task trajectory and obstacle avoidance behavior from the input data. The design of the neural network selects a three-layer fully connected neural network, and the number of neurons in each hidden layer is 80. The calculation process of each layer of the network is as follows:

[0103] From the input layer to the hidden layer 1, the input data D(x,z) will be passed to the first layer of the network. For the calculation of the first layer, the network will perform the following operations:

[0104] h 1 =σ(W 1 D(x,z)+b 1 )

[0105] Among them, W 1 is the weight matrix of the first layer, the number of neurons is 80, b 1 is the bias vector of the first layer, σ(·) is the activation function, and the tanh function is used as the activation function;

[0106] Similarly, the calculation process from hidden layer 1 to hidden layer 2 is:

[0107] h2=σ(W2h1+b2)

[0108] Among them, W 2 is the second weight matrix, the number of neurons in this layer is also 80, b2 is the bias vector of the second layer, and h 2 is the output of the second layer and also the input of the output layer;

[0109] The calculation process from the second hidden layer to the output layer is as follows:

[0110] n = W3h2 + b3

[0111] Among them, W 3 is the weight matrix of the output layer, b 3 is the bias vector of the output layer, and n is the final output of the network;

[0112] Output of the neural network: The output of the neural network is a vector n, whose dimension is where n is the dimension of the state space. This vector contains the parameters required for the modulation matrix;

[0113] Construction of the modulation matrix: According to the output vector n of the neural network, a lower triangular matrix N is constructed. Specifically, the element N ij of the matrix N is generated in the following way:

[0114] If i > j or i = j, then N ij = nk, where k is the corresponding index in the vector n;

[0115] If i < j, then N ij = 0;

[0116] Positive definite transformation of the modulation matrix: To ensure that the modulation matrix M is positive definite (this is a key condition for system stability), the idea of Cholesky decomposition is adopted for design. Specifically, the modulation matrix M is generated in the following way:

[0117]

[0118] Among them, ∈ is a small positive number, and I is the identity matrix. This step ensures that the modulation matrix M is symmetric and positive definite, thus meeting the stability requirements of the system;

[0119] Step 6: Construct the loss function, transform the parameter learning problem into a constrained optimization problem, and obtain the optimal model parameters through the method of nonlinear programming;

[0120] Step 7: Use the output of the neural dynamic system model as the robot control strategy to complete the specified task.

[0121] As Figure 5 shown and combined with Table 1 shown below, the obstacle avoidance performance data comparison of the present invention with other methods is obtained;

[0122] Table 1. Data comparison of obstacle avoidance performance of the present invention and other methods

[0123]

[0124] Compared with the existing DMP-based obstacle avoidance algorithm, the method proposed in the present invention is more stable and has higher obstacle avoidance capability. By calculating the root mean square error of the copied trajectory of the two methods, the method proposed in the present invention better retains the characteristics of the original trajectory and achieves an error of 9.73cm, while the DMP-based method reduces the trajectory accuracy during the obstacle avoidance process, resulting in an error of 10.65cm. The present invention can avoid obstacles at multiple different obstacle positions, while the DMP-based method will produce collisions at some obstacles; on the other hand, the method proposed in the present invention can teach multiple trajectories and comprehensively learn the characteristics of multiple trajectories, while the DMP-based method can only learn a single trajectory and cannot comprehensively learn the characteristics of multiple expert demonstrations.

[0125] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A robot teaching obstacle avoidance system based on neural network and dynamic system model, characterized in that: include: Demonstrate data acquisition module, decision function building module, modulation matrix building module, energy function building module, and control strategy output module; The demonstration data acquisition module is used to collect human demonstration data and obstacle information to generate a demonstration data set; The decision function building module is used to fuse state and environmental information to build a decision function, dynamically allocate the weights of task tracking and obstacle avoidance behaviors using the decision function, and generate an adaptive control strategy based on the robot state and environmental information; The modulation matrix building module is used to build a modulation matrix through a multi-task hybrid neural network. The modulation matrix dynamically adjusts the direction and amplitude of the robot's movement and learns to reproduce complex trajectory characteristics and obstacle avoidance behaviors. The energy function construction module is used to extract nonlinear feature components from the features encoded by the robot state through a neural network, weight the features and add quadratic terms to construct an energy function, and make the trajectory converge to a unique target point based on the energy function; The control strategy output module is used to construct a loss function, obtain the optimal parameters of the neural dynamic system model through a nonlinear programming method, and use the output of the neural dynamic system model as a robot control strategy to complete a specified task.

2. A robot teaching obstacle avoidance method based on neural network and dynamic system model, characterized in that: A robot teaching obstacle avoidance system based on a neural network and a dynamic system model as claimed in claim 1 is provided, comprising the following steps: Collect human demonstration data and obstacle information to generate a demonstration data set; Encode the current robot state into a high-dimensional feature vector; Extract nonlinear feature components from the encoded features through a neural network, sum the features by weights and add quadratic terms to construct an energy function; The decision function is constructed by integrating the state and environment information. The robot's state and environment information are multiplied by the corresponding weights to obtain the output of the decision function. The output of the decision function is input into the multi-task hybrid neural network to generate a parameter vector, a lower triangular matrix is ​​constructed based on the parameter vector, and a modulation matrix is ​​constructed through cholosky decomposition; Construct a loss function, transform the parameter learning problem into a constrained optimization problem, and obtain the optimal model parameters through nonlinear programming methods; The output of the neural dynamic system model is used as the robot control strategy to complete the specified task.

3. The robot teaching obstacle avoidance method based on neural network and dynamic system model according to claim 2 is characterized in that: Collect human demonstration data and obstacle information to generate a demonstration data set, which is specifically expressed as: Among them, x t,n is the position state of the robot at time t in the nth demonstration, is the speed state, x obj Obstacle information in the environment.

4. The robot teaching obstacle avoidance method based on neural network and dynamic system model according to claim 2 is characterized in that: The nonlinear feature components are extracted from the encoded features through a neural network, the features are weighted and summed, and quadratic terms are added to construct an energy function, which includes: The nonlinear feature components are extracted from the encoded features and expressed as: Among them, a k and b k is a learnable feature parameter, σ is an activation function, c(x) represents the encoded feature, g k (x) represents the nonlinear characteristic component; All nonlinear characteristic components g k (x) is weighted and summed by the weight vector ω to obtain the neural network part P1(x): Construct the quadratic term P2(x): P2(x)=κxTx(κ>0) Among them, κ is a positive number; The energy function is expressed as: V(x)=P1(x)-P1(0)+P2(x).

5. The robot teaching obstacle avoidance method based on neural network and dynamic system model according to claim 4 is characterized in that: The encoded feature c(x) is specifically expressed as: Among them, ∈>0 is the adjustment parameter, and x represents the robot state.

6. The robot teaching obstacle avoidance method based on neural network and dynamic system model according to claim 4 is characterized in that: The parameter constraints of the energy function are: Among them, ω represents the weight vector.

7. The robot teaching obstacle avoidance method based on neural network and dynamic system model according to claim 2 is characterized in that: The decision function is constructed by integrating the state and environment information. The robot's state and environment information are multiplied by the corresponding weights to obtain the output of the decision function, which includes: Calculate the distance between the robot and the obstacle, and dynamically assign the weights of the control task tracking parameter α1 and the obstacle avoidance parameter α2 through the Sigmoid function; Multiply the robot's state x and environment information z by the corresponding weights α1 and α2 respectively to get the output of the decision function, which is specifically expressed as: D(x,z)=[α1x,α2z].

8. The robot teaching obstacle avoidance method based on neural network and dynamic system model according to claim 2 is characterized in that: The environmental information z is expressed as: Among them, x represents the current state of the robot, x obj represents the center position of the obstacle, represents the speed of the robot, Represents the gradient of the energy function.

9. The robot teaching obstacle avoidance method based on neural network and dynamic system model according to claim 2, characterized in that: The lower triangular matrix is ​​constructed based on the parameter vector, and the modulation matrix is ​​constructed by cholosky decomposition. The specific expression is: Where D(x,z) represents the output of the decision function, ∈>0 is the adjustment parameter, and I is the unit matrix.

Citation Information

Patent Citations

  • Robot simulation learning method based on dynamic system model

    CN109702744A

  • Robot autonomous learning and adaptive trajectory planning method oriented to man-machine cooperation

    CN119077728A

  • Method for realizing double-arm cooperative operation task based on teaching learning

    CN112207835A

  • Mechanical arm navigation obstacle avoidance method and system, computer equipment and storage medium

    CN114603564A

  • Robot path guiding method and device based on reinforcement learning optimization and medium

    CN115599104A