A Control Method for Robotic Arm Rehabilitation Training Based on Bayesian Optimization of Asymptotic Impedance Parameters

CN122559984APending Publication Date: 2026-08-14NINGBO INST OF MATERIALS TECH & ENG CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-21
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

(1)用户实时运动意图与最优阻抗参数之间的映射关系复杂、非线性、时变,难以用精确的数学模型描述,呈现黑盒特性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122559984A_ABST
    Figure CN122559984A_ABST
Patent Text Reader

Abstract

This invention relates to a robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters. Through a closed-loop architecture and synergistic effect of real-time motion intention – Bayesian optimization to find the optimal impedance parameter set – and real-time adjustment of damping and stiffness based on an intention-weighted asymptotic function, the robotic arm actively reduces damping and stiffness to reduce resistance when the user actively exerts force; when the user's intention weakens, the robotic arm appropriately increases damping to maintain stability, avoiding abrupt parameter changes and ensuring continuous variation of impedance parameters in an arbitrary graph. This achieves an adaptive balance between the robotic arm's compliance and response speed during rehabilitation training. Finally, the end-effector torque is calculated by combining gravity and friction compensation, resulting in smooth, non-jamming robotic arm movement and reducing the risk of secondary injury to the patient due to control abrupt changes. Simultaneously, the five-dimensional impedance parameter optimization is transformed into a black-box optimization problem, and the Bayesian optimization algorithm can approximate the globally optimal parameters with a limited number of interactive samples, avoiding the dependence of model predictive control or reinforcement learning on precise system models or massive trial-and-error data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotic arm control technology, and in particular to a robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters. Background Technology

[0002] As interactive robots become increasingly prevalent in human life, they are expected to perform tasks and achieve efficient collaboration in human-robot interaction scenarios, adjusting control based on human behavior. To this end, numerous control strategies have been developed. Collaborative robots need to possess safety, flexibility, and interactivity, and be able to dynamically adjust their control strategies according to the intentions and behaviors of human users. Impedance control is one of the mainstream methods for achieving robot compliance, which establishes the dynamic relationship between the robot's end effector force and positional deviation by simulating a mass-spring-damped system.

[0003] However, the parameters of traditional impedance controllers (especially damping and stiffness) are usually fixed, making it difficult to achieve an ideal balance between "response sensitivity" (ease of actuation) and "motion stability / trajectory tracking accuracy" (disturbance resistance). To address this, researchers have proposed adaptive variable impedance control methods, attempting to adjust parameters in real time based on the user's actual motion intentions. Existing techniques include model predictive control (MPC), fuzzy logic, and reinforcement learning. However, these methods generally suffer from the following problems: (1) The mapping relationship between the user's real-time motion intention and the optimal impedance parameters is complex, nonlinear, and time-varying, and is difficult to describe with a precise mathematical model, exhibiting black box characteristics; (2) MPC has stringent requirements for model accuracy; fuzzy rules rely on expert experience and have poor generalization; reinforcement learning requires a lot of interactive trial and error, and direct training on physical robots is costly, risky, and difficult to achieve personalized and rapid adaptation.

[0004] (3) Many methods only use trajectory tracking error as the evaluation index, and fail to fully consider the multi-dimensional collaborative optimization needs such as interaction force, motion speed, and task efficiency. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a robotic arm rehabilitation training control method that can efficiently, safely, and individually optimize impedance parameters.

[0006] This invention provides a robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters. The user interacts with a virtual interactive environment by manipulating a single-degree-of-freedom robotic arm, and the method includes the following steps: S1, real-time acquisition of the angle, speed, acceleration and interaction force of the rotary joint at the end of the robotic arm; S2, based on the coupling of speed and interactive force, obtains the user's real-time motion intention; S3, Construct a variable impedance controller based on real-time motion intent. The variable impedance controller includes a five-dimensional impedance parameter set, which includes minimum damping, transition damping, maximum damping, minimum stiffness, and maximum stiffness. S4, construct a target optimization function to characterize the user's exercise efficiency and rehabilitation training task completion time; S5 generates a prior dataset through Latin hypercube sampling; The Bayesian optimization algorithm is adopted, with the five-dimensional impedance parameter set as the variable to be optimized and the negative value of the objective optimization function as the performance function. A surrogate model is constructed using a Gaussian process, and the parameter search is guided by the prior dataset. The surrogate model is updated in a finite number of iterations, and the optimal impedance parameter set that minimizes the objective optimization function is output. S6, based on real-time motion intent, the Gaussian function is used to solve the damping function weight coefficients; S7, construct a progressive damping adjustment function and stiffness adjustment function based on motion intention weighting, and dynamically adjust the damping coefficient and stiffness coefficient by combining the damping function weighting coefficient and the optimal impedance parameter set; S8, based on the adjusted damping coefficient and stiffness coefficient, calculates the torque mapped to the end joint of the robotic arm by combining gravity and friction compensation.

[0007] Compared with existing technologies, this application has the following advantages: Through a closed-loop architecture and synergistic effect of real-time motion intention-Bayesian optimization to find the optimal impedance parameter set-based adjustment of damping and stiffness in real time based on an intention-weighted asymptotic function, the robotic arm actively reduces the damping and stiffness coefficients to reduce resistance when the user actively exerts force; when the user's intention weakens, the robotic arm appropriately increases damping to maintain stability, avoiding abrupt parameter changes and ensuring continuous variation of impedance parameters in an arbitrary graph. This achieves an adaptive balance between the compliance and response speed of the robotic arm in rehabilitation training. Finally, the combined calculation of end torque using gravity and friction compensation ensures smooth and seamless robotic arm movement, reducing the risk of secondary injury to the patient due to control abrupt changes. Simultaneously, the optimization of five-dimensional impedance parameters is transformed into a black-box optimization problem, and the Bayesian optimization algorithm can approximate the globally optimal parameters with a limited number of interactive samples, avoiding the dependence of model predictive control or reinforcement learning on precise system models or massive trial-and-error data.

[0008] In one possible implementation, the user's real-time motion intent obtained in step S2 based on the coupling of velocity and interaction force is as follows: ; In the formula, Indicates speed, It represents interactive force.

[0009] Compared with existing technologies, in actual motion, the magnitude of real-time motion intention is closely related to speed. This application uses human power to represent real-time motion intention, which makes it easier to understand the inflow and outflow of energy. When the real-time motion intention is to accelerate in a certain direction, the value of the real-time motion intention is larger; when the real-time motion is to decelerate in a certain direction, the value of the real-time motion intention is smaller or negative. At this time, the variable impedance controller increases the damping according to the change of the user's real-time motion intention, and conversely, decreases the damping to suppress motion.

[0010] In one possible implementation, the objective optimization function constructed in step S4 is: ; In the formula, This represents the average speed at which the robotic arm completes rehabilitation training tasks in a virtual interactive environment; This indicates the time it takes for the robotic arm to complete rehabilitation training tasks in a virtual interactive environment; This represents the weighting of interactivity and performance on rehabilitation training tasks in a virtual interactive environment. , Indicates the index score of a task in a virtual interactive environment; Indicates the maximum speed. Indicates the average speed deviation. Indicates time deviation.

[0011] Compared with existing technologies, this method uses the user's task performance and interaction performance in the virtual interactive environment to jointly drive optimization, with the index score as the weighting adjustment factor. Thus, when the training task score is high, the weight of the performance index increases accordingly, and the optimization focuses more on improving the interaction performance; when the training task score is low, the optimization is biased towards optimizing the task completion time.

[0012] In one possible implementation, step S5 specifically includes: S51, a preset range of values ​​for minimum damping, transition damping, maximum damping, minimum stiffness, and maximum stiffness is defined in the five-dimensional impedance parameter set. The five-dimensional space spanned by these ranges is defined as the parameter search space. Within this search space, N initial parameter samples are generated using the Latin hypercube sampling method. These N initial parameter samples are sequentially deployed to the variable impedance controller. The target optimization function value for one round of rehabilitation training is collected for each initial parameter sample. Each initial parameter sample and its corresponding target optimization function value constitute a data point, thus constructing a priori dataset. ,in, This represents the five-dimensional impedance parameter set of the variable impedance controller. This represents the negative value of the corresponding objective optimization function; the prior dataset is used as the training dataset. S52, based on the current training dataset, takes the five-dimensional impedance parameter set as the variable to be optimized and the negative value of the objective optimization function as the performance function. Gaussian process is used to construct a surrogate model between the five-dimensional impedance parameter set and the performance function. The Gaussian process uses a kernel function to measure the similarity between different data points in the parameter search space and estimates the hyperparameters of the kernel function by maximizing the log marginal likelihood. S53, using the expected improvement function as the acquisition function, based on the current agent model, calculates the expected improvement value for any data point in the parameter search space, and selects the data point with the largest expected improvement value as the target data point to be tested in the next round; S54, deploy the five-dimensional impedance parameter set in the target data point to the variable impedance controller, collect the actual target optimization function value of the user under the target data point, and add the target parameter combination and the actual target optimization function value as new data points to the training dataset, return to S52, until the preset iteration round is reached or the convergence condition is met; Select the set of five-dimensional impedance parameters that minimizes the value of the objective optimization function from the current training dataset as the optimal impedance parameter set.

[0013] Compared to existing methods, this approach treats the complex mapping relationship between five-dimensional impedance parameters and rehabilitation training effects as a black box, eliminating the need to establish an analytical mathematical model between the patient's real-time movement intention and the optimal parameters. It employs Latin hypercube sampling to generate a priori dataset within the boundaries of the five-dimensional parameters, ensuring that each parameter dimension is uniformly covered within its domain and avoiding initial sampling concentrated in local areas. This achieves efficient, stable, and personalized optimization of five-dimensional impedance parameters, significantly reducing the cost of parameter optimization in rehabilitation training and providing a reliable foundation of optimal parameters for subsequent real-time intention response control.

[0014] In one possible implementation, the kernel function in step S52 employs an anisotropic Matern kernel: ; In the formula, Represents any two sets of five-dimensional impedance parameters. Represents the signal variance. This represents the Euclidean distance between any two sets of five-dimensional impedance parameters. , This represents the length scale parameter. Represents the target data point. This represents a data point in the current training dataset.

[0015] Compared with existing technologies, Gaussian processes automatically learn the importance of each parameter dimension (length scale) through anisotropic Matern kernel functions, enabling them to mine parameter-performance relationships from data and avoiding the requirement of precise modeling of system dynamics in traditional methods.

[0016] In one possible implementation, the desired improvement function in step S53 is: ; In the formula, Represents the Gaussian process in a five-dimensional impedance parameter set The predicted mean at the location; Represents the Gaussian process in a five-dimensional impedance parameter set The standard deviation of the forecast at that location; This represents the optimal objective function value in the current training dataset. The preset exploration coefficient, The cumulative distribution function represents the standard normal distribution. The probability density function representing the standard normal distribution; Indicates the amount of standardization improvement. ,in, , This represents the performance function value of each data point in the current training dataset. Represents the covariance matrix between data points. Indicates the noise variance. This represents the transpose of the covariance vector between the target data point and the data points in the current training dataset.

[0017] Compared with existing technologies, this method adopts the expected improvement function and uses the predicted mean (utilizing known optimal regions) and predicted standard deviation (exploring uncertain regions) of the Gaussian process output to set the exploration coefficient, effectively balancing local fine-grained search and global exploration within a finite number of iterations.

[0018] In one possible implementation, the damping function weighting coefficient in step S6 is: ; In the formula, Represents an exponential function. This represents a constant that affects the rise and fall of the weighting coefficients of the damping function. Range parameter representing the real-time motion intent.

[0019] In one possible implementation, the asymptotic damping adjustment function constructed in S7 is: ; ; ; In the formula, These represent the damping functions that affect the speed decrease. and the damping function as velocity increases The constant of the growth rate The minimum parameter representing the speed of the rotary joint at the end of the robotic arm. This represents a parameter within the speed range of the rotary joint at the end of the robotic arm. ; Indicates minimum damping. Indicates transition damping, Indicates maximum damping; The stiffness adjustment function is: ; In the formula, This represents a constant value that affects the rate of change of the stiffness adjustment function; This represents the real-time motion intention value in a static state. Indicates minimum stiffness. This indicates the maximum stiffness.

[0020] In one possible implementation, the torque mapped to the end-effector rotary joint in step S8 is calculated as follows: ; In the formula These represent acceleration error, velocity error, and angular error, respectively. These are the inertia matrix, damping matrix, and stiffness matrix, respectively. This represents the gravity compensation feedforward term. This represents the feedforward term for friction compensation. Attached Figure Description

[0021] Figure 1 This is a system block diagram of the interaction between the robotic arm and the variable impedance controller in this application; Figure 2 This is a system flowchart for this application; Figure 3 This is a schematic diagram of the virtual interactive environment interface of this application; Figure 4 This is a user operation scenario diagram used in the experimental verification of this application; Figure 5(a) shows the trend of real-time motion intention as a function of force and velocity in the experimental verification of this application; Figure 5(b) shows the trend of the damping function weight coefficient changing with the real-time motion intention in the experimental verification of this application; Figure 5(c) shows the trend of the damping coefficient changing with the real-time motion intention in the experimental verification of this application; Figure 5(d) shows the trend of stiffness coefficient changing with real-time motion intention in the experimental verification of this application; Figure 6 This is a graph showing the trend of average speed as a function of the number of iterations in the experimental verification of this application. Figure 7(a) is a comparison of the average speed of the subjects during the trial of the two algorithms in the comparative experiment of this application; Figure 7(b) is a comparison of the interaction forces of the subjects after the two algorithm experiments in the comparative experiment of this application. Detailed Implementation

[0022] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of this application and are not intended to limit the scope of protection of the embodiments of this application. Those skilled in the art can make adjustments as needed to adapt to specific application scenarios.

[0023] In the description of the embodiments of this application, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in the embodiments of this application based on the specific circumstances.

[0024] In the embodiments of this application, unless otherwise expressly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "on top of," and "over" the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0025] The present application will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0026] See Figures 1-3 As shown, this application discloses a robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters. Users interact with a virtual interactive environment by manipulating a single-degree-of-freedom robotic arm. The robotic arm in this application is a seven-degree-of-freedom collaborative robotic arm, with only the end-effector rotational joint (single degree of freedom) activated, and the other six joints locked in the initial posture shown in Table 1. A six-dimensional force sensor is installed at the end of the robotic arm to collect the user's interaction torque. The control cycle is set to 1ms, and high-speed communication is achieved using the EtherCAT protocol. The system is based on the Speedgoat real-time simulation system, equipped with a real-time operating system, and communicates with the robotic arm actuator and force sensor via EtherCAT to achieve high-frequency torque control and status feedback. A control algorithm is developed based on Simulink, and a virtual rehabilitation training environment (parkour game) and human-computer interaction interface are developed using Python. The development computer exchanges data with the real-time control platform via the UDP / IP protocol. The virtual interactive environment in this embodiment is a horizontally scrolling parkour game. The user rotates the end joints of a robotic arm to drive a virtual rocket on the screen to move horizontally and collect coins. The joint angle is linearly mapped to the horizontal position of the rocket, and the joint angular velocity is linearly mapped to the rocket's forward speed (which can be increased to a maximum of 56.7% of the base speed). The interface displays the joint angle, rotation speed, forward speed, and current impedance parameters in real time.

[0027] The method described in this application includes the following steps: S1, real-time acquisition of the angle, speed, acceleration and interaction force of the rotary joint at the end of the robotic arm; In this embodiment, the angle, velocity, and acceleration of the rotating joint are collected in real time by a joint encoder (acceleration can be obtained through differentiation, but in this embodiment, acceleration is not used as the real-time motion intention to avoid noise); the interactive force applied by the user is collected in real time by a six-dimensional force sensor at the end; all data are sampled at a frequency of 1kHz and preprocessed by a low-pass filter with a cutoff frequency of 20Hz.

[0028] S2, based on the coupling of speed and interactive force, obtains the user's real-time motion intention; To avoid acceleration noise, the product of velocity and interaction force is used as a quantitative representation of the user's real-time motion intent: ; In the formula, Indicates speed, Indicates interactive force; Real-time motion intent This represents the instantaneous mechanical power input by the user, which is used when the user intends to accelerate during real-time motion. Larger positive values; when the real-time motion intends to decelerate or resist, The value is small or negative; It is sensitive to changes in intent and has a clear physical meaning.

[0029] S3, Construct a variable impedance controller based on real-time motion intent, the variable impedance controller including a five-dimensional impedance parameter set; the five-dimensional impedance parameter set includes minimum damping. Transition Damping Maximum damping Minimum stiffness and maximum stiffness In this embodiment, the minimum damping represents a strong user's intention to move, with a value range of [0.001, 0.3], the transition damping ranges from [0.3, 0.5], and the maximum damping represents a weak user's intention to move, with a value range of [0.5, 0.8]. The minimum stiffness ranges from [0.03, 1], and the maximum stiffness ranges from [0.5, 0.8].

[0030] S4, construct a target optimization function to characterize the user's exercise efficiency and rehabilitation training task completion time; The objective optimization function in this embodiment is used to evaluate the performance of a single round of rehabilitation training task (collecting coins in a game), and is intended to be minimized: ; In the formula, This represents the average speed at which the robotic arm completes the current round of tasks in the virtual interactive environment; This indicates the time it takes for the robotic arm to complete the current round of tasks in the virtual interactive environment; This indicates the weight given to balancing interactivity and task performance in a virtual interactive environment. , This represents the index score of a task in a virtual interactive environment. Indicates the maximum speed. Indicates the average speed deviation. Indicates time deviation.

[0031] S5 generates a prior dataset through Latin hypercube sampling; A Bayesian optimization algorithm is employed, using a five-dimensional impedance parameter set as the variable to be optimized and the negative value of the objective function as the performance function. A Gaussian process is used to construct a surrogate model, and the parameter search is guided by a prior dataset. The surrogate model is updated in a finite number of iterations, outputting the optimal impedance parameter set that minimizes the objective function. Specifically, this includes: S51, a preset range of values ​​for minimum damping, transition damping, maximum damping, minimum stiffness, and maximum stiffness is defined in the five-dimensional impedance parameter set. The five-dimensional space spanned by these ranges is defined as the parameter search space. Within this search space, N initial parameter samples (N=16) are generated using the Latin hypercube sampling method. These N initial parameter samples are then sequentially deployed to the variable impedance controller. The target optimization function value for one round of rehabilitation training is collected for each initial parameter sample. Each initial parameter sample and its corresponding target optimization function value constitute a data point, thus constructing a priori dataset. , This represents the number of parameter samples, where, This represents the five-dimensional impedance parameter set of the variable impedance controller. This represents the negative value of the corresponding objective optimization function; the prior dataset is used as the training dataset.

[0032] S52, based on the current training dataset, takes the five-dimensional impedance parameter set as the variable to be optimized and the negative value of the objective optimization function as the performance function. Gaussian process is used to construct a surrogate model between the five-dimensional impedance parameter set and the performance function. The Gaussian process uses a kernel function to measure the similarity between different data points in the parameter search space and estimates the hyperparameters of the kernel function by maximizing the log marginal likelihood. When training the surrogate model with the current training dataset, the surrogate model outputs the predicted mean and predicted standard deviation of any five-dimensional impedance parameter set; The kernel function used in this embodiment is an anisotropic Matern kernel: ; In the formula, Represents any two sets of five-dimensional impedance parameters. Represents the signal variance. This represents the Euclidean distance between any two sets of five-dimensional impedance parameters. , This represents the length scale parameter. Represents the target data point. This represents a data point in the current training dataset.

[0033] S53, using the expected improvement function as the acquisition function, based on the current agent model, calculates the expected improvement value for any data point in the parameter search space, and selects the data point with the largest expected improvement value as the target data point to be tested in the next round; The desired improved function in this application embodiment is: ; In the formula, Represents the Gaussian process in a five-dimensional impedance parameter set The predicted mean at the location; Represents the Gaussian process in a five-dimensional impedance parameter set The standard deviation of the forecast at that location; This represents the optimal objective function value in the current training dataset. The preset exploration coefficient, The cumulative distribution function represents the standard normal distribution. The probability density function representing the standard normal distribution; Indicates the amount of standardization improvement. ,in, , This represents the performance function value of each data point in the current training dataset. Represents the covariance matrix between data points. Indicates the noise variance. This represents the transpose of the covariance vector between the target data point and the data points in the current training dataset.

[0034] S54, deploy the five-dimensional impedance parameter set in the target data point to the variable impedance controller, the user performs a round of rehabilitation training under the target data point, and collect the actual target optimization function value of the user under the target data point, and add the target parameter combination and the actual target optimization function value as new data points to the training dataset, return to S52, until the preset number of iterations is reached or the convergence condition is met; Select the set of five-dimensional impedance parameters that minimizes the value of the objective optimization function from the current training dataset as the optimal impedance parameter set.

[0035] S6, based on real-time motion intent, uses a Gaussian function to solve for the damping function weight coefficients: ; In the formula, Represents an exponential function. This represents a constant that affects the rise and fall of the weighting coefficients of the damping function. It approaches 0 when real-time motion intent is low and approaches 1 when real-time motion intent is high. Range parameter representing the real-time motion intent.

[0036] S7, construct an asymptotic damping adjustment function and stiffness adjustment function based on motion intention weighting, and dynamically adjust the damping coefficient and stiffness coefficient by combining the damping function weighting coefficient and the optimal impedance parameter set; where: The asymptotic damping adjustment function is: ; ; ; In the formula, These represent the damping functions that affect the speed decrease. and the damping function as velocity increases The constant of the growth rate The minimum parameter representing the speed of the rotary joint at the end of the robotic arm. This represents a parameter within the speed range of the rotary joint at the end of the robotic arm. ; The stiffness adjustment function is: ; In the formula, This represents a constant value that affects the rate of change of the stiffness adjustment function; This represents the real-time motion intention value in a static state, used to eliminate the influence of static conditions.

[0037] S8, based on the adjusted damping and stiffness coefficients, calculates the torque mapped to the rotary joint at the end of the robotic arm using combined gravity and friction compensation: ; In the formula These represent acceleration error, velocity error, and angle error, respectively, calculated from the expected trajectory and the measured angle, angular velocity, and angular acceleration. These are the inertia matrix, damping matrix, and stiffness matrix, respectively. This represents the gravity compensation feedforward term. This represents the feedforward term for friction compensation; Torque is sent to the robotic arm to achieve flexible interactive control of the robotic arm.

[0038] Experimental verification and results To verify the performance of the proposed Bayesian optimized asymptotic impedance model in a virtual training environment, a multi-person comparative experiment was designed. The experiment was conducted on the same hardware platform, using single-axis rotation of the robotic arm's end effector to correspond to the user's motion in the virtual environment. The robotic arm's seven joints operated in torque mode. Figure 3 The control method is input, and the other six joints are in position mode and locked. The joint angles of each joint in the initial posture of the robotic arm are shown in Table 1.

[0039] Table 1: Key Angles for End-of-Arm Rotation Eight participants were recruited for the experiment, including four men and four women, aged 20 to 27 years, all with no history of sports injury. Each participant completed multiple rounds of training tasks in a virtual rehabilitation environment before the formal experiment, with a familiarization training time of no less than 30 minutes, and was required to successfully complete the game tasks three times consecutively. To eliminate the influence of fatigue on the experimental results, there was a minimum interval of 3 hours between each participant's training and the formal data collection. The user operation scenario is as follows: Figure 4 As shown, the subject receives visual feedback through a monitor and operates the human-computer interaction actuator at the end of the robotic arm to perform movements.

[0040] The virtual interactive environment displays the current joint angle, rotational speed, forward speed, and current impedance parameters in real time. To ensure experimental fairness, all analysis data comes from 16 successfully completed formal experiments; data from repeated practice sessions before the formal experiments were used as a priori dataset.

[0041] Before the formal Bayesian optimization iterative experiment, prior data were stratified and sampled using the Latin hypercube method within the critical range of the optimization parameters. Table 2 shows the stability boundary values ​​of the optimization parameters, which were designed based on previous research on the impedance characteristics of the robotic arm in wrist-joint human-machine interaction. To ensure that the minimum damping and stiffness coefficients are non-negative to avoid system instability, the feasible range of each parameter was uniformly divided into 16 equally spaced intervals. By generating random dimensional permutations, 16 sets of parameter samples were constructed. Each set of samples satisfies the strict condition that each interval has exactly one sample point in each of the five dimensions, thus ensuring that the complete range of each parameter within its domain is uniformly covered.

[0042] Table 2: Boundary values ​​of the five-dimensional impedance parameter set The resulting set of 16 prior parameters comprehensively covered the critical ranges of each parameter, providing sufficient and balanced initial training data for the subsequent surrogate model based on Gaussian processes. The parameter dataset and the corresponding objective function estimates were calculated and used to form the training dataset for optimization. The number of iterations in the optimization process corresponded to the number of rounds in the training task; each experiment underwent 16 iterations, generating 16 sets of optimized parameters. To further verify the stability and advancement of this algorithm, each subject was required to repeat the game three times, and the mean and standard deviation of the relevant data were calculated and compared. The performance of the Bayesian impedance parameter optimization method based on real-time human motion intention and the impedance parameter method based on reinforcement learning were compared under the same conditions. The reinforcement method uses a Gaussian process to learn the robot system dynamics, and uses a neural network to represent and adjust the impedance adaptation strategy. The impedance parameter under this method... , Both methods use the same value function. To reduce subject fatigue, a rest period of at least two minutes is set between each set of experiments.

[0043] To quantify the performance improvement of the controller, several objective evaluation metrics were defined: task completion time (time from the start of a single round to collecting the fifth coin); average speed of movement in a single round, representing motion efficiency; and the cumulative L2 norm of the interaction force over the task completion time, reflecting the user's operational effort. All data were collected in real-time by the robot system. Motion speed was calculated using angle differentiation, and interaction force was collected by a six-axis force sensor at the end of the robotic arm. The data sampling frequency was 1 kHz, and preprocessing was performed using a low-pass filter with a cutoff frequency of 20 Hz to eliminate noise. Paired-samples t-tests were used to compare the two algorithms in the comparative experiment, and effect sizes were calculated to assess the significance of each metric. Performance improvement was expressed as a percentage to measure the relative improvement efficiency of the proposed algorithm.

[0044] The output of the gradually varying impedance parameter module in the adaptive impedance controller adjusts according to the change in the intention function value. The output of the damping function weights and impedance parameters is controlled using real-time motion intention, as shown in Figure 5(a). The numerical changes of velocity and interaction force follow the same trend. When the values ​​of velocity and interaction force are large... The value is also relatively large, and the speed response lags behind the interaction force. Therefore, acceleration is not used as a variable in the intention function because its real-time noise is too large and the signal lag is too late. The interaction force can better reflect human intention; as shown in Figure 5(b), the weight of the damping function Will follow The value changes when When the value approaches 0, it indicates a low real-time motion intention; at this point, the damping function weight coefficient... The value is less than 0.5, meaning the damping rising function has a significant impact; when the intended value increases, It will rise accordingly, and at 1.4×10 4 Up to 2.0×10 4 Between the time points The value has a significant peak, at which point Approaching 1, the damping decrease function plays a dominant role; as shown in Figures 5(c) and 5(d), the changes in damping and stiffness parameters will vary with... The value increases as it decreases, and with The value increases and decreases. The real-time performance of the above parameters meets the expected variable impedance effect. By reducing motion damping and increasing response sensitivity, the system assists the user's movement and improves system compliance.

[0045] In the multi-person trial, data collection began after each subject had practiced repeatedly before the trial to ensure they had a certain level of proficiency with the robotic arm and the developed parkour game. For example... Figure 6 As shown, the average speed of the robotic arm changes in each round when eight subjects complete the task three times in a row. As the number of iterations increases, the angular velocity of the driven joint rotation shows an upward trend.

[0046] Since the rotational angular velocity is directly related to the speed of the car in the virtual interactive environment, the increase in angular velocity will simultaneously accelerate the task progress, thereby increasing the task difficulty. This makes the changes in the user's task completion status between different rounds more significant, further highlighting the impact of the parameter optimization of the variable impedance controller on the interactive performance. The user's speed gradually increases, which helps to stimulate the user's enthusiasm for participating in the game.

[0047] Comparative experiment The method proposed in this application (Method 1) was compared with the method for optimizing impedance parameters using reinforcement learning (Method 2). Table 3 lists the average task completion time per round and the sequence number of the last coin picked up by the user under both algorithms.

[0048] Table 3: Multi-person training performance Experimental results show that the average task completion time and average pickup point for all users in Method 1 are lower than those in Method 2. The average time to complete one round of coin pickup is shortened by 0.68 seconds, a relative improvement of 8%. Meanwhile, the average final pickup point for users under Bayesian optimization control is 7.90, and that of the reinforcement learning algorithm is 8.27, representing an average of about 0.37 virtual coins earlier than the reinforcement learning control, indicating higher task efficiency and allowing users to complete the current task in an earlier round. These results demonstrate that the proposed control strategy can effectively improve human-machine collaboration efficiency and user motion response sensitivity. The experiment was conducted after the subjects were familiar with the operating scenario and had sufficient energy; therefore, the performance difference can be attributed to the optimization effect of the control strategy itself, rather than to increased cognitive ability. Figure 7 shows the average speed during the experiment for eight users and the cumulative L2 norm of the interaction force after completing the experiment.

[0049] As shown in Figure 7, the average speed generated by the proposed method (Method 1) is higher than that of the traditional algorithm (Method 2), while the average interaction force is lower than that of Method 2. As shown in Figure 7(b), the average interaction force of Method 2 is 3.63 N, while that of Method 1 is 3.04 N, a decrease of 0.59 N. The average force after Bayesian optimization is 16.25% lower than that of the reinforcement learning method. In Figure 7(a), the average speed of the eight subjects under Method 2 is 49.36 ° / s, while that under Method 1 is 61.2 ° / s, an increase of 24%. The interaction force of all subjects under Method 1 is lower than that under Method 2, while the speed is higher than that under Method 2, indicating that the optimization strategy has consistent effectiveness. Paired-samples t-test and ANOVA showed that, assuming normal distribution for speed and interaction force, the mean speed variance for Method 1 was 18.519 ° / s, and for Method 2 it was 20.712 ° / s. The mean variances of interaction force were 0.518 N and 0.557 N, respectively. Method 1 demonstrated stronger stability than Method 2 over 48 rounds. Furthermore, as shown in Figure 7, five out of the eight participants exhibited statistically significant differences in both speed and interaction force (p<0.05), validating the advancement of the intention-driven Bayesian optimization asymptotic impedance control method. After the experiment, all eight participants completed the NASA-TLX psychological assessment scale, with results showing that Method 1 had a lower overall workload score than Method 2. These test results indicate that under this configuration model, users can achieve ideal response speeds with less effort, stimulating their enthusiasm for game training, while the reduced interaction force prevents fatigue during prolonged exercise.

[0050] The algorithm proposed in this application is based on a variable impedance controller design. To avoid parameter jumps and system instability, piecewise functions and negative forms are avoided in the controller. The parameter design in human-computer interaction relies on human subjective perception, and the dimensions of the impedance parameters designed in advance may not be the most suitable for the current dynamic system. That is, it is difficult to find the optimal values ​​for all parameters in the variable impedance module. The focus is biased towards the specific stiffness and damping coefficient values ​​and the performance in the specific virtual rehabilitation environment. Using Gaussian Bayes optimization to handle the relationship between the two is a better solution. Although the interaction force is not included in the optimization function, the experimental results show the coupling relationship between the impedance parameters and the interaction force, further proving the robustness of the system.

[0051] It should be noted that although the experimental verification was conducted on a single-joint platform, the proposed algorithm framework has good scalability. From the perspective of control theory, the asymptotic impedance controller of this application is essentially based on impedance modeling in joint space, and its mathematical form can be naturally extended to the Cartesian space of the robotic arm's end effector, with the impedance parameter matrix extended as follows: The inertia, damping, and stiffness matrices correspond to the dynamic characteristics of the end effector in Cartesian space. At this point, the user's real-time motion intent is expanded to a six-dimensional vector; the optimization parameter set will also expand from a five-dimensional scalar to high-dimensional matrix parameters. However, the Bayesian optimization framework itself has good scalability to parameter dimensions, and the Gaussian process surrogate model can automatically learn the importance of different dimensions through anisotropic kernel functions. It is hoped that improving the acquisition function can still effectively guide sequential sampling in high-dimensional space.

[0052] Furthermore, although the algorithm improves the compliance and task efficiency of human-computer interaction through intent-driven asymptotic impedance adjustment and Bayesian optimization, its real-time performance may be limited by the computational complexity of Bayesian optimization in a high-dimensional parameter space when facing more complex multi-degree-of-freedom interaction tasks. Simultaneously, the superposition and coupling of sensor noise in high-dimensional space during multi-degree-of-freedom interaction may exacerbate the cumulative error in intent recognition and increase the ambiguity of intent judgment, thus posing a challenge to the accuracy of impedance parameter adjustment. Future work needs to verify the algorithm's scalability on a full-scale platform and combine multimodal perception fusion and transfer learning strategies to further enhance the model's adaptability in diverse interaction scenarios.

[0053] With the continuous development of robotics technology, the application scenarios of interactive collaborative robotic arms are becoming more diversified. Gaming scenarios are being applied in industries such as medical rehabilitation, skills training, sports, and entertainment. These applications often require robotic arms to meet safety requirements while possessing adaptive interaction algorithms to improve compliance. Therefore, the demand for smooth and stable variable impedance control methods is increasingly urgent. The controller needs to calculate personalized adaptation parameters in real time, requiring a high control frequency to maintain stability and reduce response time. The interaction algorithm needs to accurately identify user intentions, appropriately increase assistance and resistance, and integrate with the scenario to improve task efficiency and stimulate user participation. Therefore, this application constructs a robotic arm and variable impedance controller, using a force sensor and real-time velocity coupling method to extract human intentions. An intention-assigned impedance parameter function is designed, and the relationship between the optimized object, the robotic arm's motion performance, and the virtual environment task performance is simulated using a Gaussian process to predict the function parameters with the highest benefit. Experimental results show that the variable impedance controller developed in this paper based on Bayesian optimization can generate real-time, continuous, and stable impedance parameters according to intentions, effectively promoting user task completion, increasing average movement speed, and reducing human-machine physical interaction forces.

Claims

1. A robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters, wherein the user interacts with a virtual interactive environment by manipulating a single-degree-of-freedom robotic arm, characterized in that... Includes the following steps: S1, real-time acquisition of the angle, speed, acceleration and interaction force of the rotary joint at the end of the robotic arm; S2, based on the coupling of speed and interactive force, obtains the user's real-time motion intention; S3, Construct a variable impedance controller based on real-time motion intent. The variable impedance controller includes a five-dimensional impedance parameter set, which includes minimum damping, transition damping, maximum damping, minimum stiffness, and maximum stiffness. S4, construct a target optimization function to characterize the user's exercise efficiency and rehabilitation training task completion time; S5 generates a prior dataset through Latin hypercube sampling; The Bayesian optimization algorithm is adopted, with the five-dimensional impedance parameter set as the variable to be optimized and the negative value of the objective optimization function as the performance function. A surrogate model is constructed using a Gaussian process, and the parameter search is guided by the prior dataset. The surrogate model is updated in a finite number of iterations, and the optimal impedance parameter set that minimizes the objective optimization function is output. S6, based on real-time motion intent, the weight coefficients of the damping function are solved using a Gaussian function; S7, construct a progressive damping adjustment function and stiffness adjustment function based on motion intention weighting, and dynamically adjust the damping coefficient and stiffness coefficient by combining the damping function weighting coefficient and the optimal impedance parameter set; S8, based on the adjusted damping coefficient and stiffness coefficient, calculates the torque mapped to the end joint of the robotic arm by combining angle, velocity, acceleration, gravity and friction compensation.

2. The robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters according to claim 1, characterized in that, In step S2, the user's real-time motion intent is obtained based on the coupling of velocity and interaction force: ; In the formula, Indicates speed, It represents interactive force.

3. The robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters according to claim 1, characterized in that, The objective function constructed in step S4 is: ; In the formula, This represents the average speed at which the robotic arm completes rehabilitation training tasks in a virtual interactive environment; This indicates the time it takes for the robotic arm to complete rehabilitation training tasks in a virtual interactive environment; This indicates the weight given to balancing interactivity and task performance in a virtual interactive environment. , Indicates the index score of a task in a virtual interactive environment; Indicates the maximum speed. Indicates the average speed deviation. Indicates time deviation.

4. The robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters according to claim 1, characterized in that, The S5 steps specifically include: S51, a preset range of values ​​for minimum damping, transition damping, maximum damping, minimum stiffness, and maximum stiffness is defined in the five-dimensional impedance parameter set. The five-dimensional space spanned by these ranges is defined as the parameter search space. Within this search space, N initial parameter samples are generated using the Latin hypercube sampling method. These N initial parameter samples are sequentially deployed to the variable impedance controller. The target optimization function value for one round of rehabilitation training is collected for each initial parameter sample. Each initial parameter sample and its corresponding target optimization function value constitute a data point, thus constructing a priori dataset. ,in, This represents the five-dimensional impedance parameter set of the variable impedance controller. This represents the negative value of the corresponding objective optimization function; the prior dataset is used as the training dataset. S52, based on the current training dataset, takes the five-dimensional impedance parameter set as the variable to be optimized and the negative value of the objective optimization function as the performance function. Gaussian process is used to construct a surrogate model between the five-dimensional impedance parameter set and the performance function. The Gaussian process uses a kernel function to measure the similarity between different data points in the parameter search space and estimates the hyperparameters of the kernel function by maximizing the log marginal likelihood. S53, using the expected improvement function as the acquisition function, based on the current agent model, calculates the expected improvement value for any data point in the parameter search space, and selects the data point with the largest expected improvement value as the target data point to be tested in the next round; S54, deploy the five-dimensional impedance parameter set in the target data point to the variable impedance controller, collect the actual target optimization function value of the user under the target data point, and add the target parameter combination and the actual target optimization function value as new data points to the training dataset, return to S52, until the preset iteration round is reached or the convergence condition is met; Select the set of five-dimensional impedance parameters that minimizes the value of the objective optimization function from the current training dataset as the optimal impedance parameter set.

5. The robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters according to claim 4, characterized in that, The kernel function in step S52 uses the anisotropic Matern kernel: ; In the formula, Represents any two sets of five-dimensional impedance parameters. Indicates the signal variance. This represents the Euclidean distance between any two sets of five-dimensional impedance parameters. , This represents the length scale parameter. Represents the target data point. This represents a data point in the current training dataset.

6. The robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters according to claim 4, characterized in that, The expected improvement function in step S53 is: ; In the formula, Represents the Gaussian process in a five-dimensional impedance parameter set The predicted mean at the location; Represents the Gaussian process in a five-dimensional impedance parameter set The standard deviation of the forecast at that location; This represents the optimal objective function value in the current training dataset. The preset exploration coefficient, The cumulative distribution function represents the standard normal distribution. The probability density function representing the standard normal distribution; Indicates the amount of standardization improvement. ,in, , This represents the performance function value of each data point in the current training dataset. Represents the covariance matrix between data points. Indicates the noise variance. This represents the transpose of the covariance vector between the target data point and the data points in the current training dataset.

7. The robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters according to claim 1, characterized in that, The weighting coefficient of the damping function in step S6 is: ; In the formula, Represents an exponential function. This represents a constant that affects the rise and fall of the weighting coefficients of the damping function. Range parameter representing the real-time motion intent.

8. The robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters according to claim 7, characterized in that, The asymptotic damping adjustment function constructed in S7 is: ; ; ; In the formula, These represent the damping functions that affect the speed decrease. and the damping function as velocity increases The constant of the growth rate The minimum parameter representing the speed of the rotary joint at the end of the robotic arm. This represents a parameter within the speed range of the rotary joint at the end of the robotic arm. ; Indicates minimum damping. Indicates transition damping, Indicates maximum damping; The stiffness adjustment function is: ; In the formula, This represents a constant value that affects the rate of change of the stiffness adjustment function; This represents the real-time motion intention value in a static state. Indicates minimum stiffness. This indicates the maximum stiffness.

9. The robotic arm rehabilitation training control method based on Bayesian optimization of asymptotic impedance parameters according to claim 8, characterized in that, In step S8, the torque mapped to the end effector joint of the robotic arm is calculated as follows: ; In the formula These represent acceleration error, velocity error, and angular error, respectively. These are the inertia matrix, damping matrix, and stiffness matrix, respectively. This represents the gravity compensation feedforward term. This represents the feedforward term for friction compensation.