Deep underground man-in-the-loop smart remote control method and device
By combining electromyographic feedback and teaching learning, an extended DMPs model was constructed, and stiffness parameters were adjusted in real time. This solved the problem of unstable control in deep underground environments using traditional teleoperation technology, and achieved efficient and stable remote control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHEASTERN UNIV CHINA
- Filing Date
- 2026-01-06
- Publication Date
- 2026-04-17
AI Technical Summary
Traditional teleoperation technology struggles to adapt to nonlinear and time-varying disturbances in complex deep underground environments, leading to unstable control and difficulty in simultaneously achieving control transparency, stability, and robustness. Existing adaptive impedance methods offer limited improvement.
By combining electromyographic feedback and teaching-based learning, an extended DMPs model is constructed by collecting electromyographic signals and motion trajectory data during the offline teaching phase. Stiffness parameters are adjusted in real time, and bidirectional force feedback is achieved using admittance controllers and impedance controllers for online correction and fusion, thereby improving the system's adaptability and interactive transparency.
It enables efficient and stable remote control in deep underground environments, improves operational dexterity and human-computer interaction performance, and ensures the system's stability and reliability in complex environments.
Smart Images

Figure CN121870752A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep underground exploration and operation technology, and in particular to a method and device for the intelligent remote control of deep underground personnel. Background Technology
[0002] With the increasing demand for deep underground engineering (such as mining, tunnel construction, and physical simulation of deep disasters), teleoperation systems need to be extended to more complex underground scenarios, especially given the complex working environments of coal mines with low illumination, limited texture, confined spaces, and uncertain structures. Deep environments are typically characterized by insufficient lighting, confined spaces, and severe environmental disturbances, making it difficult for operators to make timely and accurate decisions based solely on vision and position mapping. The uncertainty and complexity of the environment significantly reduce the performance and efficiency of teleoperated robotic arms. In practice, for example, the DARPA Underground Challenge in the United States requires robots to work in extreme underground environments such as darkness, dust, and limited communication, severely testing the transparency, stability, and reliability of remote operation.
[0003] Traditional teleoperation techniques primarily employ position control or fixed impedance control strategies. Fixed impedance control assumes known environmental characteristics and struggles to adapt to nonlinear and time-varying disturbances in deep environments. When environmental changes or contact conditions alter, such controllers cannot adjust parameters in real time, potentially leading to control instability or even mission failure. While methods such as adaptive impedance and switching impedance exist to improve control performance, performance enhancements in complex interactive tasks remain limited, failing to significantly improve mission efficiency or reduce operator workload. Traditional force-position control performs poorly in unpredictable interactions, struggling to simultaneously achieve control transparency, stability, and robustness.
[0004] Therefore, this invention provides a method and apparatus for dexterous remote control of a human in a deep underground environment. It fully integrates electromyographic feedback and teaching learning to intelligently adjust robot movements, enabling real-time capture of the operator's electromyographic intentions, fusion of teaching skills learning, and adaptive control—a dexterous teleoperation system. This simulates and enhances dexterous remote control of a human in a remote environment, improving the transparency of interaction between the master operator and the slave robot in complex remote control tasks. Summary of the Invention
[0005] This invention establishes a method and device for dexterous remote control of personnel in deep underground environments, providing operators with human-environment capabilities for carrying out deep resource and energy development tasks in dangerous or remote environments. By enhancing the adaptability of the master end to the slave end environment, this invention significantly improves operational dexterity and human-machine interaction performance. The system is divided into various modules and a controller, each responsible for the functions required at each step of the pre-teach-to-real-time task process, as well as the information flow processing that supports the continuous operation of the system. This technology can also be applied to extreme environments such as the deep sea, nuclear waste disposal, and disaster relief, aiming to extend human perception and control capabilities and safely transfer complex tasks to dangerous areas.
[0006] The technical solution of the present invention is as follows: A method for intelligent remote control of deep underground people in a loop, specifically including the following steps: Step 1: Add an offline teaching phase before the real-time task phase; Step 2: Offline teaching phase, perform multiple teaching tasks, collect electromyographic signals and record motion trajectory data, perform variable impedance estimation on the electromyographic signals, perform stiffness mapping of the end of the arm, and obtain stiffness data; Step 3: Construct an extended DMPs model based on the multiple stiffness data and motion trajectory data from Step 2, perform weighted fusion, and output a comprehensive model that can generalize position stiffness. ; Step 4: In the real-time task phase, in the bidirectional force feedback master-slave system, the stiffness parameters in both the admittance controller at the master end and the impedance controller at the slave end adopt the comprehensive model from Step 3. ; Step 5: During the real-time task phase, simultaneously acquire real-time stiffness data and calculate the integrated model used in Step 4. The tracking error e of the medium stiffness data is used to update the weights using gradient increments. This allows for the acquisition of online fused stiffness output, integration of reference trajectories, and the generation of a comprehensive real-time correction model. ; Step Six: During the real-time task, the bidirectional force feedback master-slave system continuously uses real-time updated... .
[0007] Step two specifically involves: Step 2.1: Perform low-pass filtering and envelope extraction on the acquired electromyographic signals to obtain the envelope amplitude. The calculation formula is as follows: Where W is the window size and the number of sampling points, Indicates a point in time t The amplitude of the electromyographic signal; The absolute values of the electromyographic signal envelope amplitude for each channel are taken and summed to obtain the comprehensive stiffness index p. The calculation formula is as follows: in, N For the number of channels, N =8; Step 2.2: Stiffness index based on muscle co-contraction model p Joint stiffness Estimation of variable impedance; Step 2.3: Adjust joint stiffness using conservative uniform transformation. End stiffness mapped to Cartesian space Stiffness data, calculated using the following formula: in, It is the Jacobian matrix of the human arm. It is an external force at the end. It is the pseudo-inverse of the Jacobian matrix.
[0008] The calculation formula for the muscle co-contraction model is as follows: in, It is the inherent joint stiffness during minimal muscle co-contraction. and All of them are positive coefficients that need to be calibrated.
[0009] Step three specifically involves: Step 3.1: Fit the residuals of the motion trajectory data to the radial basis function to demonstrate the trajectory. Analyzing the demonstration data backwards The weights are solved by local weighted regression. ; Step 3.2: Extend the DMPs framework into a dual-transformation system, simultaneously encoding trajectory data and stiffness data. The calculation formula is as follows: in, It is a scaling ratio of motion speed. p It is a stiffness index. s They share the same standard system. It is a trajectory forcing term. It is the location of the trajectory target. It is the trajectory spring coefficient. This is the current location. This is the current speed. It is a stiffness forcing term. It is the target value for stiffness. It is the stiffness convergence coefficient. It is the rate of change of stiffness; Step 3.3: The weighted average output model of DMPs for multiple demonstration trajectories is calculated using the following formula: in, It is the first i The score for this demonstration It consists of trajectory and stiffness data from a single demonstration.
[0010] The bidirectional force feedback master-slave system is specifically as follows: The controller of the master teleoperated device is an admittance controller. The torque input required at the master end is the sum of the operator's operating force on the end effector and the feedback force from the slave end actuator. The specific controller of the master teleoperated device is as follows: The controller of the slave actuator uses a basic impedance control model. The input of the slave actuator comes from the sum of the virtual force of the master actuator, the inverse dynamics compensation of the slave actuator, and the zero-space controller. The first interface between the operator and the master teleoperated device has an admittance controller that reproduces the interactive dynamics at the slave end effector. The second interface between the slave and the environment is represented by the impedance of the operator input obtained by the master teleoperated device.
[0011] Step five specifically involves: Step 5.1: Extract the comprehensive model from Step 4. Stiffness data in The operator's real-time electromyographic intentions are obtained and estimated into stiffness data. The formula for calculating the tracking error is as follows: in, It is tracking error; Step 5.2: To minimize the tracking error, a quadratic cost function is constructed, calculated as follows: Online reduction The gradient descent is applied to the fusion weights, and the gradient components are calculated using the following formula: in, It is the objective function. It is an online weight fusion. This is the initial value of the offline weights; The continuous-time gradient flow for each weight is calculated using the following formula: in, The learning rate; Step 5.3: Online correction is performed in small increments to maintain the overall passivity of the system. The bidirectional force feedback master-slave system employs critical damping. For the adaptive attenuation strategy, the actual control updates are based on discrete step lengths. The calculation formula is as follows: Obtain online fused stiffness output: In conjunction with the offline reference trajectory, control commands are generated, ultimately resulting in a comprehensive online real-time correction model. The output expression is as follows: .
[0012] A smart remote control device using a smart remote control method for deep underground human circumference includes: The master-end teleoperation device includes a six-degree-of-freedom robotic arm and an end effector handle, which is used by the operator to perform operations at the master end; The slave execution device is a composite robot consisting of a six-degree-of-freedom robotic arm, an end-effector tool, and a deep-earth all-terrain vehicle. It is used to replicate the operation output from the operator at the master end and execute the operation tasks at the slave end. Human upper limb motion acquisition device, used to acquire electromyographic signals of the operator's upper arm; Force sensing devices, including six-dimensional force sensors at the ends of the master and slave robotic arms, are used to collect the operating force of the master operator and the contact force between the slave and the environment. Terminal interaction equipment is used to transmit and process force signals and chassis control signals from the master and slave ends.
[0013] The beneficial effects of this invention are as follows: This invention can be applied to various scenarios involving dangerous conditions in deep underground environments. The impedance characteristics of the master-slave system can dynamically match and follow the operator's real-time intentions, ensuring system robustness and stability while enhancing the tactile feedback of the mechanical properties of the remote environment, achieving higher interactive transparency and operational efficiency. It enables agile remote control under extreme conditions in deep underground environments; the system architecture is clear and the functions are complete, ensuring stable and reliable operational performance in complex environments. Attached Figure Description
[0014] Figure 1 A schematic diagram of a method for the intelligent remote control of people in deep underground environments.
[0015] Figure 2 A flowchart illustrating the output of the fusion model method for the teaching demonstration module.
[0016] Figure 3 This is a schematic diagram of a bidirectional force feedback master-slave system.
[0017] Figure 4 This is a schematic diagram illustrating the specific implementation process of the present invention. Detailed Implementation
[0018] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0019] A dexterous remote control system for deep underground human-enclosed environments includes a master-end teleoperation device, a slave-end execution device, a human upper limb motion acquisition device, a force sensing device, and a terminal interaction device. The terminal interaction device includes a teaching demonstration module, an impedance control mapping module, a bidirectional force feedback interaction module, and an online calibration module. This system simulates and enhances dexterous human manipulation in remote environments, improves the transparency of interaction between the master-end operator and the slave robot in complex remote control tasks, and can promptly capture the master-end operator's control intentions to improve the effectiveness of real-time dynamic control. (See schematic diagram below.) Figure 1 As shown.
[0020] The teaching demonstration module is driven by the instructor through mechanical coupling, with the slave arm synchronously following and recording motion and stiffness data. At the same time, tactile feedback is generated through virtual spring-damping. During the teaching phase, it ensures that the instructor can naturally adjust electromyography and stiffness, while allowing the robot to accurately obtain trajectory and impedance information, providing high-quality demonstration data for subsequent skill learning.
[0021] The online calibration module continuously calibrates the fusion model, allowing operators to instantly sense changes in robot stiffness during cutting or loading, thus obtaining more realistic force feedback. It can quickly adapt to the environment without offline parameter adjustment, significantly shortening the system's deployment and iteration cycle. When faced with sudden working conditions such as changes in material hardness or path inflection points, the online calibration mechanism can dynamically adjust the fusion model to suppress overshoot or oscillation caused by sudden changes, ensuring smooth and continuous robot movements.
[0022] The impedance control mapping module transforms the fused stiffness index into force-motion control commands that the robot can execute. The mapping can maintain energy conservation between the joint and the end effector space, and match the end effector force-displacement relationship with the joint stiffness configuration, thereby ensuring the physical rationality of the control. The robot can both track the reference trajectory and dynamically adjust the interaction force according to the environmental stiffness.
[0023] The bidirectional force feedback interaction module uses an admittance controller for the master device. The required torque input at the master end is the sum of the operator's operating force on the end effector and the feedback force from the slave device. The slave device controller uses a basic impedance control model, with its input derived from the virtual force of the master device, the slave device's inverse dynamics compensation, and the zero-space controller. The first interface between the operator and the master device has an admittance controller that reproduces the interactive dynamics at the slave end effector. The second interface between the slave and the environment simulates the impedance of the operator input acquired by the master device.
[0024] Furthermore, in the aforementioned dexterous remote control device for deep underground human-in-the-loop system, the master-end remote control device is a six-degree-of-freedom robotic arm and an end effector handle; the slave-end execution device is a composite robot consisting of a six-degree-of-freedom robotic arm, an end effector tool, and a deep-earth all-terrain vehicle; the human upper limb motion acquisition device is a MYO armband, which incorporates eight surface electromyography (EMG) signal sensors and a nine-axis inertial sensor; and the force sensing device is a six-dimensional force sensor installed at the ends of the master and slave robotic arms.
[0025] Furthermore, the above offline teaching process is as follows: Electromyography (EMG) signals were collected. The stiffness of the human arm joint is influenced by synergistic muscle contraction (i.e., muscle activation). To minimize the deviation between the stiffness expression calculated from EMG signals and the end-effector stiffness indirectly obtained through measurement, a Frobenius norm-based minimization method was used to estimate the parameters in the model. After mapping the end-effector stiffness, an extended DMPs framework was used in the offline teaching phase to simultaneously model the motion trajectory and stiffness changes. A transformation system with two shared phase variables was used, combined with weighted fusion of multiple teaching data, to achieve collaborative generalization and autonomous generation of the motion trajectory and stiffness strategy in the new scenario, resulting in a comprehensive output. The offline teaching flowchart is as follows: Figure 2 As shown.
[0026] Furthermore, the above-mentioned bidirectional force feedback master-slave system is as follows: Theoretically, the master and slave ends have completely identical mechanical structures and dynamics (isomorphic). When they execute the same impedance / admittance controller locally, the slave end uses the same stiffness and damping parameters. For the slave end, there are only two external inputs: one is the virtual force from the master end, and the other is the contact force between the slave end and the environment. The contact force does not directly participate in the local control law calculation of the slave end, but is transmitted back to the master end via communication. The virtual force from the master end serves as a virtual coupling between the master and slave ends, and the coupling stiffness is designed to be constant instead of using the stiffness of the master end. A schematic diagram of the bidirectional force feedback is shown below. Figure 3 As shown.
[0027] Furthermore, the algorithm for the above online correction fusion model is as follows: In real-time tasks, it is desirable for the fused stiffness output by the robotic arm to closely follow the operator's immediate electromyographic intentions. To minimize the tracking error, a quadratic cost function is constructed to describe the deviation between the "operator's desired stiffness" and the "model output stiffness". Gradient descent is applied to the fusion weights, and normalized projection is performed after each update. Normalization after each weight update ensures that the weights are non-negative and sum to one, so that the fusion result always falls within the convex hull of the historical taught stiffness curve, avoiding "outrageous" stiffness and obtaining online fused stiffness output.
[0028] The flowchart illustrates a method for simulating and enhancing human dexterity in a remote environment within a dexterous remote control device for deep underground human environments. Figure 4 As shown, the steps are as follows: Step 1: Select various unstructured objects such as soil, rocks, and stones commonly found in deep underground or hazardous environments as the targets for the teaching task, which involves cutting, pushing, moving, or manipulating them. While performing these teaching tasks, the operator wears a MYO electromyography armband, and the positional trajectory and electromyography (EMG) signals of the operator's forearm are recorded for each teaching process.
[0029] Step 2: The acquired raw electromyographic signals are processed by low-pass filtering (cutoff frequency of about 5 Hz), rectification, envelope extraction, etc., to calculate feature values reflecting the degree of muscle activation. These feature values are then mapped to operational stiffness reference values through a human-like model, and the stiffness values and position trajectories are time-stamped together.
[0030] Step 3: The extended discrete DMP (Dynamic Motion Element) framework is used to learn the motion trajectory and impedance (stiffness) strategy simultaneously. The output DMP parameter set contains the complete reference position trajectory and stiffness strategy as the overall fused output.
[0031] The specific implementation method of the extended discrete DMP framework includes the following steps: Step 3.1: Use two transformation systems to represent the changes in motion and stiffness over time, respectively. They share the same gauge phase system s to ensure that the motion output and stiffness output are synchronized in time or phase. Each teaching iteration determines the transformation parameters for trajectory and stiffness by fitting a Gaussian function.
[0032] Step 3.2: Assign a quality score to the teaching data output in Step 3.1, then calculate the weighted fusion weight according to the score, and take the weighted average of the DMP outputs from multiple teaching sessions to obtain the final reference signal.
[0033] Step 4: In the real-time task, the fused output of the teaching output serves as the stiffness and position reference in the master-slave control; the admittance controller between the operator and the master, the torque input required by the master is the sum of the operator's operating force on the end effector and the feedback force of the slave device, the master output to the slave simulates the impedance of the operator input obtained by the master device, and a virtual force is applied at the slave to drive the slave haptic enhancement.
[0034] Step 5: The slave device controller uses a basic impedance control model, with the slave device's input derived from the virtual force of the master device, the slave device's inverse dynamics compensation, and the sum of the zero-space controller. The interaction force FS between the slave end actuator and the environment is transmitted back to the operator through the admittance gain module.
[0035] Step 6: Simultaneously with the real-time task, the stiffness index was obtained in real time, and the stiffness tracking error was obtained. Through the closed-loop design of optimization-gradient-normalization, the overall online correction model was output to ensure that the robot stiffness output not only inherits the offline multi-teaching experience, but also responds to the operator's subjective intentions in real time.
[0036] The specific implementation method of the optimization-gradient-normalization closed-loop design includes the following steps: Step 6.1: For the tracking error, construct a quadratic cost function to reduce it online. Apply gradient descent to the fusion weights and correct it online with small increments to avoid compromising the overall passivity of the system. This system uses critical damping and adopts an adaptive attenuation strategy.
[0037] Step 6.2: In actual control, update the weights using discrete step lengths. After each update, perform normalized projection. Normalize the weights after each weight update to ensure that the weights are non-negative and sum to one, dynamically matching the operator's intentions, and outputting the overall online real-time correction model.
[0038] Step 7: Theoretically, the master and slave ends have completely identical mechanical structures and dynamics (isomorphic), and the slave end uses the same stiffness and damping parameters. The online correction model is output in real time, and stiffness information is transmitted back to both the master and slave ends. The master and slave platforms execute the output stiffness for task control, acquire stiffness errors in real time, and continue online correction to form a real-time correction closed loop.
Claims
1. A method for the dexterous remote control of a person in a deep underground loop, characterized in that, The specific steps are as follows: Step 1: Add an offline teaching phase before the real-time task phase; Step 2: Offline teaching phase, perform multiple teaching tasks, collect electromyographic signals and record motion trajectory data, perform variable impedance estimation on the electromyographic signals, perform stiffness mapping of the end of the arm, and obtain stiffness data; Step 3: Construct an extended DMPs model based on the multiple stiffness data and motion trajectory data from Step 2, perform weighted fusion, and output a comprehensive model that can generalize position stiffness. ; Step 4: In the real-time task phase, in the bidirectional force feedback master-slave system, the stiffness parameters in both the admittance controller at the master end and the impedance controller at the slave end adopt the comprehensive model from Step 3. ; Step 5: During the real-time task phase, simultaneously acquire real-time stiffness data and calculate the integrated model used in Step 4. The tracking error e of the medium stiffness data is used to update the weights using gradient increments. This allows for the acquisition of online fused stiffness output, integration of reference trajectories, and the generation of a comprehensive real-time correction model. ; Step Six: During the real-time task, the bidirectional force feedback master-slave system continuously uses real-time updated... .
2. The method for skillfully remotely controlling a person in a deep underground loop according to claim 1, characterized in that, Step two specifically involves: Step 2.1: Perform low-pass filtering and envelope extraction on the acquired electromyographic signals to obtain the envelope amplitude. The calculation formula is as follows: ; Where W is the window size and the number of sampling points, This represents the amplitude of the electromyographic signal at time point t; The absolute values of the electromyographic signal envelope amplitude for each channel are taken and summed to obtain the comprehensive stiffness index p. The calculation formula is as follows: ; Where N is the number of channels, N=8; Step 2.2: Calculate joint stiffness based on the muscle co-contraction model and the stiffness index p. Estimation of variable impedance; Step 2.3: Adjust joint stiffness using conservative uniform transformation. End stiffness mapped to Cartesian space Stiffness data, calculated using the following formula: ; in, It is the Jacobian matrix of the human arm. It is an external force at the end. It is the pseudo-inverse of the Jacobian matrix.
3. The method for skillfully remotely controlling a person in a deep underground loop according to claim 2, characterized in that, The calculation formula for the muscle co-contraction model is as follows: ; ; in, It is the inherent joint stiffness during minimal muscle co-contraction. and All of them are positive coefficients that need to be calibrated.
4. The method for skillfully remotely controlling a person in a deep underground loop according to claim 2, characterized in that, Step three specifically involves: Step 3.1: Fit the residuals of the motion trajectory data to the radial basis function to demonstrate the trajectory. Analyzing the demonstration data backwards The weights are solved by local weighted regression. ; Step 3.2: Extend the DMPs framework into a dual-transformation system, simultaneously encoding trajectory data and stiffness data. The calculation formula is as follows: ; ; in, It represents the scaling ratio of motion velocity, p is the stiffness index, and s indicates sharing the same standard system. It is a trajectory forcing term. It is the location of the trajectory target. It is the trajectory spring coefficient. This is the current location. This is the current speed. It is a stiffness forcing term. It is the target value for stiffness. It is the stiffness convergence coefficient. It is the rate of change of stiffness; Step 3.3: The weighted average output model of DMPs for multiple demonstration trajectories is calculated using the following formula: ; ; in, This is the score for the i-th demonstration. It consists of trajectory and stiffness data from a single demonstration.
5. The method for skillfully remotely controlling a person in a deep underground loop according to claim 1, characterized in that, The bidirectional force feedback master-slave system is specifically as follows: The controller of the master teleoperated device is an admittance controller. The torque input required at the master end is the sum of the operator's operating force on the end effector and the feedback force from the slave end actuator. The specific controller of the master teleoperated device is as follows: The controller of the slave device uses a basic impedance control model. The input of the slave device comes from the virtual force of the master device, the inverse dynamics compensation of the slave device, and the sum of the zero-space controller. The first interface between the operator and the master teleoperation device has an admittance controller that reproduces the interactive dynamics at the end effector of the slave end; the second interface between the slave end and the environment is characterized by simulating the impedance of the operator input acquired by the master teleoperation device.
6. A method for the skillful remote control of a deep underground person in a loop according to claim 5, characterized in that, Step five specifically involves: Step 5.1: Extract the comprehensive model from Step 4. Stiffness data in The operator's real-time electromyographic intentions are obtained and estimated into stiffness data. The formula for calculating the tracking error is as follows: ; in, It is tracking error; Step 5.2: To minimize the tracking error, a quadratic cost function is constructed, calculated as follows: ; Online reduction The gradient descent is applied to the fusion weights, and the gradient components are calculated using the following formula: ; in, It is the objective function. It is an online weight fusion. This is the initial value of the offline weights; The continuous-time gradient flow for each weight is calculated using the following formula: ; in, The learning rate; Step 5.3: Online correction is performed in small increments to maintain the overall passivity of the system. The bidirectional force feedback master-slave system employs critical damping. For the adaptive attenuation strategy, the actual control updates are based on discrete step lengths. The calculation formula is as follows: ; Obtain online fused stiffness output: ; In conjunction with the offline reference trajectory, control commands are generated, ultimately resulting in a comprehensive online real-time correction model. The output expression is as follows: 。 7. A smart remote control device using the smart remote control method for deep underground human circumference as described in any one of claims 1-6, characterized in that, include: The master-end teleoperation device includes a six-degree-of-freedom robotic arm and an end effector handle, which is used by the operator to perform operations at the master end; The slave execution device is a composite robot consisting of a six-degree-of-freedom robotic arm, an end-effector tool, and a deep-earth all-terrain vehicle. It is used to replicate the operation output from the operator at the master end and execute the operation tasks at the slave end. Human upper limb motion acquisition device, used to acquire electromyographic signals of the operator's upper arm; Force sensing devices, including six-dimensional force sensors at the ends of the master and slave robotic arms, are used to collect the operating force of the master operator and the contact force between the slave and the environment. Terminal interaction equipment is used to transmit and process force signals and chassis control signals from the master and slave ends.