Space manipulator multi-mode control method and system based on intention estimation

By combining depth cameras and eye-tracking devices with hidden Markov models and virtual gripper force fields, the problems of insufficient intent capture and poor environmental adaptability in the teleoperation of space robotic arms were solved, achieving efficient and safe operation of space robotic arms.

CN121893239APending Publication Date: 2026-04-21CHINA ACAD OF AEROSPACE SCI & TECH INNOVATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA ACAD OF AEROSPACE SCI & TECH INNOVATION
Filing Date
2025-11-28
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

There are problems in teleoperation of space robotic arms, such as insufficient capture of implicit intentions, poor adaptability to three-dimensional environment, and low efficiency of human-machine collaboration. In particular, it is difficult to balance operation efficiency and safety in complex task scenarios.

Method used

By acquiring multimodal information through depth cameras and eye-tracking devices, and combining it with hidden Markov models and virtual gripper force fields, the robot can achieve real-time recognition of operator intentions and path planning, generate safe movement trajectories for the robotic arm, and provide tactile feedback.

Benefits of technology

It significantly improves operational precision and efficiency, increasing efficiency by 11.92% and accuracy by 40.33%, while reducing the cognitive load and psychological stress on operators. It is suitable for complex path planning and delicate operation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121893239A_ABST
    Figure CN121893239A_ABST
Patent Text Reader

Abstract

The invention provides a space manipulator multi-mode control method and system based on intention estimation. The method comprises the steps that multi-mode information is obtained; preprocessing and dynamic probability modeling are conducted on the fixation point data, the target area intention of an operator is recognized, a two-dimensional fixation point coordinate is determined, and the two-dimensional fixation point coordinate and space depth information are fused and then mapped into a three-dimensional target coordinate under a mechanical arm base coordinate system; according to the three-dimensional target coordinates, a three-dimensional voxel grid map is combined for safe path planning, a mechanical arm movement track is generated, a virtual clamp force field containing boundary constraining force, path guiding force, position deviation force and speed damping force is constructed, guiding of a force feedback master controller is achieved through the virtual clamp force field, and the mechanical arm movement track is generated. Further, the operation of the operator is restrained; the force feedback controller generates a motion instruction according to the position and posture information, collected in real time, of the operation finger; and the position and posture of the end effector are calculated and controlled in real time according to the motion instruction, a control signal is driven to the mechanical arm, and then the mechanical arm is controlled to move towards the to-be-grabbed target according to the intention of the operator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of space robotics and intelligent teleoperation technology, specifically relating to a multimodal control method and system for a space robotic arm based on intent estimation. Background Technology

[0002] With the rapid growth in demand for on-orbit space services and the expansion of complex mission scenarios, space robotic arms, with their high-precision operation and autonomous adaptability, have become the core execution unit for tasks such as satellite maintenance, extravehicular equipment assembly, and space debris cleanup.

[0003] However, the space environment has strong unstructured characteristics, such as target dynamic drift, communication delay fluctuations, and drastic changes in lighting. Traditional robotic arm teleoperation relies on manual control throughout the process, and faces the following key challenges: (1) It is difficult to balance operation efficiency and safety. Operators need to infer the three-dimensional spatial target pose through two-dimensional image information. Path planning failure is easily caused by visual misjudgment or singular points in the robotic arm's movement, resulting in long task time and increased collision risk; (2) The level of human-machine collaboration intelligence is insufficient. Existing systems mostly adopt a single control mode (such as handle or keyboard commands), lacking the ability to dynamically perceive the operator's implicit intentions, making it difficult to achieve the complementary advantages of human situational decision-making and precise machine execution; (3) The adaptability to complex working conditions is limited. Signal delays cause operation instructions to lag. Traditional virtual fixture technology relies on preset path constraints and cannot dynamically adjust the guidance strategy according to real-time environmental changes and operator intentions, resulting in reduced operation smoothness.

[0004] To address the aforementioned issues, existing research primarily optimizes teleoperation performance through visual enhancement algorithms, motion trajectory pre-compensation, or static force feedback. However, two major shortcomings remain: First, intent perception is largely based on explicit trigger signals (such as specific buttons or voice commands), making it difficult to capture the operator's implicit attention to non-preset targets. Second, human-computer interaction strategies are rigid and fail to combine real-time intent estimation with environmental conditions to generate an adaptive guiding force field, resulting in a failure to simultaneously achieve efficiency improvements and safety assurance. Summary of the Invention

[0005] The technical problem solved by this invention is to address the problems of insufficient implicit intent capture, poor three-dimensional environment adaptability, and low human-machine collaboration efficiency in the teleoperation of space robotic arms in the prior art. This invention proposes a multimodal control method and system for space robotic arms based on intent estimation.

[0006] The solution of this invention is: a multimodal manipulation method for a spatial robotic arm based on intent estimation, comprising:

[0007] Multimodal information is acquired by using a depth camera to obtain color images and spatial depth information of the slave control terminal's working area and transmitting them to the master control terminal for real-time display. The color image displayed to the operator on the master control terminal is collected by an eye-tracking device to collect the operator's eye movement data. The force feedback controller collects the position and posture information of the operator's fingers on the master control terminal in real time. The slave control terminal's working area is an area that includes the robotic arm and the target to be grasped.

[0008] The gaze point data is preprocessed and dynamic probability modeled to identify the operator's target area intention, determine the two-dimensional gaze point coordinates, and then the two-dimensional gaze point coordinates are fused with spatial depth information and mapped to three-dimensional target coordinates in the robot arm base coordinate system.

[0009] Based on the three-dimensional target coordinates, a safe path is planned using a three-dimensional voxel mesh map to generate the robotic arm's motion trajectory. A virtual gripper force field is constructed, which includes boundary constraints, path guidance forces, position deviation forces, and velocity damping forces. The virtual gripper force field guides the force feedback controller, thereby constraining the operator's actions.

[0010] The force feedback controller generates motion commands based on the real-time (e.g., 2kHz) acquisition of the operator's finger position and posture information (the device's 7 degrees of freedom position and posture, and real-time rendering of force feedback for three degrees of freedom);

[0011] The position and attitude of the end effector are calculated and controlled in real time according to the motion command, and the control signal is sent to the robotic arm to control the robotic arm to move toward the target to be grasped according to the operator's intention.

[0012] Preferably, preprocessing and dynamic probability modeling of the gaze point data to identify the operator's target region intent includes:

[0013] Preprocess the eye-tracking data to remove outliers;

[0014] The floating region of the gaze point is determined based on the 3σ range of the normal distribution, and the physical size of the floating region is mapped to the pixel range.

[0015] The workspace of the robotic arm is rasterized into a discrete set of states according to pixel range;

[0016] The preprocessed eye-tracking coordinates were used as the observation sequence. By initializing the initial attention probability π, the inter-region transition probability A, and the state generation observation probability B, a hidden Markov model of eye-tracking gaze behavior and target region was constructed.

[0017] The Viterbi algorithm is used to decode the sequence of the highest probability target regions representing the operator's intention (this is a continuous reasoning process, inferring the next state from the previous state, and finally obtaining a complete sequence of the highest probabilities. The last and latest state of this sequence is the target region with the highest probability at present), and a dynamic update mechanism is used to adjust the transition probability A between regions in real time.

[0018] Dynamic adjustment involves adjusting the sigma value based on the latest target state (i.e., the speed and position of movement), thereby adjusting the transition probability between regions.

[0019] Preferably, the dynamic update mechanism dynamically adjusts the mean of the normal distribution based on the real-time gaze point position and dynamically adjusts the standard deviation σ based on the gaze movement speed;

[0020] The relationship between standard deviation σ and eye movement speed v is as follows:

[0021]

[0022] The unit of movement speed is m / s.

[0023] Preferably, the operator's actions are constrained in the following manner:

[0024] Based on the color images and depth information acquired by the depth camera, 3D point cloud data is constructed and a voxel mesh map is generated;

[0025] The coordinates of the two-dimensional gaze point are fused with the depth information of the corresponding area, and then converted into the coordinates of the three-dimensional target point in the robot arm's base coordinate system through the coordinate transformation matrix of the depth camera and the robot arm.

[0026] Complete the intrinsic parameter calibration and distortion calibration of the depth camera, as well as the hand-eye calibration of the robotic arm (eye-outside-the-hand means: the camera is not mounted on the end of the robotic arm; hand-eye calibration means: determining the geometric relationship (coordinate transformation matrix) between the end of the robotic arm and the camera, and establishing a precise spatial mapping relationship between the vision system and the robotic arm).

[0027] Path search is performed on the voxel grid map and the path is smoothed to generate a collision-free motion trajectory for the robotic arm.

[0028] Based on the planned path, an admittance-type virtual fixture is generated, and the composite force field composed of boundary constraint force, path guiding force, position deviation force and velocity damping force is calculated.

[0029] The virtual fixture's composite force field is mapped to the force feedback controller via a tactile interface, providing tactile guidance to the operator and assisting in completing precise operations.

[0030] Preferably, a heuristic path search algorithm is used to search the three-dimensional voxel grid map, and the path is smoothed based on fifth-order polynomial interpolation.

[0031] Preferably, the composite force field calculation of the virtual fixture force field satisfies the following relationship:

[0032] The boundary constraint force increases linearly with the distance of the force feedback controller's end off the path, preventing the operator from exceeding the safe or permissible operating range.

[0033] Guiding forces include centripetal forces pointing towards the center of the path and axial forces pointing towards the path points; by providing force feedback along the task path, the operator is guided to complete specific operational tasks.

[0034] The positional deviation force is proportional to the displacement difference between the force feedback controller and the end effector of the robotic arm;

[0035] The speed damping force is related to the moving speed of the force feedback controller and the preset damping coefficient, and is used to suppress the operator's rapid or excessive movements.

[0036] A multimodal manipulation system for a space robotic arm based on intent estimation includes:

[0037] The multimodal information acquisition module includes an eye-tracking device, a force feedback controller, and a depth camera. The depth camera is used to acquire color images and spatial depth information of the working area of ​​the slave control terminal and transmit them to the vision interface module for real-time display. When the operator gazes at the color image displayed by the vision interface module, the eye-tracking device collects the operator's eye movement data, and the force feedback controller is installed on the operator's hand to collect the operator's manual position and speed data at the master control terminal.

[0038] The vision interface module acquires color images and spatial depth information of the working area of ​​the slave device through a depth camera;

[0039] The human intent estimation module preprocesses and dynamically probabilistically models the operator's eye movement data collected in real time by the eye-tracking device, identifies the operator's intent in the target area, determines the two-dimensional gaze point coordinates, and maps the two-dimensional gaze point coordinates with depth information to the three-dimensional target coordinates in the robotic arm's base coordinate system.

[0040] The path planning module calculates the robotic arm's trajectory within the operating space based on the 3D target coordinates provided by the intent estimation module and the 3D voxel mesh map of the current environment, and outputs it to the virtual fixture module.

[0041] The virtual clamp module transmits the virtual clamp force field, which includes boundary constraint force, path guiding force, position deviation force and velocity damping force, to the force feedback controller in conjunction with the tactile interface.

[0042] The force feedback controller generates motion commands based on the real-time collected information on the position and posture of the operating fingers;

[0043] The motion control module receives motion commands from the force feedback controller, calculates the position and attitude of the end effector in real time, and drives control signals to the robotic arm module.

[0044] The robotic arm module is responsible for completing precise operation tasks according to the target instructions issued by the motion control module, realizing the precise movement of the end effector in three-dimensional space.

[0045] Preferably, the human intention estimation module processing steps include:

[0046] Preprocessing and dynamic probability modeling of the gaze point data to identify the operator's target region intent includes:

[0047] Preprocess the eye-tracking data to remove outliers;

[0048] The floating region of the gaze point is determined based on the 3σ range of the normal distribution, and the physical size of the floating region is mapped to the pixel range.

[0049] The workspace of the robotic arm is rasterized into a discrete set of states according to pixel range;

[0050] The preprocessed eye-tracking coordinates were used as the observation sequence. By initializing the initial attention probability π, the inter-region transition probability A, and the state generation observation probability B, a hidden Markov model of eye-tracking gaze behavior and target region was constructed.

[0051] The Viterbi algorithm is used to decode the sequence of the highest probability target regions representing the operator's intention, and a dynamic update mechanism is used to correct the inter-region transition probability A in real time.

[0052] The advantages of this invention compared to the prior art are:

[0053] 1. This invention is based on a dynamic eye-tracking intention estimation algorithm using a hidden Markov model. Combined with adaptive probability adjustment and virtual fixture force field guidance, it significantly improves operational accuracy. Experimental verification shows that the efficiency is improved by 11.92% and the accuracy is improved by 40.33%. It is especially suitable for complex path planning and fine operation tasks.

[0054] 2. This invention integrates eye-tracking, depth vision and force feedback technologies to construct a hierarchical human-computer collaboration framework, achieving seamless mapping from two-dimensional intentions to three-dimensional actions. Furthermore, it optimizes task phase switching through dynamic weight allocation, enhancing the system's adaptability in target localization and fine operation.

[0055] 3. This invention significantly reduces the cognitive load and psychological pressure on operators by combining the composite force feedback (boundary constraint force, path guidance force, position deviation force, and velocity damping force) of the virtual fixture with real-time visual prompts. User experiments have confirmed that it significantly improves operating comfort and is more suitable for long-term, high-precision teleoperation tasks. Attached Figure Description

[0056] Figure 1 Schematic diagram of a multimodal teleoperation hierarchical framework design;

[0057] Figure 2 Schematic diagram of the visual focus position coordinate acquisition process;

[0058] Figure 3 Flowchart of the intent estimation algorithm based on Hidden Markov Model;

[0059] Figure 4 Schematic diagram of the transformation process between coordinate systems;

[0060] Figure 5 Flowchart for calculating boundary constraints and guiding forces of virtual fixtures;

[0061] Figure 6 Space robotic arm human-machine hybrid multimodal teleoperation system architecture. Detailed Implementation

[0062] The present invention will be further described below with reference to the embodiments.

[0063] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0064] Figure 1 This is a schematic diagram of a hierarchical framework design for a multimodal teleoperation method for a space robotic arm based on intent estimation, as described in this invention, comprising the following parts:

[0065] The first layer is the perception layer, which uses a depth camera to acquire color images and spatial depth information of the working area of ​​the slave terminal and transmits them to the master terminal for real-time display; the color image displayed by the operator on the master terminal is collected by an eye-tracking device to collect the operator's eye movement data; the force feedback controller collects the position and posture information of the operator's fingers on the master terminal in real time; the working area of ​​the slave terminal is the area containing the robotic arm and the target to be grasped.

[0066] The second layer is the task planning and intent recognition layer. It identifies the operator's intent, such as selecting a target and determining a region, through eye-tracking gaze points. Combining path planning algorithms and the constraint model of the virtual fixture, the operator's intent is mapped into an operation path and region constraints. Furthermore, this invention can consider dynamically allocating the weights of eye-tracking interaction and virtual fixture according to different stages of the task. In the initial target localization stage, eye-tracking interaction is given priority; while in the fine operation stage (when the robotic arm end approaches the target at a distance of less than 15cm), virtual fixture constraints are given priority.

[0067] The third layer is the control and feedback layer, which controls the posture of the robotic arm by combining a hybrid control model that integrates eye tracking, force feedback master controller and virtual gripper, and provides tactile feedback when the virtual gripper guides, allowing the operator to perceive the constraints and guidance of the operation, and displays the position and target of eye tracking in real time on the screen interface for visual feedback.

[0068] The fourth layer is the hardware implementation layer. The eye tracker is used to acquire eye movement data, the depth camera is used to acquire depth information and color images, the robotic arm is used to perform tasks, the master controller (force feedback controller) is used to control the robotic arm and provide force feedback to the operator, and the computer is used to process multimodal data and generate control commands in real time.

[0069] The following describes the implementation details of each part of the multimodal teleoperation method for space robotic arms based on intent estimation, including the following steps:

[0070] S1. Intent estimation based on operator gaze point using a hidden Markov model:

[0071] S11. Preprocessing of eye movement signals:

[0072] S111. Data Acquisition: Preferably, the present invention can use the Tobii Eye Tracker 4C eye tracker as the hardware for acquiring eye movement data. After connecting the eye tracker to a computer, eye movement data is acquired at a frequency of 90Hz using a timer. Finally, the visual focus coordinates can be obtained. The process of acquiring the visual focus position coordinates of eye movement fixation data is as follows: Figure 2 As shown;

[0073] S112. Median Filtering: A non-linear filtering method is used to sort the values ​​of each data point and its neighborhood window, and then take the median to replace the original value, effectively removing burst noise and outliers.

[0074] Specifically, after selecting a suitable data window size, it is applied to the i-th data point x. i The processing procedure is as follows:

[0075]

[0076] In the formula, xj The data is after median filtering. ω represents the data value after mean filtering, and ω2 is the size of the mean filtering window;

[0077] S113. Mean filtering: The median-filtered data is further subjected to a sliding window weighted average to smooth the high-frequency jitter of continuous data;

[0078] Specifically, after selecting a suitable data window size, it applies to the j-th data point x. j The processing procedure is as follows:

[0079]

[0080] In the formula, x j The data is after median filtering. ω represents the data value after mean filtering, and ω2 is the size of the mean filtering window;

[0081] S114. Outlier Removal: Combine threshold judgment method to remove jittery data that deviates continuously from the center point by more than 2 standard deviations;

[0082] S12. Probabilistic Dynamic Modeling and Analysis of Eye-Motion Information:

[0083] S121. Fixation region definition: The fixation region is determined based on the 3σ range of a normal distribution (covering 99.7% of the data), and the physical size is mapped to the pixel range;

[0084] Specifically, by collecting fixation data through actual use of an eye tracker, the standard deviation σ in the left and right X-axis and up and down Y-axis directions is calculated, and the range of 3σ is used as the average range of jitter in the up, down, left and right directions when the operator fixates.

[0085] Specifically, based on the screen's physical size and resolution, and combined with the standard deviation of the normal distribution, the physical range of the gaze point in the x and y directions can be mapped to pixel coordinates to obtain the pixel region. The calculation formula is as follows:

[0086]

[0087] In the formula, X pixel and Y pixel These represent the length and width of the mapped pixel region, respectively. width and .com height These represent the length and width of the computer, respectively; img width and img height These are the actual length and width of the image, respectively. width and src height These are the length and width of the screen, respectively;

[0088] S122. Gridded working area: The working space of the robotic arm is divided into discrete grids according to the pixel range, and the state space is transformed from continuous coordinates to discrete region numbers;

[0089] Specifically, the image resolution used in this invention is 720×1280. The area is rasterized according to the given size, discretizing the continuous human intention space into 60×71 partitions, which are represented sequentially by rows from the top-left coordinate zero point as regions {S0, S1, S2, S3, ..., S...}. 4259 This division method can limit the target selection range of the robotic arm to a smaller spatial area, thereby reducing the uncertainty of the target position;

[0090] Specifically, the rasterization formula is as follows:

[0091]

[0092] S123. Dynamically adjust normal distribution parameters: dynamically adjust the mean of the normal distribution based on the real-time gaze point position, and dynamically adjust the standard deviation σ based on the gaze movement speed (calculated using Euclidean distance). The faster the speed, the larger σ is, and the wider the distribution of the transition probability.

[0093] Specifically, μ is adjusted according to the eye gaze position. μ is the center point of the operator's current gaze. The coordinates (x, y) of this point can be obtained through an eye-tracking device and used as the mean μ of a normal distribution. By updating the mean μ in real time, the system can track changes in the center point of the operator's gaze area, making the model react more quickly to changes in its gaze trajectory.

[0094] Specifically, adjusting σ based on movement speed: When focusing on a specific target, the distribution of fixation points is relatively concentrated with a small range of variation. In this case, reducing the value of σ can make the two-dimensional normal distribution very narrow and sharp, concentrated in a small area of ​​fixation, thus increasing the probability of shifting to the current area or a small surrounding area. Conversely, when quickly browsing an area, the distribution of fixation points is more dispersed with a larger range of variation. In this case, adjusting the value of σ to a larger value can make the two-dimensional normal distribution very wide and flat, thus increasing the probability of shifting to other more distant areas, while the probability of shifting to an area near the previous fixation point is smaller compared to a smaller value of σ.

[0095] Specifically, based on actual measurement data and experience from eye trackers, this invention establishes the following relationship between the standard deviation σ and the eye movement speed v:

[0096]

[0097] In this example, a 15-inch 1080p resolution monitor is used, and the moving speed of a person is measured in m / s when the person is approximately 50cm away from the monitor.

[0098] S13. Eye-tracking intention estimation algorithm based on Hidden Markov Model:

[0099] S131. Model Construction: The working area of ​​the robotic arm is discretized into a set of hidden states. Preprocessed eye-tracking coordinates are used as the observation sequence. The relationship between eye-tracking fixation and the target area is modeled by initializing three elements: π, A, and B (representing the initial attention probability, inter-region transition probability, and probability of generating the current eye-tracking coordinate, respectively). The transition probability is calculated using a dynamic normal distribution, and the observation probability is determined by the conditional probability of generating the current eye-tracking coordinates given the current state.

[0100] q1 is the initial state, o1 is the initial observation, and (x, y) are the two-dimensional coordinates of the gaze point.

[0101] S132. Decoding process: Based on the Viterbi algorithm, the most likely hidden state (target region) sequence is decoded by initializing local probabilities, recursively calculating the maximum probability path and backtracking the optimal sequence; at the same time, a dynamic update mechanism is introduced to adjust the transition matrix in real time based on the movement speed, thereby enhancing the model's adaptability to temporal changes;

[0102] The intent estimation algorithm based on Hidden Markov Model described in this invention is as follows: Figure 3 As shown, the algorithm includes the following steps:

[0103] (1) Discretize the human intention space into a finite number of regions of 60×71;

[0104] (2) Given the mean and standard deviation μ1 and σ1 of the normal distribution, initialize the probability transition matrix A;

[0105] (3) Initialize the initial probability, N = 60 × 71 = 4260.

[0106] (4) Repeat the following:

[0107] (a) Calculate the eye movement speed v based on the operator's gaze point data;

[0108] (b) Calculate σ1 using... Calculate P = (q2|q1), and update A;

[0109] (c) Initialize probability P = (q1|o1), update B;

[0110] (d) Calculate o based on the operator's gaze point (x, y). t ;

[0111] (e) Update the observation sequence set O;

[0112] (f) Update P(q) according to time. t+1 =qj |O t )=P(q t+1 =q j |q t =q i )P(q t =q i |O t ,λ), and update A;

[0113] (g) According to Calculate the most possible state q at the next time step. t+1 ;

[0114] (h) when If the probability exceeds the set probability threshold and the cumulative number of outputs for the most likely state reaches the threshold, the most likely state is confirmed as the operator's intention for the next moment.

[0115] (5) End;

[0116] S14. The method of the present invention can be experimentally verified and evaluated as needed:

[0117] S141. Experimental Environment Deployment: Set up an Ubuntu system environment, integrate ROS and OpenCV frameworks; display a three-color standard target on the screen;

[0118] Preferably, the equipment used in the experiment includes a laptop computer, a Tobii Eye Tracker 4C eye tracker, and a D435 depth camera;

[0119] Preferably, the selected target size is 9.5cm × 9.5cm, and the target center size is 3.7cm × 3.7cm;

[0120] Specifically, the eye tracker is fixed directly below the laptop screen, which displays a color image of the desktop captured by a D435 camera. Three targets of the same size are placed on the desktop, labeled A, B, and C from left to right.

[0121] S142. Real-time test: Set 3-5 target points and ask the operator to look at these targets in a fixed order for 5 seconds. Record and calculate the difference between the gaze movement time Δt and the system response time Δτ. The smaller the difference, the better the real-time performance and the closer the estimation of eye movement speed and position is to the real situation.

[0122] S143. Recognition accuracy test: Set 3-5 target points and have the operator look at these targets in a fixed order for 5 seconds. Record and include the correctness of the gaze state area matching and the misjudgment rate of the saccade state.

[0123] S144. Data visualization presentation: Generate gaze point heatmap and trajectory map, and display the terminal recognition status in real time;

[0124] S2. Human-computer interaction strategy based on eye-tracking intention control: The Hidden Markov Model predicts and estimates the operator's attention area, intention, and operation priority. Based on this, it matches the 2D intention information with the depth information acquired by the depth camera, linking the planar intention with the 3D environment to realize human-computer interaction using eye-tracking information. The specific steps include:

[0125] S21. 3D Scene Perception: Acquire color images and depth information of the working area through a depth camera, construct 3D point cloud data and generate a voxel mesh map to provide environmental perception support for path planning;

[0126] Specifically, by utilizing the typical linear field of view characteristics in visual information and combining the depth data captured by the depth camera, the position estimated by the intention on the screen is mapped to the real three-dimensional environment, generating three-dimensional coordinates corresponding to the environment, thereby realizing the association between the intention information and the real three-dimensional spatial position.

[0127] Specifically, by capturing color images through a depth camera and using them as image feedback from the control end, and by normalizing the pixel coordinates with the top left corner of the image as the origin, the complex mapping between the sensor, the robotic arm and the world coordinate system can be reduced, thereby achieving the goal of directly expressing the intent information in the world coordinate system.

[0128] Specifically, by utilizing coordinate transformation relationships, the estimated state information region of interest can be mapped to the pixel coordinate system. The formula for converting the estimated most likely state at the next moment into the target position of most interest to the operator in the pixel coordinate system is as follows:

[0129]

[0130] Where S represents the state, and u and v are the pixel values ​​corresponding to the state in the pixel coordinate system;

[0131] S22. Two-dimensional to three-dimensional conversion: The two-dimensional gaze point coordinates obtained in step S1 are fused with the depth information of the corresponding area and converted into three-dimensional target point coordinates in the robot arm base coordinate system through a coordinate transformation matrix;

[0132] Specifically, the camera imaging model established in this invention includes a pixel coordinate system (u,v), an image coordinate system (x,y), and a camera coordinate system (X). c ,Y c Z c ), world coordinate system (X) w ,Y w Z w The four coordinate systems and their transformation process are as follows: Figure 4 As shown;

[0133] S23. Hand-eye calibration: Complete the intrinsic parameter calibration and distortion calibration of the depth camera, as well as the hand-eye calibration of the robotic arm with the eye outside the hand, and establish a precise spatial mapping relationship between the vision system and the robotic arm;

[0134] Preferably, the present invention uses the camera_calibration tool of ROS to calibrate the intrinsic and extrinsic parameters of the camera;

[0135] Preferably, in hand-eye calibration, the present invention uses ArUco codes as calibration plates to provide position and orientation information in space;

[0136] S3. Human-machine hybrid path planning and virtual fixture force guidance:

[0137] S31. Construction of 3D Voxel Mesh Map: This invention uses point cloud data acquired by a depth camera as input and performs voxel filtering. After defining a 2cm×2cm×2cm cube grid, the point cloud data is divided into multiple 3D voxel meshes;

[0138] S32. Safe path planning: The A* algorithm is used to search for paths in the 3D voxel map, and the path is smoothed by fifth-order polynomial interpolation to generate a collision-free motion trajectory for the robotic arm.

[0139] Preferably, the present invention uses the A* algorithm for path planning. This algorithm combines heuristics and cost search, and efficiently finds the optimal path by evaluating path cost and target prediction.

[0140] Preferably, the present invention uses Manhattan distance to measure the distance between two points in space relative to each other in the grid space;

[0141] Preferably, the present invention uses Manhattan distance to calculate the heuristic function, and the actual cost function g(n) and the heuristic cost function h(n) are as follows:

[0142] g(n) = D*(|x i -x b |+|y i -y b |+|z i -z b |) (7)

[0143] h(n) = D*(|x i -x e |+|y i -y e |+|z i -z e |) (8)

[0144] In the formula, D is the unit cost, (x i ,y i,z i ),(x b ,y b ,z b ),(x e ,y e ,z e ) represent the three-dimensional coordinates of the current node, the start point, and the end point, respectively;

[0145] S33. Virtual Fixture Construction: Generate admittance-type virtual fixtures based on the planned path, and calculate the composite force field composed of boundary constraint force, path guiding force, position deviation force and velocity damping force;

[0146] Specifically, the purpose of boundary constraints is to prevent operators from exceeding the safe or permissible operating range, and their function is to restrict the operator's movement;

[0147] Specifically, guiding force aims to guide the operator to complete a specific operational task by providing force feedback along the task path;

[0148] Specifically, the deviation force is a feedback force generated by the displacement difference between the position of the master controller and the position of the end effector of the robotic arm, which is used to reflect the relative offset between the two.

[0149] Specifically, speed damping force improves handling stability and reduces the possibility of misoperation by suppressing rapid or excessive motion;

[0150] Specifically, when the end of the main controller deviates from V pi When the current end position of the master controller is at the desired trajectory V, pi As the vertical distance d increases, it generates a gradually increasing attractive force from the end of the main controller to the virtual path in the direction of the virtual path, up to a maximum of 10N, to prompt the operator to return to the desired trajectory V. pi In this invention, this attractive force is named the boundary constraint force F. eb ;

[0151] Specifically, this invention designs two guiding forces: one is a centripetal force F pointing towards the center of the path. ec Secondly, the axial force F pointing towards a specific path point. ep These two forces combined constitute the guiding force of the main controller, which is used to assist in guiding the operation process;

[0152] Specifically, the boundary constraint force Feb, guiding force Fec, and Fep are calculated as follows: Figure 5 As shown, there are four scenarios: when the master controller moves to the middle of the path between adjacent path points (i.e., ... Figure 5In the context of `num_zeros==1`, the current `Fec` is only subject to the radial attraction of the current path segment, with a magnitude proportional to the vertical distance from the end of the hand controller to the path; the current `Fep` is the axial attraction pointing towards the next path point along the current path segment direction, with a magnitude proportional to the size of the current path segment vector; if the vertical distance `d` from the end of the hand controller to the path is greater than the radius `r` of the virtual fixture, then `Feb` uses radial attraction, with a magnitude proportional to its distance from the boundary (`dr`); otherwise, `Feb=0`. When the master controller moves to the corner side of the cylinder formed by the two adjacent paths as axes (i.e., ... Figure 5 In the given condition, num_zeros = 2. The current Fec is subjected to the combined radial attraction of the current and next path segments, and its magnitude is proportional to the resultant force of the radial attraction from the end of the hand controller to both paths calculated separately. The current Fep is subjected to the combined axial attraction of the current and next path segments, and its magnitude is proportional to the resultant force of the axial attraction from both paths calculated separately. If the maximum vertical distance d from the end of the current hand controller to the two paths is greater than the radius r of the virtual fixture, then Feb adopts the direction of the resultant force of Fec and Fep, and its magnitude is proportional to its distance from the boundary (dr). Otherwise, Feb = 0. When the master controller moves to the corner of the cylinder formed by the two adjacent paths as axes (i.e., ... Figure 5 In the context of `num_zeros==0`, the current `Fec` is subject to the combined radial attraction of the current and next path segments, and its magnitude is proportional to the resultant force of the radial attraction from the end of the controller to both ends of the path when calculated individually. The current `Fep` is subject to the combined axial attraction of the current and next path segments, and its magnitude is proportional to the resultant force of the axial attraction from both ends of the path when calculated individually. If the distance `d` from the end of the current controller to the nearest endpoint of the path is greater than the radius `r` of the virtual fixture, then `Feb` adopts the direction from the end of the controller to the nearest endpoint of the path, and its magnitude is proportional to its distance from the boundary (`dr`). Otherwise, `Feb=0`. When the master controller moves to a path point or the path between it (i.e., ... Figure 5 In the YES branch where distance==0, the current Fep is the axial attraction pointing to the next path point along the current path segment direction, and its magnitude is proportional to the magnitude of the current path segment vector; Feb and Fec are both 0);

[0153] Specifically, the deviation force generated by the positional deviation of the robotic arm's end effector mapped to the main controller can be described by the following formula:

[0154] F deviation =k d *(x c -x m ,y c -y m ,z c -z m(9)

[0155] In the formula, k d The proportionality coefficient of the deviation force, (x) c ,y c ,z c ) is the location of the main controller, (x m ,y m ,z m () represents the position of the robotic arm's end effector mapped in the main controller;

[0156] Specifically, to improve system stability and handling, and to prevent oscillations or unstable movements that may occur due to inertia or external disturbances, this invention introduces a velocity damping force. The calculation of the damping force is as follows:

[0157] F damping =μv (10)

[0158] In the formula, v represents the moving speed, and μ is the damping coefficient;

[0159] S34. Real-time force feedback: The virtual clamp force field is mapped to the main controller workspace through the tactile interface, providing tactile guidance for the operator and assisting in completing precise operations.

[0160] This invention provides a human-machine hybrid multimodal teleoperation system for a space robotic arm, the architecture of which is as follows: Figure 6 As shown, it includes a visual interface module, a human intent estimation module, a path planning module, a virtual gripper module, a motion control module, and a robotic arm module;

[0161] Specifically, the vision interface module is responsible for both 3D depth information acquisition and processing. It acquires color images and spatial depth information of the working area of ​​the control terminal through a depth camera. The image information can be used to extract features of the operating area to assist the subsequent intent recognition process, while the depth information is used to generate a 3D point cloud and construct an environmental map within the workspace, providing environmental modeling support for the path planning module, thereby improving the overall environmental perception capability of the system.

[0162] Specifically, the human intention estimation module receives working area image data from the visual interface module, and at the same time, based on the gaze point position information collected by the eye tracker, it analyzes the gaze trajectory through a hidden Markov model to predict the operator's current operation intention. The system performs two-dimensional coordinate positioning of the predicted target area and transmits it as input to the path planning module, thereby realizing real-time perception and translation of the operator's intention.

[0163] Specifically, the path planning module, based on the target location information provided by the intent estimation module and combined with the current 3D map of the environment, calculates the optimal motion path within the operating space using a heuristic path search algorithm (such as the A* algorithm) and outputs it to the virtual fixture module. This path ensures both target reachability and path rationality, enhancing the system's motion efficiency.

[0164] Specifically, the virtual gripper module is used to enhance the operator's perception and control capabilities for remote operation. It constructs multiple auxiliary force feedback mechanisms such as guiding force field, deviation force field and damping force field, and transmits them to the main controller handle in conjunction with the tactile interface, so that the operator can obtain real-time force feedback information based on path guidance, which significantly improves the intuitiveness of the interaction and the accuracy of the operation.

[0165] Specifically, the motion control module acts as a bridge between the main controller input, path planning, and actual execution. It receives motion commands from the main controller after processing by the virtual fixture module, calculates the position and attitude of the end effector in real time, and drives control signals to the robotic arm module, enabling the tracking and execution of path and force feedback inputs from the main controller. This module also possesses force feedback capabilities, allowing for dynamic adjustment of the control strategy based on actual interaction conditions.

[0166] Specifically, the robotic arm module, as the end effector of the system, is responsible for completing precise operation tasks according to the target instructions issued by the motion control module. The robotic arm drives the motor through an internal controller and combines inverse kinematics calculations to achieve precise movement of the end effector in three-dimensional space, thereby completing multimodal interactive tasks with the target.

[0167] Example

[0168] The present invention has completed the deployment of a space robotic arm human-machine hybrid multimodal teleoperation system verification experimental platform, characterized in that it consists of two parts, a master control end and a slave control end, forming a complete teleoperation system;

[0169] Specifically, the main control unit mainly includes equipment such as eye trackers, main controllers, and computers, as well as the operator responsible for operating them;

[0170] Specifically, the slave device consists of a depth camera, a robotic arm, a target object, and a propulsion area, and is mainly responsible for the actual execution of the task;

[0171] Specifically, the experiment compared two teleoperation methods: a traditional teleoperation method and an intention-based human-machine hybrid multimodal teleoperation method.

[0172] Specifically, in a traditional teleoperation system, the operator directly controls the robotic arm through the main controller, causing the pusher plate at its end to push the target object to complete the movement task.

[0173] Specifically, in the space robotic arm human-machine hybrid multimodal teleoperation system verification experimental platform described in this invention, the operation process combines intent recognition and virtual gripper guidance: First, the operator needs to look at the surface of the target object. The system generates a virtual gripper based on the intent recognition result to guide the end of the robotic arm to approach the target object. Then, the operator looks at the target position, and the system adjusts the virtual gripper again based on the intent recognition result to guide the end of the robotic arm to move accurately to the target red box area, thereby completing the operation task.

[0174] In the embodiments of this invention, eight sets of experiments were conducted. In all eight sets of experiments, the push time value of the traditional teleoperation system was higher than that of the system proposed in this invention. The experimental results show that the eye-tracking intention control and virtual gripper human-machine hybrid strategy can effectively improve the operation efficiency.

[0175] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solutions of the present invention by utilizing the methods and techniques disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the protection scope of the technical solutions of the present invention.

Claims

1. A multimodal manipulation method for a spatial robotic arm based on intent estimation, characterized in that... include: Acquire multimodal information by using a depth camera to obtain color images and spatial depth information of the working area of ​​the slave terminal and transmit them to the master terminal for real-time display; The color image displayed to the operator on the master control terminal is collected by an eye-tracking device to capture the operator's eye movement data; the force feedback controller collects the position and posture information of the operator's fingers on the master control terminal in real time; the working area of ​​the slave control terminal is the area that includes the robotic arm and the target to be grasped. The gaze point data is preprocessed and dynamic probability modeled to identify the operator's target area intention, determine the two-dimensional gaze point coordinates, and then the two-dimensional gaze point coordinates are fused with spatial depth information and mapped to three-dimensional target coordinates in the robot arm base coordinate system. Based on the three-dimensional target coordinates, a safe path is planned using a three-dimensional voxel mesh map to generate the robotic arm's motion trajectory. A virtual gripper force field is constructed, which includes boundary constraints, path guidance forces, position deviation forces, and velocity damping forces. The virtual gripper force field guides the force feedback controller, thereby constraining the operator's actions. The force feedback controller generates motion commands based on the real-time collected information on the position and posture of the operating fingers; The position and attitude of the end effector are calculated and controlled in real time according to the motion command, and the control signal is sent to the robotic arm to control the robotic arm to move toward the target to be grasped according to the operator's intention.

2. The method according to claim 1, characterized in that, Preprocessing and dynamic probability modeling of the gaze point data to identify the operator's target region intent includes: Preprocess the eye-tracking data to remove outliers; The floating region of the gaze point is determined based on the 3σ range of the normal distribution, and the physical size of the floating region is mapped to the pixel range. The workspace of the robotic arm is rasterized into a discrete set of states according to pixel range; The preprocessed eye-tracking coordinates were used as the observation sequence. By initializing the initial attention probability π, the inter-region transition probability A, and the state generation observation probability B, a hidden Markov model of eye-tracking gaze behavior and target region was constructed. The Viterbi algorithm is used to decode the sequence of the highest probability target regions representing the operator's intention, and a dynamic update mechanism is used to correct the inter-region transition probability A in real time.

3. The method according to claim 2, characterized in that, The dynamic update mechanism described above dynamically adjusts the mean of the normal distribution based on the real-time gaze point position and dynamically adjusts the standard deviation σ based on the gaze movement speed. The relationship between standard deviation σ and eye movement speed v is as follows: The unit of movement speed is m / s.

4. The method according to claim 1, characterized in that, The operator's actions are constrained in the following ways: Based on the color images and depth information acquired by the depth camera, 3D point cloud data is constructed and a voxel mesh map is generated; The coordinates of the two-dimensional gaze point are fused with the depth information of the corresponding area, and then converted into the coordinates of the three-dimensional target point in the robot arm's base coordinate system through the coordinate transformation matrix of the depth camera and the robot arm. Complete the intrinsic parameter calibration and distortion calibration of the depth camera, as well as the hand-eye calibration of the robotic arm with the eye outside the hand, and establish a precise spatial mapping relationship between the vision system and the robotic arm; Path search is performed on the voxel grid map and the path is smoothed to generate a collision-free motion trajectory for the robotic arm. Based on the planned path, an admittance-type virtual fixture is generated, and the composite force field composed of boundary constraint force, path guiding force, position deviation force and velocity damping force is calculated. The virtual fixture's composite force field is mapped to the force feedback controller via a tactile interface, providing tactile guidance to the operator and assisting in completing precise operations.

5. The method according to claim 4, characterized in that, A heuristic path search algorithm is used to search the 3D voxel grid map, and the path is smoothed based on fifth-order polynomial interpolation.

6. The method according to claim 4, characterized in that, The composite force field calculation of the virtual fixture force field satisfies the following relationship: The boundary constraint force increases linearly with the distance of the force feedback controller's end off the path, preventing the operator from exceeding the safe or permissible operating range. Guiding forces include centripetal forces pointing towards the center of the path and axial forces pointing towards the path points; by providing force feedback along the task path, the operator is guided to complete specific operational tasks. The positional deviation force is proportional to the displacement difference between the force feedback controller and the end effector of the robotic arm; The speed damping force is related to the moving speed of the force feedback controller and the preset damping coefficient, and is used to suppress the operator's rapid or excessive movements.

7. A multimodal control system for a spatial robotic arm based on intent estimation, characterized in that... include: A multimodal information acquisition module, the module including an eye-tracking device, a force feedback controller, and a depth camera; The depth camera is used to acquire color images and spatial depth information of the working area of ​​the slave control terminal, and transmits them to the vision interface module for real-time display. When the operator gazes at the color image displayed by the vision interface module, the eye-tracking device collects the operator's eye movement data. The force feedback controller is installed on the operator's hand and collects the operator's manual position and speed data at the master control terminal. The vision interface module acquires color images and spatial depth information of the working area of ​​the slave device through a depth camera; The human intent estimation module preprocesses and dynamically probabilistically models the operator's eye movement data collected in real time by the eye-tracking device, identifies the operator's intent in the target area, determines the two-dimensional gaze point coordinates, and maps the two-dimensional gaze point coordinates with depth information to the three-dimensional target coordinates in the robotic arm's base coordinate system. The path planning module calculates the robotic arm's trajectory within the operating space based on the 3D target coordinates provided by the intent estimation module and the 3D voxel mesh map of the current environment, and outputs it to the virtual fixture module. The virtual clamp module transmits the virtual clamp force field, which includes boundary constraint force, path guiding force, position deviation force and velocity damping force, to the force feedback controller in conjunction with the tactile interface. The force feedback controller generates motion commands based on the real-time collected information on the position and posture of the operating fingers; The motion control module receives motion commands from the force feedback controller, calculates the position and attitude of the end effector in real time, and drives control signals to the robotic arm module. The robotic arm module is responsible for completing precise operation tasks according to the target instructions issued by the motion control module, realizing the precise movement of the end effector in three-dimensional space.

8. The system according to claim 7, characterized in that: The human intent estimation module processing steps include: Preprocessing and dynamic probability modeling of the gaze point data to identify the operator's target region intent includes: Preprocess the eye-tracking data to remove outliers; The floating region of the gaze point is determined based on the 3σ range of the normal distribution, and the physical size of the floating region is mapped to the pixel range. The workspace of the robotic arm is rasterized into a discrete set of states according to pixel range; The preprocessed eye-tracking coordinates were used as the observation sequence. By initializing the initial attention probability π, the inter-region transition probability A, and the state generation observation probability B, a hidden Markov model of eye-tracking gaze behavior and target region was constructed. The Viterbi algorithm is used to decode the sequence of the highest probability target regions representing the operator's intention, and a dynamic update mechanism is used to correct the inter-region transition probability A in real time.