Hand three-dimensional occupation sensing method for underwater man-machine cooperation
By constructing an optical scattering dataset and a dynamic sparse visual converter, a spatiotemporal graph convolutional network, and a sliding window Kalman filter, the problems of accuracy and real-time performance of three-dimensional reconstruction of hands in underwater human-machine collaboration are solved, high-precision collision avoidance safety boundary generation is achieved, and the safety and efficiency of underwater human-machine collaboration are improved.
Patent Information
- Application Number
- CN202510764346.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-01
- Publication Date
- 2025-10-17
AI Technical Summary
In underwater human-machine collaboration scenarios, existing technologies find it difficult to accurately capture hand semantic information in low light and turbid waters with suspended particles. Dynamic occlusion and turbulent disturbances lead to joint positioning deviations, and traditional methods find it difficult to achieve real-time, high-precision three-dimensional reconstruction and collision avoidance control under complex movements.
By constructing a dataset based on the optical scattering physics model, designing a dynamic sparse visual converter and a spatiotemporal graph convolutional network, and combining sliding window Kalman filtering and Bayesian optimization, robust estimation and real-time 3D reconstruction of hand joints are achieved, and a lightweight skinning algorithm is used to generate a collision avoidance safe space.
In complex underwater environments, it achieved millimeter-level reconstruction accuracy of hand posture and 30fps real-time tracking capability, generated centimeter-level precision collision avoidance safety boundaries, and improved the safety and operational efficiency of underwater human-computer interaction.
Smart Images

Figure CN120807629A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of underwater machine vision, and specifically relates to pose estimation of underwater gestures combined with a spatio-temporal graph convolution network and a visual transformer, and adoption of a lightweight visual transformer to enable the network to capture complex relationships between hand joints. BACKGROUND
[0002] In underwater human-robot collaboration scenarios, real-time reconstruction of the three-dimensional hand-occupancy space of a humanoid robot operator is a core challenge for achieving autonomous collision avoidance and intention interaction. Current monocular vision-based solutions face a double dilemma in turbid waters: low light and suspended particles cause a sharp drop in image signal-to-noise ratio, making it difficult for traditional feature extraction methods to capture effective hand semantic information; dynamic occlusion and turbulent disturbance-induced limb deformation disrupt the spatio-temporal continuity of the hand topology, and existing visual transformer models based on fixed attention weights are prone to joint positioning deviations. In addition, the stringent requirements for real-time performance and spatial occupancy accuracy in underwater scenarios make the cumulative error of the skeletal transformation matrix under complex motion prone to occur in traditional linear blend skinning algorithms.
[0003] To address the above problems, the present patent proposes a hand three-dimensional occupancy perception method for underwater human-robot collaboration. First, an underwater augmented dataset based on an optical scattering physical model is constructed, and 2D skeletal key point labeling is optimized through joint kinematic constraints to address the texture degradation problem caused by turbid media. At the network architecture level, a dynamic sparse visual transformer and a spatio-temporal graph convolution network are designed for collaborative inference: the dynamic sparse visual transformer focuses on the effective hand region through a deformable sparse attention window, enhancing the robustness to local occlusion; the spatio-temporal graph convolution network models the rigid motion constraints of finger joints through an adaptive adjacency matrix, suppressing abnormal topological connections caused by water flow disturbance. An uncertainty-aware sliding window Kalman filter is introduced innovatively, which triggers kinematic prior reinforcement when the inter-frame displacement is >0.5m / s, fuses the MANO parameter probability distribution output by the neural network with the robot's motion prior, and constructs a Bayesian optimization model in the time dimension, reducing the dynamic estimation error of the skinning parameters. Finally, through skeletal-driven deformation and a lightweight linear blend skinning engine, real-time three-dimensional occupancy space mapping of the underwater hand is achieved, providing a centimeter-precision dynamic collision avoidance safety boundary for underwater human-robot collaboration. SUMMARY
[0004] The purpose of the application is: for the problems of low precision of hand three-dimensional reconstruction, poor motion continuity and insufficient real-time performance caused by low light scattering, suspended particle interference, dynamic occlusion and water flow disturbance in underwater human-robot cooperation scene, a hand three-dimensional occupation real-time perception method for underwater human-robot cooperation is proposed. Through physical constraint optimization data enhancement, dynamic sparse attention and occlusion sensitive graph convolution multi-scale feature fusion, uncertainty guided iterative optimization and Bayesian time smoothing technology, the hand posture reconstruction accuracy of millimeter level and 30fps real-time tracking ability in complex underwater environment are realized, and the collision avoidance safety envelope space is generated, which aims to solve the core problems of intention recognition delay and collision avoidance control error in underwater human-robot interaction, and provides high reliability physical interaction basis for underwater cooperative robot. Specifically, the following steps are included:
[0005] Step 1, through the physical constraint optimization of 21 key points of hand, the semantic degradation caused by low light and suspended particles is actively compensated, the key point displacement constraint between adjacent frames is defined to detect the kinematic feasibility of hand key points, ν max = 0.5m / s is the maximum physiological speed of hand;
[0006] Step 2, design a dynamic sparse visual converter, combine the adversarial domain adaptation strategy, learn the probability distribution mapping of hand MANO parameters from low signal-to-noise ratio RGB images, and focus on the effective area through deformable sparse attention to resist dynamic occlusion interference;
[0007] Step 3, construct an occlusion-sensitive spatio-temporal graph convolution network, extract local joint features constrained by rigid motion, and perform multi-scale fusion with global semantic features of dynamic sparse visual converter to generate spatio-temporally consistent hand topology representation;
[0008] Step 4, based on the uncertainty distribution of predicted parameters, dynamically activate the lightweight refinement network module, and combine the human hand motion prior to reassign weights and iterative optimization in high error areas;
[0009] Step 5, use sliding window time optimization Kalman filter (when the inter-frame displacement > 0.5m / s, trigger kinematic prior reinforcement), fuse the probabilistic MANO parameters output by the neural network with the hand motion trajectory prior, and generate a robust skin transformation matrix in the Bayesian framework;
[0010] Step 6, the generated transformation matrix is deformed through the skeleton driving and the lightweight linear mixed skinning engine, and the hand three-dimensional occupation point cloud and collision avoidance safety space are generated in real time, realizing the occupation estimation of the diver's hand.
[0011] The beneficial effects of the present application are: by optimizing the labeling with physical constraints and dynamic sparse attention mechanism, the hand joint positioning error in turbid water area is reduced. The generalization ability of the model to complex underwater environment is significantly improved. The deformable sparse attention and the occlusion sensitive graph convolution cooperate to maintain the visibility of the joints under the condition of occlusion and fast water flow. The generated three-dimensional collision avoidance safety space coverage rate is improved, and the speed adaptive dynamic inflation layer is combined to improve the success rate of robot collision avoidance path planning. The present technology provides a centimeter-level dynamic hand occupancy perception capability for underwater human-robot collaboration, which significantly improves the safety and operation efficiency of underwater robot collaborative operation. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the example embodiments of the present application, the advantages of the above and / or additional aspects of the present application will become apparent and easy to understand in the description of the embodiments in conjunction with the following drawings. Obviously, the drawings introduced are only a part of the drawings of the embodiments described by the present application, and not all the drawings. Those skilled in the art can obtain other drawings from these drawings without creating creative labor.
[0013] Figure 1 is a flowchart of a hand three-dimensional occupancy real-time perception method for underwater human-robot collaboration according to an embodiment of the present application; DETAILED DESCRIPTION
[0014] In order to more clearly illustrate the technical solutions of the example embodiments of the present application, the advantages of the above and / or additional aspects of the present application will become apparent and easy to understand in the description of the embodiments in conjunction with the following drawings. Obviously, the drawings introduced are only a part of the drawings of the embodiments described by the present application, and not all the drawings. Those skilled in the art can obtain other drawings from these drawings without creating creative labor.
[0015] Figure 1 The hand three-dimensional occupancy real-time perception method for underwater human-robot collaboration provided by the embodiment of the present application is a flowchart. The present embodiment uses RGB data as input to perform real-time reconstruction of diver's hand, which is suitable for underwater human-robot interaction.
[0016] As Figure 1 shown, the method of the present embodiment specifically includes the following steps:
[0017] Step 1, through the network approach to obtain high-quality hand motion data, at the same time, use the camera to collect diver gesture video sequence under different water quality, light intensity, construct the original data set containing optical scattering characteristics. The improved biomechanical constraint enhanced OpenPose model is used to extract the initial 2D skeleton key points. A deformable convolution layer is added to the network of the model, and the convolution kernel offset is dynamically adjusted by the water quality parameters (turbidity, attenuation coefficient): Δpn=Φ(NTU,c)·W n, where Δpn is the offset of the nth convolution kernel, Φ is the water quality parameter mapping function, W n is the learnable weight matrix.
[0018] To address the texture degradation caused by low light scattering underwater, a physical optimization model based on rigid body kinematics is established. The biomechanical length ratio constraints between 21 key points of the hand are defined, and the bone length error term is constructed. where ε is the set of anatomically defined bone edges, is the theoretical bone length (such as the metacarpal length ratio threshold ±10%); the finger joint flexion angle is calculated by vector cross product, and the angle error term is defined as: is the physiological range of activity. Where k is the set of joints. Construct the hand key point adjacency matrix A∈R 21 ×21 , maintaining topological coherence through graph Laplace regularization: E topo =tr(P T LP), where P = [p1, ..., p 21 ] T is the key point coordinate matrix, L=DA is the graph Laplace matrix, D is the degree matrix. Comprehensive optimization objective function: min p (λ1E length +λ2E angle +λ3E topo ) is solved by the Levenberg-Marquardt algorithm, with weight coefficients λ1:λ2:λ a =3:2:1.
[0019] To address the sudden noise caused by suspended particle occlusion, a sliding window timing optimizer is designed. Kinematic feasibility testing of hand key points is performed by defining key point displacement constraints between adjacent frames: Among them, v max =0.5m / s is the maximum physiological motion speed, and Δt is the frame interval time.
[0020] Step 2: Input low signal-to-noise ratio RGB image I∈R H×W×3 , dynamic region segmentation is performed through turbidity-aware block encoder. Based on the local signal-to-noise ratio (SNR) and turbidity parameter β, the image is divided into a set of adaptively sized blocks. The block size satisfies: Size(P i ·(1+α·SNR(P i ) -1 )), where α is a learnable parameter. Larger blocks are used in high turbidity areas to suppress noise interference. For each block P i , by linear projection W e ∈R d×3 Generate initial embedding e i ∈Rd And introduce gradient reversal layer driven adversarial training to force the encoder to ignore the water domain difference.
[0021] We design deformable sparse attention to focus on the effective hand region and resist occlusion interference by dynamically querying offsets and calculating sparse correlations. For each query position q i ∈R d , we predict k position offsets Δp ij ∈R 2 : Δp ij =MLP offset (q i )·σ(β·NTU), where σ(·) is the Sigmoid function, and β is the softening coefficient, limiting the offset amplitude when the turbidity is high, to avoid excessive deviation from the target region. At the offset position {p i + Δp ij}, we extract the key-value pair (K j , V j ), and filter the top-k high-response regions to calculate the attention weight: where k is dynamically adjusted according to the local occlusion rate (for every 10% increase in occlusion rate, k decreases by 2).
[0022] Based on the Gaussian parameterized output, we model the joint distribution of hand pose θ ∈ R4 8 , shape β ∈ R 10 , and translation t ∈ R 3 , and output the probability distribution of MANO parameters (hand pose parameters, shape parameters, and global rotation and translation parameters) through the probabilistic decoder.
[0023] Step 3: Each hand joint is considered as a node in the graph, denoted as V = {v1, v2, …, v N}, where N is the number of joints. For example, the wrist joint, the joints of each finger, etc. are all independent nodes. The edges between nodes are defined according to the anatomical relationship and relative position between hand joints, denoted as For example, the finger joints are connected in order, and the finger joints are also connected to the wrist joint.
[0024] The space-time graph G = (V, E) contains 21 hand joint nodes, where the nodes. For the non-rigid motion of underwater gestures affected by water flow, we design a rigid motion invariant convolution kernel: where R ij ∈ SO(3) is the local rotation matrix from node i to j, T ij ∈ R 3 is the translation vector, and W rigid is the rotation equivariant learnable weight, Normalization factor.
[0025] The global features extracted by the visual transformer are hierarchically fused with the local features extracted by the spatio-temporal graph convolutional network.
[0026] Step 4, Estimate the uncertainty of the prediction by multiple forward propagation. Perform kinematic feasibility check (such as joint angle limit, finger bending direction) on the predicted poses generated by Monte Carlo sampling, eliminate the results that do not conform to anatomy, and correct the uncertainty distribution of the prediction. Finally, the uncertainty is quantified by the variance of these results: where M is the number of samples (usually set to 50), v valid is the prediction subset verified by hand inverse kinematics (IK), is the prediction result of the m-th forward propagation, and Var(y) is the variance of the prediction result, which is used to measure the uncertainty. A lightweight residual network RefineNet is designed, whose activation weight is controlled by an uncertainty threshold: Only high-uncertainty parameters (ω i > 0) are refined: Δθ i = ω i · RefineNet(f i ), where f i is the spatio-temporal feature of the corresponding joint, and σ i 2 is calculated by the effective sample set V valid .
[0027] The traditional loss function uses the same weight for all regions, which cannot effectively handle high-error regions. This method introduces an uncertainty weight to dynamically adjust the loss function: where, is the loss of the i-th sample, and ω i is the uncertainty-based weight. By adaptive loss weighting, the model will focus on high-error regions (i.e., regions with high uncertainty), thereby improving the overall prediction accuracy.
[0028] Step 5, Model the temporal MANO parameters as a hidden Markov process and construct the state space equation. Define the state vector as the joint representation of the hand pose, shape, and translation parameters: x t = [θ t , β t , t t ] T ∈ R 61 . Based on the continuity assumption of hand motion, a kinematic prior model is designed: χ t = Fχ t-1 + Bu t + ω twhere F encodes the inertial properties of joint motion, u is the control input t provided by inertial sensor data or motion trajectory prediction network, process noise ω t Uncertainty of reaction water flow disturbance. Add velocity constraint judgment in state transition equation: where Δt is the inter-frame time, if v t > 0.5 m / s, trigger kinematics prior compensation.
[0029] Bayesian optimization using sliding window Kalman filter. Within the window, the probabilistic MANO parameters output by the dynamic sparse visual converter as observations, construct the observation equation: where the observation matrix H is the identity matrix, and the noise covariance R t =∑ t Quantify the confidence of neural network prediction.
[0030] Convert the optimized state estimation to skeletal driving parameters. For each bone b, calculate its local rotation matrix and translation vector where J b (β) is the joint position function defined by the MANO model. Calculate the global transformation matrix step by step through the skeletal hierarchy:
[0031]
[0032] Step 6, load the MANO parameterized model, parse the skeletal hierarchy and skinning weights. Extract the parent-child relationship parent(b) and joint initial position from the predefined 21-skeletal topology: Load the skinning weight matrix where W i,b represents the weight of vertex i affected by bone b, satisfying∑ b W i,b =1.
[0033] Apply the optimized skinning transformation matrix to the vertex initial position to generate real-time occupancy point cloud: Based on real-time point cloud, generate the safety space required for robot collision avoidance. For each point construct a safety sphere with radius r i Adaptively adjust with hand motion speed, use QuickHull algorithm to generate point cloud convex hull, and calculate its axial bounding box (intersection with directional bounding box, as the final collision avoidance safety space.
[0034] Using the rendering technique by instantiation, the safety envelope space is superimposed in real time to the robot operation interface as an array of semi-transparent red spheres.
[0035] The technical scheme of the embodiment designs a hand three-dimensional occupation sensing method for underwater human-robot collaboration, designs a dynamic sparse visual converter, utilizes a deformable sparse attention focusing effective area, and aligns clear water and turbid water area feature distribution through an adversarial domain adaptation strategy; secondly, a MANO parameter probability distribution is generated based on Monte Carlo sampling, high error areas are iteratively optimized in combination with inverse kinematics verification and biomechanical energy function, adaptive fine adjustment is realized by using a variance weighted loss function; a sliding window Kalman filter is used to fuse neural network prediction and motion trajectory prior in a Bayesian framework, generate a spatio-temporal smooth skin transformation matrix, and suppress trajectory jitter caused by water flow disturbance; a lightweight linear mixed skin engine is used to drive the hand vertex deformation result of the bone transformation matrix, real-time rendering is performed on the down-sampled point cloud, and a speed-adaptive dynamic inflation layer is combined to construct a collision avoidance safety envelope space. The scheme realizes the deep integration of physical driving, probability modeling and lightweight calculation, solves the hand pose estimation problem in the underwater complex environment, and provides the collaborative robot with a dynamic collision avoidance sensing capability of centimeter-level precision and millisecond-level response.
[0036] It should be noted that the above only describes the preferred embodiments of the present application and the applied technical principles. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and those skilled in the art can make various obvious changes, readjustments and substitutions without departing from the scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.
Claims
1. A three-dimensional hand occupancy perception method for underwater human-machine collaboration. It is characterized by: Step 1: Optimize the annotation of 21 skeletal key points of the hand through physical constraints, actively compensate for the semantic degradation caused by low light and suspended particles, define the displacement constraints of key points between adjacent frames to perform kinematic feasibility detection on the key points of the hand, v max =0.5m / s is the maximum physiological movement speed of the hand; Step 2: Design a dynamic sparse visual converter and combine it with an adversarial domain adaptation strategy to learn the probability distribution mapping of the hand MANO parameters from low signal-to-noise ratio RGB images. Use deformable sparse attention to focus on the effective area to resist dynamic occlusion interference. Step 3: Build an occlusion-sensitive spatiotemporal graph convolutional network to extract local joint features constrained by rigid body motion, and perform multi-scale fusion with the global semantic features of the dynamic sparse visual transformer to generate a spatiotemporally consistent hand topology representation; Step 4: Based on the uncertainty distribution of the prediction parameters, the lightweight refinement network module is dynamically activated, and the weights of the high-error areas are redistributed and iteratively optimized in combination with the human hand motion prior. Step 5: Using a sliding window timing-optimized Kalman filter (kinematic prior reinforcement is triggered when the inter-frame displacement is greater than 0.5 m / s), the probabilistic MANO parameters output by the neural network are fused with the hand motion trajectory prior to generate a robust skinning transformation matrix in a Bayesian framework. Step 6: The generated transformation matrix is applied to the skeleton-driven deformation and lightweight linear blend skinning engine to generate a 3D hand occupancy point cloud and collision avoidance safety space in real time, thus achieving accurate estimation of the diver's hand occupancy.
2. The method for three-dimensional hand occupancy perception for underwater human-machine collaboration according to claim 1, characterized in that: Step 1 is achieved by: High-quality hand motion data was acquired through a network. A camera was used to capture video sequences of divers' hand gestures under varying water quality and light intensity, constructing a raw dataset that incorporates optical scattering characteristics. Initial 2D skeletal keypoints were extracted using an improved OpenPose model enhanced with biomechanical constraints. To address texture degradation caused by low light scattering underwater, a physical optimization model based on rigid body kinematics was established. Biomechanical length ratio constraints were defined between 21 key points of the hand to construct a bone length error term. The knuckle flexion angle was calculated through vector cross product, and the angle error term was defined: is the range of physiological activities. Where k is the set of joints. A sliding window timing optimizer is designed to address the sudden noise caused by suspended particle occlusion.
3. The method for three-dimensional hand occupancy perception for underwater human-machine collaboration according to claim 1, characterized in that: Step 2 is achieved by: Input low signal-to-noise ratio RGB image I∈R H×W×3 , dynamic region segmentation is performed through turbidity-aware block encoder. Based on the local signal-to-noise ratio (SNR) and turbidity parameter β, the image is divided into a set of adaptively sized blocks. The block size satisfies: Size(P i ·(1+α·SNR(P i ) -1 ), where α is a learnable parameter and larger blocks are used in high turbidity areas to suppress noise interference. A deformable sparse attention algorithm is designed to focus on the effective hand area and resist occlusion interference through dynamic query offset and sparse correlation calculation. Based on the Gaussian parameterized output, the probability distribution of MANO parameters (hand posture parameters, shape parameters, and global rotation and translation parameters) is output through a probabilistic decoder.
4. The method for three-dimensional hand occupancy perception for underwater human-machine collaboration according to claim 1, characterized in that: Step 3 is achieved by: Each hand joint is regarded as a node in the graph, denoted as V = {v1, V2, ..., v N }, where N is the number of joints. For example, the wrist joints, the knuckles of each finger, etc. are all treated as independent nodes. The edges between the nodes are defined according to the anatomical relationship and relative position between the hand joints, which are expressed as For example, the finger joints are connected in sequence, and the finger joints are also connected to the wrist joints. A spatiotemporal graph convolutional network is constructed to extract local motion features of hand joints. The spatiotemporal graph G = (V, E) contains 21 hand joint points and the time window T = 5 frames. Aiming at the non-rigid motion of underwater gestures caused by water flow, a rigid motion invariant convolution kernel is designed: Among them, R ij ∈SO(3) is the local rotation matrix from node i to j, T ij ∈R 3 is the translation vector, W rigid are rotationally equivariant learnable weights, is the normalization factor. The global features extracted by the visual transformer are hierarchically fused with the local features extracted by the spatiotemporal graph convolutional network.
5. The method for three-dimensional hand occupancy perception for underwater human-machine collaboration according to claim 1, characterized in that: Step 4 is achieved by: The uncertainty of the prediction is estimated through multiple forward propagations. The predicted poses generated by Monte Carlo sampling are then subjected to kinematic feasibility checks (e.g., joint angle limits and finger bending directions), and anatomically inconsistent results are eliminated to correct the uncertainty distribution of the prediction. Through uncertainty estimation and dynamic network activation weights. The traditional loss function uses the same weight for all regions and cannot effectively handle high error regions. This method dynamically adjusts the loss function by introducing uncertainty weights: in, is the loss of the i-th sample, ω i is the uncertainty-based weight. By adaptively weighting the loss, the model focuses on areas with high error (i.e., areas with greater uncertainty), thereby improving overall prediction accuracy.
6. The method for three-dimensional hand occupancy perception for underwater human-machine collaboration according to claim 1, characterized in that: Step 5 is achieved by the following steps: The time-series MANO parameters are modeled as a hidden Markov process, and the state space equation is constructed. The state vector is defined as the joint expression of hand posture, shape and translation parameters: t =[θ t , β t , t t ] T ∈R 61 Add velocity constraint judgment to the state transfer equation. If the inter-frame displacement is greater than 0.5m / s, kinematic prior compensation is triggered. A sliding window Kalman filter is used for Bayesian optimization. Within the window [tk, t], the probabilistic MANO parameters output by the dynamic sparse visual converter are used as observations to construct the observation equation. The optimized state estimate Convert to bone-driven parameters and calculate the global transformation matrix level by level through the bone hierarchy.
7. The method for three-dimensional hand occupancy perception for underwater human-machine collaboration according to claim 1, characterized in that: Step 6 is achieved by: Loads a MANO parameterized model and analyzes the bone hierarchy and skin weights. Apply the optimized skin transformation matrix to the initial vertex positions to generate a real-time occupancy point cloud. The safe space required for the robot to avoid collisions is generated based on real-time point clouds. Through instanced rendering technology, the safe envelope space is superimposed on the robot operation interface in real time as an array of translucent red spheres.