Multi-mode dexterous hand integrating visual touch sense and palm eye and application of multi-mode dexterous hand

By integrating a multimodal dexterous hand that combines vision, touch, and palm-eyes, combined with multispectral depth photometric stereo 3D reconstruction and dynamic spatial propagation network, the multimodal perception problem of existing robotic arms in high-precision scenarios is solved, and high-precision and robust perception effects are achieved.

CN120697060AActive Publication Date: 2025-09-26HUAZHONG UNIV OF SCI & TECH

Patent Information

Application Number
CN202510843035.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-26
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Existing rigid manipulators have difficulty meeting the multimodal perception requirements in high-precision, high-dynamic scenarios. Tactile perception accuracy is insufficient, vision fails in occluded environments, ToF camera depth maps are highly sparse, and traditional methods have large errors at low input density.

Method used

A multimodal dexterous hand that integrates vision, touch, and palm-eye is used. It combines vision-tactile sensors and palm-eye modules, and realizes multimodal perception through RGB cameras, NIR cameras, TOF cameras, and processors. It combines multispectral depth photometric stereo 3D reconstruction and dynamic spatial propagation networks to improve perception accuracy and robustness.

Benefits of technology

It achieves high-precision and high-robust multimodal perception, reduces the normal vector estimation error, and significantly reduces the root mean square error and mean relative error of depth estimation, making it suitable for scenarios such as precision assembly and minimally invasive surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120697060A_ABST
    Figure CN120697060A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of robots, and particularly discloses a multi-mode dexterous hand fusing visual tactile sense and palmtop eye and application thereof.The multi-mode dexterous hand comprises a hand body, a visual tactile sensor, a palmtop eye module and a processor, the visual tactile sensor comprises a skeleton, an elastic body, RGB and an NIR camera, the skeleton is of a finger-shaped structure, and the skeleton and the hand body form a hand structure; the elastomer covers the outer side of the framework, and a diffuse reflection coating covers the surface of the elastomer; the RGB camera and the NIR camera are installed in the framework and used for obtaining an RGB image and an NIR image when the first object makes contact with the elastic body; the handheld eye module is installed in the middle of the hand body, comprises a second RGB camera and a TOF camera and is used for taking a second object RGB image and sparse depth; and the processor determines the first object surface three-dimensional point cloud and the second object depth map to realize pose alignment of the two. According to the invention, touch-vision-depth multi-mode perception can be realized, and the precision and robustness of the manipulator are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of robotics, and more particularly, relates to a multimodal dexterous hand integrating vision, touch, and palm-eye, and its application. Background Art

[0002] With the growing demand for dexterous robotic manipulation in fields such as industrial automation and medical surgery, existing rigid manipulators are unable to meet the multimodal, high-precision perception requirements in high-precision, high-dynamic scenarios due to insufficient tactile perception accuracy and low depth estimation resolution.

[0003] Existing robotic arms often rely on a single modality, either vision or touch, making it difficult to balance global environmental understanding with local contact characteristics: vision fails in occluded environments, while touch cannot perceive untouched areas. Furthermore, the depth maps output by existing ToF cameras are highly sparse, requiring post-processing algorithms to complete them. However, traditional methods (such as Kalman filtering and SGM (Semi-Global Matching)) suffer from a sharp increase in root mean square error (RMS) at low input density. Summary of the Invention

[0004] In response to the above-mentioned defects or improvement needs of the prior art, the present invention provides a multimodal dexterous hand that integrates vision, touch and palm-eye and its application, with the aim of realizing multimodal perception of touch, vision and depth and improving the working accuracy and robustness of the manipulator.

[0005] To achieve the above objectives, according to one aspect of the present invention, a multimodal dexterous hand that integrates vision, touch, and eye-in-the-palm is proposed, comprising a hand body, several vision-tactile sensors, an eye-in-the-palm module, and a processor, wherein:

[0006] The visual-tactile sensor includes a skeleton, an elastic body, an RGB camera, and a NIR camera. The skeleton is a finger-shaped structure mounted on the hand body. The skeletons of each visual-tactile sensor and the hand body together form a hand structure. The elastic body covers the outside of the skeleton, and the surface of the elastic body is covered with a diffuse reflection coating. The RGB camera and the NIR camera are mounted within the skeleton and are used to respectively obtain an RGB image and an NIR image when a first object contacts the elastic body.

[0007] The palm eye module is installed in the middle of the hand body, and the palm eye module includes a second RGB camera and a TOF camera, and the second RGB camera and the ToF camera are used to obtain an RGB image and a sparse depth of a second object respectively;

[0008] The processor is used to reconstruct a three-dimensional point cloud of the surface of the first object based on the RGB image and NIR image of the first object when it is in contact with the elastic body, and to obtain a depth map of the second object based on the RGB image and sparse depth of the second object; and to align the postures of the first object and the second object based on the three-dimensional point cloud of the surface of the first object and the depth map of the second object.

[0009] According to another aspect of the present invention, a multimodal dexterous hand application integrating visual-tactile and palm-eye is provided, comprising the following steps:

[0010] When the dexterous hand grasps the first object, it uses the RGB camera and the NIR camera to capture RGB and NIR images of the first object when it contacts the sensor. RGB and NIR pixel values ​​and pixel coordinates are obtained based on the RGB and NIR images. A predicted normal vector for the first object is obtained based on the RGB and NIR pixel values ​​and pixel coordinates. A depth field is obtained based on the predicted normal vector and the image boundary depth, thereby reconstructing a three-dimensional point cloud of the first object's surface.

[0011] The second RGB camera and the ToF camera respectively obtain an RGB image and a sparse depth of the second object, and the sparse depth is completed by the RGB image to obtain a depth map of the second object;

[0012] Position alignment of the first object and the second object is achieved according to the three-dimensional point cloud of the first object surface and the depth map of the second object.

[0013] As a further preferred method, the first object predicted normal vector is obtained based on the RGB, NIR pixel values ​​and pixel coordinates, specifically:

[0014] Inputting RGB, NIR pixel values ​​and pixel coordinates into the normal vector prediction model to obtain the first object predicted normal vector;

[0015] The method for obtaining the normal vector prediction model includes:

[0016] A calibration object with characteristic protrusions is placed on the sensor surface, and the actual image is acquired using an RGB camera or NIR camera. Based on the calibration object's CAD model, an image of the contact surface between the calibration object and the sensor is rendered to obtain a calibration object rendered image. The corresponding camera parameters are obtained when the characteristic patterns of the actual image and the calibration object rendered image overlap. Based on the camera parameters, the scene of the undeformed elastic body is rendered to obtain the background depth and background normal vector, and the image boundary depth is determined.

[0017] A probe with a sphere at its end is placed in contact with the sensor at different locations. RGB and NIR images of different surface areas are captured using RGB and NIR cameras, i.e., contact images. The RGB and NIR images are used as calibration images. Based on the sensor CAD model, the contact surface image of the sphere and the sensor is rendered to obtain a sphere-rendered image. The scene depth is obtained when the calibration image and the sphere-rendered image completely overlap.

[0018] The mask area is obtained by subtracting the scene depth from the background depth, and the RGB, NIR pixel values ​​and pixel coordinates of the mask area in the contact image are extracted; the surface normal vector is calculated based on the depth of the mask area;

[0019] The neural network model is trained with background depth, background normal vector, RGB, NIR pixel values ​​and pixel coordinates as input and surface normal vector as output, and the trained neural network model is used as the normal vector prediction model.

[0020] As a further preferred method, a depth field is obtained based on the predicted normal vector and the image boundary depth, including:

[0021] The 3D reconstruction problem of the object is described as the Poisson equation. The Poisson equation is discretized using central difference to obtain the depth constraint of each pixel of the contact image. The depth constraints of all pixels are combined into a discrete linear system.

[0022] The image boundary depth is used as the depth prior z prior and transform the depth prior z into prior As an additional constraint, the discrete linear system is expanded into an augmented sparse linear system;

[0023] The augmented sparse linear system is solved using the least squares method to obtain the depth field.

[0024] As a further preferred embodiment, the augmented sparse linear system is specifically expressed as:

[0025]

[0026] Where A is the sparse matrix of the encoded depth coefficients, z is the vector containing the depth of all pixels; b is the divergence term of the Poisson equation, which is obtained by calculating the depth gradient by predicting the normal vector; I prior is a diagonal matrix.

[0027] As a further preferred embodiment, the sparse depth is supplemented by the RGB image to obtain a second object depth map, specifically: the RGB image and the sparse depth of the second object are aligned and then input into the depth correction model to obtain the second object depth map;

[0028] The depth correction model includes ResNet34-Unet and dynamic spatial propagation network, where:

[0029] ResNet34-Unet obtains the initial depth map V based on the aligned RGB image and sparse depth 0 , unweighted affinity matrix Initial similarity matrix Attention weight and confidence mask C;

[0030] The dynamic spatial propagation network consists of multiple submodules based on the initial depth map V 0 , each submodule updates the depth map in turn, and the final corrected depth map is obtained after all submodules complete the update;

[0031] Where, for the t-th iteration, t = 1, 2…N, N is the number of submodules:

[0032] Based on attention weight and the initial similarity matrix Calculate the dynamic similarity matrix W t ;

[0033] By the dynamic similarity matrix W t and Generate affinity matrix A t :

[0034] Based on the affinity matrix A t And confidence mask C for the current depth map V t-1 Update and get the updated depth map V t .

[0035] As a further preferred embodiment, the current depth map V t-1 The updated calculation formula is:

[0036] V t =(D -1 A t )V t-1 (1-C)+V sparse ·C

[0037] Among them, (D -1 A t )V t-1 is the depth propagation term, which is activated in the area with lower confidence mask value, and D is the affinity matrix A t The rows and matrices of V sparse =V 0 It is a sparse depth constraint term and is activated in areas with higher confidence mask values.

[0038] As a further preferred embodiment, the depth map is not updated in the edge area of ​​the object, that is, the edge area V t=V t-1 .

[0039] As a further preference, the attention weight is determined as follows:

[0040] For the current depth map V at the tth iteration t-1 The pixel at position (i, j) in Represents the attention weight of the pixel to its neighboring pixels with a distance of k. As k increases, the attention weight Gradually decrease.

[0041] As a further preference, the unweighted affinity matrix Neighborhood-based To update:

[0042]

[0043] in, express The value of position (i, j) in the neighborhood The method of determining is:

[0044]

[0045] Among them, (p,q) is the neighborhood offset, is the offset field predicted by the deformable convolutional network, and is the parameter of the deformable convolutional network.

[0046] In general, the above technical solutions conceived by the present invention have the following technical advantages compared with the existing technology:

[0047] This invention replaces rigid fingertips with finger-shaped visual tactile sensors and embeds a palm-mounted eye module equipped with an RGB camera and a Time of Flight (ToF) camera in the palm. This collaborative approach overcomes the limitations of single-modal information and enables multimodal perception through touch, vision, and depth. This invention fills a technological gap in multimodal perception for robotic arms, providing a highly robust solution for applications requiring precise perception, such as precision assembly and minimally invasive surgery.

[0048] 2. In order to address the problems of difficulty in obtaining the true value of normal vectors, limited normal vector reconstruction accuracy, and cumulative errors in normal vector integrals in the three-dimensional reconstruction of complex surface visual-tactile sensing, the present invention proposes a multispectral depth photometric stereo three-dimensional reconstruction method based on contact surface rendering and boundary priors. The true value of the surface normal is obtained by rendering, the constraints on the solution of photometric stereo normal vectors are added by introducing NIR spectra, and the cumulative errors of normal vector integrals are reduced by introducing boundary priors, ultimately achieving high-precision and high-robust surface visual-tactile normal estimation and depth reconstruction.

[0049] 3. This paper proposes a dynamic spatial propagation network based on adaptive weights, dynamic paths, and diffusion suppression strategy, which integrates dense RGB color information with sparse depth information to achieve high-resolution depth estimation in the operating area. The details are as follows:

[0050] 1) By weighting the static affinity matrix using the similarity matrix, dynamic decoupling of the propagation process is achieved, breaking through the static affinity limitations of traditional spatial propagation networks;

[0051] 2) Propose an adaptive weighting strategy to decouple neighborhoods by distance, assigning higher weights to close neighbors while attenuating the weights of distant neighbors with distance, thus enhancing the depth accuracy of texture and edge details;

[0052] 3) Propose a diffusion suppression operation to stop depth updating in edge areas, avoid cross-edge diffusion, and preserve the clarity of object contours;

[0053] 4) A dynamic path strategy is proposed, which predicts the neighborhood offset through deformable convolution and dynamically generates the propagation path, allowing the neighborhood to break through the fixed grid, expanding the path to accelerate filling in flat areas, and shrinking it to self-loop triggering in edge areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 Schematic diagram of the overall structure of the multimodal dexterous hand according to an embodiment of the present invention;

[0055] Figure 2 This is a schematic structural diagram of a palm-eye module according to an embodiment of the present invention;

[0056] Figure 3 This is a schematic diagram of the structure of a visual-tactile sensor according to an embodiment of the present invention;

[0057] Figure 4 Schematic diagram of a multispectral depth photometric stereo 3D reconstruction method according to an embodiment of the present invention;

[0058] Figure 5 This is a flow chart of visual-tactile sensor calibration according to an embodiment of the present invention;

[0059] Figure 6 This is a diagram showing the experimental results of a three-dimensional reconstruction according to an embodiment of the present invention;

[0060] Figure 7 This is the overall framework diagram of the depth correction model of an embodiment of the present invention. DETAILED DESCRIPTION

[0061] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0062] The embodiment of the present invention provides a multimodal dexterous hand that integrates vision, touch and palm-eye, such as Figure 1 and Figure 2 As shown, it includes a hand body, several visual and tactile sensors and a palm eye module, wherein:

[0063] The visual tactile sensor surface has a complex curved structure similar to the shape of a human finger, which can produce continuous deformation when in contact with an object. Specifically, it includes a skeleton, an elastic body, an RGB camera, and a NIR camera. The skeleton is a finger-shaped structure installed on the hand body. The skeleton of each visual tactile sensor and the hand body together form the hand structure; the elastic body covers the outside of the skeleton and is coated with a diffuse reflection coating. The RGB camera and NIR camera are installed in the skeleton and are used to obtain RGB images and NIR images respectively when the first object contacts the elastic body.

[0064] The palm-eye module is mounted in the middle of the hand body and includes a second RGB camera and a time-of-flight camera for acquiring an RGB image and sparse depth of a second object, respectively. Furthermore, the module also includes a time-of-flight laser and a working status indicator.

[0065] Furthermore, it also includes RGB three-color light source, NIR light source and dichroic prism, such as Figure 3 As shown; the RGB three-color light source and NIR light source are specifically 3 RGB LEDs and 1 NIR LED, which are evenly distributed on the inner surface of the skeleton. The resistance value of the series resistor is changed to make the LED have appropriate brightness and achieve uniform lighting. The spectroscopic prism realizes the synchronous perception of the two cameras. It is used to split the incident light into two identical and mutually perpendicular beams, which enter the RGB camera and NIR camera respectively to achieve multispectral imaging. The RGB camera and NIR camera are equipped with filters of different bands, which filter out visible light and near-infrared light in a specific wavelength range, so that the camera can independently perceive visible light and near-infrared light and capture the two types of light respectively; in order to establish a controlled lighting environment, the camera's automatic white balance and automatic exposure correction functions are turned off. In addition, a transparent acrylic is provided on the skeleton to provide a window for the camera. The camera's field of view can be adjusted by changing the size of the acrylic.

[0066] Furthermore, the skeleton is prepared by 3D printing using photosensitive resin, and its shape is coordinated with the finger-shaped surface, which is used for forming and positioning the elastomer and serves as a container for other components.

[0067] Furthermore, the silicone elastomer is made by mixing Smooth-On's Solaris curing silicone and Slacker softener. The higher the proportion of softener, the softer the elastomer. When making finger-shaped elastomers, the mixture is first placed in a vacuum pump to remove bubbles to avoid affecting imaging; then the mixture is poured into a mold with a finger-shaped inner surface and waited for solidification. In order to avoid the appearance of a parting line on the surface of the elastomer that affects imaging, an upper and lower parting method is adopted. To improve the smoothness of the elastomer surface, the mold is made of acrylic machine and polished. The diffuse reflective coating is made by mixing a silicone mixture, silver powder and diluent, and is sprayed onto the surface of the elastomer using a spray gun. The silicone mixture ensures the adhesion of the coating to the elastomer, the silver powder gives the coating diffuse reflective properties, and the diluent prevents the coating from solidifying too quickly.

[0068] An embodiment of the present invention provides an application of a multimodal dexterous hand that integrates vision, touch, and eye-in-the-hand, including multispectral depth photometric stereo 3D reconstruction based on vision-tactile sensors, depth map acquisition based on an eye-in-the-hand module, and the combined use of the two. Specifically, the dexterous hand grasps a first object, and uses the vision-tactile sensor to acquire RGB and NIR images of the first object in contact with an elastic body to create a 3D point cloud of the first object's surface. The eye-in-the-hand module then acquires an RGB image and sparse depth map of a distant second object, thereby obtaining a depth map of the second object. The first and second objects can be aligned based on the 3D point cloud of the first object's surface and the depth map of the second object.

[0069] (1) Multi-spectral depth photometric stereo 3D reconstruction based on visual-tactile sensors aims to obtain the true value of the surface normal through rendering, and to increase the constraints of the photometric stereo normal vector solution by introducing the NIR spectrum, that is, to solve the surface normal vector of the object by the light intensity reflected from the surface of the object under the illumination of multiple light sources. The three RGB light sources provide three normal vector constraint equations, and the introduction of the NIR light source provides the fourth equation to form an overdetermined equation group, thereby reducing the influence of random errors; by introducing the boundary prior to reduce the normal vector integral cumulative error, finally achieve high-precision and high-robust surface visual-tactile normal vector estimation and depth reconstruction. Figure 4 As shown, the details are as follows:

[0070] (1.1) The affine transformation parameters of the RGB camera and NIR camera are obtained through a special structure calibration object to achieve alignment of the acquired RGB and NIR images.

[0071] A calibration object with raised features that conform to the surface of the elastomer is placed on the sensor surface. Pressing the object to ensure full contact with the elastomer, the RGB and NIR images are simultaneously captured. Feature point matching is used to determine the affine transformation parameters of the RGB image relative to the NIR image, including scaling, rotation, and translation.

[0072] (1.2) Calibrate the camera's internal and external parameters by rendering, such as Figure 5 As shown, the background depth, background normal vector and image boundary depth are obtained.

[0073] Place a specially structured calibration object on the sensor surface and acquire the actual image using an RGB camera or NIR camera (since the images acquired by the two cameras are aligned, either image can be used).

[0074] Based on the calibration object CAD model, an image of the contact surface between the calibration object and the sensor is rendered to obtain a calibration object rendered image. In this embodiment, the OpenGL-based 3D rendering engine Pyrender is used to create a rendering scene, the trimesh mesh is converted to a Pyrender mesh and added to the rendering scene, and a renderer with a size of 640×480 pixels (the same as the captured image) is created.

[0075] Initialize the internal and external parameters of the camera, where the external parameters include the rotation matrix R and the translation vector T, and the internal parameters include the focal length of the image in the X-axis and Y-axis directions (f x ,f y ), principal point coordinates (c x ,c y ); by shooting multi-angle images of known geometric patterns (such as chessboards) in advance, the Zhang Zhengyou calibration method is used to calculate the intrinsic parameters of the two cameras; the camera extrinsic parameters are used to calculate the extrinsic parameter matrix, and a camera node is created in the rendering scene. In this embodiment, the camera extrinsic parameters are initialized, the three-dimensional coordinates are initialized to [0,0,10], and the pitch / yaw / roll angles are initialized to 0. The initial extrinsic parameter matrix of the camera is calculated based on the initialization parameters. The specific steps are: converting the angle to radians, constructing the X, Y, and Z axis rotation matrices, combining the rotation matrix R in the order of Z, Y, and X, adding the translation component T to construct a complete extrinsic parameter matrix, and converting the coordinate system from OpenCV to OpenGL. The RGB camera intrinsic parameters (f x1 ,f y1 ,c x1 ,c y With the NIR camera internal reference (f x2 ,f y2 ,c x2 ,c y2 ) is unified as (f x ',f y ',c x ',c y '), by the camera internal parameter (f x ',f y ',c x ',c y '), create a camera node with external parameters (R, T) and add it to the rendering scene, add a light source to the scene, and import the actual calibration object image taken.

[0076] Read pixels and depth from the rendered scene, overlay the actual image with the rendered image, adjust the camera parameters (extrinsic parameters), update the camera's extrinsic parameter matrix in real time and synchronize it to the rendered scene until the feature patterns of the actual image and the rendered image coincide, and obtain the camera parameters at this time.

[0077] The sensor surface CAD model is read, the calibrated camera parameters are imported, and the scene with the elastomer undeformed is rendered to obtain the background depth and background normal vector. A certain width of the image edge (10 pixels in this implementation) is extracted as the image boundary, and the image boundary depth is obtained.

[0078] (1.3) The true value of the normal vector is obtained by rendering the image of a small ball with a known radius, and then the normal vector prediction model is trained.

[0079] The sensor is fixed on the base, and computer numerical control technology is used to control the probe with a small sphere of known radius at the end to contact the sensor at different positions. The RGB camera and NIR camera are used to evenly collect contact images of different areas of the surface, and the RGB image or NIR image is used as the calibration image.

[0080] Based on the CAD model of the sensor's outer surface, render an image of the contact surface between the sphere and the sensor. Enter the calibrated camera parameters to add a camera node. Enter a known radius to create a small sphere mesh and add a material. Combine the mesh and material into a renderable object, add it to the scene, and create a sphere node.

[0081] The calibration image and the rendered image are superimposed and displayed, and the 3D coordinates of the sphere are adjusted so that the calibration image and the rendered image completely overlap. The scene depth at this time is read from the renderer. The contact area mask is obtained by subtracting the scene depth from the background depth, and the RGB, NIR pixel values ​​and pixel coordinates of the mask area in the contact image are extracted.

[0082] The true value of the surface normal vector is determined based on the depth of the mask area, and the training set is constructed by combining the corresponding RGB, NIR pixel values ​​and pixel coordinates; specifically:

[0083] The central difference is used to calculate the gradient of the depth map (scene depth) in the x and y directions of the image coordinate system:

[0084]

[0085] Then the true value n of the surface normal vector is calculated as follows:

[0086]

[0087] With background depth, background normal vector, RGB, NIR pixel values ​​and pixel coordinates as input and surface normal vector as output, the neural network model is trained through the training set, and the trained neural network model is used as the normal vector prediction model.

[0088] In this embodiment, the specific method for preparing the training set is as follows: processing the RGB image through affine transformation parameters to align it with the NIR image data; normalizing the X-axis coordinate to [-1, 1], the Y-axis coordinate to H / W (H is the image height, W is the image height), and adding the Z-axis coordinate to 0. Processing the image through the ball contact area mask retains the pixel values ​​and pixel coordinates of the contact area. By generating random numbers, loading contact data with a 50% probability and background data with a 50% probability, the model learns to distinguish between contact and non-contact states. When using NIR training, the RGB and NIR data are spliced, and the input channels are expanded to 6.

[0089] The neural network model adopts the MLP model, and the input of the MLP network is the difference between the foreground pixel and the background pixel, and the position code. If only RGB training is used, the number of input channels is 3; if NIR training is enabled, the number of input channels is 6; if position coding is enabled, the number of input channels is 6; if both NIR and position coding are enabled, the number of input channels is 9. In this embodiment, the specific structure of the MLP model is shown in Table 1, where the 1×1 convolution layer is used for cross-channel feature interaction and fusion of multimodal information; the batch normalization layer standardizes the activation value to accelerate training convergence; the ReLU activation function layer introduces nonlinear factors to improve the model's expressiveness; the DropOut layer randomly blocks 40% of neurons to prevent overfitting of tactile noise; and the output layer compresses the features to 3 channels, representing the components of the three directions of the normal vector.

[0090] Table 1 MLP model

[0091]

[0092] The model training configuration is as follows: GPU-accelerated training is used, with FP32 precision, an initial learning rate of 0.01, a batch size of 16, 300 epochs, and a weight decay coefficient of 0.0001. Weight decay is applied to standard convolutional layers to prevent overfitting; to avoid interfering with batch normalization statistics calculations and because the bias term has a minimal impact on model complexity, weight decay is not applied to batch normalization layers and bias terms. The Adam optimizer is used because its adaptive learning rate feature is suitable for scenarios with large gradient variations in tactile data. The learning rate is reduced to 0.1 every 100 epochs to achieve rapid initial convergence and fine-tune the results later. num_workers is set to 8 to fully utilize multi-core CPUs; persistent_workers is set to True to avoid the overhead of repeated creation and destruction. The L1 loss function is selected, and the gradient error calculated from the normal vector is monitored.

[0093] (1.4) Integrate the depth prior and solve the depth field through normal vector integration to achieve 3D reconstruction.

[0094] When the dexterous hand grasps the first object to be tested, the RGB image and NIR image of the first object when it contacts the sensor are collected through the RGB camera and the NIR camera, and the RGB and NIR pixel values ​​and pixel coordinates are obtained based on the RGB and NIR images; based on the RGB and NIR pixel values ​​and pixel coordinates as well as the background depth and background normal vector, the predicted normal vector of the object surface is obtained through the normal vector prediction model.

[0095] Then, based on the predicted normal vector and the image boundary depth, the depth field is obtained, which specifically includes:

[0096] Assuming the surface depth is z(x,y), the relationship between its gradient (p,q) and normal vector n(x,y,z) is as follows:

[0097]

[0098] The 3D reconstruction problem can be described as solving the following Poisson equation:

[0099]

[0100] The above Poisson equation is discretized using central difference. For pixel (i, j), its depth divergence can be expressed by the neighborhood depth:

[0101]

[0102] For all valid pixels in the image, the constraints of all pixels are combined into a sparse matrix format to obtain the following discrete linear system, which is obtained:

[0103] Az=b

[0104] Where A is a sparse matrix encoding the depth coefficient (1 or -1), z is a vector containing the depth of all pixels, and b is the divergence term of the Poisson equation, which is obtained by calculating the depth gradient from the normal vector.

[0105] The image boundary depth is used as the depth prior z prior , using the weight λ as an additional constraint, the above sparse matrix is ​​expanded to obtain the augmented sparse linear system:

[0106] A total z=b total

[0107] in I prior is a diagonal matrix whose diagonal elements are 1 at positions corresponding to pixels with valid depth priors.

[0108] Use the least squares method to solve the weighted linear equations, that is, to find:

[0109] min z ||A total zb total || 2

[0110] This is equivalent to solving:

[0111] A total T A total z=A total T b total

[0112] Use the sparse matrix solvers sparseqr or spsolve for fast solutions. sparseqr is based on the QR decomposition of Householder reflection and has high stability and is suitable for large-scale problems. spsolve is based on the LU decomposition method and has low memory usage but is slower, making it suitable for small-scale applications.

[0113] Assume that the camera internal parameter is focal length (f x ,f y ), principal point (c x ,c y ), depth z i,j The corresponding three-dimensional points are:

[0114]

[0115] Z i,j =z i,j

[0116] This completes the reconstruction of the three-dimensional point cloud of the object surface.

[0117] It should be noted that each visual-tactile sensor performs the above three-dimensional reconstruction process separately, thereby reflecting the appearance of the object in different directions.

[0118] (2) Based on the depth map acquisition of the palm-eye module, the second RGB camera and the ToF (Time of Flight) camera respectively obtain the RGB image and sparse depth of the second object, and design a dynamic spatial propagation network based on adaptive weights, dynamic paths, and diffusion suppression strategies to construct a depth correction model. Through this depth correction model, the RGB image is combined with the sparse depth to obtain a depth map.

[0119] Accurate and dense depth measurement is very important for robot perception, but the accuracy and density of depth maps are often difficult to balance. Considering the high resolution of RGB images and the accuracy of direct depth measurement, the image-guided depth completion (IGDC) method generates a high-density, high-precision depth map through sparse depth measurement and the corresponding dense RGB image. The present invention proposes a dynamic spatial propagation network for image-guided sparse depth completion, which decomposes pixel affinity through a nonlinear propagation model, combines adaptive weights, dynamic paths and diffusion suppression strategies to achieve dynamic propagation and optimization of depth values. The details are as follows:

[0120] like Figure 7 As shown, the deep correction model includes ResNet34-Unet and dynamic spatial propagation network:

[0121] (2.1) ResNet34-Unet obtains the initial depth map V based on the aligned RGB image and sparse depth 0 , unweighted affinity matrix Initial similarity matrix Attention weight and confidence mask C.

[0122] Specifically, a ToF camera acquires sparse depth measurements, while an RGB camera captures a high-resolution color image, and the two are aligned. The RGB image provides scene color and texture information, guiding depth propagation; the sparse depth map, containing a small number of known depth values, serves as the seed point for depth propagation.

[0123] Input the aligned RGB image and sparse depth into ResNet34-Unet to generate the initial depth map V 0 , attention weight ((i, j) is the pixel coordinate, k is the distance between the pixel and its neighbor), unweighted affinity matrix Initial similarity matrix And the confidence mask C. The initial depth map is a rough estimate of the depth value, the elements in the attention weight represent the attention that should be given to the neighbors at the iteration step t and the specific distance k, the affinity matrix represents the connectivity of the pixel neighborhood, and the initial similarity matrix represents the static similarity between the pixel and the neighborhood based on the RGB image prediction; the confidence mask marks the sparse depth reliable area, 1 means retaining the original value, and 0 means that the calculation needs to be propagated.

[0124] (2.2) The dynamic spatial propagation network consists of N submodules, based on the initial depth map V 0 , each sub-module updates the depth map in turn. That is, after N iterations, when all sub-modules have completed the update, the depth completion is completed and the final corrected depth map is obtained.

[0125] For the t-th submodule, i.e., at the t-th iteration (t=1, 2…N), affinity is decomposed through the nonlinear propagation model, and then a deep update is performed, specifically:

[0126] Unweighted affinity matrix The details are as follows:

[0127]

[0128] Among them, the affinity matrix is ​​a 0-1 matrix, marking the neighborhood The pixel pairs within control the propagation path.

[0129] Attention weight Reflects the importance of neighbor pixels at different distances k to the depth propagation of the current pixel (i, j); based on attention weight and the initial similarity matrix Calculate the dynamic similarity matrix W t ; Specifically based on the initial similarity matrix elements Calculate the dynamic similarity matrix W t element

[0130]

[0131] in is the similarity weight between pixel (i, j) and its neighbor with offset (a, b) at iteration step t.

[0132] By the dynamic similarity matrix W t right Perform weighting to generate affinity matrix A t , A t Integrate the propagation path and propagation intensity as the core matrix of deep propagation:

[0133]

[0134] Based on the affinity matrix A t And confidence mask C for the current depth map V t-1 Update and get the updated depth map V t :

[0135] V t =(D -1 A t )V t-1 (1-C)+V sparse ·C

[0136] Where D is the affinity matrix A t The rows and matrices (diagonal elements are At The sum of the corresponding row elements) normalizes the depth propagation weight of each pixel; (D -1 A t )V t-1 is the depth propagation item, which is activated in the area with low confidence mask value and is used to interpolate and update the depth value according to the neighborhood information; V sparse =V 0 It is a sparse depth constraint term, which is activated in areas with higher confidence mask values ​​and is used to re-inject high-precision initial depth values.

[0137] Furthermore, the attention weight Here’s how to update:

[0138] At the tth iteration, for the current depth map V t-1 For the pixel at position (i, j) in the network, k represents the distance between the pixel and its neighboring pixels. ResNet34-Unet is trained to generate a series of attention weights related to k: for close neighbors (such as k = 1), a higher attention weight is output to make the neighboring pixels propagate depth first; for distant neighbors (such as k ≥ 2), a lower attention weight is output, and the attention weight decays with increasing distance, reducing its propagation contribution, thereby avoiding the introduction of irrelevant depth information and causing blur.

[0139] This adaptive weight strategy ensures that neighboring pixels with high similarity are filled first, and distant neighboring pixels with low similarity contribute less, ensuring accurate local details and avoiding blur.

[0140] Furthermore, edge diffusion suppression is added to the dynamic spatial propagation network:

[0141] The depth update process includes a diffusion suppression step to process pixels in the edge area of ​​the object, where the weight of the target pixel connected to itself is significantly higher than the weight of the connection with its neighbor pixels, so that the affinity matrix A t The corresponding row of the pixel is approximately a self-loop matrix, and the propagation operator D formed based on the self-loop matrix -1 A t Approximately the unit matrix, then V t =V t-1 , thereby suppressing the cross-edge propagation and update of depth.

[0142] Specifically, during the iterative update of the depth map, the dynamic spatial propagation network uses its attention module to dynamically identify edge areas based on the rich gradient information provided by the RGB image (especially the significant gradient changes at the object boundaries). For the identified edge pixels, the network will learn to output specific attention in the current or subsequent iterations. The specific attention is manifested as: for the edge pixel itself (i.e., the neighborhood distance k = 0), its attention weight is significantly improved and approaches a larger value, while for all other neighboring pixels (k>0), their attention weights is suppressed and approaches 0, which makes the overall affinity matrix A t On the row corresponding to the edge pixel, it degenerates into a self-loop matrix with non-zero values ​​only in the diagonal position. At this time, since the diagonal elements of the row and matrix D are equal to the self-loop weight, the normalized propagation operator D - 1 A t On this line, it becomes a unit vector. Therefore, in the edge area without initial sparse points (C=0), the above depth update formula becomes:

[0143] V t =(D -1 A t )·V t-1 =V t-1

[0144] That is, the depth no longer spreads.

[0145] Furthermore, dynamic path updates are adopted in dynamic spatial communication networks:

[0146] The dynamic path strategy optimizes the depth propagation path by dynamically adjusting the connectivity of pixel neighborhoods, reducing invalid calculations and improving edge preservation capabilities. Neighborhood offset field Predicted by ResNet34-Unet, neighborhood Dynamic update via:

[0147]

[0148] Where (p,q) is the offset predicted by the deformable convolution. By dynamically predicting the neighborhood offset through deformable convolution, the neighborhood can be expanded beyond the fixed grid, adaptively capturing long-range dependencies or edge structures. In flat areas, the offset expands the neighborhood, including more distant neighbors and accelerating depth filling. In edge regions, the offset shrinks the neighborhood to only the center pixel.

[0149] (3) Based on the three-dimensional point cloud of the surface of the first object and the depth map of the second object, the spatial posture alignment of the first object and the second object is achieved, providing key perception support for the precise operation of the dexterous hand (such as precision assembly, grasping control or path planning). For example: During precision assembly, the three-dimensional point cloud of the surface of the electronic component is reconstructed by the visual tactile sensor, and the sparse depth of the environment is obtained by the palm eye module and completed into a high-resolution depth map, so as to achieve sub-millimeter posture alignment between the assembly component and the target substrate. During minimally invasive surgery, the surface of the precision instrument is reconstructed by the visual tactile sensor, and the depth map of the affected area is completed by the palm eye module under the interference of liquid, so as to generate a safe operation path that avoids the vital parts.

[0150] To address the issues of insufficient tactile perception accuracy and low depth estimation resolution in existing rigid robotic arms, this paper proposes a multimodal three-finger dexterous hand that integrates visual-tactile and eye-in-the-palm sensing. By replacing the rigid fingertips with finger-shaped visual-tactile sensors, a multispectral depth-photometric stereo 3D reconstruction method is proposed, enabling the robotic hand to possess high-spatial-resolution tactile perception and fine-grained perception of surface textures. By embedding an eye-in-the-palm module equipped with a monocular RGB camera and a Time-of-Flight (ToF) camera in the palm of the robotic hand, a dynamic spatial propagation network combining adaptive weights, dynamic paths, and a diffusion suppression strategy is proposed. This network fuses dense RGB color information with sparse depth information to achieve high-resolution depth estimation of the operating area.

[0151] This invention achieves high-precision, high-resolution tactile perception (normal vector estimation error of 0.013, surface texture spatial resolution of 0.01mm); when the sparse input is 0.05%, the root mean square error of depth estimation is reduced to 0.089, and the average relative error is reduced to 0.012. Due to its high-precision and highly robust multimodal visual and tactile perception capabilities, this invention has applications in precision assembly, minimally invasive surgery, human-machine collaboration, and other fields.

[0152] The following are specific embodiments:

[0153] Example 1

[0154] The 3D reconstruction method proposed in this invention is used to perform 3D reconstruction on four types of objects: a screwdriver, a fingertip, a grid pen holder, and a color ring resistor. Figure 6 As shown in the figure, the normal vector and the 3D point cloud reconstruction results show that, compared with using only the RGB channel, the slight texture changes of the test object can be more clearly displayed after adding NIR imaging, which proves the effectiveness of the method proposed in this invention in improving the accuracy of visual and tactile 3D reconstruction of curved surfaces.

[0155] Example 2

[0156] Application of multimodal three-finger dexterous hand in precision electronic component assembly:

[0157] In semiconductor chip assembly scenarios, micro-capacitor components measuring 0.5mm x 0.3mm need to be accurately placed on circuit board pads. This operation requires the robot to have submillimeter positioning accuracy and to sense the surface texture of the component and the three-dimensional structure of the circuit board to avoid assembly errors caused by visual obstruction or tactile noise. The operation process is as follows:

[0158] 1) Visual-tactile sensor initialization and calibration.

[0159] Structural configuration: Finger-shaped multispectral visual and tactile sensors are installed on the fingertips of three fingers. Each sensor integrates an RGB camera (640×480 pixels), a NIR camera, a spectroscopic prism and a silicone elastomer (Shore hardness 20A). A palm-mounted eye module (including a ToF camera, depth resolution 160×120, sparse input density 0.5%) is embedded in the palm.

[0160] Calibration process: By rendering the CAD model of the calibration object, interactively adjust the camera's extrinsic parameters (rotation matrix, translation vector) to align the rendered image with the actual calibration image features and obtain the camera's intrinsic parameters (focal length, principal point). A small ball probe with a radius of 3mm touches the sensor surface, and the true normal vector is obtained through rendering matching (error ≤ 0.013 rad). The MLP model is then trained to predict the normal vector of the contact area.

[0161] 2) Component grasping and surface perception.

[0162] Contact sensing: Three fingers close together to grasp the capacitive element, causing the silicone elastomer to deform (maximum deformation 0.1mm), and the RGB / NIR camera synchronously captures the contact image.

[0163] 3D Reconstruction: Surface normals are calculated using multispectral photometric stereo, and NIR spectra are integrated to improve normal accuracy in areas with weak textures (such as smooth component surfaces). A boundary prior (10-pixel width and depth at the image edge) is introduced to solve the Poisson equation, reconstructing the component surface point cloud and determining the component's pose and surface curvature distribution.

[0164] 3) Deep completion of dynamic spatial propagation networks.

[0165] Input data: The PalmEye ToF camera provides sparse depth of the circuit board pad area (input density 0.05%), and the RGB camera outputs a 1920×1080 pixel color image.

[0166] Depth completion process:

[0167] Adaptive weighting strategy: Decouple pixel neighborhoods by distance into near-neighbor regions (e.g., a 3×3 area) and far-neighbor regions (e.g., a 5×5 area). A learnable attention mechanism assigns dynamic weights to regions of varying distances and iteration steps. This mechanism tends to prioritize far-neighbor regions for rapid infill in the early stages of an iteration, while assigning higher weights to near-neighbor regions in the later stages of an iteration to refine details. It also effectively learns to suppress noise interference from distant regions (e.g., a circuit board background).

[0168] Dynamic Path Strategy: Leveraging the mechanism of deformable convolution, a set of sparse neighborhood offsets (e.g., offsets within ±2 pixels) is dynamically predicted for each pixel. By only interacting with a few (e.g., 3-5) key neighbors, this strategy significantly reduces computational complexity in 6 iterations.

[0169] Diffusion suppression: In the identified object boundary area (such as the junction of a component and a pad), the propagation weights of the pixels in this area are locally degenerated into a self-circular structure through attention or path adjustment. This causes the propagation matrix to form unit vectors on the rows corresponding to the boundary pixels, thereby terminating the cross-edge propagation of depth and ensuring the boundary clarity of the final depth map.

[0170] 4) Precise assembly execution.

[0171] Path planning: Based on the completed dense depth map (resolution 1920×1080), an obstacle avoidance trajectory is generated (avoiding component pins and capacitors around pads), with an end-effector positioning accuracy of ≤±0.05mm.

[0172] Posture adjustment: Using the visual-tactile sensor to reconstruct the posture of components and pads, the three-finger grasping angle and position are adjusted in real time. Through multispectral image matching and depth completion data, submillimeter posture alignment (error ≤ 0.1mm) is achieved.

[0173] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multimodal dexterous hand that integrates vision, touch, and palm-eye, characterized by: It includes a hand body, several visual and tactile sensors, a palm eye module and a processor, among which: The visual-tactile sensor includes a skeleton, an elastic body, an RGB camera, and a NIR camera. The skeleton is a finger-shaped structure mounted on the hand body. The skeletons of each visual-tactile sensor and the hand body together form a hand structure. The elastic body covers the outside of the skeleton, and the surface of the elastic body is covered with a diffuse reflection coating. The RGB camera and the NIR camera are mounted within the skeleton and are used to respectively obtain an RGB image and an NIR image when a first object contacts the elastic body. The palm eye module is installed in the middle of the hand body, and the palm eye module includes a second RGB camera and a TOF camera, and the second RGB camera and the ToF camera are used to obtain an RGB image and a sparse depth of a second object respectively; The processor is used to reconstruct a three-dimensional point cloud of the surface of the first object based on the RGB image and NIR image of the first object when it is in contact with the elastic body, and to obtain a depth map of the second object based on the RGB image and sparse depth of the second object; and to align the postures of the first object and the second object based on the three-dimensional point cloud of the surface of the first object and the depth map of the second object.

2. An application of the multimodal dexterous hand integrating vision, touch and palm-eye as claimed in claim 1, characterized in that: The steps include: When the dexterous hand grasps the first object, it uses the RGB camera and the NIR camera to capture RGB and NIR images of the first object when it contacts the sensor. RGB and NIR pixel values ​​and pixel coordinates are obtained based on the RGB and NIR images. A predicted normal vector for the first object is obtained based on the RGB and NIR pixel values ​​and pixel coordinates. A depth field is obtained based on the predicted normal vector and the image boundary depth, thereby reconstructing a three-dimensional point cloud of the first object's surface. The second RGB camera and the ToF camera respectively obtain an RGB image and a sparse depth of the second object, and the sparse depth is completed by the RGB image to obtain a depth map of the second object; Position alignment of the first object and the second object is achieved according to the three-dimensional point cloud of the first object surface and the depth map of the second object.

3. The application of the multimodal dexterous hand integrating vision, touch and palm-eye as claimed in claim 2, characterized in that: Based on the RGB, NIR pixel values ​​and pixel coordinates, the first object predicted normal vector is obtained, specifically: Inputting RGB, NIR pixel values ​​and pixel coordinates into the normal vector prediction model to obtain the first object predicted normal vector; The method for obtaining the normal vector prediction model includes: A calibration object with characteristic protrusions is placed on the sensor surface, and the actual image is acquired using an RGB camera or NIR camera. Based on the calibration object's CAD model, an image of the contact surface between the calibration object and the sensor is rendered to obtain a calibration object rendered image. The corresponding camera parameters are obtained when the characteristic patterns of the actual image and the calibration object rendered image overlap. Based on the camera parameters, the scene of the undeformed elastic body is rendered to obtain the background depth and background normal vector, and the image boundary depth is determined. A probe with a sphere at its end is placed in contact with the sensor at different locations. RGB and NIR images of different surface areas are captured using RGB and NIR cameras, i.e., contact images. The RGB and NIR images are used as calibration images. Based on the sensor CAD model, the contact surface image of the sphere and the sensor is rendered to obtain a sphere-rendered image. The scene depth is obtained when the calibration image and the sphere-rendered image completely overlap. The mask area is obtained by subtracting the scene depth from the background depth, and the RGB, NIR pixel values ​​and pixel coordinates of the mask area in the contact image are extracted; the surface normal vector is calculated based on the depth of the mask area; The neural network model is trained with background depth, background normal vector, RGB, NIR pixel values ​​and pixel coordinates as input and surface normal vector as output, and the trained neural network model is used as the normal vector prediction model.

4. The application of the multimodal dexterous hand integrating vision, touch and palm-eye as claimed in claim 3, characterized in that: Based on the predicted normal vector and the image boundary depth, the depth field is obtained, including: The 3D reconstruction problem of the object is described as the Poisson equation. The Poisson equation is discretized using central difference to obtain the depth constraint of each pixel of the contact image. The depth constraints of all pixels are combined into a discrete linear system. The image boundary depth is used as the depth prior z prior and transform the depth prior z into prior As an additional constraint, the discrete linear system is expanded into an augmented sparse linear system; The augmented sparse linear system is solved using the least squares method to obtain the depth field.

5. The application of the multimodal dexterous hand integrating vision, touch and palm-eye as claimed in claim 4, characterized in that: The augmented sparse linear system is specifically expressed as: Where A is the sparse matrix of the encoded depth coefficients, z is the vector containing the depth of all pixels; b is the divergence term of the Poisson equation, which is obtained by calculating the depth gradient by predicting the normal vector; I prior is a diagonal matrix.

6. The application of the multimodal dexterous hand integrating vision, touch and palm-eye according to any one of claims 2 to 5, characterized in that: The sparse depth is complemented by the RGB image to obtain a depth map of the second object, specifically: the RGB image and the sparse depth of the second object are aligned and input into the depth correction model to obtain the depth map of the second object; The depth correction model includes ResNet34-Unet and dynamic spatial propagation network, where: ResNet34-Unet obtains the initial depth map V based on the aligned RGB image and sparse depth 0 , unweighted affinity matrix Initial similarity matrix Attention weight and confidence mask C; The dynamic spatial propagation network consists of multiple submodules based on the initial depth map V 0 , each submodule updates the depth map in turn, and the final corrected depth map is obtained after all submodules complete the update; Where, for the t-th iteration, t = 1, 2…N, N is the number of submodules: Based on attention weight and the initial similarity matrix Calculate the dynamic similarity matrix W t ; By the dynamic similarity matrix W t and Generate affinity matrix A t : Based on the affinity matrix A t And confidence mask C for the current depth map V t-1 Update and get the updated depth map V t .

7. The application of the multimodal dexterous hand integrating vision, touch and palm-eye as claimed in claim 6, characterized in that: For the current depth map V t-1 The updated calculation formula is: In t =(D -1 And t )V t-1 ·(1-C)+V sparse ·C Among them, (D -1 A t )V t-1 is the depth propagation term, which is activated in the area with lower confidence mask value, and D is the affinity matrix A t The rows and matrices of V sparse =V 0 It is a sparse depth constraint term and is activated in areas with higher confidence mask values.

8. The application of the multimodal dexterous hand integrating vision, touch and palm-eye as claimed in claim 7, characterized in that: The depth map is not updated in the edge area of ​​the object, that is, the edge area V t =V t-1 .

9. The application of the multimodal dexterous hand integrating vision, touch and palm-eye as claimed in claim 6, characterized in that: Attention weight is determined as follows: For the current depth map V at the tth iteration t-1 The pixel at position (i, j) in Represents the attention weight of the pixel to its neighboring pixels with a distance of k. As k increases, the attention weight Gradually decrease.

10. The application of the multimodal dexterous hand integrating vision, touch and palm-eye as claimed in claim 6, characterized in that: Unweighted affinity matrix Neighborhood-based To update: in, express The value of the (i, j) position in the neighborhood range The method of determining is: Among them, (p,q) is the neighborhood offset, is the offset field predicted by the deformable convolutional network, and is the parameter of the deformable convolutional network.

Citation Information

Patent Citations

  • Non-contact control method of bionic manipulator based on learning of hand motion gestures

    CN106346485A

  • Real-time three-dimensional reconstruction method for object grabbed by mechanical arm

    CN113313815A

  • Mechanical arm sensing method based on multi-modal data fusion

    CN117103277A

  • Visual and tactile fusion-based robot grabbing pose optimization method

    CN119681901A

  • System and method for tracking hand motion using strong coupling fusion of image sensor and inertial sensor

    KR102456872B1

Cited By

  • Finger type visual tactile sensor calibration method and device and server

    CN121527200A

  • Calibration method and device of finger-type tactile sensor and server

    CN121527200B