Multi-modal data acquisition method, device and system for human hand teaching process of robot
By acquiring six-DOF pose data of human finger joints and tactile information of fingertips, and using binocular cameras and visual matching algorithms to determine coordinate system transformation parameters, the problem of lack of quantification of tactile feedback information and data misalignment in existing technologies is solved, and high-quality multimodal data acquisition is achieved, thereby improving the data acquisition capability of robot operation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF AUTOMATION CHINESE ACAD OF SCI
- Filing Date
- 2026-01-06
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies struggle to achieve high-quality multimodal data acquisition during robot dexterity operations, especially in complex environments where tactile feedback information lacks quantitative recording and tactile sensor and human hand pose data cannot be aligned in the same coordinate system.
By acquiring six-degree-of-freedom pose data of the finger joints and tactile information of the fingertips, and using a binocular camera module to collect deformation images of the tactile sensor in real time, combined with a binocular vision matching algorithm and a light refraction tracking model, the coordinate transformation parameters of the tactile sensing coordinate system under the human wrist coordinate system are determined, thereby achieving the unification of tactile information and pose data.
It achieves spatial alignment of high-resolution tactile information and human hand pose data in a unified coordinate system, generating a high-quality multimodal dataset and improving the data acquisition capability and efficiency of robot dexterity tasks.
Smart Images

Figure CN121848447A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot tactile sensing and imitation learning technology, and in particular to a method, apparatus, system and computing device for acquiring multimodal data during human hand teaching process. Background Technology
[0002] In the field of dexterous manipulation in robotics, imitation learning allows robots to learn new skills by mimicking human instruction on specific tasks. This has shown great potential in enabling robotic systems to handle diverse tasks in unstructured and complex environments. The key to achieving dexterous manipulation through imitation learning lies in obtaining high-quality expert instruction data.
[0003] Currently, researchers have developed various data acquisition devices and corresponding methods to collect expert teaching data. The main technical solutions and their limitations are as follows: The first type is data acquisition based on visual observation or teleoperation. This type of solution utilizes exoskeletons, motion capture devices, data gloves, augmented reality, or video demonstrations. While these methods have advantages in terms of natural movement, operators often rely solely on visual observation to remotely control the robot without direct physical interaction with the experimental environment. On the one hand, visual feedback is easily affected by occlusion or lighting conditions in the experimental environment, reducing the effectiveness of the data; on the other hand, the lack of tactile feedback from the experimental object during the teaching process significantly reduces the quality and efficiency of this type of teaching data in dexterity tasks with high precision requirements, frequent contact, and complex processes. The second type is data acquisition devices that introduce real-time tactile feedback. To address the above problems, some technologies introduce tactile feedback during the teaching process, allowing operators to directly interact with the manipulated object. However, existing devices often only serve a "cue" function, relying entirely on the subjective touch of the human fingertips for tactile feedback. They neglect the quantitative recording of tactile feedback during interaction, resulting in the final collected data remaining limited to the pose information of the robotic arm or human hand. Due to the lack of quantification and representation of tactile data as a contact modality, robots cannot use it as policy input when conducting subsequent imitation learning.
[0004] Furthermore, existing tactile sensing technologies also have shortcomings at the hardware level. Traditional thin-film pressure sensor arrays typically achieve an accuracy of only about 10 millimeters, and often can only calculate the normal force perpendicular to the surface, making it difficult to acquire complex mechanical information such as the three-dimensional deformation of the contact surface, the precise contact position, and the tangential force. More importantly, although data gloves for acquiring pose and knuckle sleeves for acquiring tactile sensation exist in existing technologies, there is a lack of an effective mechanism to unify the high-dimensional point cloud data of tactile sensors with the positional framework of human hand movements. If the tactile information relative to the fingertip coordinate system cannot be accurately aligned to a unified coordinate system with the wrist as the origin, even if the hand moves in space, the internal grasping data cannot remain stable and consistent, making it difficult to generate multimodal datasets for training high-quality robot strategies.
[0005] Therefore, how to provide a multimodal data acquisition scheme that enables physical contact interaction and spatially aligns high-resolution tactile information with human hand pose data in a unified coordinate system is an urgent problem to be solved. Summary of the Invention
[0006] This disclosure provides a multimodal data acquisition method for a robot's hand teaching process. The multimodal data acquisition method includes: acquiring six-degree-of-freedom pose data of each finger joint of a human hand in a human wrist coordinate system; acquiring tactile information from a tactile sensor located at the fingertip, wherein the tactile information includes at least: three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system of the tactile sensor, the multiple sampling points being used to characterize the deformation of the elastic sensing module of the tactile sensor; transforming the three-dimensional spatial coordinates of the multiple sampling points to the human wrist coordinate system based on predetermined coordinate transformation parameters of the tactile sensing coordinate system in the human wrist coordinate system to obtain coordinate-transformed tactile information; and associating the six-degree-of-freedom pose data with a time synchronization relationship with the coordinate-transformed tactile information to obtain the multimodal data.
[0007] According to an embodiment of this disclosure, acquiring tactile information from a tactile sensor located at the fingertip includes: using a binocular camera module to acquire in real time a deformation image of the inner surface of the elastic sensing module of the tactile sensor; identifying the two-dimensional image coordinates of the plurality of sampling points in the deformation image; and obtaining the three-dimensional spatial coordinates of the plurality of sampling points in the tactile sensing coordinate system based on the two-dimensional image coordinates using a binocular vision matching algorithm and a light refraction tracking model.
[0008] According to embodiments of this disclosure, the coordinate system transformation parameters of the tactile sensing coordinate system under the human wrist coordinate system are determined through the following steps: when a person wears the tactile sensor and grasps a calibration object with a predetermined geometric shape, the pose representation of the calibration object under the human wrist coordinate system is obtained, and the three-dimensional spatial coordinates of multiple sampling points under the tactile sensing coordinate system of the tactile sensor are obtained; the three-dimensional spatial coordinates of the multiple sampling points when in contact with the calibration object are compared with the three-dimensional spatial coordinates when not in contact with the object, and the multiple sampling points are divided into contact areas and non-contact areas according to a preset three-dimensional deformation threshold; a geometric optimization objective function is constructed, which characterizes the geometric consistency between the transformed coordinates obtained by the three-dimensional spatial coordinates of the sampling points as the contact area after transformation by the coordinate system transformation parameters to be solved and the surface of the calibration object determined based on the pose representation; and the coordinate system transformation parameters are obtained by minimizing the geometric optimization objective function.
[0009] According to an embodiment of this disclosure, the predetermined geometry is a sphere. The geometry optimization objective function includes the mean value of the difference between the distance from the center of the sphere to the transformed coordinates of the sampling points in the contact area and the radius of the sphere.
[0010] According to an embodiment of this disclosure, the geometric optimization objective function further includes: the variance of the distance from the center of the sphere to the transformed coordinates of the sampling points in the contact area.
[0011] According to embodiments of this disclosure, the geometric optimization objective function further includes: Wherein, the ReLU function is a non-linear activation function, and R is the radius of the sphere. The Euclidean distance from the center of the sphere to the transformed coordinates of the sampling point in the non-contact area after being converted to the coordinate system of the human wrist.
[0012] This disclosure also provides a multimodal data acquisition device for a robot's hand teaching process. The multimodal data acquisition device includes: a data glove worn on the palm and knuckles of a human hand to detect the motion state of each finger and knuckle to obtain six-degree-of-freedom pose data of each finger and knuckle in the wrist coordinate system; a tactile flexible fingertip sleeve including an elastic sensing module, worn on the fingertip as a tactile sensor to acquire tactile information of the fingertip, wherein the tactile information includes at least the three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system of the tactile flexible fingertip sleeve, the multiple sampling points being used to characterize the deformation of the elastic sensing module; and a data processing module communicatively connected to the data glove and the tactile flexible fingertip sleeve, configured to execute the acquisition method described above.
[0013] According to embodiments of this disclosure, the elastic sensing module includes: a contact elastomer having an inner surface and a plurality of tactile markers distributed on the inner surface, the contact elastomer being configured to deform upon contact with an object, and the plurality of tactile markers being configured to be identified to generate the plurality of sampling points. The tactile flexible fingertip sleeve further includes: a base frame defining a wearing hole for accommodating a human fingertip, the elastic sensing module being disposed on the base frame; a flexible wearing ring disposed on the inner wall of the wearing hole, configured to adapt to fingertips of different sizes and provide frictional fixation during wear; a light source illumination module disposed on the base frame for illuminating the inner surface of the contact elastomer; and a binocular camera module fixedly mounted on the base frame, facing the inner surface of the contact elastomer, for acquiring deformation images of the inner surface including the plurality of tactile markers.
[0014] This disclosure also provides a multimodal data acquisition system for the human hand teaching process of a robot. The multimodal data acquisition system includes: a finger joint pose module configured to acquire six-degree-of-freedom pose data of each finger joint in the human wrist coordinate system; a fingertip tactile sensing module configured to acquire tactile information of the fingertips, wherein the tactile information includes at least: three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system of the fingertip tactile sensing module, the multiple sampling points being used to characterize the deformation of the elastic sensing module of the fingertip tactile sensing module; a wearable pose calibration module configured to pre-determine coordinate system transformation parameters of the tactile sensing coordinate system in the human wrist coordinate system; and a multimodal data alignment module configured to transform the three-dimensional spatial coordinates of the multiple sampling points to the human wrist coordinate system based on the coordinate system transformation parameters, to obtain coordinate-transformed tactile information, and to associate the six-degree-of-freedom pose data with time synchronization with the coordinate-transformed tactile information to obtain the multimodal data.
[0015] This disclosure also provides a computing device, the computing device comprising: a processor; and a memory storing a computer program, wherein when the computer program is executed by the processor, the method for acquiring multimodal data of the manual teaching process as described above is implemented.
[0016] Using the acquisition device, method, and system provided in this disclosure, a human hand wears a data glove on the palm and tactile flexible fingertip sleeves on the fingertips. Based on the human wrist coordinate system, 6D pose data of each finger joint is acquired in real time, and tactile information (such as deformation, contact position, force, etc.) from the fingertip sensors is simultaneously acquired based on the tactile sensing framework. This disclosure transforms the tactile information from the tactile sensing coordinate system of each fingertip to the human wrist coordinate system by determining relevant parameters of the hand's pose (coordinate system transformation parameters), achieving spatial alignment of tactile information and finger joint 6D pose data in a unified coordinate system. This method overcomes the shortcomings of existing technologies in expert teaching data acquisition, which lacks contact interaction and tactile modal data. It can provide high-quality multimodal data containing high-resolution contact surface deformation and precise contact position, thereby significantly improving the data acquisition capability and efficiency of expert teaching in robot dexterity tasks. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the structure of a multimodal data acquisition device for the manual teaching process.
[0018] Figure 2 This is a flowchart illustrating the multimodal data acquisition method during the manual teaching process.
[0019] Figure 3 It is a visualization of the 6D pose data of each finger joint in different hand gesture states, based on the human wrist coordinate system.
[0020] Figure 4 This is a flowchart illustrating the process of acquiring real-time contact deformation of a tactile flexible fingertip sleeve.
[0021] Figure 5 This is a flowchart illustrating the process of acquiring multimodal data based on fingertip tactile sensation and human hand pose, determining relevant parameters of human hand wear pose, and aligning multimodal data in space.
[0022] Figure 6 It is a visualization of the representation of tactile information in the human wrist coordinate system under various fingertip tactile sensing coordinate systems based on coordinate system transformation parameters.
[0023] Figure 7 This is a schematic diagram of a multimodal data acquisition system for human hand teaching of robots.
[0024] Figure 8 This is a schematic diagram of the structure of the computing device provided by the present invention. Detailed Implementation
[0025] The following detailed embodiments are provided to aid the reader in gaining a comprehensive understanding of the methods, apparatus, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but may be changed as will become clear upon understanding this disclosure, except for operations that must occur in a specific order. Furthermore, for clarity and conciseness, descriptions of features known in the art may be omitted.
[0026] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Rather, the examples described herein are provided only to illustrate some of the many feasible ways of implementing the methods, apparatus, and / or systems described herein, which will become clear upon understanding the disclosure of this application.
[0027] As used herein, the term “and / or” includes any one of the associated listed items and any combination of any two or more.
[0028] Although terms such as “first,” “second,” and “third” may be used herein to describe various components, assemblies, regions, layers, or parts, these components, assemblies, regions, layers, or parts should not be limited by these terms. Rather, these terms are used only to distinguish one component, assembly, region, layer, or part from another. Thus, without departing from the teaching of the examples described herein, the first component, first assembly, first region, first layer, or first part referred to as the first component, first assembly, first region, first layer, or first part may also be referred to as the second component, second assembly, second region, second layer, or second part.
[0029] The terminology used herein is for the purpose of describing various examples only and is not intended to limit disclosure. Unless the context clearly indicates otherwise, the singular form is intended to include the plural form as well. The terms “comprising,” “including,” and “having” indicate the presence of the described features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0030] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains upon understanding this disclosure. Unless expressly defined herein, terms (such as those defined in a general dictionary) shall be interpreted as having a meaning consistent with their meaning in the context of the relevant field and in this disclosure, and shall not be interpreted in an idealized or overly formalistic manner.
[0031] Furthermore, in the description of the examples, detailed descriptions of well-known related structures or functions will be omitted when it is believed that such detailed descriptions would lead to a vague interpretation of this disclosure.
[0032] Figure 1 This is a schematic diagram of the structure of a multimodal data acquisition device for the manual teaching process.
[0033] like Figure 1 As shown, the data acquisition device mainly consists of a data glove 100 for wearing on the palm of the hand, a tactile flexible fingertip sleeve 200 for wearing on the fingertips, and a data processing module. Figure 1 (Not shown in the text) consists of...
[0034] Data glove 100 is designed for wearing on the palm and knuckles of the human hand. In practice, wearing guidelines must be followed to ensure that the IMU (Inertial Measurement Unit) sensors at each node of the data glove are positioned at the center of the corresponding knuckle. This data glove is used to detect the motion state of each finger and knuckle of the human hand. After pose calibration and binding to the hand skeleton as described later, it is used to acquire real-time six-degree-of-freedom (6D) pose data of each finger and knuckle in the wrist coordinate system.
[0035] The tactile flexible fingertip sleeve 200 is worn on the fingertips of different types and sizes of the human hand, and acts as a tactile sensor to synchronously acquire tactile information from the fingertips. The tactile flexible fingertip sleeve consists of: a base frame 201, a flexible wear ring 202, a light source illumination module 203, a binocular camera module 204, and an elastic sensing module 205.
[0036] The base frame 201 is used to fix and connect the binocular camera module 204 and the light source illumination module 203 and elastic sensing module 205 of the tactile sensor. The bottom of the base frame 201 has a hole (i.e., a wear hole) for embedded wear by human fingertips. Its bottom projection plane is elliptical, with the lengths of the major axis and minor axis being 19mm and 16mm respectively, and the depth being 24.5mm.
[0037] A flexible wearable ring 202 is disposed on the inner wall of the wear hole to adapt to the fingertips of different types and sizes of human hands. Since the base frame 201 has a rigid structure and fixed dimensions, direct wear may result in insecure fixation or pressure on the fingers. The fingertip achieves a size-adaptive insertion by pressing the flexible wearable ring 202 (e.g., a soft silicone elastic ring). After wear, as the finger pressure decreases, the elastic ring returns to its natural deformation state to provide frictional fixation, thereby protecting the fingertip skin and joint skeleton, and adapting to prolonged expert teaching tasks. The flexible wearable ring 202 is made of silicone, the same material as the contact surface of the tactile sensor, possessing good elasticity and toughness. Its shape is similar to the wear hole at the bottom of the base frame 201, but its length along the major and minor axes is reduced by 2 to 4 millimeters (depending on the fingertip size).
[0038] During the manufacturing process, the A and B components of the transparent silicone are first mixed in a 1:1 ratio and vacuumed to eliminate air bubbles and ensure resilience. Then, Sil-Poxy adhesive is applied to the inner wall of the elliptical hole at the bottom of the base frame 201 to increase adhesion. The mixed silicone is then poured in and shaped using a mold. After curing, the mold is removed, thus completing the bonding between the flexible wearable ring 202 and the base frame 201.
[0039] The light source illumination module 203 is mounted on the base frame 201 and includes a ring-shaped illumination lamp and a light shield. It is used to illuminate the inner surface of the elastic sensing module 205 (i.e., to provide an illumination source for the interior of the tactile sensor). The power supply voltage is 5V, and multiple sensors can be powered in series or parallel.
[0040] A light shield is used to control the brightness of the internal lighting source and reduce reflected light spots. To avoid the formation of tiny light spots on the surface of the acrylic support shell during light reflection within the sensor, which would reduce imaging quality, this embodiment preferably uses a light shield with parameters of 75° incident angle, 2mm thickness, and 1mm width to ensure the acquisition quality of binocular tactile images.
[0041] The binocular camera module 204 is fixedly installed at a designated position on the base frame 201, facing the inner surface of the elastic sensing module 205.
[0042] The binocular camera module 204 is used to directly observe the deformation of the elastic surface of the tactile sensor, acquiring deformation images of the inner surface containing multiple tactile markers (described further later with reference to the elastic sensing module 205). Before use, the focal length needs to be adjusted by observing the imaging of the tactile markers, and the Zhang calibration method is used to extract corner coordinates by taking multiple sets of calibration checkerboard images, and to calculate camera intrinsic parameters, radial distortion parameters, tangential distortion parameters, and binocular transformation parameters.
[0043] The elastic sensing module 205 includes a transparent support frame, a contact elastomer, and multiple tactile markers distributed on its inner surface. The contact elastomer undergoes geometric deformation upon contact with an object. The tactile markers are identified to generate sampling points, thereby reflecting the deformation state. In this embodiment, there are a total of 285 tactile markers, distributed as follows: 51 markers in the central cylindrical region and 117 markers in each of the two hemispherical regions. The geometric deformation of the elastomer is reflected by detecting its real-time three-dimensional coordinate values.
[0044] In another embodiment, the elastic sensing module 205 is also provided with a silver external coating to filter external light to reduce environmental interference, and works with the internal light source illumination module 203 to enhance the robustness of binocular imaging, while protecting the sensor surface from contamination and scratches.
[0045] In this embodiment, the tactile flexible fingertip sleeve 200 achieves higher resolution through the cooperation of the binocular camera module 204 and the elastic sensing module 205. Specifically, unlike traditional thin-film pressure sensor arrays which are typically limited by the physical unit spacing (e.g., achieving only about 10 mm of accuracy), this embodiment employs a vision-based tactile sensing principle. By capturing the continuous displacement of high-density tactile markers (e.g., 285 markers) on the inner surface of the elastic sensing module 205 through the binocular camera module 204, the device can achieve fine sensing within 10 mm, with a sensing resolution controllable to about 1 mm. Based on this high-resolution hardware architecture, the tactile flexible fingertip sleeve 200 can provide richer tactile information than traditional sensors. In addition to the three-dimensional spatial coordinates (point cloud) of multiple sampling points representing deformation, the acquired tactile information can further include multi-dimensional mechanical information (including vector information of the magnitude and direction of forces, such as the direction and angle of normal force, tangential force, and oblique force on the contact surface) and contact state characteristics (the three-dimensional deformation topology of the contact surface and the precise contact position) calculated based on the displacement vectors of the sampling points. This flexible sensing characteristic enables the device to capture minute deformation feedback that rigid sensors cannot record when handling fragile or easily deformable objects, thereby generating high-quality flexible interactive data.
[0046] As per the above reference Figure 1 The structure of the tactile flexible fingertip sleeve 200 is described in detail. However, specific embodiments are not limited thereto.
[0047] Furthermore, despite Figure 1 Not shown, the data acquisition device also includes a data processing module. The data processing module is communicatively connected to the data glove 100 and the tactile flexible fingertip sleeve 200. The data processing module is configured to receive data acquired by the data glove 100 and the tactile flexible fingertip sleeve 200, and process it to obtain multimodal data. The specific operation of the data processing module will be described in further detail below.
[0048] Figure 2 This is a flowchart illustrating the multimodal data acquisition method during the manual teaching process.
[0049] In step S100, the six-degree-of-freedom pose data of each finger joint of the human hand in the coordinate system of the human wrist are obtained.
[0050] In this embodiment, a communication connection is established between the data glove 100 and the data processing module (hereinafter referred to as the host computer) via a wireless receiver, and the hand posture is calibrated after a successful connection. The specific calibration process includes: first, having the hand wearing the data glove draw an "8" in the air with minimal movement to calibrate the magnetic field; then, maintaining a specific hand gesture posture, for example, with the thumb and index finger at a 45° angle and all other fingers flat, holding the position vertically downward and horizontally forward for several seconds to bind the hand skeletal posture.
[0051] After calibration, based on multiple IMU inertial sensors embedded inside the data glove, real-time six-degree-of-freedom (6D) pose data of the human hand (including 1 wrist node, 5 palm nodes, 2 thumb nodes, and 3 nodes each for the other four fingers) is acquired as the hand assumes different hand gestures. The visualization effect of the real-time 6D pose data of each finger joint based on the wrist coordinate system will be discussed in [reference needed]. Figure 3 To provide a more detailed description.
[0052] In step S200, tactile information from a tactile sensor located at the fingertip is acquired. The tactile information includes at least the three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system of the tactile sensor. The multiple sampling points are used to characterize the deformation of the elastic sensing module of the tactile sensor.
[0053] In this embodiment, a calibrated binocular camera module 204 (e.g., resolution set to 1280 x 720, frequency set to 60Hz) is used to acquire in real-time deformation images of the inner surface of the elastic sensing module 205 of the tactile flexible fingertip sleeve 200, which serves as a tactile sensor. As an example, when an object contacts the inner surface of the elastic sensing module 205, the contact elastomer undergoes three-dimensional geometric deformation, causing the tactile markers transferred on the inner surface to undergo two-dimensional coordinate changes under camera observation.
[0054] The two-dimensional image coordinates of each tactile marker point in the deformed image are identified using image processing algorithms (including Gaussian filtering, image differencing, binarization, and morphological operations). Subsequently, based on a binocular vision matching algorithm and a ray refraction tracing model, the three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system (e.g., a three-dimensional tactile point cloud) are obtained from the two-dimensional image coordinates.
[0055] The specific process for obtaining the 3D tactile point cloud coordinates of the sensor contact surface based on the principle of binocular vision will refer to... Figure 4 To provide a more detailed description.
[0056] In step S300, based on the predetermined coordinate transformation parameters of the tactile sensing coordinate system in the human wrist coordinate system, the three-dimensional spatial coordinates of multiple sampling points are transformed to the human wrist coordinate system to obtain the coordinate-transformed tactile information.
[0057] In this embodiment, the core lies in using "coordinate system transformation parameters" to achieve spatial unification of the data. These transformation parameters are predetermined constant parameters used to describe the pose relationship between the tactile sensing coordinate system of the fingertip sleeve and the coordinate system of the human hand's end joints.
[0058] During the pre-calibration process, a calibration object with a predetermined geometric shape (e.g., a sphere with a known radius R) is repeatedly grasped by a wearable hand device. An optimization function is constructed based on geometric constraints, and the coordinate system transformation parameters are obtained through optimization. Based on these parameters, the tactile point cloud data relative to the fingertip sensor obtained in step S200 can be transformed into a unified human wrist coordinate system. The specific calibration and optimization process for pre-determining the coordinate system transformation parameters (hand-wearing pose-related parameters) will be discussed in [reference needed]. Figure 5 To provide a more detailed description.
[0059] In step S400, the six-DOF pose data with time synchronization is associated with the tactile information after coordinate transformation to obtain multimodal data.
[0060] In this embodiment, the 6D pose data of each finger joint of the human hand obtained in step S100 and the tactile information (i.e., the tactile three-dimensional point cloud in the coordinate system of the human wrist) obtained in step S300 are synchronized and aligned in time series.
[0061] Through this association, the system can output multimodal data that includes "fingertip tactile sensation - hand pose", so that tactile information is no longer isolated fingertip perception, but comprehensive data that is strictly aligned with the overall hand movement state in space, thereby improving the quality of expert teaching data in robot operation tasks.
[0062] The visualization effect of representing tactile information from various fingertip tactile sensing coordinate systems in the human wrist coordinate system based on coordinate system transformation parameters will refer to... Figure 6 To provide a more detailed description.
[0063] Figure 3 It is a visualization of the 6D pose data of each finger joint in different hand gesture states, based on the human wrist coordinate system.
[0064] like Figure 3The image shows the effect of acquiring and visualizing the pose of each finger joint in real time when a person's hand is in an open gesture posture. This process specifically corresponds to step S100 as described above, and can be achieved using methods such as those described in the reference diagram. Figure 1 The data gloves 100 described perform data acquisition.
[0065] In a specific embodiment of the present invention, in order to achieve Figure 3 The data glove 100, as shown in the fine-grained pose capture, is configured with a specific sensor node distribution.
[0066] Specifically, based on the human hand model and the sequential numbering of the corresponding finger joints of each IMU, the data glove has a total of 20 nodes, distributed as follows: 1 wrist node, used to establish the coordinate system reference of the human wrist; 5 palm nodes, used to capture the overall shape of the palm; 2 thumb nodes, used to capture the movement of the thumb; and 3 nodes for each of the other four fingers, used to capture the movement of each finger joint of the index, middle, ring, and little fingers.
[0067] During data acquisition, the IP address and port number are first specified in the host computer software to establish communication and begin continuous data broadcasting. Subsequently, the system acquires the 6D pose data of the aforementioned 20 nodes in their respective coordinate systems in real time. For example... Figure 3 As shown in the visualization interface, regardless of whether the fingers are flat or bent, the system can calculate and display their posture in the human wrist coordinate system in real time. To ensure the smoothness of motion capture and the real-time nature of the data, the data update frequency is set to 60Hz in this embodiment.
[0068] Figure 4 This is a flowchart illustrating the process of acquiring real-time contact deformation of a tactile flexible fingertip sleeve.
[0069] like Figure 4 The diagram illustrates the specific process for acquiring the coordinates of the three-dimensional tactile point cloud of the sensor contact surface based on the principle of binocular vision. This process is detailed in step S200 as described above.
[0070] In a specific embodiment of the present invention, the process includes the following steps: The deformation images of the inner surface (e.g., the inner surface of the contact elastomer) of the elastic sensing module 205 of the tactile sensor (e.g., the tactile flexible fingertip sleeve 200) are acquired in real time using a binocular camera module.
[0071] In this embodiment, firstly, a calibrated binocular camera module 204 is used, with its capture resolution set to 1280 x 720 and sampling frequency set to 60 Hz, to acquire tactile images of the inner surface of the elastic sensing module 205 in real time. When an object contacts the tactile sensor, the elastic sensing module 205 undergoes three-dimensional geometric deformation, causing the tactile markers transferred on its inner surface to correspondingly change their two-dimensional coordinates under camera observation. The system performs Gaussian filtering, image differencing, binarization, and morphological operations on the acquired binocular tactile images to extract the two-dimensional image coordinates of each tactile marker on the sensor contact surface in real time.
[0072] Next, the two-dimensional image coordinates of multiple sampling points are identified in the deformation image. Tactile markers distributed on the inner surface of the contact elastomer are configured to be identified to generate multiple sampling points.
[0073] In this embodiment, to address the matching challenge caused by the inconsistent arrangement order of tactile markers in the left and right images, this disclosure employs a partitioned sorting strategy. Based on the two-dimensional coordinates of the tactile markers, they are divided into three regions: a left annular region, a central rectangular region, and a right annular region, corresponding to the left hemisphere, central cylindrical surface, and right hemisphere of the tactile sensor in three-dimensional space, respectively. Processing is performed in the order of "central rectangle - left annular - right annular." The rectangular region follows a left-to-right, top-to-bottom order; each annular region has a defined center point, following an order of increasing center distance and increasing central angle. After sorting the tactile markers in the binocular images according to the above rules, feature point matching is performed between the left and right images to obtain two-dimensional matching point pairs of the tactile markers.
[0074] Then, based on the binocular vision matching algorithm and the light refraction tracking model, the three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system are obtained according to the two-dimensional image coordinates.
[0075] Specifically, based on the known true coordinates of the tactile sensor during transfer, the relative positional relationship of the three-dimensional coordinates of the surface tactile markers is obtained. The parameters of the light refraction tracing model are self-calibrated to determine the refractive index and the representation of the refractive surface coordinate system in the camera coordinate system. Finally, using two-dimensional matching point pairs as input, the three-dimensional spatial coordinates of the tactile markers (i.e., sampling points) in the tactile sensing coordinate system are calculated.
[0076] Figure 5 This is a flowchart illustrating the process of acquiring multimodal data based on fingertip tactile sensation and human hand pose, determining relevant parameters of human hand wear pose, and aligning multimodal data in space.
[0077] like Figure 5The diagram illustrates how to determine the constant parameters of the human hand's pose using geometric optimization methods, specifically how to pre-obtain coordinate system transformation parameters and ultimately achieve spatial alignment of multimodal data. This process is the core algorithm implementation of step S300.
[0078] To predetermine the coordinate system transformation parameters of the tactile sensing coordinate system within the human wrist coordinate system, when a person wears the tactile sensor and grasps a calibration object with a predetermined geometric shape, the pose representation of the calibration object within the human wrist coordinate system is acquired, and the three-dimensional spatial coordinates of multiple sampling points within the tactile sensing coordinate system of the tactile sensor are obtained. The three-dimensional spatial coordinates of the multiple sampling points when in contact with the calibration object are compared with their three-dimensional spatial coordinates when not in contact with the object, and the multiple sampling points are divided into contact and non-contact regions according to a preset three-dimensional deformation threshold. For example, in an embodiment, the data processing module compares the real-time obtained three-dimensional tactile point cloud coordinates with the coordinates corresponding to the initial state of the non-contact object, and calculates the Euclidean distance change between the two. A three-dimensional deformation threshold is set; if the distance change of a sampling point is greater than this threshold, it is classified as a contact region (e.g., ...). Figure 4 and Figure 5 The yellow part of the point cloud); conversely, it is divided into non-contact areas (e.g., Figure 4 and Figure 5 (The red part of the point cloud).
[0079] In a specific embodiment of the present invention, the process mainly includes the following stages: when a person wears a tactile sensor and grasps a calibration object with a predetermined geometric shape, the pose representation of the calibration object in the coordinate system of the person's wrist is obtained, and the three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system of the tactile sensor are obtained. Figure 5 The tactile point cloud shown in the figure has three-dimensional coordinates SP1, where the contact portion is indicated by the reference numeral SP. c (to indicate).
[0080] In a specific embodiment, during the calibration data acquisition and initialization phase, the operator wears tactile flexible fingertips and data gloves to grasp a rigid calibration object with a predetermined geometric shape. In this embodiment, to simplify calculations and utilize geometric symmetry, a calibration object with a known radius is selected. The sphere. During this process, two sets of data are acquired simultaneously: the pose representation of the calibration object (sphere) in the human wrist coordinate system obtained based on the data glove (i.e., the sphere center coordinates WP2).
[0081] In order to solve for the unknown "representation of the fingertip tactile sensing coordinate system in the human hand end-joint coordinate system" (i.e., the human hand wearing posture WT) JThis disclosure provides a composite geometric optimization objective function. The geometric optimization objective function characterizes: the three-dimensional spatial coordinates of the sampling points in the coordinate system to be solved (optimized wearing pose transformation parameters JT). s The transformed coordinates obtained after the transformation (e.g.) Figure 5 The transformed coordinates of the sampling points in the contact area, as shown in WP1, are represented by WP. c The objective function aims to achieve geometric consistency between the object and the surface determined based on the pose representation of the calibrated object. The objective function was constructed through a three-dimensional progressive optimization to address the problem of uneven or overlapping point cloud distribution under a single constraint.
[0082] First dimension: Distance to mean constraint Based on the geometric fact that "the points in the contact area should be located on the surface of the sphere," the first part of the objective function is used to characterize the transformed coordinates of the sampling points in the contact area (e.g., ...). Figure 5 The WP shown in c The distance from the center of the sphere and the radius of the sphere The mean of the differences between them.
[0083] Second dimension: Distance variance constraint This embodiment introduces a variance term into the objective function, which is the variance of the distance from the center of the sphere to the transformed coordinates of the sampling points in the contact area. By assigning a specific weight to this term, the fluctuation of the distance from each fingertip contact area to the center of the sphere can be evaluated and minimized, resulting in a more uniform distribution of the optimized sphere relative to each fingertip, which is closer to the real grasping situation. If the variance term is missing, the optimization algorithm may obtain a solution where the "average distance" is equal to the radius, but the spatial position is severely offset. For example, the sphere model may be incorrectly placed too close to some fingertips and too far away from others, causing the physical meaning of the contact depth of each fingertip to be distorted.
[0084] Third dimension: Spatial constraints of non-contact areas This disclosure additionally introduces a penalty term to constrain the geometric location of the non-contact region. This term is constructed using the ReLU (Rectified Linear Unit) function:
[0085] in, Let be the radius of the sphere. This represents the Euclidean distance from the center of the sphere to the sampling point in the non-contact area after transformation to the human wrist coordinate system. When the non-contact point is located outside the sphere ( When a non-contact point erroneously enters the sphere ( ), the function value is 0, and no penalty is applied; when a non-contact point erroneously enters the sphere ( ), the function value is 0, and no penalty is applied. When this occurs, a positive error penalty is generated, thus "pushing" the optimization algorithm to move the non-contact region outside the sphere's boundaries. In complex spatial optimization, if this term is missing, anomalous solutions may occur where the point cloud of the non-contact region mathematically penetrates into the rigid sphere. This term applies a one-way penalty through the ReLU function: that is, it encourages non-contact points to remain outside the sphere ( The function value is 0), and once a non-contact point mistakenly enters the sphere ( The function value increases sharply, thus forcing the optimization result to satisfy the physical law of non-intrusion of rigid bodies.
[0086] According to embodiments of this disclosure, coordinate system transformation parameters can be obtained by minimizing the geometric optimization objective function. In these embodiments, the solution and application of the coordinate system transformation parameters are based on the aforementioned comprehensive geometric optimization objective function, which includes the mean, variance, and ReLU penalty term. The function value is minimized through an iterative optimization algorithm, thereby obtaining the optimal coordinate system transformation parameters.
[0087] After obtaining the aforementioned constant coordinate system transformation parameters, the system can then use these parameters to transform the real-time acquired fingertip tactile 3D point cloud data during subsequent formal teaching processes. Combined with real-time human hand joint pose Transform to the human wrist coordinate system Down:
[0088] This achieves strict alignment between tactile information and human hand pose data in the same spatial coordinate system.
[0089] Figure 6 It is a visualization of the representation of tactile information in the human wrist coordinate system under various fingertip tactile sensing coordinate systems based on coordinate system transformation parameters.
[0090] like Figure 6 As shown, this visualization demonstrates the spatial alignment effect after uniformly transforming the tactile point clouds distributed across different fingertips to the human wrist coordinate system based on the solved coordinate system transformation parameters. This visualization intuitively verifies the final output of step S400.
[0091] To further explain Figure 5 The necessity of the three constraints in the objective function of geometric optimization. Figure 6 By comparing the point cloud distribution under different optimization conditions, the impact of each constraint on the final alignment accuracy is intuitively demonstrated.
[0092] Reference Figure 6In (a) and (b), if the objective function lacks distance variance constraints and spatial constraints on the non-contact area, and optimization relies solely on the mean distance, it can be observed that the grasped calibration sphere is spatially too close to some fingertips and too far from others. This indicates that the mean alone cannot constrain the physical consistency of the contact depth. Figure 6 (c) introduces a distance variance constraint, ensuring a uniform distribution of the contact areas of each fingertip relative to the center of the sphere, further recreating the physical state of multi-finger coordinated grasping. However, although the 3D haptic point cloud is uniformly distributed on the contact surface, anomalies occur where the 3D haptic point cloud in non-contact areas (such as the sides of the fingertips) overlaps or intersects with the grasped sphere model. This is because mathematical optimization, in the absence of exclusive constraints, cannot automatically satisfy the physical law of non-intrusion of rigid bodies. (Refer to...) Figure 6 Figure (d) shows the perfect alignment after incorporating all three constraints. The contact portion of the fingertip tactile point cloud precisely fits the surface of the sphere, while the non-contact portion is strictly outside the geometric model of the sphere, proving the physical accuracy of the coordinate system transformation parameters.
[0093] In summary, Figure 6 This not only provides a visual representation of the multimodal data acquisition results, but also directly verifies the effectiveness of the alignment algorithm for multimodal data used in robot hand teaching proposed in this disclosure.
[0094] Figure 7 This is a schematic diagram of a multimodal data acquisition system for human hand teaching of robots.
[0095] like Figure 7 As shown, this embodiment provides a multimodal data acquisition system for the human hand teaching process of a robot. The multimodal data acquisition system includes: a human finger joint pose module 701, a fingertip tactile sensing module 702, a wearable pose calibration module 703, and a multimodal data alignment module 704.
[0096] The finger joint pose module 701 is configured to acquire six-DOF pose data of each finger joint in the wrist coordinate system. In its specific implementation, this module corresponds to the data glove 100 described above, responsible for receiving and processing data from the IMU sensor and outputting the finger joint movement state relative to the wrist. The finger joint pose module 701 is configured to execute step S100 as described above; redundant descriptions are omitted here.
[0097] The fingertip tactile sensing module 702 is configured to acquire tactile information from a human fingertip. The tactile information includes at least the three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system of the fingertip tactile sensing module. These multiple sampling points are used to characterize the deformation state of the elastic sensing module within the fingertip tactile sensing module. The fingertip tactile sensing module 702 corresponds to the aforementioned tactile flexible fingertip sleeve 200 and is configured to perform step S200 as described above; redundant descriptions are omitted here.
[0098] The wearable pose calibration module 703 is configured to pre-determine the coordinate system transformation parameters of the tactile sensing coordinate system in the human wrist coordinate system. This module performs the aforementioned reference... Figure 5 The method for determining coordinate system transformation parameters is described, but redundant descriptions are omitted here.
[0099] The multimodal data alignment module 704 is configured to perform spatial transformation and temporal correlation of the data. The multimodal data alignment module 704 is configured to perform steps S300 and S400 as described above, and redundant descriptions are omitted here.
[0100] Figure 8 This is a schematic diagram of the structure of the computing device provided by the present invention.
[0101] Reference Figure 8 The computing device 800 according to embodiments of the present disclosure may include a processor 810, a communication interface 820, and a memory 830. The processor 810 may include (but is not limited to) a central processing unit (CPU), a digital signal processor (DSP), a microcomputer, a field-programmable gate array (FPGA), a system-on-a-chip (SoC), a microprocessor, an application-specific integrated circuit (ASIC), etc. The memory 830 may store computer programs to be executed by the processor 810. The memory 830 includes high-speed random access memory and / or a non-volatile computer-readable storage medium. When the processor 810 executes the computer program stored in the memory 830, the multimodal data acquisition method for the human hand teaching process of a robot, as described above, can be implemented.
[0102] In addition, the communication interface 820 is used to enable communication between the computing device and other devices (such as data gloves, fingertip tactile sensors, or external servers). The communication bus 840 is used to enable communication between the processor 810, the communication interface 820, and the memory 830, ensuring smooth transmission of instructions and data between the components.
[0103] The multimodal data acquisition method for a robot's hand teaching process according to embodiments of this disclosure can be programmed into a computer program and stored on a computer-readable storage medium. When the computer program is executed by a processor, the multimodal data acquisition method for a robot's hand teaching process as described above can be implemented. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or ultra-fast digital (XD) cards), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store computer programs and any associated data, data files, and data structures in a non-transitory manner and to provide the computer programs and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer programs. In one example, the computer programs and any associated data, data files, and data structures are distributed across a networked computer system, such that the computer programs and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.
[0104] While some embodiments of this disclosure have been shown and described, those skilled in the art will understand that modifications may be made to these embodiments without departing from the principles and spirit of this disclosure, which are defined by the claims and their equivalents.
Claims
1. A multimodal data acquisition method for the human hand teaching process of a robot, characterized in that, The multimodal data acquisition method includes: Acquire the six-degree-of-freedom pose data of each finger joint in the human wrist coordinate system; Acquire tactile information from a tactile sensor located at the fingertip of a person, wherein the tactile information includes at least the three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system of the tactile sensor, and the multiple sampling points are used to characterize the deformation of the elastic sensing module of the tactile sensor; Based on predetermined coordinate transformation parameters of the tactile sensing coordinate system under the human wrist coordinate system, the three-dimensional spatial coordinates of the multiple sampling points are transformed to the human wrist coordinate system to obtain coordinate-transformed tactile information; and The six-degree-of-freedom pose data with time synchronization is associated with the coordinate-transformed tactile information to obtain the multimodal data.
2. The multimodal data acquisition method according to claim 1, characterized in that, The acquisition of tactile information from a tactile sensor located at the fingertip includes: The deformation image of the inner surface of the elastic sensing module of the tactile sensor is acquired in real time using a binocular camera module. Identify the two-dimensional image coordinates of the plurality of sampling points in the deformed image; and Based on the binocular vision matching algorithm and the light refraction tracking model, the three-dimensional spatial coordinates of the multiple sampling points in the tactile sensing coordinate system are obtained according to the two-dimensional image coordinates.
3. The multimodal data acquisition method according to claim 2, characterized in that, The following steps are used to determine the coordinate transformation parameters of the tactile sensing coordinate system in the human wrist coordinate system: When a person wears the tactile sensor on their hand and grasps a calibration object with a predetermined geometric shape, the pose representation of the calibration object in the coordinate system of the person's wrist is obtained, and the three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system of the tactile sensor are obtained. The three-dimensional spatial coordinates of the multiple sampling points when they are in contact with the calibrated object are compared with the three-dimensional spatial coordinates when they are not in contact with the object, and the multiple sampling points are divided into contact areas and non-contact areas according to a preset three-dimensional deformation threshold. A geometric optimization objective function is constructed, which is used to characterize the geometric consistency between the three-dimensional spatial coordinates of the multiple sampling points after transformation by the coordinate system transformation parameters to be solved and the surface of the calibration object determined based on the pose representation. as well as The coordinate system transformation parameters are obtained by minimizing the geometric optimization objective function.
4. The multimodal data acquisition method according to claim 3, characterized in that, The predetermined geometry is a sphere, and The geometric optimization objective function includes the mean value of the difference between the distance of the transformed coordinates of the sampling points in the contact area from the center of the sphere and the radius of the sphere.
5. The multimodal data acquisition method according to claim 4, characterized in that, The geometric optimization objective function also includes: the variance of the distance from the center of the sphere to the transformed coordinates of the sampling points in the contact area.
6. The multimodal data acquisition method according to claim 4, characterized in that, The geometric optimization objective function also includes: , Wherein, the ReLU function is a non-linear activation function, and R is the radius of the sphere. The Euclidean distance from the center of the sphere to the transformed coordinates of the sampling point in the non-contact area after being converted to the coordinate system of the human wrist.
7. A multimodal data acquisition device for the human hand teaching process of a robot, characterized in that, The multimodal data acquisition device includes: Data gloves, worn on the palm and knuckles of a human hand, are used to detect the movement state of each finger and knuckle of a human hand in order to obtain six-degree-of-freedom pose data of each finger and knuckle of a human hand in the wrist coordinate system. A tactile flexible fingertip sleeve, including an elastic sensing module, is worn on the fingertip to act as a tactile sensor to acquire tactile information from the fingertip. The tactile information includes at least the three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system of the tactile flexible fingertip sleeve, wherein the multiple sampling points characterize the deformation of the elastic sensing module; and A data processing module, communicatively connected to the data glove and the tactile flexible fingertip sleeve, is configured to perform the acquisition method as described in any one of claims 1 to 6 to obtain the multimodal data based on data acquired by the data glove and the tactile flexible fingertip sleeve.
8. The multimodal data acquisition device according to claim 7, characterized in that, The elastic sensing module includes: a contact elastomer having an inner surface and a plurality of tactile markers distributed on the inner surface; the contact elastomer is configured to deform upon contact with an object; and the plurality of tactile markers are configured to be identified to generate the plurality of sampling points. The tactile flexible fingertip sleeve also includes: A base frame is defined with wear holes for accommodating a person's fingertips, and the elastic sensing module is disposed on the base frame; A flexible wear ring, disposed on the inner wall of the wear hole, is configured to fit the fingertips of people of different sizes and provide frictional fixation force when worn; A light source illumination module, disposed on the base frame, is used to illuminate the inner surface of the contact elastomer; and A binocular camera module is fixedly mounted on the base frame, facing the inner surface of the contact elastomer, and is used to acquire deformation images of the inner surface containing the plurality of tactile markers.
9. A multimodal data acquisition system for the human hand teaching process of a robot, characterized in that, The data acquisition system includes: The finger joint pose module is configured to acquire six-degree-of-freedom pose data of each finger joint in the wrist coordinate system. A fingertip tactile sensing module is configured to acquire tactile information from a human fingertip, wherein the tactile information includes at least the three-dimensional spatial coordinates of multiple sampling points in the tactile sensing coordinate system of the fingertip tactile sensing module, and the multiple sampling points are used to characterize the deformation of the elastic sensing module of the fingertip tactile sensing module. The wearable pose calibration module is configured to pre-determine the coordinate system transformation parameters of the tactile sensing coordinate system in the coordinate system of the human wrist; and The multimodal data alignment module is configured to transform the three-dimensional spatial coordinates of the multiple sampling points to the coordinate system of the human wrist based on the coordinate system transformation parameters, thereby obtaining coordinate-transformed tactile information, and to associate the six-degree-of-freedom pose data with time synchronization with the coordinate-transformed tactile information to obtain the multimodal data.
10. A computing device, characterized in that, The computing device includes: processor; and A memory storing a computer program that, when executed by a processor, implements the multimodal data acquisition method for a robot's hand teaching process according to any one of claims 1 to 6.